Ship detection and cargo segmentation method and system based on shared network mechanism

By adopting a shared network mechanism method in ship detection and cargo segmentation tasks, and using the improved YOLOv5 network model for high-level feature extraction and feature fusion, the problem of repeated feature extraction and computing resource consumption is solved, efficient and accurate detection and segmentation effects are achieved, and suitable for resource-constrained equipment.

CN120107905APending Publication Date: 2025-06-06JIANGSU UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510172352.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art has problems such as repeated feature extraction, large computing resource consumption, excessive memory usage, low overall accuracy and inability to operate effectively on equipment with limited resources in ship detection and cargo segmentation tasks.

Method used

Using a method based on the shared network mechanism, high-level feature extraction is performed through the shared backbone network of the improved YOLOv5 network model, and feature fusion is performed by combining the path aggregation network to reduce repeated calculations and improve computing efficiency and resource utilization.

Benefits of technology

It significantly improves the accuracy of detection and segmentation, reduces computing resource consumption and processing time, improves inference speed, reduces memory usage, enables the model to run efficiently on resource-constrained devices, and improves the accuracy of cargo estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107905A_ABST
    Figure CN120107905A_ABST
Patent Text Reader

Abstract

The invention discloses a ship detection and cargo segmentation method and system based on a shared network mechanism. The method comprises the following steps: preprocessing input ship image information; performing high-level feature extraction on the preprocessed image information through a shared backbone network of an improved YOLOv5 network model; carrying out feature fusion through a path aggregation network; performing branch processing through a target detection branch and an image segmentation branch according to the fused shared features; calculating an estimated value of the cargo quantity by combining a detection result of the ship and a cargo segmentation result; and outputting a result based on a head network of YOLOv5. According to the method, the detection and segmentation precision is remarkably improved, and the calculation efficiency and the resource utilization rate are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of shipping, and relates to a ship cargo capacity detection technology, and in particular to a ship detection and cargo segmentation method and system based on a shared network mechanism. Background Art

[0002] As the most common means of freight transportation between countries, sea transportation is of great significance to ports for being able to calculate the cargo capacity of ships in a timely manner. In the field of target detection and image segmentation, especially in ship detection and cargo segmentation tasks, the current mainstream technical method is to use YOLOv5 for target detection and segmentation, but the existing target detection and image segmentation tasks usually use independent network structures for ship detection and cargo segmentation, which leads to repeated calculations of a large number of network layers. For example, in traditional detection and segmentation models, the detection network and the segmentation network extract image features separately, which consumes a lot of computing resources and time. Specifically, the traditional model requires multiple convolution calculations on the image, which increases the computing overhead by about 30% to 50%, resulting in a decrease in inference speed and affecting the ability of real-time processing.

[0003] In addition, when dealing with complex ship and cargo segmentation tasks, existing object detection models often ignore the synergy between tasks. Ship detection and cargo segmentation are closely related tasks that rely on similar image features, but existing technologies often process them separately through independent networks and fail to effectively utilize the shared feature information between the two. For example, when extracting ship position features, the traditional detection branch fails to provide accurate contextual information for the segmentation task, resulting in a decrease in segmentation accuracy. The overall accuracy of traditional methods in ship and cargo segmentation tasks is often less than 15% to 20%.

[0004] Existing detection and segmentation models require a lot of memory when running, especially in complex scenes. Since detection and segmentation tasks need to process different parts of the image separately, redundant calculations at the network layer lead to excessive memory usage. On devices with limited resources, the memory usage of traditional methods may be as high as 6GB or more, which makes it impossible to run effectively on some embedded systems or low-power devices, limiting the feasibility of its widespread application. Summary of the invention

[0005] Purpose of the invention: In order to overcome the deficiencies in the prior art, a method and system for ship detection and cargo segmentation based on a shared network mechanism is provided, which not only significantly improves the accuracy of detection and segmentation, but also greatly improves the computing efficiency and resource utilization.

[0006] Technical solution: To achieve the above purpose, the present invention provides a ship detection and cargo segmentation method based on a shared network mechanism, comprising the following steps:

[0007] S1: preprocessing the input ship image information;

[0008] S2: The preprocessed image information is passed through the shared backbone network of the improved YOLOv5 network model to extract high-level features;

[0009] S3: Combine feature maps at different levels, perform feature fusion through a path aggregation network, and fuse high-level semantic features with low-level detail features through upsampling and downsampling operations;

[0010] S4: Based on the fused shared features, branch processing is performed through the target detection branch and the image segmentation branch. For the target detection branch, based on the prediction head of YOLOv5, the shared features are used to locate and classify the ship, and the detection results including the bounding box coordinates, category labels and confidence scores are output; for the image segmentation branch, based on the lightweight segmentation network Mask R-CNN, the shared features are used to perform pixel-level segmentation of the cargo, generate the segmentation mask of the cargo, and use upsampling and skip connection technology to combine low-level detail features and high-level semantic features to complete cargo segmentation;

[0011] S5: Calculate the estimated value of the cargo volume by combining the ship's inspection results and the cargo segmentation results;

[0012] S6: Output the results based on the head network of YOLOv5.

[0013] Furthermore, the preprocessing in step S1 includes size adjustment, normalization, and data enhancement, which are specifically as follows:

[0014] The resizing is to adjust the input ship image to a fixed resolution to meet the input requirements of the improved YOLOv5 network, and use bilinear interpolation or nearest neighbor interpolation to adjust the image clarity;

[0015] Image normalization is to standardize the pixel values ​​of the image to the [0,1] interval, by subtracting the mean and dividing by the standard deviation to improve the training stability and convergence speed of the model. The calculation formula is:

[0016]

[0017] Where I is the original image, μ and σ are the mean and standard deviation of the image respectively;

[0018] Data augmentation methods include random rotation, translation, scaling, cropping, and color jittering.

[0019] Furthermore, the step S2 specifically includes:

[0020] A1: Improve CSPDarknet, integrate some structures of the cross-stage, perform residual connections and dense connections, and use the improved CSPDarknet as the shared backbone network;

[0021] A2: Utilizes a multi-layer convolutional structure with a shared backbone network to extract high-level semantic features of the input image, including a combination of multiple convolutional layers, batch normalization layers, and SiLU activation functions to gradually extract image features, including the edge, texture, and shape of the image.

[0022] Furthermore, the target detection branch in step S4 specifically includes multiple convolutional layers and prediction layers, and detects ships of different sizes through an anchor mechanism and a multi-scale detection strategy.

[0023] Furthermore, in step S5, the proportion of cargo to the hull is calculated, and the ship type coefficient is adjusted, and the cargo volume is estimated by the proportion calculation module in the improved YOLOv5 network model. The specific calculation process of the estimated value of the cargo volume is:

[0024] B1: Mask generation;

[0025] According to the results of target detection and image segmentation, the Mask M of the hull is generated respectively. ship and MaskM of goods cargo , Mask is a binary image, representing the hull and cargo areas respectively;

[0026] B2: Area calculation;

[0027] The pixel area covered by the hull Mask and cargo Mask is calculated through matrix operation to obtain the hull area A ship and cargo area A cargo , the calculation expression is as follows:

[0028]

[0029] Among them, H and W are the height and width of the image respectively;

[0030] B3: Proportional calculation;

[0031] Calculate the ratio of cargo area to hull area P, the formula is as follows:

[0032]

[0033] B4: Ship type coefficient adjustment;

[0034] The coefficient C is preset according to different ship types type , calculate the estimated value Q of the cargo volume, the formula is as follows:

[0035] Q=P×Ctype ;

[0036] Among them, C type It is a coefficient set according to different ship types and is used to correct the estimation of cargo volume.

[0037] Furthermore, in step B1, the detected ship area is filled with 1, and the remaining areas are filled with 0; the segmented cargo area is filled with 1, and the remaining areas are filled with 0.

[0038] Furthermore, the ship types in step B4 include bulk carriers, container ships, and oil tankers.

[0039] Furthermore, in step S6, the detection information of the ship, the segmentation mask of the cargo and the estimated value of the cargo quantity are integrated to generate a comprehensive detection and segmentation result, including corresponding the bounding box coordinates, category labels and confidence scores to the specific ship instance, and associating the corresponding cargo segmentation mask and cargo quantity estimate.

[0040] Furthermore, in step S6, the final detection and segmentation results are output to the port management system and the logistics scheduling system through the interface or storage system to provide real-time ship cargo volume data support, including generating a visual detection result image, a cargo segmentation mask image, and a numerical report of cargo volume estimation.

[0041] The present invention also provides a ship detection and cargo segmentation system based on a shared network mechanism, the system comprising a network interface, a memory and a processor; wherein:

[0042] The network interface is used to receive and send signals during the process of sending and receiving information with other external network elements;

[0043] The memory is used to store computer program instructions that can be executed on the processor;

[0044] The processor is used to execute the steps of a ship detection and cargo segmentation method based on a shared network mechanism when running the computer program instructions.

[0045] Beneficial effects: Compared with the prior art, the present invention has the following advantages:

[0046] (1) By introducing a shared backbone network structure, the repeated feature extraction process in the ship detection and cargo segmentation tasks is avoided, the computational overhead is reduced by about 30% to 50%, and the network's computing resource consumption and processing time are significantly reduced. Shared feature extraction increases the overall inference speed of the model by about 1.5 times, ensuring that the system has real-time processing capabilities and can perform ship detection and cargo segmentation more efficiently in practical applications. The shared backbone network reduces the redundant calculation of the model and reduces the overall memory usage by about 25%, allowing the model to run efficiently on resource-constrained devices.

[0047] (2) By adopting an efficient proportion calculation method and combining it with the adjustment of ship type coefficient, the accuracy of cargo volume estimation has been improved by 15% to 20%. This method can accurately reflect the distribution characteristics of cargo under different ship types and meet diverse practical needs. By introducing the adjustment of ship type coefficient, the model can adapt to the cargo distribution and storage methods of different types of ships (such as bulk carriers, container ships, oil tankers, etc.), further improving the accuracy and applicability of cargo volume estimation. Accurate cargo volume estimation provides reliable data support for port management and logistics scheduling, improves the scientificity and efficiency of resource allocation, and reduces operating costs.

[0048] (3) Through weight pruning and channel pruning, the model parameters are reduced by more than 30%, significantly reducing the storage requirements of the model. The inference speed of the pruned model is increased by about 2 times, while maintaining high detection and segmentation accuracy, ensuring the efficiency and practicality of the model. The lightweight backbone network ShuffleNet and the efficient convolution module GhostModule are adopted, and the overall computational workload is reduced by 40%, significantly improving the operation efficiency of the model. The optimized network structure reduces the computational overhead while maintaining rich feature expression capabilities, ensuring high accuracy of detection and segmentation tasks. Through multi-threading and asynchronous computing technology, the inference speed is increased by about 1.3 times; the optimization of hardware platforms such as GPU and FPGA further improves the inference efficiency of the model on different hardware, reduces the cache miss rate, and improves the data transmission efficiency, so that the model can still run stably and efficiently under high load. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 is a flow chart of the method of the present invention;

[0050] Figure 2 It is a schematic diagram of the structure of the improved YOLOv5 network model in the present invention;

[0051] Figure 3 Schematic diagram of the structure of the GhostConv module in the present invention;

[0052] Figure 4 It is a structural diagram of the C3Ghost module in the present invention;

[0053] Figure 5 It is a structural schematic diagram of the CBS module in the present invention;

[0054] Figure 6 It is a schematic diagram of the structure of the GhostBottleneck module in the present invention;

[0055] Figure 7 It is a schematic diagram of the structure of the BasicRFB module in the present invention. DETAILED DESCRIPTION

[0056] The present invention is further explained below in conjunction with the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, various equivalent forms of modifications to the present invention by those skilled in the art all fall within the scope defined by the claims attached to this application.

[0057] Embodiment 1:

[0058] like Figure 1 As shown, this embodiment provides a ship detection and cargo segmentation method based on a shared network mechanism, comprising the following steps:

[0059] S1: preprocessing the input ship image information;

[0060] Preprocessing includes size adjustment, normalization, and data enhancement, as follows:

[0061] The resizing is to adjust the input ship image to a fixed resolution to meet the input requirements of the improved YOLOv5 network, and use bilinear interpolation or nearest neighbor interpolation to adjust the image clarity;

[0062] Image normalization is to standardize the pixel values ​​of the image to the [0,1] interval, by subtracting the mean and dividing by the standard deviation to improve the training stability and convergence speed of the model. The calculation formula is:

[0063]

[0064] Where I is the original image, μ and σ are the mean and standard deviation of the image respectively;

[0065] Data enhancement: Due to the limited size of the ship dataset, there may be problems such as insufficient data for different ship types. Therefore, data enhancement techniques such as random rotation, translation, scaling, cropping, and color jittering are used to increase the diversity of training samples and improve the generalization ability of the model.

[0066] S2: Shared feature extraction: The preprocessed image information is extracted through the shared backbone network of the improved YOLOv5 network model at a high level to reduce repeated calculations and improve computational efficiency; specifically, it includes:

[0067] A1: Improve CSPDarknet, integrate the cross-stage part (CSP) structure, perform residual connection and dense connection, and use the improved CSPDarknet as the shared backbone network; improve the efficiency and expression ability of feature extraction, enhance the transmission and reuse of features through residual connection and dense connection, and improve the depth and width of the network;

[0068] A2: High-level feature extraction: Utilize the multi-layer convolutional structure of the shared backbone network to extract high-level semantic features of the input image, including a combination of multiple convolutional layers, batch normalization layers, and SiLU activation functions to gradually extract image features, including the edge, texture, and shape of the image.

[0069] Extract image features step by step: The following step-by-step description of how each convolution layer, batch normalization layer, SiLU activation function, and their combination extract image features at different stages:

[0070] The input image undergoes a preliminary convolution operation. This convolution layer uses a 3*3 convolution kernel to capture the basic local features of the image, such as edges and simple shapes. Features at this stage may include the outline of an object or the transition area of ​​different colors;

[0071] The convolution result will pass through the batch normalization layer, which aims to standardize the data and solve the problem of gradient disappearance or explosion during training, thereby accelerating model training. Subsequently, the SiLU activation function is used for nonlinear transformation to enhance the network's ability to learn complex features.

[0072] At this stage, the network continues to use convolutional layers to extract more complex features, such as texture and shape. The size of the convolution kernel is increased and the stride is set higher, further deepening the network's level of abstraction. By stacking multiple convolutional layers, the network is able to understand more image features, such as texture details, edge contours of objects, etc.

[0073] The network may have extracted deeper features through multiple convolutional layers. These features may represent more complex image information, such as the structure, spatial relationship, and more detailed shapes of objects. As the depth of the network increases, the features are gradually abstracted from low-level edges, textures, etc. to higher-level semantic understanding.

[0074] After being processed by multiple convolutional layers and activation functions, the network will output multi-dimensional feature maps. These feature maps combine different levels of image information: edges, textures, shapes, and may even include the location information of objects. By combining features at different levels, the network can gradually achieve a holistic understanding of the image.

[0075] The final output is a high-level semantic feature map, which combines all the features extracted previously and can represent the semantic content of the image. Through these high-level features, the network can make more accurate predictions for classification, detection or other high-level tasks.

[0076] S3: Combine feature maps at different levels, perform feature fusion through a path aggregation network, and fuse high-level semantic features with low-level detail features through upsampling and downsampling operations;

[0077] S4: Task branch processing: According to the fused shared features, branch processing is performed through the target detection branch and the image segmentation branch respectively;

[0078] 1. For the target detection branch, based on the prediction head of YOLOv5, shared features are used to locate and classify ships, and the detection results including bounding box coordinates, category labels and confidence scores are output; specifically, multiple convolutional layers and prediction layers are used to achieve accurate detection of ships of different sizes through anchor mechanism and multi-scale detection strategy;

[0079] Specific implementation method:

[0080] 1. Feature extraction and shared features: In the network structure of YOLOv5, the features of the input image are first extracted through the backbone network (such as CSPDarknet). The network gradually extracts high-level feature maps through multiple convolutional layers and activation functions. These feature maps are shared with the prediction head at different scales for target positioning and classification.

[0081] 2. Prediction head and multi-scale output: YOLOv5 uses a multi-scale output design, that is, it performs object detection at different levels (small, medium, and large) to adapt to objects of different sizes. This is achieved through the structure of the prediction head, which extracts the location, category label, and confidence score of the target from the feature map.

[0082] At each scale, YOLOv5 predicts each grid cell. Each cell is responsible for detecting the target at that location in the image. The prediction results of each cell include:

[0083] Bounding box coordinates: expressed as (x, y) (the coordinate offset of the center of the object relative to the cell), w (the width of the object), and h (the height of the object).

[0084] Category label: The probability distribution of the category to which the target belongs.

[0085] Confidence score: c n Indicates the confidence that the nth cell detects the target.

[0086] The specific formula is:

[0087] Prediction=(x,y,w,h,c 1 ,c 2 ,…,c n )

[0088] 3. Anchor mechanism: The anchor mechanism is a key step in YOLOv5 object detection. The network predefines some fixed anchor boxes (AnchorBoxes) during the training phase. These anchor boxes have different aspect ratios. Each anchor box is used to predict the bounding box of the target.

[0089] For each anchor box, the network outputs the relative offset to the anchor box, and the bounding box position of the target is adjusted by the anchor box. For each cell, the network predicts the offsets of multiple anchor points.

[0090] The specific bounding box coordinate adjustment formula is usually:

[0091] Δx=σ(tx)·p w +p x

[0092] Δy=σ(ty)·p h +p y

[0093] Δw=p w ·exp(tw)

[0094] Δh=p h ·exp(th)

[0095] Among them: (p x ,p y ) is the center coordinate of the anchor point, p w ,p h are the width and height of the anchor; tw, th are the relative offsets predicted by the network; σ(·) is an activation function used to limit the predicted value between 0 and 1; these formulas adjust each anchor box to the bounding box predicted by the network through the anchor mechanism, thereby achieving accurate target positioning.

[0096] 4. Multi-scale detection: YOLOv5 adopts a multi-scale detection strategy to detect objects at different layers of the network (e.g., low-level feature maps, mid-level feature maps, and high-level feature maps). Low-level feature maps are suitable for detecting small objects, while high-level feature maps are suitable for detecting large objects. Through multi-scale detection, the network can achieve better detection results on objects of different sizes.

[0097] In the implementation of YOLOv5, this strategy is implemented through multiple feature pyramids and different prediction head layers. The specific multi-scale feature maps will process objects of different sizes respectively. The output of each scale can be optimized through the corresponding loss function, thereby improving the overall detection accuracy of the model.

[0098] 5. Final detection results: Finally, YOLOv5 will merge the prediction results of all scales, and use post-processing steps (such as non-maximum suppression, NMS) to remove duplicate boxes, retain the best prediction box, and output the final target detection results.

[0099] Second, for the image segmentation branch, based on the lightweight segmentation network MaskR-CNN, shared features are used to perform pixel-level segmentation of goods, generate segmentation masks of goods, and use upsampling and skip connection technology to combine low-level detail features and high-level semantic features to complete high-precision segmentation of goods;

[0100] The core task of MaskR-CNN is to generate pixel-level segmentation masks for each object region. For each candidate box, Mask R-CNN uses a small fully convolutional network (FCN) to generate a segmentation mask.

[0101] This branch is independent of the shared feature network of the object detection part, but relies on the ROIAlign operation to extract the features of each candidate region and generate the segmentation mask of the target through the convolution operation. The mask is a binary image in which each pixel value is 1 for the target part and 0 for the background.

[0102] The specific generation steps are as follows:

[0103] 1. ROIAlign: First, for each candidate region, ROIAlign is used to accurately extract the features of the region from the shared convolutional feature map. The ROIAlign operation not only crops the feature map, but also avoids the accuracy loss caused by the pooling operation.

[0104] The process of ROIAlign can be expressed by the following formula:

[0105]

[0106] Among them, feature(i,j) represents the value of a specific position extracted from the feature map, and mask(i,j) is the corresponding segmentation mask.

[0107] 2. Upsampling and skip connections: In order to solve the problem of detail loss of low-level features, Mask R-CNN uses upsampling and skip connection techniques. Skip connections combine low-level detail features with high-level semantic features so that the network can generate more refined segmentation masks.

[0108] This process mainly combines the detailed features from the lower layers with the semantic features extracted from the higher layers through skip connections. Skip connections help maintain the spatial information of the image while improving the resolution and accuracy of the image.

[0109] The upsampling process is usually performed through transposed convolution, which can be expressed as follows:

[0110]

[0111] Among them, x m,n is the value of the input feature map, w m,n is the weight of the convolution kernel, y i,j is the output image after upsampling.

[0112] 3. Mask generation: Through convolution and upsampling operations, MaskR-CNN generates a pixel-level segmentation mask for each candidate box. These masks are combined with the bounding box of the target detection part, and finally output a complete detection result containing the target category label, location (bounding box coordinates) and segmentation mask.

[0113] 4. Loss function: In order to optimize the segmentation network, MaskR-CNN designed a comprehensive loss function consisting of three parts:

[0114] Object detection loss: mainly includes bounding box regression loss (smooth L1 loss) and classification loss (cross entropy loss).

[0115] Segmentation loss: Use the cross entropy loss function to calculate the difference between the predicted segmentation mask and the true mask. The loss formula is:

[0116]

[0117] Among them, mask ture (i,j) is the true label of the target mask, mask pred (i,j) is the segmentation mask predicted by the network. Total loss: The weighted sum of all loss terms gives the final loss:

[0118] L total =Lbox +L cls +L mask

[0119] Among them, L box and L cls are the losses for bounding box regression and classification, respectively.

[0120] S5: Proportion calculation module: Combine the ship detection results and cargo segmentation results to calculate the proportion of cargo to the ship body, adjust the ship type coefficient, and estimate the cargo volume through the proportion calculation module in the improved YOLOv5 network model. The specific calculation process of the estimated value of the cargo volume is as follows:

[0121] 1) Mask generation;

[0122] According to the results of target detection and image segmentation, the Mask M of the hull is generated respectively. ship and MaskM of goods cargo , Mask is a binary image, representing the hull and cargo areas respectively;

[0123] The detected ship area is filled with 1 and the rest of the area is filled with 0; the segmented cargo area is filled with 1 and the rest of the area is filled with 0.

[0124] 2) Area calculation;

[0125] The pixel area covered by the hull Mask and cargo Mask is calculated through matrix operation to obtain the hull area A ship and cargo area A cargo , the calculation expression is as follows:

[0126]

[0127] Among them, H and W are the height and width of the image respectively;

[0128] 3) Proportional calculation;

[0129] Calculate the ratio of cargo area to hull area P, the formula is as follows:

[0130]

[0131] 4) Adjustment of ship type coefficient;

[0132] Ship types include bulk carriers, container ships, and oil tankers; coefficients C preset according to different ship types type , calculate the estimated value Q of the cargo volume, the formula is as follows:

[0133] Q=P×C type ;

[0134] Among them, C typeIt is a coefficient set according to different ship types and is used to correct the estimation of cargo volume.

[0135] S6: Output the results based on the head network of YOLOv5:

[0136] The ship detection information, cargo segmentation mask and cargo quantity estimate are integrated to generate comprehensive detection and segmentation results, including mapping bounding box coordinates, category labels and confidence scores to specific ship instances, and associating corresponding cargo segmentation masks and cargo quantity estimates.

[0137] The final detection and segmentation results are output to the port management system, logistics scheduling system, etc. through the interface or storage system, providing real-time ship cargo volume data support, including the generation of visual detection result images, cargo segmentation mask images and numerical reports of cargo volume estimation for relevant departments to view and analyze.

[0138] The following combination Figure 2 to Figure 7 The improved yolov5 network model provided by the present invention is described in detail:

[0139] like Figure 2 As shown, the improved yolov5 network model includes a backbone network, a neck network connected to the backbone network, and a head network connected to the neck network;

[0140] The backbone network includes an image input module, a first CBS module, a first GhostConv module, a first C3Ghost module, a second GhostConv module, a second C3Ghost module, a third GhostConv module, a third C3Ghost module, a fourth GhostConv module and a fourth C3Ghost module which are connected in sequence; the input end of the image input module serves as the image input end of the backbone network; the first ship feature output end of the second C3Ghost module and the second ship feature output end of the third C3Ghost module are connected to out3 and out4 in sequence; the third ship feature output end of the fourth C3Ghost module is connected to the BasicRFB module;

[0141] The backbone network in this invention uses a modular structure and efficient convolution methods (GhostConv and C3Ghost), as well as receptive field expansion (BasicRFB module). The network can fully extract high-level and low-level features of the image while maintaining low computational overhead, and achieve accurate understanding and processing of complex images. The core goal of this innovative design is to improve the network's computational efficiency, feature extraction capabilities, and ability to adapt to different scales.

[0142] The neck network includes the first splicing + convolution layer * 5 modules, the second splicing + convolution layer * 5 modules, the third splicing + convolution layer * 5 modules, the fourth splicing + convolution layer * 5 modules, the first convolution + upsampling module, the second convolution + upsampling module, the first downsampling module and the first convolution + downsampling module;

[0143] out3 is connected with the channel attention mechanism and input to the first splicing + convolution layer * 5 module through a two-layer pyramid structure, and is sequentially connected to the first downsampling module, the fourth splicing + convolution layer * 5 module, the first convolution + downsampling module, the third splicing + convolution layer * 5 module, the second convolution + upsampling module, the second splicing + convolution layer * 5 module, the first convolution + upsampling module, and finally reconnected to the initial first splicing + convolution layer * 5 module to form a positive pyramid and an inverted pyramid splicing structure;

[0144] out4 is connected with the channel attention mechanism and input into the second concatenation + convolution layer * 5 module in the neck network;

[0145] The present invention achieves deep fusion of information by splicing low-level and high-level features and further processing them through convolutional layers. Detailed information is restored through upsampling, and contextual information of a wider range is obtained through downsampling, which enhances the processing capability of multi-scale features. Each module enhances the nonlinear representation capability and improves the recognition accuracy of complex targets by stacking multiple convolutional layers.

[0146] The BasicRFB module is connected to the first concatenation + convolution layer * 3 modules, and then connected to out5 and the channel attention mechanism in turn to access the second convolution + upsampling module in the neck network;

[0147] In target detection and image segmentation tasks, the BasicRFB module is usually used to expand the network's receptive field and enhance its perception of different scales and contextual information. Its core function is to improve the network's performance in processing multi-scale targets through specific module design, making it more efficient in capturing image details and global information. The BasicRFB module's functions include:

[0148] 1. Expand the receptive field: The receptive field refers to the area of ​​the input image that a layer or a neuron in the network can perceive. Shallower convolutional layers in the network usually have smaller receptive fields, which means that they can only capture local, smaller features (such as edges and textures).

[0149] 2. Multi-scale feature extraction: In an image, objects or targets of different sizes may require receptive fields of different sizes to effectively capture them. The BasicRFB module extracts multi-scale features in an image by using multiple convolutional layers with different convolution kernel sizes and strides.

[0150] 3. Combining local and global information

[0151] Local information: Through the operation of small convolution kernels, the BasicRFB module can focus on details, such as local features such as the contour and texture of the object.

[0152] Global information: Through larger convolution kernels or operations that expand the receptive field, the BasicRFB module can capture a wider range of contextual information in the image, such as the relationship between objects and the background or the spatial relationship between objects.

[0153] Combining local and global information: The BasicRFB module can improve detection accuracy while maintaining sensitivity to details by combining local and global information. Combining local and global information is particularly important for complex backgrounds, multiple targets, and situations where the target and background are similar.

[0154] 4. Improve feature expression capabilities: Through the combination of convolution kernels of different sizes, the BasicRFB module can extract features of different scales and integrate these features together. The network no longer relies solely on features of a single scale, but can effectively integrate feature information of multiple scales to improve the overall feature representation capabilities.

[0155] The first splicing + convolutional layer * 5 module, the third splicing + convolutional layer * 5 module and the fourth splicing + convolutional layer * 5 module in the neck network are used as output ends and are connected to the first prediction head, the second prediction head and the third prediction head in the head network in sequence;

[0156] like Figure 3 As shown, the first GhostConv module, the second GhostConv module, the third GhostConv module, the fourth GhostConv module, the fifth GhostConv module, the sixth GhostConv module, the seventh GhostConv module and the eighth GhostConv module all adopt the GhostConv structure, and each module includes a first convolution module, an identity mapping channel and a plurality of feature channels. The input end of the first convolution module serves as the input end of the first ship image, and its output end is respectively connected to one end of the identity mapping channel and one end of each feature channel. The other end of the identity mapping channel is connected to the other end of each feature channel in turn. In this embodiment, the convolution kernel size of the first convolution module is 1×1.

[0157] The first convolution module performs convolution operation on the input image to generate the first ship feature map; the identity mapping channel extracts features from the feature map to obtain the intrinsic feature map; each feature channel extracts features from the first ship feature map respectively, and splices the extracted results to form a Ghost feature map. Subsequently, the intrinsic feature map is connected with the Ghost feature map, and then activated by batch normalization and Mish activation function in turn, and finally the GhostConv ship feature map is obtained.

[0158] The calculation expression of the Mish activation function is as follows:

[0159] f(x″)=x″*tanh(SoftPlus(x″));

[0160]

[0161] Where f(x") represents the Mish activation function, x" represents the batch normalization result after the intrinsic feature map and the Ghost feature map are concatenated, tanh(·) represents the tanh activation function, SoftPlus(·) represents the SoftPlus activation function, e represents the exponential basis constant, and log(·) represents the logarithmic function.

[0162] like Figure 4 As shown, the first to eighth C3Ghost modules all adopt the C3Ghost structure, and each module includes a second CBS module, several GhostBottleneck modules connected in sequence, a third CBS module, a fifth Concat splicing module and a fourth CBS module. Specifically, the input end of the second CBS module is connected to the input end of the third CBS module and serves as the feature input end of the C3Ghost module. The output end of the second CBS module is connected to the input end of the first GhostBottleneck module, and the output end of the third CBS module is connected to the first input end of the fifth Concat splicing module. At the same time, the output end of the nth GhostBottleneck module is connected to the second input end of the fifth Concat splicing module. The spliced ​​feature map is output through the fifth Concat splicing module and passed to the input end of the fourth CBS module, and finally output by the fourth CBS module as the feature output end of the C3Ghost module, where n is a positive integer, indicating the number of GhostBottleneck modules in each C3Ghost module;

[0163] In each C3Ghost module, the second CBS module first performs preliminary processing on the input features and passes the processed features to the GhostBottleneck module for further feature extraction. The GhostBottleneck module efficiently extracts and processes features through a lightweight GhostConv structure, reducing computational overhead. The third CBS module further processes the features from the second CBS module to prepare for feature splicing. The output of the nth GhostBottleneck module and the output of the third CBS module are spliced ​​in the fifth Concat splicing module to integrate multi-level feature information. The spliced ​​feature map is processed by the fourth CBS module, and finally outputs high-quality features for subsequent detection and segmentation tasks.

[0164] like Figure 5 As shown in the figure, the first CBS module, the second CBS module, the third CBS module and the fourth CBS module all adopt the CBS structure, and each module includes a second convolution module, a batch normalization (BN) layer and a SiLU activation module connected in sequence. Specifically, the second convolution module performs a convolution operation on the input image to generate a second feature map; then, the BN layer normalizes the feature map to extract and standardize features; finally, the SiLU activation module performs nonlinear activation on the normalized feature map through the SiLU activation function to form the final CBS feature map.

[0165] Each CBS module operates sequentially according to the above process: First, the second convolution module is responsible for performing preliminary convolution extraction on the input image and generating the corresponding feature map. Then, the BN layer normalizes the feature map of the convolution output to improve the training stability and performance of the model. Finally, the SiLU activation function introduces nonlinear transformation to further enhance the expressiveness of the feature map and generate an activated CBS feature map. This modular design not only optimizes the feature extraction process, but also improves the expressiveness and accuracy of the overall network through normalization and activation steps, thereby enhancing the performance of the improved YOLOv5 in ship detection and cargo segmentation tasks.

[0166] like Figure 6As shown in the figure, the GhostBottleneck structure consists of multiple GhostBottleneck modules, each of which contains the ninth GhostConv module, the tenth GhostConv module and the first Add module in sequence. It is worth noting that the design of the ninth GhostConv module and the tenth GhostConv module is the same as the first GhostConv module. First, the input image is feature extracted by the ninth GhostConv module, and then further processed by the batch normalization (BN) layer and nonlinearly activated by the Mish activation function. This activated feature map is then input to the tenth GhostConv module for further feature extraction and BN processing.

[0167] In the tenth GhostConv module, the image after feature extraction and BN processing is activated again by the Mish activation function. Finally, the output feature map of the tenth GhostConv module and the output feature map of the ninth GhostConv module are added element by element through the first Add module to generate the final GhostBottleneck feature map. This design not only enhances the expressive power of the features, but also improves the overall stability and learning efficiency of the network through the residual connection between modules, thereby optimizing the performance of the GhostBottleneck module in the feature extraction process.

[0168] like Figure 7As shown, the BasicRFB module in the BasicRFB structure consists of a third convolution module, a fourth convolution module, a fifth convolution module, a sixth convolution module, a seventh convolution module, an eighth convolution module, a ninth convolution module, a tenth convolution module, an eleventh convolution module, a Shortcut module, a sixth Concat splicing module, a second Add module, and a ReLU module. Specifically, the input ends of the third convolution module, the fourth convolution module, the fifth convolution module, and the Shortcut module are all connected to the output end of the fourth C3Ghost module as the input source of the BasicRFB module. Subsequently, the output of the third convolution module is passed to the sixth convolution module, the output of the fourth convolution module is connected to the seventh convolution module, and the output of the fifth convolution module is connected to the eighth convolution module. Then, the output of the sixth convolution module is passed to the ninth convolution module, and the output of the seventh convolution module is connected to the tenth convolution module. The outputs of the eighth, ninth, and tenth convolution modules are jointly input to the sixth Concat splicing module, and the spliced ​​feature map is further passed to the eleventh convolution module. The output of the eleventh convolution module is connected to the first input of the second Add module, while the output of the Shortcut module is directly connected to the second input of the second Add module. After the second Add module adds the two feature maps, it is activated by the ReLU module, and the final output is connected to the first splicing + convolution layer * 3 module.

[0169] In the BasicRFB module, the fifth convolution module and the eighth convolution module extract features from the input image in turn to generate the first RFB feature; the third convolution module, the sixth convolution module and the ninth convolution module process the input image in turn to extract the second RFB feature; the fourth convolution module, the seventh convolution module and the tenth convolution module generate the third RFB feature in turn; the Shortcut module extracts independent features from the input image to obtain the fourth RFB feature. Subsequently, the sixth Concat module concatenates the first, second and third RFB features, and further extracts them through the eleventh convolution module to form the fifth RFB feature. The second Add module adds the fourth RFB feature to the fifth RFB feature, and activates it through the ReLU activation function module to finally generate the BasicRFB feature map. In this embodiment, the third to fifth convolution modules all use 1×1 convolution kernels, the sixth convolution module uses 3×3 convolution kernels, the seventh convolution module uses 5×5 convolution kernels, the eighth convolution module uses 3×3 convolution kernels with a dilation rate of 1, the ninth convolution module uses 3×3 convolution kernels and sets the dilation rate to 3, the tenth convolution module also uses 3×3 convolution kernels, but the dilation rate is 5, and the eleventh convolution module uses 1×1 convolution kernels again. The image input to the BasicRFB module is derived from the ship feature map output by the fourth C3Ghost module. The generated BasicRFB feature map is then passed to the first splicing + convolution layer * 3 module, and finally input to the second convolution + upsampling module in the neck network through the channel attention mechanism module.

[0170] Embodiment 2:

[0171] This embodiment provides a ship detection and cargo segmentation system based on a shared network mechanism, the system comprising a network interface, a memory and a processor; wherein the network interface is used to realize the reception and transmission of signals in the process of sending and receiving information between other external network elements; the memory is used to store computer program instructions that can be run on the processor; the processor is used to execute the steps of the above-mentioned consensus method when running the computer program instructions.

[0172] The present embodiment also provides a computer storage medium, which stores a computer program, and the method described above can be implemented when the processor executes the computer program. The computer readable medium can be considered to be tangible and non-temporary. Non-limiting examples of non-temporary tangible computer-readable media include non-volatile memory circuits (such as flash memory circuits, erasable programmable read-only memory circuits or mask read-only memory circuits), volatile memory circuits (such as static random access memory circuits or dynamic random access memory circuits), magnetic storage media (such as analog or digital tapes or hard drives) and optical storage media (such as CDs, DVDs or Blu-ray discs), etc. The computer program includes processor executable instructions stored on at least one non-temporary tangible computer-readable medium. The computer program may also include or rely on stored data. The computer program may include a basic input / output system (BIOS) that interacts with the hardware of a special-purpose computer, a device driver that interacts with a specific device of a special-purpose computer, one or more operating systems, user applications, background services, background applications, etc.

[0173] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0174] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0175] Embodiment 3:

[0176] In order to verify the effectiveness and effect of the solution of the present invention, in this embodiment, a ship movement video dataset in some complex environments is used as a test set, and the ship detection network model is tested based on the test set. During the test, a comparative test is performed with the yolov5 network model and the yolov8 network model. The test results are shown in Table 1:

[0177] Table 1 Target detection test results

[0178]

[0179] In Table 1, mAP is the mean average accuracy, Recall is the recall rate, Accuracy is the accuracy, IoU is the intersection over union ratio of the model, FPS is the detection frame rate of the model, Parameters is the number of parameters of the model, and GFLOPS is the number of floating point operations per second. It can be seen from Table 1 that the improved yolov5 has greatly improved various indicators compared with the yolov5 model, which can effectively improve the accuracy and speed of ship detection, and compared with the latest model yolov8, the performance has also been greatly improved.

Claims

1. A ship detection and cargo segmentation method based on a shared network mechanism, characterized in that: The steps include: S1: preprocessing the input ship image information; S2: The preprocessed image information is passed through the shared backbone network of the improved YOLOv5 network model to extract high-level features; S3: Combine feature maps at different levels and perform feature fusion through a path aggregation network to fuse high-level semantic features with low-level detail features; S4: Based on the fused shared features, branch processing is performed through the target detection branch and the image segmentation branch. For the target detection branch, based on the prediction head of YOLOv5, the shared features are used to locate and classify the ship, and the detection results including the bounding box coordinates, category labels and confidence scores are output; for the image segmentation branch, based on the lightweight segmentation network Mask R-CNN, the shared features are used to perform pixel-level segmentation of the cargo, generate the segmentation mask of the cargo, and use upsampling and skip connection technology to combine low-level detail features and high-level semantic features to complete cargo segmentation; S5: Calculate the estimated value of the cargo volume by combining the ship's inspection results and the cargo segmentation results; S6: Output the results based on the head network of YOLOv5.

2. A ship detection and cargo segmentation method based on a shared network mechanism according to claim 1, characterized in that: The preprocessing in step S1 includes size adjustment, normalization, and data enhancement, which are specifically as follows: The resizing is to adjust the input ship image to a fixed resolution to meet the input requirements of the improved YOLOv5 network, and use bilinear interpolation or nearest neighbor interpolation to adjust the image clarity; Image normalization is to standardize the pixel values ​​of the image to the [0,1] interval, by subtracting the mean and dividing by the standard deviation to improve the training stability and convergence speed of the model. The calculation formula is: Where I is the original image, μ and σ are the mean and standard deviation of the image respectively; Data augmentation methods include random rotation, translation, scaling, cropping, and color jittering.

3. The ship detection and cargo segmentation method based on a shared network mechanism according to claim 1 is characterized in that: The step S2 specifically includes: A1: Improve CSPDarknet, integrate some structures of the cross-stage, perform residual connections and dense connections, and use the improved CSPDarknet as the shared backbone network; A2: Utilizes a multi-layer convolutional structure with a shared backbone network to extract high-level semantic features of the input image, including a combination of multiple convolutional layers, batch normalization layers, and SiLU activation functions to gradually extract image features, including the edges, textures, and shapes of the image.

4. A ship detection and cargo segmentation method based on a shared network mechanism according to claim 3, characterized in that: The shared backbone network in step A1 includes an image input module, a first CBS module, a first GhostConv module, a first C3Ghost module, a second GhostConv module, a second C3Ghost module, a third GhostConv module, a third C3Ghost module, a fourth GhostConv module and a fourth C3Ghost module connected in sequence; the input end of the image input module serves as the image input end of the backbone network; the first ship feature output end of the second C3Ghost module and the second ship feature output end of the third C3Ghost module are connected to out3 and out4 in sequence; the third ship feature output end of the fourth C3Ghost module is connected to the BasicRFB module.

5. A ship detection and cargo segmentation method based on a shared network mechanism according to claim 4, characterized in that: The composition of the shared backbone network and the extraction process of high-level features in step A2 specifically include: In the shared backbone network, the first GhostConv module, the second GhostConv module, the third GhostConv module, the fourth GhostConv module, the fifth GhostConv module, the sixth GhostConv module, the seventh GhostConv module and the eighth GhostConv module all adopt the GhostConv structure, and each module includes a first convolution module, an identity mapping channel and a plurality of feature channels; the input end of the first convolution module is used as the input end of the first ship image, and the output end thereof is respectively connected to one end of the identity mapping channel and one end of each feature channel; the other end of the identity mapping channel is sequentially connected to the other end of each feature channel; The first convolution module performs a convolution operation on the input image to generate the first ship feature map; the identity mapping channel extracts features from the feature map to obtain an intrinsic feature map; each feature channel extracts features from the first ship feature map respectively, and concatenates the extracted results to form a Ghost feature map; then, the intrinsic feature map is connected to the Ghost feature map, and then activated by batch normalization and Mish activation function in turn, and finally a GhostConv ship feature map is obtained; The calculation expression of the Mish activation function is as follows: f(x″)=x″*tanh(SoftPlus(x″)); Where f(x″) represents the Mish activation function, x″ represents the batch normalization result after the intrinsic feature map and the Ghost feature map are concatenated, tanh(·) represents the tanh activation function, SoftPlus(·) represents the SoftPlus activation function, e represents the exponential basis constant, and log(·) represents the logarithmic function; The first to eighth C3Ghost modules all adopt the C3Ghost structure, and each module includes a second CBS module, a plurality of sequentially connected GhostBottleneck modules, a third CBS module, a fifth Concat splicing module and a fourth CBS module; specifically, the input end of the second CBS module is connected to the input end of the third CBS module and serves as the feature input end of the C3Ghost module; the output end of the second CBS module is connected to the input end of the first GhostBottleneck module, and the output end of the third CBS module is connected to the first input end of the fifth Concat splicing module; at the same time, the output end of the nth GhostBottleneck module is connected to the second input end of the fifth Concat splicing module; the spliced ​​feature map is output through the fifth Concat splicing module and passed to the input end of the fourth CBS module, and finally output by the fourth CBS module as the feature output end of the C3Ghost module, where n is a positive integer, indicating the number of GhostBottleneck modules in each C3Ghost module; In each C3Ghost module, the second CBS module first performs preliminary processing on the input features and passes the processed features to the GhostBottleneck module for further feature extraction; the GhostBottleneck module uses a lightweight GhostConv structure to efficiently extract and process features and reduce computational overhead; the third CBS module further processes the features from the second CBS module to prepare for feature concatenation; the output of the nth GhostBottleneck module and the output of the third CBS module are concatenated in the fifth Concat concatenation module to integrate multi-level feature information; the concatenated feature map is processed by the fourth CBS module, and finally outputs high-quality features for subsequent detection and segmentation tasks; The first CBS module, the second CBS module, the third CBS module and the fourth CBS module all adopt the CBS structure, and each module includes a second convolution module, a batch normalization layer and a SiLU activation module connected in sequence; specifically, the second convolution module performs a convolution operation on the input image to generate a second feature map; then, the BN layer normalizes the feature map to extract and normalize the features; finally, the SiLU activation module performs nonlinear activation on the normalized feature map through the SiLU activation function to form the final CBS feature map; Each CBS module operates in sequence according to the process: first, the second convolution module is responsible for performing preliminary convolution extraction on the input image to generate the corresponding feature map; then, the BN layer standardizes the feature map output by the convolution to improve the training stability and performance of the model; finally, the SiLU activation function introduces nonlinear transformation to further enhance the expression ability of the feature map and generate an activated CBS feature map; The GhostBottleneck structure consists of multiple GhostBottleneck modules, each of which contains the ninth GhostConv module, the tenth GhostConv module, and the first Add module in sequence; first, the input image is extracted through the ninth GhostConv module, and then further processed by the batch normalization layer and nonlinearly activated by the Mish activation function; this activated feature map is then input to the tenth GhostConv module for further feature extraction and BN processing; In the tenth GhostConv module, the image after feature extraction and BN processing is activated again by the Mish activation function; finally, the output feature map of the tenth GhostConv module and the output feature map of the ninth GhostConv module are added element by element through the first Add module to generate the final GhostBottleneck feature map; The BasicRFB module in the BasicRFB structure consists of a third convolution module, a fourth convolution module, a fifth convolution module, a sixth convolution module, a seventh convolution module, an eighth convolution module, a ninth convolution module, a tenth convolution module, an eleventh convolution module, a Shortcut module, a sixth Concat module, a second Add module and a ReLU module; specifically, the input ends of the third convolution module, the fourth convolution module, the fifth convolution module and the Shortcut module are all connected to the output end of the fourth C3Ghost module as the input source of the BasicRFB module; then, the output of the third convolution module is passed to the sixth convolution module, and the output of the fourth convolution module is passed to the sixth convolution module. The output of the fifth convolution module is connected to the eighth convolution module; then, the output of the sixth convolution module is passed to the ninth convolution module, and the output of the seventh convolution module is connected to the tenth convolution module; the outputs of the eighth, ninth and tenth convolution modules are input to the sixth Concat module, and the concatenated feature map is further passed to the eleventh convolution module; the output of the eleventh convolution module is connected to the first input of the second Add module, and the output of the Shortcut module is directly connected to the second input of the second Add module; the second Add module adds the two parts of the feature map, activates them through the ReLU module, and finally connects the output to the first concatenation + convolution layer * 3 module; In the BasicRFB module, the fifth convolution module and the eighth convolution module extract features from the input image in turn to generate the first RFB feature; the third convolution module, the sixth convolution module and the ninth convolution module process the input image in turn to extract the second RFB feature; the fourth convolution module, the seventh convolution module and the tenth convolution module generate the third RFB feature in turn; the Shortcut module extracts independent features from the input image to obtain the fourth RFB feature; then, the sixth Concat module concatenates the first, second and third RFB features, and further extracts them through the eleventh convolution module to form the fifth RFB feature; the second Add module adds the fourth RFB feature to the fifth RFB feature, and activates them through the ReLU activation function module to finally generate the BasicRFB feature map.

6. A ship detection and cargo segmentation method based on a shared network mechanism according to claim 1, characterized in that: In step S4, the target detection branch specifically includes multiple convolutional layers and prediction layers, and detects ships of different sizes through an anchor mechanism and a multi-scale detection strategy; the specific implementation method includes: B1: Prediction head and multi-scale output: YOLOv5 uses a multi-scale output design, that is, object detection is performed at different levels to adapt to objects of different sizes; this is achieved through the structure of the prediction head, which extracts the location, category label, and confidence score of the object from the feature map; At each scale, YOLOv5 makes predictions for each cell; each cell is responsible for detecting the target at that location in the image; the prediction results for each cell include: Bounding box coordinates: expressed as (x, y)-the coordinate offset of the target center relative to the cell, w-the width of the target, h-the height of the target; category label: the probability distribution of the category to which the target belongs; confidence score: c n Indicates the confidence that the nth cell detects the target; B2: Anchor mechanism: The network predefines some fixed anchor boxes with different aspect ratios during the training phase; each anchor box is used to predict the bounding box of the target; For each anchor box, the network outputs the relative offset with the anchor box, and the bounding box position of the target is adjusted by the anchor box; for each cell, the network predicts the offset of multiple anchor points; The specific bounding box coordinate adjustment formula is usually: Δx=σ(tx)·p w +p x Δy=σ(ty)·p h +p y Δw=p w ·exp(tw) Δh=p h ·exp(th) Among them: (p x ,p y ) is the center coordinate of the anchor point, p w ,p h are the width and height of the anchor point; tw, th are the relative offsets predicted by the network; σ(·) is an activation function used to limit the predicted value between 0 and 1; these formulas adjust each anchor box to the bounding box predicted by the network through the anchor mechanism, thereby achieving accurate target positioning; B3: Multi-scale detection: YOLOv5 adopts a multi-scale detection strategy to detect targets at different layers of the network; the feature maps of the low layers are suitable for detecting small targets, while the feature maps of the high layers are suitable for detecting large targets; In the implementation of YOLOv5, the multi-scale detection strategy is implemented through multiple feature pyramids and different prediction head layers. The specific multi-scale feature maps will process objects of different sizes respectively; the output of each scale can be optimized by the corresponding loss function. B4: Final detection result: Finally, YOLOv5 will merge the prediction results of all scales, and use post-processing steps to remove duplicate boxes, retain the best prediction box, and output the final target detection result.

7. The method for ship detection and cargo segmentation based on a shared network mechanism according to claim 1, characterized in that: In step S4, for the image segmentation branch, the generation of the segmentation mask of the goods specifically includes: C1: ROI Align: First, for each candidate region, ROI Align is used to extract the features of the region from the shared convolutional feature map; The process of ROI Align is expressed by the following formula: Among them, feature(i,j) represents the value of a specific position extracted from the feature map, and mask(i,j) is the corresponding segmentation mask; C2: Upsampling and skip connections: Mask R-CNN uses upsampling and skip connections; skip connections combine low-level detail features with high-level semantic features so that the network can generate more refined segmentation masks; The upsampling process is performed through transposed convolution, which is expressed by the following formula: Among them, x m,n is the value of the input feature map, w m,n is the weight of the convolution kernel, y i ,j is the output image after upsampling; C3: Mask Generation: Through convolution and upsampling operations, Mask R-CNN generates a pixel-level segmentation mask for each candidate box; these masks are combined with the bounding box of the object detection part to finally output a complete detection result containing the object category label, location and segmentation mask.

8. The method for ship detection and cargo segmentation based on a shared network mechanism according to claim 1, characterized in that: The specific calculation process of the estimated value of the cargo volume in step S5 is: D1: Mask generation; According to the results of target detection and image segmentation, the Mask M of the hull is generated respectively. ship And the goods Mask M cargo , Mask is a binary image, representing the hull and cargo areas respectively; D2: Area calculation; The pixel area covered by the hull Mask and cargo Mask is calculated through matrix operation to obtain the hull area A ship and cargo area A cargo , the calculation expression is as follows: Among them, H and W are the height and width of the image respectively; D3: Proportional calculation; Calculate the ratio of cargo area to hull area P, the formula is as follows: D4: Ship type coefficient adjustment; The coefficient C is preset according to different ship types type , calculate the estimated value Q of the cargo volume, the formula is as follows: Q=P×C type ; Among them, C type It is a coefficient set according to different ship types and is used to correct the estimation of cargo volume.

9. The method for ship detection and cargo segmentation based on a shared network mechanism according to claim 1, characterized in that: In step S6, the detection information of the ship, the segmentation mask of the cargo and the estimated value of the cargo quantity are integrated to generate a comprehensive detection and segmentation result, including corresponding the bounding box coordinates, category labels and confidence scores to the specific ship instance, and associating the corresponding cargo segmentation mask and cargo quantity estimate.

10. A ship detection and cargo segmentation system based on a shared network mechanism, characterized in that: The system comprises a network interface, a memory and a processor; wherein, The network interface is used to receive and send signals during the process of sending and receiving information with other external network elements; The memory is used to store computer program instructions that can be executed on the processor; The processor is used to execute the steps of a ship detection and cargo segmentation method based on a shared network mechanism according to any one of claims 1 to 9 when running the computer program instructions.