A training method and an application method for a ship detection network based on federated learning
By building a ship detection network and a local detection network, the dual-branch attention enhancement and dynamic non-monotonic focusing method are used to solve the problems of data inhomogeneity and quality disparity in federated learning, and the accuracy of the model and data utilization efficiency are improved.
Patent Information
- Application Number
- CN202411249309.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-06
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-09-06
AI Technical Summary
The existing federated learning methods are difficult to reasonably utilize training data with unbalanced distribution and uneven quality during the training process, resulting in reduced model accuracy.
Build a ship detection network and a local detection network of multiple clients, and use dual-branch attention enhancement and feature fusion detection, combine dynamic non-monotonic focusing method to determine the predicted total loss, and perform adaptive weighted aggregation of local parameters, and iteratively update local and global parameters until the network performance is no longer improved.
On the premise of protecting data privacy, the accuracy of the ship detection network is effectively improved, and the training data of each client is reasonably utilized to balance training data of different quality and distribution.
Smart Images

Figure CN119251513B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly relates to a method for training a ship detection network based on federated learning and an application method thereof. Background Art
[0002] Visual supervision of waterway traffic plays an important role in the modern maritime field. With the rapid development of the shipping industry, the demand for waterway traffic monitoring systems is increasing day by day. An effective waterway traffic supervision system can not only improve the efficiency of ship traffic management, but also prevent accidents, ensure maritime safety, and support key tasks such as maritime law enforcement. With the rapid development of deep learning technology, waterway traffic supervision based on images and videos has been significantly improved. The powerful feature extraction ability of deep learning algorithms makes it possible to automatically identify ships and other waterway traffic targets. However, in large-scale waters, in the traditional centralized machine learning training process, it is necessary to obtain data from various departments. For the maritime field, the centralized machine learning method is difficult to meet the privacy and security requirements of each maritime supervision department.
[0003] In order to ensure the data privacy and security of each maritime supervision department (client) during the training process, federated learning, as a distributed machine learning method, has gradually attracted the attention of researchers. Federated learning allows multiple waterway traffic supervision departments to use their respective data sets to train their respective models and share the network parameters obtained from the training. On the premise of ensuring data privacy, they can jointly cooperate to train machine learning models. However, in current federated learning, due to the uneven distribution of data from different sources and the uneven quality of data, existing federated learning methods are difficult to reasonably utilize these training data with unbalanced distribution and uneven quality during the training process, thereby leading to a decrease in model accuracy.
[0004] Therefore, the technical problem that existing federated learning methods are difficult to reasonably utilize training data with unbalanced distribution and uneven quality during the training process, resulting in a decrease in model accuracy, needs to be improved. Summary of the Invention
[0005] In view of this, it is necessary to provide a method for training a ship detection network based on federated learning and an application method thereof, which are used to solve the technical problem that existing federated learning methods are difficult to reasonably utilize training data with unbalanced distribution and uneven quality during the training process, resulting in a decrease in model accuracy.
[0006] To solve the above technical problems, on the one hand, the present invention provides a method for training a ship detection network based on federated learning, including:
[0007] Constructing a ship detection network and multiple local detection networks of multiple clients;
[0008] Input the obtained ship image data into each local detection network, perform dual-branch attention enhancement and feature fusion detection on the ship image data to obtain ship detection output, determine the predicted total loss based on the dynamic non-monotonic focusing method, and update the local parameters of the local detection network according to the predicted total loss;
[0009] Adaptive weighted aggregation of the local parameters of each local detection network to obtain the global parameters of the ship detection network, update the local parameters according to the global parameters to obtain a new round of local detection networks, and iteratively update the local parameters and global parameters until the network performance no longer improves.
[0010] In a possible implementation, the ship image data includes the local ship image data obtained by each client, and the local ship image data is respectively input into the local detection network of the corresponding client.
[0011] In a possible implementation, performing dual-branch attention enhancement and feature fusion detection on the ship image data to obtain ship detection output includes:
[0012] Perform feature segmentation on the ship image data to obtain a number of input feature images;
[0013] Perform dual-branch attention enhancement on the input feature images to obtain enhanced feature images;
[0014] Perform multi-scale feature extraction and prediction output on the enhanced feature images to obtain ship detection output.
[0015] In a possible implementation, performing feature segmentation on the ship image data to obtain a number of input feature images includes:
[0016] Perform channel-dimensional feature segmentation on the ship image data after feature convolution to obtain a number of input feature images.
[0017] In a possible implementation, performing dual-branch attention enhancement on the input feature images to obtain enhanced feature images includes:
[0018] Perform height average pooling and width average pooling on the input feature images respectively to obtain a height feature map and a width feature map;
[0019] Pass the height feature map and the width feature map through a 1×1 convolution, sigmoid activation function, weighted merging, and group normalization in sequence to obtain the first-branch feature;
[0020] Perform 3×3 convolution on the input feature images to obtain the second-branch feature;
[0021] The first branch feature is successively passed through 2D average pooling and the softmax activation function, and then fused with the second branch feature through matrix multiplication to obtain the first attention feature. The second branch feature is successively passed through 2D average pooling and the softmax activation function, and then fused with the first branch feature through matrix multiplication to obtain the second attention feature;
[0022] The first attention feature and the second attention feature are merged and then passed through the sigmoid activation function to obtain the cross-space attention weight. The input feature image is weighted with the cross-space attention weight to obtain the enhanced feature image.
[0023] In a possible implementation, determining the predicted total loss based on the dynamic non-monotonic focusing method includes:
[0024] Calculating the loss of the ship detection output based on the double-layer distance attention mechanism to obtain the initial bounding box loss;
[0025] Determining the exponential moving average according to the iteration round, determining the outlier degree of the anchor box according to the exponential moving average, determining the non-monotonic focusing coefficient according to the outlier degree of the anchor box, and determining the bounding box regression loss according to the non-monotonic focusing coefficient and the initial bounding box loss;
[0026] Determining the predicted probability loss and the classification loss according to the ship detection output;
[0027] Determining the predicted total loss according to the bounding box regression loss, the predicted probability loss and the classification loss.
[0028] In a possible implementation, adaptively weighted aggregation of the local parameters of each local detection network to obtain the global parameters of the ship detection network includes:
[0029] Determining the aggregation weight according to the amount of valid data in the ship image data of each client;
[0030] Adapting and weighted aggregating the local parameters according to the aggregation weight to obtain the global parameters of the ship detection network.
[0031] On the other hand, the present invention also provides a method for applying a ship detection network, including:
[0032] Obtaining the ship picture to be detected;
[0033] Inputting the ship picture to be detected into the trained ship detection network to obtain the ship detection result;
[0034] Wherein, the trained ship detection network is determined according to the above-mentioned ship detection network training method based on federated learning.
[0035] On the other hand, the present invention also provides an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above-mentioned method for training a ship detection network based on federated learning and / or the above-mentioned method for applying a ship detection network are implemented.
[0036] On the other hand, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned method for training a ship detection network based on federated learning and / or the above-mentioned method for applying a ship detection network are implemented.
[0037] The beneficial effects of the present invention are as follows: The method for training a ship detection network based on federated learning provided by the embodiments of the present invention first constructs a ship detection network and multiple local detection networks of multiple clients; then inputs the obtained ship image data into each local detection network, performs dual-branch attention enhancement and feature fusion detection on the ship image data to obtain a ship detection output, determines the predicted total loss based on the dynamic non-monotonic focusing method, and updates the local parameters of the local detection network according to the predicted total loss; finally, adaptively weights and aggregates the local parameters of each local detection network to obtain the global parameters of the ship detection network, updates the local parameters according to the global parameters to obtain a new round of local detection networks, and iteratively updates the local parameters and global parameters until the network performance no longer improves. In the process of federated learning of the present invention, more comprehensive feature information is learned through dual-branch attention enhancement, training data of different qualities is balanced through dynamic non-monotonic focusing, and training data with different distributions is balanced through adaptive weighted aggregation of local parameters. It is possible to reasonably utilize the training data of each client on the premise of protecting the data privacy of each client, and effectively improve the accuracy of the ship detection network obtained by federated learning training. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0039] Figure 1 It is a schematic flowchart of an embodiment of the method for training a ship detection network based on federated learning provided by the present invention;
[0040] Figure 2 It is a schematic flowchart of the ship detection output in the embodiment of the present invention;
[0041] Figure 3 It is a schematic flowchart of the dual-branch attention enhancement in the embodiment of the present invention;
[0042] Figure 4 The network structure diagram of the dual-branch attention for the embodiment of the present invention;
[0043] Figure 5 The flow schematic diagram of the loss calculation for the embodiment of the present invention;
[0044] Figure 6 The flow schematic diagram of the adaptive weighted aggregation of network parameters for the embodiment of the present invention;
[0045] Figure 7 The flow schematic diagram of an embodiment of the method for applying the ship detection network provided by the present invention;
[0046] Figure 8 The structural schematic diagram of an embodiment of the electronic device provided by the present invention. Detailed implementation manners
[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0048] In the description of the embodiments of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone.
[0049] The descriptions such as "first" and "second" involved in the embodiments of the present invention are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Therefore, the technical features defined with "first" and "second" can explicitly or implicitly include at least one such feature.
[0050] Referring to "embodiment" in this article means that the specific features, structures, or characteristics described in conjunction with the embodiment can be included in at least one embodiment of the present invention. The appearance of this phrase in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0051] The present invention provides a ship detection network training method, application method, electronic device, and medium based on federated learning, which will be described separately below.
[0052] Figure 1 The flowchart of an embodiment of the ship detection network training method based on federated learning provided by the present invention is shown as Figure 1 follows. The ship detection network training method based on federated learning includes:
[0053] S101. Construct a ship detection network and multiple local detection networks of multiple clients;
[0054] S102. Input the obtained ship image data into each local detection network, perform double-branch attention enhancement and feature fusion detection on the ship image data to obtain a ship detection output, determine the total prediction loss based on the dynamic non-monotonic focusing method, and update the local parameters of the local detection network according to the total prediction loss;
[0055] S103. Adaptively weighted aggregate the local parameters of each local detection network to obtain the global parameters of the ship detection network, update the local parameters according to the global parameters to obtain a new round of local detection networks, and iteratively update the local parameters and global parameters until the network performance no longer improves.
[0056] Compared with the prior art, the ship detection network training method based on federated learning provided by the embodiment of the present invention first constructs a ship detection network and multiple local detection networks of multiple clients; then inputs the obtained ship image data into each local detection network, performs double-branch attention enhancement and feature fusion detection on the ship image data to obtain a ship detection output, determines the total prediction loss based on the dynamic non-monotonic focusing method, and updates the local parameters of the local detection network according to the total prediction loss; finally, adaptively weighted aggregate the local parameters of each local detection network to obtain the global parameters of the ship detection network, update the local parameters according to the global parameters to obtain a new round of local detection networks, and iteratively update the local parameters and global parameters until the network performance no longer improves. In the process of federated learning, the present invention can learn more comprehensive feature information through double-branch attention enhancement, balance training data of different qualities through dynamic non-monotonic focusing, and balance training data of different distributions through adaptive weighted aggregation of local parameters, so as to reasonably utilize the training data of each client on the premise of protecting the data privacy of each client, and effectively improve the accuracy of the ship detection network obtained by federated learning training.
[0057] In some embodiments of the present invention, the ship image data includes local ship image data obtained by each client, and the local ship image data is respectively input into the local detection network of the corresponding client.
[0058] Specifically, the ship image data used for model training in the embodiments includes the local ship image data of each maritime supervision department's client, which is collected by camera devices installed on different clients for shooting waterway ships. To ensure data privacy and security between clients during the training process, each local ship image data only conducts local data feature learning in the local detection network corresponding to each client itself.
[0059] In some embodiments of the present invention, Figure 2 is a schematic flowchart of the ship detection output of the embodiments of the present invention. As Figure 2 shown, the ship detection output is obtained by performing dual-branch attention enhancement and feature fusion detection on the ship image data, including:
[0060] S201. Feature-split the ship image data to obtain a number of input feature images;
[0061] S202. Perform dual-branch attention enhancement on the input feature images to obtain enhanced feature images;
[0062] S203. Perform multi-scale feature extraction and prediction output on the enhanced feature images to obtain the ship detection output.
[0063] Specifically, the embodiments construct a cross-space detection network based on the detection framework of YOLOv8. The cross-space detection network consists of four components, including input, backbone, neck, and head.
[0064] Among them, the input component part is used to process the input image to increase data and enhance the generalization ability of the model, such as cropping, transformation, and scaling, etc.
[0065] In the backbone part, feature splitting and dual-branch attention enhancement are added in the embodiments. Among them, feature splitting reshapes some channels into the batch dimension and divides them into multiple sub-features according to the channel dimension, so that multi-semantic features are distributed in each group, which can effectively prevent the disappearance of details caused by convolutional dimensionality reduction, and can achieve local cross-channel interaction in parallel sub-networks without reducing channels. Dual-branch attention enhancement then learns features for each group of features through two parallel branches. After the channel weights in each branch are updated, the outputs of the two parallel branches are aggregated using cross-space interaction to capture the enhanced feature images containing pixel-level feature information.
[0066] The neck and head are based on the framework structure of YOLOv8. The neck performs multi-scale feature extraction on the features and fuses features of different scales, and finally the head of the network outputs the regression result and class probability.
[0067] In some embodiments of the present invention, feature-splitting the ship image data to obtain a number of input feature images includes:
[0068] After convolving the features of the ship image data, feature splitting is performed in the channel dimension to obtain a number of input feature images.
[0069] Specifically, after the ship image data is input into the local detection network, for each ship image, in the embodiment, the input ship image is first subjected to convolution processing through a CBS convolution block to obtain an input feature , where the CBS convolution block consists of a convolution layer, a normalization layer, and an activation function layer. Then, the input feature is split into sub-features , which are used as the input feature images of the model.
[0070] In some embodiments of the present invention, Figure 3 is a schematic flow diagram of the dual-branch attention enhancement in the embodiments of the present invention. As shown in Figure 3 , the input feature image is subjected to dual-branch attention enhancement to obtain an enhanced feature image, including:
[0071] S301. Respectively perform height average pooling and width average pooling on the input feature image to obtain a height feature map and a width feature map;
[0072] S302. The height feature map and the width feature map are sequentially passed through a 1×1 convolution, a sigmoid activation function, weighted merging, and group normalization to obtain a first-branch feature;
[0073] S303. Perform 3×3 convolution on the input feature image to obtain a second-branch feature;
[0074] S304. The first-branch feature is sequentially passed through 2D average pooling and a softmax activation function, and then matrix multiplication is used to fuse the second-branch feature to obtain a first attention feature. The second-branch feature is sequentially passed through 2D average pooling and a softmax activation function, and then matrix multiplication is used to fuse the first-branch feature to obtain a second attention feature;
[0075] S305. The first attention feature and the second attention feature are merged and then passed through a sigmoid activation function to obtain a cross-space attention weight. The input feature image is feature-weighted according to the cross-space attention weight to obtain an enhanced feature image.
[0076] Specifically, in order to learn more comprehensive feature information, the embodiment adds a dual-branch attention module to the backbone part of the local detection network. Figure 4 is the network structure diagram of the dual-branch attention in the embodiments of the present invention. Combining with Figure 4 to see. In the dual-branch attention enhancement, the embodiment extracts attention in two branches of 1×1 and 3×3.
[0077] For the 1×1 branch, first divide the input feature image into two directions, namely the height and width directions, and distribute them for global average pooling to obtain two feature maps, which are expressed by the formula:
[0078]
[0079]
[0080] Among them, represents 's height feature map, represents 's width feature map.
[0081] Then connect the height feature map and the width feature map with the global receptive field, send the result to the shared 1×1 convolution, and after activation by the Sigmoid function, obtain the attention weights and respectively, which are expressed by the formula:
[0082]
[0083]
[0084] Among them, represents activation by the Sigmoid function, represents the 1×1 convolution operation.
[0085] Then, according to the attention weights, perform weighted calculation on the input feature image to obtain the feature maps with attention weights in the height and width directions, which are expressed by the formula:
[0086]
[0087] Subsequently, the following embodiments use group normalization to replace the conventional normalization method to process the feature image to obtain the first branch feature , so as to reduce the dependence on the batch size, and then perform 2D average pooling to obtain the global information , which is expressed by the formula:
[0088]
[0089]
[0090] Finally, apply the softmax activation function and matrix multiplication (MM) operation to fuse the features of the 3×3 branch to obtain the -dimensional first attention feature , which is expressed by the formula:
[0091]
[0092] wherein represents the second branch feature obtained from the 3×3 branch represents the softmax activation function
[0093] For the 3×3 branch, in the embodiment, the second branch feature is obtained by performing 3×3 convolution on the input feature image The formula is expressed as:
[0094]
[0095] wherein represents the 3×3 convolution operation
[0096] Then, a 2D average pooling operation is performed on the second branch feature to obtain a pooling result with a dimension of The formula is expressed as:
[0097]
[0098] Finally, after passing the pooling result through the softmax activation function, a matrix multiplication operation is performed with the first branch feature of the 1×1 branch to obtain a second attention feature with a dimension of The formula is expressed as:
[0099]
[0100] After completing the attention extraction of the two branches, in the embodiment, the output results of the two branches and are combined, and then passed through the sigmoid activation function to obtain the cross - spatial attention weight. Finally, the input feature image is weighted according to the cross - spatial attention weight to obtain the enhanced feature image of the final output The formula is expressed as:
[0101]
[0102] In summary, in the dual-branch attention module of the present invention, attention is extracted from two branches of 1×1 and 3×3 respectively. Among them, the 1×1 branch performs global average pooling on the features in both the height and width directions, and after concatenation, it passes through a shared 1×1 convolution and is activated by the sigmoid function. The first-branch features are obtained through weighted merging and group normalization. The 3×3 branch performs 3×3 convolution on the input feature image to obtain the second-branch features, and the features obtained from both branches are averaged and pooled and then fused with the other branch through matrix multiplication after passing through the activation function to obtain the attention features of both branches. In this way, the cross-space attention weights of the input feature image are determined, enabling the detection network to learn more comprehensive feature information and effectively improving the training effect of the ship detection network.
[0103] In some embodiments of the present invention, Figure 5 is a schematic flowchart of the loss calculation of the embodiments of the present invention. As Figure 5 shown, determining the total prediction loss based on the dynamic non-monotonic focusing method includes:
[0104] S501. Calculate the initial bounding box loss for the ship detection output based on the double-layer distance attention mechanism;
[0105] S502. Determine the exponential moving average according to the iteration round, determine the outlier degree of the anchor box according to the exponential moving average, determine the non-monotonic focusing coefficient according to the outlier degree of the anchor box, and determine the bounding box regression loss according to the non-monotonic focusing coefficient and the initial bounding box loss;
[0106] S503. Determine the prediction probability loss and the classification loss according to the ship detection output;
[0107] S504. Determine the total prediction loss according to the bounding box regression loss, the prediction probability loss, and the classification loss.
[0108] Specifically, the total prediction loss of the local detection network includes the bounding box regression loss , the prediction probability loss and the classification loss . The total prediction loss is obtained by the weighted sum of each loss, and the formula is expressed as:
[0109]
[0110] Among them, , and correspond to the weight values of the three losses respectively.
[0111] Considering that the quality of training data is uneven, in order to balance training data of different qualities and improve the training effect of the local detection network, in the embodiment, the bounding box regression loss is optimized by a non-monotonic focusing coefficient. For the bounding box regression loss , the embodiment first adopts a double-layer distance attention mechanism to obtain the initial bounding box loss, which is expressed by the formula:
[0112]
[0113] where and represent the abscissa and ordinate of the center point of the predicted box, and represent the abscissa and ordinate of the center point of the ground truth box, and are the width and height of the minimum detection box, is the intersection over union loss.
[0114] Based on the initial bounding box loss, in order to improve the detection accuracy, the embodiment uses the outlier degree of the anchor box and constructs a non-monotonic focusing coefficient by using it, which is expressed by the formula:
[0115]
[0116]
[0117] where is the exponential moving average with momentum , is the gradient gain, which is controlled by the hyperparameters and to map the outlier degree and the gradient gain . The smaller the value, the higher the quality of the anchor box. In addition, in order to avoid ignoring low-quality anchor boxes in the early training, a momentum is designed. The embodiment designs a momentum to delay the time when the exponential moving average approaches the true value. During the training process, when the batch size is , the momentum is:
[0118]
[0119] where is the number of training iterations.
[0120] Then the bounding box regression loss can be expressed as:
[0121]
[0122] In this way, in the middle and late stages of training, small gradient gains can be assigned to low-quality anchor boxes to reduce harmful gradients.
[0123] For the prediction probability loss , it can be expressed by the formula:
[0124]
[0125] where, is the predicted probability of the model for the target class. If the sample is a positive sample, then , if the sample is a negative sample, then , is the balance factor for adjusting the balance between positive and negative samples, represents the parameter for adjusting the balance between easy and difficult samples, is the weighting term for difficult-to-distinguish samples, aiming to increase the loss weight of difficult-to-distinguish samples.
[0126] For the classification loss , it can be expressed by the formula:
[0127]
[0128] where, is the predicted IoU-aware classification score, is the target score, representing a soft label related to the IOU. When , there is no hyperparameter, that is, there is no attenuation. While when , negative samples have a hyperparameter term, which will reduce the contribution of negative samples, is a hyperparameter to prevent over-suppression, but generally still reduces the contribution of negative samples.
[0129] In some embodiments of the present invention, Figure 6 is a schematic flow diagram of the adaptive weighted aggregation of the network parameters in the embodiments of the present invention. As Figure 6 shown, the local parameters of each local detection network are adaptively weighted and aggregated to obtain the global parameters of the ship detection network, including:
[0130] S601. Determine the aggregation weight according to the amount of valid data in the ship image data of each client;
[0131] S602. Adaptively weight and aggregate the local parameters according to the aggregation weight to obtain the global parameters of the ship detection network.
[0132] Specifically, considering the problems of unbalanced data distribution and non-independent data distribution among different clients, in the process of aggregating local parameters of each local detection network to obtain global parameters, the embodiment proposes an adaptive weighted aggregation model. The adaptive weighted aggregation model sets weights according to the effective quantity in each client's dataset to ensure that clients with a larger amount of effective data in the dataset receive more attention. For clients with a smaller effective dataset, these datasets are considered less common in the real world, and the present invention reduces their influence on the global model. The formula for adaptive weighted aggregation is expressed as:
[0133]
[0134] Wherein, represents the total number of all local datasets, represents the client the amount of effective data in the dataset. represents the parameters of the previous round distributed by the aggregator to the client
[0135] The embodiment continuously performs local training on the local detection network to obtain local parameters, then adaptively weights and aggregates the local parameters to obtain the global parameters of the ship detection network, and then updates the local parameters of each local detection network through the global parameters, and iterates this training process to reasonably utilize the data of each client for model training on the premise of ensuring the data privacy of each client until the prediction effect of the trained model converges, and a trained complete ship detection network is obtained.
[0136] In summary, in order to reasonably utilize unbalanced and uneven-quality training data in the process of federated learning training and improve the model accuracy, the present invention first constructs a ship detection network and multiple local detection networks of multiple clients; then inputs the obtained ship image data into each local detection network, performs double-branch attention enhancement and feature fusion detection on the ship image data to obtain ship detection outputs, determines the total prediction loss based on the dynamic non-monotonic focusing method, and updates the local parameters of the local detection network according to the prediction loss; finally, adaptively weights and aggregates the local parameters of each local detection network to obtain the global parameters of the ship detection network, updates the local parameters according to the global parameters to obtain a new round of local detection networks, and iteratively updates the local parameters and global parameters until the network performance no longer improves. In the process of federated learning, the present invention can learn more comprehensive feature information through double-branch attention enhancement, balance training data of different qualities through dynamic non-monotonic focusing, and balance training data of different distributions through adaptive weighted aggregation of local parameters, and can reasonably utilize the training data of each client on the premise of protecting the data privacy of each client, effectively improving the accuracy of the ship detection network obtained by federated learning training.
[0137] An embodiment of the present invention also provides a method for applying a ship detection network. In combination with Figure 7 viewing, Figure 7 The flowchart of an embodiment of the method for applying a ship detection network provided by the present invention is as Figure 7 shown. The method for applying a ship detection network includes:
[0138] S701. Obtain a picture of the ship to be detected;
[0139] S702. Input the picture of the ship to be detected into the trained ship detection network to obtain a ship detection result;
[0140] Among them, the trained ship detection network is determined according to the above-mentioned method for training a ship detection network based on federated learning.
[0141] In the embodiment of the present invention, first, effectively obtain a picture of the ship to be detected, and then use the above-mentioned trained ship detection network to effectively detect the picture of the ship to be detected, and then the ship detection result can be output.
[0142] As Figure 8 shown, the present invention also correspondingly provides an electronic device 800. The electronic device 800 includes a processor 801, a memory 802, and a display 803. Figure 8 Only some components of the electronic device 800 are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.
[0143] The processor 801 may be a central processing unit (CPU), a microprocessor, or other data processing chips in some embodiments, and is used to run the program code stored in the memory 802 or process data, such as the method for training a ship detection network based on federated learning and / or the method for applying a ship detection network in the present invention.
[0144] In some embodiments, the processor 801 may be a single server or a server group. The server group may be centralized or distributed. In some embodiments, the processor 801 may be local or remote. In some embodiments, the processor 801 may be implemented on a cloud platform. In one embodiment, the cloud platform may include a private cloud, a public cloud, a hybrid cloud, a community cloud, a distributed cloud, an internal cloud, a multi-cloud, etc., or any combination of the above.
[0145] The memory 802 can be an internal storage unit of the electronic device 800 in some embodiments, such as the hard disk or memory of the electronic device 800. The memory 802 can also be an external storage device of the electronic device 800 in other embodiments, such as a plug-in hard disk equipped on the electronic device 800, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.
[0146] Furthermore, the memory 802 can include both the internal storage unit of the electronic device 800 and the external storage device. The memory 802 is used to store the application software installed on the electronic device 800 and various types of data.
[0147] The display 803 can be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. in some embodiments. The display 803 is used to display the information of the electronic device 800 and to display a visual user interface. The components 801 - 803 of the electronic device 800 communicate with each other through the system bus.
[0148] In one embodiment, when the processor 801 executes the ship detection network training program in the memory 802, the following steps can be implemented:
[0149] Construct a ship detection network and multiple local detection networks for multiple clients;
[0150] Input the obtained ship image data into each local detection network, perform dual-branch attention enhancement and feature fusion detection on the ship image data to obtain a ship detection output, determine the predicted total loss based on the dynamic non-monotonic focusing method, and update the local parameters of the local detection network according to the predicted total loss;
[0151] Adaptive weighted aggregation of the local parameters of each local detection network to obtain the global parameters of the ship detection network, update the local parameters according to the global parameters to obtain a new round of local detection networks, and iteratively update the local parameters and global parameters until the network performance no longer improves.
[0152] In one embodiment, when the processor 801 executes the ship detection network application program in the memory 802, the following steps can be implemented:
[0153] Obtain a ship picture to be detected;
[0154] Input the ship picture to be detected into the trained ship detection network to obtain a ship detection result.
[0155] It should be understood that when the processor 801 executes the ship detection network training program and / or the ship detection network application program in the memory 802, in addition to the above functions, other functions can also be realized. For specific details, reference can be made to the descriptions of the corresponding method embodiments above.
[0156] Furthermore, the type of the electronic device 800 mentioned in the embodiments of the present invention is not specifically limited. The electronic device 800 can be a mobile phone, a tablet computer, a personal digital assistant (PDA), a wearable device, a laptop computer, or other portable electronic devices. Exemplary embodiments of portable electronic devices include, but are not limited to, portable electronic devices equipped with IOS, android, microsoft, or other operating systems. The above portable electronic devices can also be other portable electronic devices, such as a laptop computer with a touch-sensitive surface (such as a touch panel). It should also be understood that in some other embodiments of the present invention, the electronic device 800 may not be a portable electronic device, but a desktop computer with a touch-sensitive surface (such as a touch panel).
[0157] Correspondingly, the embodiments of the present application further provide a computer-readable storage medium. The computer-readable storage medium is used to store computer-readable programs or instructions. When the programs or instructions are executed by a processor, the steps or functions in the ship detection network training method and / or the ship detection network application method provided by the above method embodiments can be realized.
[0158] Those skilled in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The computer program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a disk, an optical disk, a read-only memory, or a random access memory, etc.
[0159] The ship detection network training method and application method based on federated learning provided by the present invention have been introduced in detail above. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for training a ship detection network based on federated learning, characterized in that, Including: Constructing a ship detection network and multiple local detection networks of multiple clients; Inputting the obtained ship image data into each of the local detection networks, performing double-branch attention enhancement and feature fusion detection on the ship image data to obtain a ship detection output, determining a predicted total loss based on the dynamic non-monotonic focusing method, and updating the local parameters of the local detection network according to the predicted total loss; Among them, the determining of the predicted total loss based on the dynamic non-monotonic focusing method includes: calculating a loss of an initial bounding box for the ship detection output based on a double-layer distance attention mechanism; determining an exponential moving average according to the iteration round, determining an anchor box outlier degree according to the exponential moving average, determining a non-monotonic focusing coefficient according to the anchor box outlier degree, and determining a bounding box regression loss according to the non-monotonic focusing coefficient and the initial bounding box loss; determining a predicted probability loss and a classification loss according to the ship detection output; determining the predicted total loss according to the bounding box regression loss, the predicted probability loss, and the classification loss; Performing adaptive weighted aggregation on the local parameters of each of the local detection networks to obtain the global parameters of the ship detection network, updating the local parameters according to the global parameters to obtain a new round of local detection networks, and iteratively updating the local parameters and the global parameters until the network performance no longer improves.
2. The method for training a ship detection network based on federated learning according to claim 1, wherein, The ship image data includes local ship image data obtained by each client, and the local ship image data is respectively input into the local detection network corresponding to the client.
3. The method for training a ship detection network based on federated learning according to claim 1, wherein The performing of double-branch attention enhancement and feature fusion detection on the ship image data to obtain a ship detection output includes: Performing feature segmentation on the ship image data to obtain a number of input feature images; Performing double-branch attention enhancement on the input feature images to obtain enhanced feature images; Performing multi-scale feature extraction and prediction output on the enhanced feature images to obtain a ship detection output.
4. The method for training a ship detection network based on federated learning according to claim 3, characterized in that The performing of feature segmentation on the ship image data to obtain a number of input feature images includes: Performing channel-dimension feature segmentation on the ship image data after feature convolution to obtain a number of input feature images.
5. The method for training a ship detection network based on federated learning according to claim 3, characterized in that The performing of double-branch attention enhancement on the input feature images to obtain enhanced feature images includes: Performing height average pooling and width average pooling on the input feature images respectively to obtain a height feature map and a width feature map; Successively passing the height feature map and the width feature map through a 1×1 convolution, a sigmoid activation function, weighted merging, and group normalization to obtain a first-branch feature; Performing 3×3 convolution on the input feature images to obtain a second-branch feature; Successively passing the first-branch feature through 2D average pooling and a softmax activation function and then performing matrix multiplication to fuse the second-branch feature to obtain a first attention feature, and successively passing the second-branch feature through 2D average pooling and a softmax activation function and then performing matrix multiplication to fuse the first-branch feature to obtain a second attention feature; After merging the first attention feature and the second attention feature, a cross - spatial attention weight is obtained through a sigmoid activation function, and the input feature image is feature - weighted according to the cross - spatial attention weight to obtain an enhanced feature image.
6. The method for training a ship detection network based on federated learning according to claim 1, characterized in that The adaptive weighted aggregation of the local parameters of each local detection network to obtain the global parameters of the ship detection network includes: Determining an aggregation weight according to the amount of valid data in the ship image data of each client; Adaptive weighted aggregation of the local parameters according to the aggregation weight to obtain the global parameters of the ship detection network.
7. A method for applying a ship detection network, characterized in that, Including: Obtaining a ship picture to be detected; Inputting the ship picture to be detected into a trained ship detection network to obtain a ship detection result; Among them, the trained ship detection network is determined according to the ship detection network training method based on federated learning described in any one of claims 1 to 6.
8. An electronic device, comprising a processor, a memory, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the ship detection network training method based on federated learning described in any one of claims 1 to 6 and / or the ship detection network application method described in claim 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the ship detection network training method based on federated learning described in any one of claims 1 to 6 and / or the ship detection network application method described in claim 7.
Citation Information
Patent Citations
Infrared ship target detection method based on improved yolov7
CN117152691A
Ship target detection method and device and readable storage medium
CN118379696A