A single-source domain target recognition generalization method, product, medium and device

By generating stylized images in a single source domain and decoupling features, and optimizing the model using an orthogonal loss function, the problem of poor target recognition performance under complex weather conditions is solved, and efficient recognition under complex and unknown weather conditions is achieved.

CN118570781BActive Publication Date: 2026-05-29SHANGHAI UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI UNIV
Filing Date
2024-05-16
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing domain-adaptive target recognition methods perform poorly under complex weather conditions and are difficult to generalize to unknown weather conditions. This is mainly because the training dataset contains relatively little complex weather data and is biased towards conventional weather, resulting in insufficient generalization ability of the model under complex weather conditions.

Method used

A single-source domain target recognition generalization method is adopted, which generates stylized images through Fourier transform. Decoupled representation learning is used to decouple image features into domain-invariant features and domain-unique features. The model is optimized through orthogonal loss function to improve the model's recognition ability under complex and unknown weather conditions.

Benefits of technology

It improves the accuracy of the model in identifying maritime targets under complex and unknown weather conditions, realizes the ability to generalize from a single domain to multiple domains, and enhances the model's recognition performance under different complex and unknown weather conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118570781B_ABST
    Figure CN118570781B_ABST
Patent Text Reader

Abstract

The application discloses a single-source domain target recognition generalization method, product, medium and equipment, relates to the field of domain adaptive target recognition, and comprises the following steps: generating a stylized image corresponding to an original image through a style feature space; encoding original image and stylized image features; jointly decoupling the original image and the stylized image features into domain-invariant features and domain-unique features; training a region candidate network using the domain-invariant features, and optimizing the network using orthogonal loss and target recognition loss functions; and inputting a complex unknown weather sea target image into the optimized network to obtain the category and position of the sea target in the complex unknown weather sea target image. The application can overcome the problems of difficulty in extracting domain-invariant features in complex weather data sets, single style generation and data generated being biased towards source domain distribution, and difficulty in method model generalization, improve the generalization ability of the model, and effectively improve the recognition ability of sea targets under different complex unknown weather conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of domain adaptive target recognition technology, and in particular to a single-source domain target recognition generalization method, product, medium, and device. Background Technology

[0002] As ships navigate, the weather changes constantly, and the training set (data collected in the marine environment) is difficult to cover various complex weather scenarios in the real environment. This limits the application of current domain adaptive target recognition methods, and the effect of maritime target recognition under complex weather conditions is still not ideal.

[0003] Existing domain-adaptive target recognition methods are primarily based on adversarial learning. The most classic domain-adaptive target recognition model is the Domain Adaptation Faster R-CNN (DA-Faster) model, which, based on the Faster R-CNN framework, has achieved significant results on cross-domain datasets. Unlike classification tasks, the DA-Faster model extracts domain-invariant features between the source and target domains by adding gradient inversion layers at both the image and instance levels. Intrinsic Feature Extraction Domain Adaptation (IFEDA) builds upon DA-Faster by extracting intrinsic features to reduce background interference, thereby improving the model's robustness and recognition accuracy under complex weather conditions.

[0004] IFEDA and DA-Faster are two different methods (or models). The IFEDA method consists of two parts: Intrinsic Feature Extraction (IFE) and Object Consistency Constraint (OCC). IFEE ensures at the instance level that the model extracts domain-invariant features while also including discriminative features of the target, reducing interference from redundant background information. The Object Consistency Constraint maintains class consistency between the target and source domains, further improving the accuracy of domain-adaptive target recognition. IFEDA's IFEDA method can only guarantee domain-invariant feature extraction at the instance level and cannot achieve domain-invariant feature extraction in complex weather datasets.

[0005] While the IFEDA method achieves high accuracy on datasets with foggy conditions, its performance deteriorates on datasets with twilight conditions. This is primarily because weather conditions are complex and variable during intelligent ship navigation, and the training set for target recognition models struggles to cover all weather conditions in the marine environment. This is due to the high cost of data acquisition, resulting in limited training data for complex navigation scenarios. Most data is collected under favorable weather conditions, such as clear skies. Consequently, models trained on existing data perform well in target recognition under normal weather conditions but poorly in complex scenarios, thus limiting the application of the IFEDA method in complex and variable weather conditions. Therefore, improving the perception capabilities of intelligent ships under complex and unknown weather conditions remains a challenge.

[0006] Because existing marine scene data primarily consists of data from normal clear weather with limited data on complex weather conditions, conventional methods like IFEDA, which generate data from highly biased datasets, inevitably tend to favor the source domain (the normal data domain). Models trained on this basis perform well in identifying targets in normal weather but poorly in complex weather. During ship navigation, marine weather conditions change rapidly. In real-world environments, models cannot access the test set during the training phase. Traditional domain generalization models (IFEDA and DA-Faster) require combining multiple source domains to enhance the backbone network's ability to acquire domain-invariant features across different domains. However, in practical scenarios, data acquisition and annotation are extremely time-consuming and labor-intensive, significantly limiting the generalization and application difficulty of these methods.

[0007] In summary, how to overcome the difficulties in extracting domain-invariant features from complex weather datasets, the single generation style, and the generation of data that is biased towards the source domain distribution, which leads to difficulties in the generalization of the model, and how to improve the generalization ability of the model (generalization ability means that the model can achieve good recognition performance when transferred to other domains), thereby effectively improving the ability to identify maritime targets under different complex and unknown weather conditions, has become an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0008] The purpose of this invention is to provide a single-source-domain target recognition generalization method, product, medium, and device that can overcome the problems of difficulty in extracting central domain invariant features from complex weather datasets, the single generation style, and the generation of data biased towards the source domain distribution, which leads to difficulties in the generalization of the method model. This improves the generalization ability of the model and thus effectively enhances the ability to identify maritime targets under different complex and unknown weather conditions.

[0009] To achieve the above objectives, the present invention provides the following solution.

[0010] On one hand, the present invention provides a single-source domain target recognition generalization method, comprising:

[0011] Two image samples are randomly selected from a single-source domain dataset, and image features are extracted through two-dimensional convolution to construct a style feature space; the single-source domain dataset is a set of marine target image samples containing various weather conditions; the two image samples randomly selected from the single-source domain dataset are marine target image samples with two different weather conditions.

[0012] A stylized image corresponding to the original image is generated using the style feature space; the original image is one of two image samples arbitrarily selected from the single-source domain dataset.

[0013] The original image and the stylized image are encoded with features by a feature encoder, respectively.

[0014] The original image and stylized image features are decoupled into domain-invariant features and domain-unique features by two feature decoders.

[0015] The region candidate network is trained using the domain-invariant features, and then optimized using an orthogonal loss function and an object recognition loss function to obtain the optimized region candidate network.

[0016] The optimized region candidate network is used to input the marine target image with complex and unknown weather conditions into the marine target image with complex and unknown weather conditions to obtain the category and location of the marine target in the marine target image with complex and unknown weather conditions.

[0017] Optionally, the style feature space is constructed by extracting the phase and amplitude of images from different domains through Fourier transform; the images from different domains are two image samples arbitrarily selected from the single-source domain dataset.

[0018] Optionally, a stylized image corresponding to the original image is generated using the style feature space, specifically including:

[0019] For each frequency domain image obtained through Fourier transform, a new amplitude value sample set corresponding to the frequency domain image is constructed to obtain the arbitrary style amplitude corresponding to each frequency domain image. Finally, the stylized images are obtained by Fourier inversion.

[0020] Optionally, the feature encoder is a variational autoencoder.

[0021] Optionally, the two feature decoders are composed of a series of convolutional layers.

[0022] Optionally, the method for calculating the orthogonal loss function includes:

[0023] The RoI features are extracted from the domain-invariant features and the domain-unique features respectively to obtain instance-level domain-invariant features and instance-level domain-unique features;

[0024] Average pooling is performed on the instance-level domain-invariant features and the instance-level domain-unique features respectively to obtain the instance-level domain-invariant features and the instance-level domain-unique features after the operation.

[0025] Calculate the vector product between the instance-level domain-invariant features after the operation and the instance-level domain-unique features after the operation;

[0026] The orthogonal loss function is calculated by summing the vector products of each pixel in the image.

[0027] Optionally, the method for calculating the target recognition loss function includes:

[0028] Obtain the regression loss and classification loss of the target bounding box, as well as the loss function of the region candidate network;

[0029] The target recognition loss function is calculated by adding the regression loss and classification loss of the target bounding box and the region candidate network loss function.

[0030] On the other hand, the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the single-source domain target recognition generalization method.

[0031] On the other hand, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the single-source domain target recognition generalization method.

[0032] In another aspect, the present invention provides a computer device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the single-source domain target recognition generalization method.

[0033] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0034] This invention discloses a single-source domain target recognition generalization method, product, medium, and device. By generating data with a different style from the source domain data on a single source domain, and using a decoupling approach, the source domain data and the generated data are decoupled to extract domain-invariant features and domain-specific features respectively. Based on the domain-invariant features, a model (i.e., a region candidate network) is trained, and the model is optimized using an orthogonal loss function. This enables the model to fully extract domain-invariant features, improving the generalization ability of the perception model on complex and unknown environmental datasets. Therefore, it has a good recognition effect on maritime targets under different complex and unknown weather conditions, thus overcoming the problems of difficulty in extracting domain-invariant features in complex weather datasets, the single generation style, and the generation of data biased towards the source domain distribution, which leads to the difficulty in generalization of the method model. This improves the model's generalization ability and effectively enhances the model's ability to recognize maritime targets under different complex and unknown weather conditions. Attached Figure Description

[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0036] Figure 1 This is a flowchart of the single-source domain target recognition generalization method provided in Embodiment 1 of the present invention;

[0037] Figure 2 This is a flowchart of the SDRL method of the present invention;

[0038] Figure 3 This is a visual comparison of the SDRL method of this invention for three complex weather conditions: sunny, foggy, dusk, and cloudy. Detailed Implementation

[0039] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0040] The purpose of this invention is to provide a single-source-domain target recognition generalization method, product, medium, and device that can overcome the problems of difficulty in extracting central domain invariant features from complex weather datasets, the single generation style, and the generation of data biased towards the source domain distribution, which leads to difficulties in the generalization of the method model. This improves the generalization ability of the model and thus effectively enhances the ability to identify maritime targets under different complex and unknown weather conditions.

[0041] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0042] Example 1

[0043] like Figure 1 As shown in the figure, the single-source domain target recognition generalization method provided in this embodiment includes:

[0044] Step 101: Select any two image samples from the single-source domain dataset, extract image features through two-dimensional convolution, and construct a style feature space; the single-source domain dataset is a set of marine target image samples containing various weather conditions; the two image samples selected from the single-source domain dataset are marine target image samples with two different weather conditions.

[0045] In step 101, the construction of the style feature space is achieved by extracting the phase and amplitude of images from different domains through Fourier transform; the images from different domains are two image samples arbitrarily selected from a single-source domain dataset.

[0046] Step 102: Generate a stylized image corresponding to the original image through the style feature space; the original image is one of two image samples arbitrarily selected from the single-source domain dataset.

[0047] Step 102 specifically includes:

[0048] For each frequency domain image obtained through Fourier transform, a new amplitude value sample set corresponding to the frequency domain image is constructed to obtain the arbitrary style amplitude corresponding to each frequency domain image. Finally, the stylized images are obtained by Fourier inversion.

[0049] Step 103: Encode the features of the original image and the stylized image respectively using a feature encoder.

[0050] In step 103, the feature encoder is a variational autoencoder.

[0051] Step 104: Decouple the features of the original image and the stylized image into domain-invariant features and domain-unique features using two feature decoders.

[0052] In step 104, the two feature decoders are composed of a series of convolutional layers.

[0053] Step 105: Train the region candidate network using domain-invariant features, and optimize the region candidate network using orthogonal loss function and target recognition loss function to obtain the optimized region candidate network.

[0054] In step 105, the calculation method for the orthogonal loss function includes:

[0055] RoI features are extracted from domain-invariant features and domain-unique features respectively to obtain instance-level domain-invariant features and instance-level domain-unique features.

[0056] Average pooling is performed on instance-level domain-invariant features and instance-level domain-unique features respectively to obtain the instance-level domain-invariant features and instance-level domain-unique features after the operation.

[0057] Calculate the product of the vectors between the instance-level domain-invariant features and the instance-level domain-unique features after the operation.

[0058] The orthogonal loss function is calculated by summing the vector products of each pixel in the image.

[0059] In step 105, the calculation method for the target recognition loss function includes:

[0060] Obtain the regression loss and classification loss of the target bounding box, as well as the loss function of the region candidate network.

[0061] The target recognition loss function is calculated by adding the regression loss and classification loss of the target bounding box and the region candidate network loss function.

[0062] Step 106: Input the image of the marine target in complex and unknown weather into the optimized region candidate network, and use the optimized region candidate network to obtain the category and location of the marine target in the image of the marine target in complex and unknown weather.

[0063] The technical solution of the present invention is illustrated below with a specific embodiment:

[0064] This invention provides a single-source domain target recognition generalization method, which is a single-source domain target recognition generalization method based on Stylization and Decoupled Representation Learning (SDRL). This invention primarily addresses the difficulties in extracting domain-invariant features from datasets under complex and unknown weather conditions, as well as the challenges of model generalization caused by a single generated style and a bias towards the source domain distribution. It employs a single-source domain generalization strategy, combined with data stylization and decoupled learning, to improve the model's generalization ability. Here, generalization refers to domain transfer; for this invention, target recognition under unknown weather conditions is the domain that needs to be transferred. The overall idea of ​​this invention is as follows:

[0065] First, a Fourier transform-based stylization method is used to generate data with a style different from the source domain data in a single source domain. Second, a decoupled representation learning method is used to decouple style-invariant features (i.e., domain-invariant features) and style-unique features (i.e., domain-unique features) from the source domain data and the generated data. Decoupled representation learning (DRL) refers to decomposing the variable factors representing features in a dataset that meet certain conditions, and these variable factors are independent of each other. Then, a vector orthogonal loss function (i.e., orthogonal loss function) is used to train the model, enabling it to fully extract domain-invariant features and improve the generalization ability of the perceptual model on unknown environment datasets (i.e., complex unknown environment datasets). Finally, classification and regression tasks are performed on the branch of the decoupled representation learning model that extracts domain-invariant features, improving the model's ability to identify maritime targets under complex unknown weather conditions. This achieves good recognition results for maritime targets under different complex unknown weather conditions, effectively improving the model's ability to identify maritime targets under various complex unknown weather conditions.

[0066] The overall process of the single-source domain target recognition generalization method based on stylization and decoupled representation learning in this invention is as follows: Figure 2 As shown, it includes the following steps:

[0067] Step 1: Select any two images from the single-source domain dataset, extract image features through two-dimensional convolution, and construct a style feature space. Specifically, the construction of the style feature space is mainly achieved through Fourier transform, extracting the phase and amplitude of images from different domains (styles). The method is shown in formulas (1) to (3):

[0068]

[0069]

[0070]

[0071] Where H and W represent the height and width of the image, f represents the input function, which takes the image in the time domain as the Fourier transform input, (w,h) represents the spatiotemporal position of the current pixel, and (u,v) represents the frequency domain position of the pixel in image x. This represents the frequency domain image obtained by performing a Fourier transform on the image, thus obtaining the image's phase. and amplitude e is the exponent of e, π is the mathematical constant pi, i represents the imaginary unit, R(u,v) represents the real part of the Fourier transform, and I(u,v) represents the imaginary part of the Fourier transform. Formula (1) is the application of the conventional two-dimensional discrete Fourier transform to images.

[0072] When constructing the style feature space, the styles of any two images selected from a single-source domain dataset should be different. A single-source domain dataset includes images of different styles. However, a single source domain is not limited to one domain; it contains images from multiple domains, including images from different weather conditions.

[0073] Step 2: Generate a stylized image corresponding to the original image using the stylization module (style feature space) to enhance the diversity of styles within the single source domain. For each frequency domain image obtained in Step 1... Construct new amplitude value sample sets for each. Obtain arbitrary style amplitude (Each frequency domain image corresponds to an arbitrary style amplitude). Finally, Fourier inversion is performed. Obtain a new style image and generate an image. The specific method is shown in formula (4):

[0074]

[0075] in, Let N represent the k-th image of the j-th source domain, N be the image width (maximum pixel value per row), M be the image height (maximum pixel value per column), and (u, v) represent the current pixel position (frequency domain position of the current image pixel). Calculate the phase according to formula (2) in step one. for The phase is calculated according to formula (2) in step one, where e is the exponent and i represents the imaginary unit. This indicates the generated image (i.e., the stylized image).

[0076] The Fourier transform can convert a two-dimensional image from the conventional spatial domain to the Fourier frequency domain. Through the Fourier transform, each two-dimensional image can be transformed into its corresponding frequency domain image. The frequency domain image can be represented by phase and amplitude. By recombining different phases and amplitudes, a new frequency domain image can be obtained. Then, by performing an inverse Fourier transform on the new frequency domain image, a new two-dimensional image can be restored.

[0077] Step 3: Encode the original image and the generated image separately using a feature encoder. Features. A variational autoencoder (VAE) is used to encode the original image x. S Compared with the stylized image Obtain the single-source domain feature set Where w, h, and c represent the width, length, and number of channels of the feature map, respectively, and F b Represents a single-source domain feature set. It is a set of feature maps, which contains three dimensions: w, h, and c.

[0078] To generate an image of a specific style, an image of that style is selected as the source image, and a new image of that style is generated based on that source image. The single-source domain feature set contains pairs of features.

[0079] Step 4: Decouple the features of the original image and the generated image into domain-invariant features and domain-unique features using two feature decoders (feature extractors). Design two feature extractors E. DIR With E DSR From F respectively b Invariant features of the decoupling domain Unique features of the domain Where E DIR With E DSR It consists of a series of convolutional layers, as shown in equations (5) and (6):

[0080] F di =E DIR (F b (5)

[0081] F ds =E DSR (F b ) = F b -F di (6)

[0082] Among them, F di Representation domain invariant features, F ds Represents a domain-specific feature, E DIR With E DSR This represents two feature decoders. Feature extractor E DIR With E DSR It consists of a series of convolutional layers, i.e., a CNN structure, which is well known to those skilled in the art.

[0083] Then, the domain-invariant features F extracted by the model are... di The target region candidate network (RPN) is used as input to extract a series of candidate regions with domain-invariant features. The RPN is the input feature map, and its structure allows for the acquisition of target candidate boxes and RoI regions.

[0084] The method of this invention decouples the features of the original image and the generated image into domain-invariant features and domain-unique features through decoupling. The model is trained based on the domain-invariant features, thus achieving good recognition results for targets under different weather conditions. This solves the problem of the method model having difficulty in generalization due to the single generation style and the generated data being biased towards the source domain distribution.

[0085] Step 5: Employ vector orthogonality to constrain domain-invariant and domain-unique features, improving the accuracy and completeness of the backbone network in extracting domain-invariant features. Based on vector decomposition theory, and adhering to the principle of independent operations between vectors, an orthogonal strategy is introduced to maintain the orthogonal relationship between the decomposed vectors, thereby promoting the extraction of domain-invariant features F. di Domain-specific features F ds Decoupling is achieved by employing an orthogonal loss function. Maintain F di With F ds The orthogonal relationship between them is used to fully extract F di Specifically, firstly, extract F respectively. di With F ds RoI features yield instance-level domain-invariant features. Unique features of instance-level domains Secondly, for instance-level domain-invariant feature A... di and instance-level domain-specific features Ad s进 Perform average pooling to obtain instance-level domain-invariant features after the operation. Unique characteristics of instance-level domains after operations Then, calculate P. di and P ds The orthogonality loss is calculated by multiplying the vectors between each pixel in the image by S. Finally, the orthogonality loss is calculated by summing the vector products of each pixel. The calculation method is as follows:

[0086]

[0087]

[0088] Where n represents the number of candidate boxes, c represents the number of channels, and w and h represent the width and length, respectively. Let represent the L2 norm, ⊙ represent the dot product operation between vectors, and S[i,j] represent the product of vectors i and j, where i and j are P... di and P ds The corresponding pixel value. Formula (7) is the product formula, since P di and P ds Since the two feature maps are the same size, the corresponding pixels can be multiplied by vectors, thus obtaining the vector product of each pixel in the image.

[0089] Step Six: Use the features generated by the domain-invariant feature extraction module for the final classification and regression of the model (object recognition model), thereby improving the generalization performance of the single-source domain object recognition model. During the training of the object recognition model, optimization is required using the object recognition loss function. (Object recognition loss function) The definition is shown in formula (9):

[0090]

[0091] in, and These represent the regression loss and classification loss of the target bounding box, respectively. The regression loss and classification loss of the target bounding box are well-known to those skilled in the art. The bounding box regression loss is generally achieved using the smooth-L1 loss, while the classification loss in this invention uses the cross-entropy loss function. The loss function of the Region Candidate Network (RPN) is well known to those skilled in the art. It is used to distinguish the foreground and background of an image and to fine-tune the bounding box.

[0092] Steps one through five encompass the training and testing of the target recognition model. The training dataset primarily consists of target image data under normal weather conditions in marine scenes, along with a small amount of target image data under complex weather conditions. The trained model is tested using marine target image data under different complex weather conditions as input. Tables 1, 2, and 3 below show the sample test results under different weather conditions. The trained target recognition model can extract domain-invariant features from the input images for recognition. The model can be applied to multiple domains (different weather conditions) for recognition. By training the model using domain-invariant features and then performing recognition using the trained model, it achieves generalization from a single model to multiple source domains; that is, it achieves good recognition performance in one domain and good recognition performance when generalized from one domain to other domains.

[0093] To verify the effectiveness of the proposed SDRL method, real-world datasets were used for training and testing. For dataset selection, the experiment employed a marine target dataset containing four different weather conditions: sunny, twilight, cloudy, and foggy. 4218 sunny weather images were used as the training set for the single-source domain generalization method, while 671 foggy, 329 twilight, and 1819 cloudy images were used as the test set. The test set was not accessed during model training. The experiment utilized six common categories from the dataset: bulk carriers, container ships, cruise ships, sailboats, other vessels, and islands / reefs. The mean average precision (mAP) with a threshold of 0.5 was used as a metric to measure the effectiveness of the single-source domain target recognition generalization method based on stylized and decoupled representation learning.

[0094] Experiments were conducted on a marine dataset under various complex weather conditions. The performance of the proposed single-source domain marine target recognition generalization model based on stylization and decoupled representation learning was compared with that of a classic single-source domain target recognition model. Tables 1, 2, and 3 show the recognition results of the model on the marine target dataset under different complex sea conditions. Table 1 shows the marine target recognition results when the source domain is clear weather and the unknown target domain is dusk. Table 2 shows the marine target recognition results when the source domain is clear weather and the unknown target domain is foggy weather. Table 3 shows the marine target recognition results when the source domain is clear weather and the unknown target domain is cloudy weather.

[0095] Table 1. Results of identification under dusk weather conditions (%)

[0096] method bulk carriers container ship cruise ship Other ships Islands and reefs sailboat mAP FasterR-CNN 51.3 71.1 82.6 29.7 42.6 85.1 60.4 IBN-Net 54.3 72.9 83.3 29.1 41.9 85.7 61.2 SW 52.6 71.4 83.2 30.7 43.9 86.7 61.4 IterNorm 50.3 69.1 83.6 31.1 42.2 85.5 60.3 ISW 52.3 71.4 82.8 31.5 42.4 86.4 61.1 IFEDA 52.1 71.7 86.0 32.7 44.7 85.9 62.1 SDRL 60.8 75.1 88.8 34.1 40.2 87.9 64.5

[0097] As shown in Table 1, the stylized and decoupled representation learning single-source domain generalization model proposed in this invention achieves a recognition accuracy of 64.5% mAP in the marine target recognition dataset during twilight weather. Compared with the benchmark model Faster R-CNN and other classic single-source domain target recognition generalization models IFEDA, IBN-Net, IterNorm, ISW, and SW, the accuracy is improved by 4.1%, 2.4%, 3.3%, 4.2%, 3.4%, and 3.1%, respectively.

[0098] Table 2. Results of identification under foggy weather (%)

[0099] method bulk carriers container cruise ship Other ships Islands and reefs sailboat mAP FasterR-CNN 75.7 80.2 82.1 38.3 59.8 85.7 70.3 IBN-Net 76.3 79.6 81.1 37.1 56.6 84.1 69.1 SW 76.9 80.8 82.2 37.4 60.1 84.3 70.3 IterNorm 75.6 78.6 81.9 38.7 56.4 85.8 69.5 ISW 75.8 80.7 83.3 40.6 60.1 88.1 71.5 IFEDA 79.3 80.9 85.1 41.7 60.2 88.3 72.6 SDRL 79.4 85.6 87.6 41.6 61.4 88.5 74.0

[0100] As shown in Table 2, the single-source domain generalization model based on stylization and decoupled representation learning proposed in this invention achieves an accuracy of 74.0% mAP in the recognition of marine targets in foggy weather. Compared with the benchmark model Faster R-CNN and other classic single-source domain target recognition generalization models IFEDA, IBN-Net, IterNorm, ISW, and SW, the accuracy is improved by 3.7%, 1.4%, 4.9%, 4.5%, 2.5%, and 3.7%, respectively.

[0101] Table 3. Recognition results under cloudy weather conditions (%)

[0102] method bulk carriers container ship cruise ship Other ships Islands and reefs sailboat mAP FasterR-CNN 66.1 80.3 80.2 46.0 40.5 79.5 65.4 IBN-Net 66.2 77.5 79.8 44.5 39.1 79.6 64.5 SW 67.1 79.7 80.6 46.2 40.4 78.8 65.5 IterNorm 67.2 77.6 78.9 46.1 37.9 73.6 63.6 ISW 67.9 80.8 80.1 45.7 41.7 79.7 66.0 IFEDA 68.6 80.7 84.1 48.3 42.7 79.8 67.4 SDRL 72.0 81.7 88.0 50.2 42.4 80.9 69.2

[0103] As shown in Table 3, the stylized and decoupled representation learning single-source domain generalization model proposed in this invention achieves an accuracy of 69.2% mAP in the identification of marine targets in cloudy weather. Compared with the benchmark model Faster R-CNN and other classic single-source domain target identification generalization models IFEDA, IBN-Net, IterNorm, ISW, and SW, the accuracy is improved by 3.8%, 1.8%, 4.7%, 5.6%, 3.2%, and 3.7%, respectively.

[0104] Table 4 Comparison results under different weather conditions (%)

[0105] method cloudy day Foggy day Dusk mAP FasterR-CNN 65.4 70.3 60.4 65.4 IBN-Net 64.5 69.1 61.2 64.9 SW 65.5 70.3 61.4 65.7 IterNorm 63.6 69.5 60.3 64.5 ISW 66.0 71.5 61.1 66.2 IFEDA 67.4 72.6 62.1 67.4 SDRL 69.2 74.0 64.5 69.2

[0106] As shown in Table 4, the stylized and decoupled representation learning single-source domain generalization model proposed in this invention achieves an accuracy of 69.2% mAP in the identification of marine targets in unknown weather scenarios. Compared with the benchmark model Faster R-CNN and other classic single-source domain target recognition generalization models IFEDA, IBN-Net, IterNorm, ISW, and SW, the accuracy is improved by 3.8%, 1.8%, 4.3%, 4.7%, 3.0%, and 3.5%, respectively.

[0107] This invention's method annotates predicted bounding boxes and category scores on randomly selected samples identified during twilight, foggy, and overcast weather conditions, qualitatively analyzing the model's performance. Specifically, the samples refer to image data containing labeled targets under complex weather conditions such as twilight and fog. The predicted bounding boxes will label bulk carriers, container ships, cruise ships, sailboats, other vessels, and islands / reefs in the image. The dataset used for training and testing the model in this invention annotates these target categories. The labeled category score refers to the score indicating whether the predicted bounding box contains a bulk carrier, container ship, cruise ship, sailboat, other vessel, or island / reef; this score is a probability value, representing the probability of belonging to a specific target category. Figure 3 This paper demonstrates the recognition results of the proposed single-source domain maritime target recognition generalization model (SDRL) based on stylization and decoupled representation learning in unknown and complex weather conditions. Figure 3 Parts (a), (d), and (g) demonstrate the recognition performance of the SDRL model. Figure 3 Parts (b), (e), and (h) demonstrate the recognition performance of the IFEDA model. Figure 3 Parts (c), (f), and (i) demonstrate the recognition performance of the Faster R-CNN model, by... Figure 3 As shown in the set of recognition results, the SDRL method proposed in this invention has better recognition performance than the models trained by the other two methods. Figure 3 Parts (a), (b), and (c) show the identification results of some samples from the marine target dataset collected in foggy weather. Figure 3 Parts (d), (e), and (f) show the identification results of some samples from the marine target dataset collected during twilight weather. Figure 3 Parts (g), (h), and (i) represent the identification results of some samples in the marine target dataset collected on cloudy days. Figure 3 Parts (a), (d), and (g) are visualizations of the SDRL model generalizing from clear weather to complex, unknown weather conditions. Figure 3 Part (a) shows the visualization results of the SDRL model generalizing from sunny weather to foggy weather, part (d) shows the visualization results of the SDRL model generalizing from sunny weather to dusk weather, and part (g) shows the visualization results of the SDRL model generalizing from sunny weather to cloudy weather. Figure 3 Parts (b), (e), and (h) are visualizations of the IFEDA model's generalization from clear weather to complex, unknown weather conditions. Figure 3 Part (b) shows the visualization results of the IFEDA model generalizing from sunny to foggy weather, part (e) shows the visualization results of the IFEDA model generalizing from sunny to dusk weather, and part (h) shows the visualization results of the IFEDA model generalizing from sunny to cloudy weather. Figure 3 Parts (c), (f), and (i) are visualizations of the Faster R-CNN model generalizing from a sunny day dataset to complex, unknown weather conditions. Figure 3 Part (c) shows the visualization results of the Faster R-CNN model generalizing from a sunny day dataset to foggy weather; part (f) shows the visualization results of the Faster R-CNN model generalizing from a sunny day dataset to dusk weather; and part (i) shows the visualization results of the Faster R-CNN model generalizing from a sunny day dataset to cloudy weather. During model training, the model cannot access the target domain data and labels, so the three complex weather conditions are unknown to the model. The visualization results show that the stylized and decoupled representation learning-based single-source domain maritime target recognition generalization model proposed in this invention has better recognition results and can effectively improve the ship recognition capability under complex and unknown weather conditions.

[0108] This invention addresses the challenge of generalization difficulties in domain-adaptive target recognition models under complex and unknown weather conditions. It proposes a single-source domain target recognition generalization method based on stylization and decoupled representation learning, further enhancing the ability of intelligent ships to identify targets under complex, variable, and even unknown weather conditions during navigation. The single-source domain target recognition generalization model based on stylization and decoupled representation learning consists of two parts: a stylization method and a decoupled representation learning method. First, a Fourier transform-based stylization method is used to generate various data with different styles from the source domain during training. Then, the decoupled representation learning method is used to extract domain-invariant and domain-unique features between the generated images and the single-source domain images. Finally, a vector orthogonal loss function is used to enhance the network's ability to acquire domain-invariant features. This invention lays the foundation for improving the perception capabilities of intelligent ships, especially their target perception capabilities in complex environments.

[0109] Example 2

[0110] A computer program product includes a computer program that, when executed by a processor, implements the single-source domain target recognition generalization method of Embodiment 1.

[0111] Example 3

[0112] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the single-source domain target recognition generalization method of Embodiment 1.

[0113] Example 4

[0114] A computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The database stores pending transactions. The I / O interfaces facilitate information exchange between the processor and external devices. The communication interface enables communication with external terminals via a network connection. When executed by the processor, the computer program implements the single-source domain target recognition generalization method described in Embodiment 1.

[0115] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A single-source domain target recognition generalization method, characterized in that, include: Two image samples are randomly selected from a single-source domain dataset, and image features are extracted through two-dimensional convolution to construct a style feature space; The single-source domain dataset is a set of marine target image samples containing various weather conditions; any two image samples selected from the single-source domain dataset are marine target image samples for two different weather conditions. A stylized image corresponding to the original image is generated using the style feature space; the original image is one of two image samples arbitrarily selected from the single-source domain dataset. The original image and the stylized image are encoded with features by a feature encoder, respectively. The original image and stylized image features are decoupled into domain-invariant features and domain-unique features by two feature decoders. The region candidate network is trained using the domain-invariant features, and optimized using an orthogonal loss function and an object recognition loss function to obtain an optimized region candidate network. The calculation method of the orthogonal loss function includes: extracting RoI features from the domain-invariant features and the domain-unique features respectively to obtain instance-level domain-invariant features and instance-level domain-unique features; performing average pooling operations on the instance-level domain-invariant features and the instance-level domain-unique features respectively to obtain the processed instance-level domain-invariant features and the processed instance-level domain-unique features; calculating the vector product between the processed instance-level domain-invariant features and the processed instance-level domain-unique features; and summing the vector products of each pixel in the image to calculate the orthogonal loss function. The optimized region candidate network is used to input the marine target image with complex and unknown weather conditions into the marine target image with complex and unknown weather conditions to obtain the category and location of the marine target in the marine target image with complex and unknown weather conditions.

2. The single-source domain target recognition generalization method according to claim 1, characterized in that, The style feature space is constructed by extracting the phase and amplitude of images from different domains through Fourier transform; the images from different domains are two image samples arbitrarily selected from the single-source domain dataset.

3. The single-source domain target recognition generalization method according to claim 2, characterized in that, Using the style feature space, a stylized image corresponding to the original image is generated, specifically including: For each frequency domain image obtained through Fourier transform, a new amplitude value sample set corresponding to the frequency domain image is constructed to obtain the arbitrary style amplitude corresponding to each frequency domain image. Finally, the stylized images are obtained by Fourier inversion.

4. The single-source domain target recognition generalization method according to claim 1, characterized in that, The feature encoder is a variational autoencoder.

5. The single-source domain target recognition generalization method according to claim 1, characterized in that, The two feature decoders are composed of a series of convolutional layers.

6. The single-source domain target recognition generalization method according to claim 1, characterized in that, The method for calculating the target recognition loss function includes: Obtain the regression loss and classification loss of the target bounding box, as well as the loss function of the region candidate network; The target recognition loss function is calculated by adding the regression loss and classification loss of the target bounding box and the region candidate network loss function.

7. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the single-source domain target recognition generalization method as described in any one of claims 1-6.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the single-source domain target recognition generalization method as described in any one of claims 1-6.

9. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the single-source domain target recognition generalization method according to any one of claims 1-6.