Multi-modal image marine target individual identification method, system and equipment based on wavelet decomposition, medium and product
By combining infrared and visible light images, multi-scale feature extraction and feature fusion are performed using wavelet decomposition and double-tree complex wavelet transformation, the recognition accuracy problem of offshore target individual recognition in complex sea conditions is solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510597229.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, when facing a signalless or "three-no" target, it is difficult to achieve individual recognition of maritime targets, and a single modal image is susceptible to environmental interference in complex sea conditions, resulting in a reduced recognition accuracy.
A multimodal image recognition method based on wavelet decomposition is adopted. By combining infrared and visible light images, the image is multi-scale feature decomposition using double-tree complex wavelet transformation, feature fusion and noise processing are performed to improve the robustness and accuracy of the recognition model.
It improves the accuracy and robustness of sea target individual recognition, can better identify target shapes and texture features in complex scenarios, and enhances the adaptability and anti-interference ability of the algorithm.
Smart Images

Figure CN120220037A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of marine target recognition, and in particular to a method, system, device, medium and product for multimodal image marine target individual recognition based on wavelet decomposition. Background Art
[0002] With the rapid development of the global marine economy and the increasing demand for maritime safety management, individual identification of maritime targets has become an extremely critical research field. Accurate identification of maritime targets is not only of great significance for improving navigation safety and combating illegal activities, but also provides strong support for marine resource management and environmental protection. The existing methods mainly use radar radiation source signals to achieve accurate individual identification, but when faced with many "three-no" targets (i.e. no name, no number, no signal), traditional identification methods based on radiation source signals show obvious limitations. These targets often turn off the signal of the Automatic Identification System (AIS), making it difficult for existing signal-based identification technology to be effective.
[0003] With the continuous advancement of imaging technology, especially the widespread application of high-resolution cameras, infrared sensors and other equipment, image-based individual recognition is gradually becoming a reality and a research hotspot. Compared with traditional signal recognition methods, image recognition can intuitively obtain the visual features of the target and has stronger adaptability and flexibility. However, single-modal image information has limitations, especially under complex sea conditions (such as lighting changes, fog, surges, etc.), single-modal images are easily affected by environmental interference, resulting in reduced recognition accuracy. Summary of the invention
[0004] The purpose of this application is to provide a method, system, device, medium and product for multimodal image marine target individual recognition based on wavelet decomposition, which can improve the robustness and accuracy of marine target individual recognition.
[0005] To achieve the above objectives, this application provides the following solutions.
[0006] In a first aspect, the present application provides a method for identifying individual targets at sea in a multimodal image based on wavelet decomposition, comprising the following steps.
[0007] Acquire a target image at sea; the target image at sea includes an infrared image and a visible light image.
[0008] An individual recognition model is constructed and trained; the individual recognition model comprises: a dual-tree complex wavelet transform decomposition module, a noise processing module, a feature fusion module, a feature encoding module and a classification module which are connected in sequence.
[0009] Input the image of the maritime target into the trained individual recognition model to obtain the identity of the maritime target, thus completing the individual recognition of the maritime target.
[0010] Among them, inputting the image of the maritime target into the trained individual recognition model to obtain the identity of the maritime target specifically includes the following.
[0011] Input the maritime target image into the dual-tree complex wavelet transform decomposition module for wavelet decomposition to obtain multiple subbands; the subbands include: low-frequency subbands and high-frequency subbands.
[0012] Input the high-frequency subbands into the noise processing module for denoising to obtain multiple denoised high-frequency subbands.
[0013] Input multiple low-frequency subbands and multiple denoised high-frequency subbands into the feature fusion module for feature fusion to obtain a feature fusion map.
[0014] Input the feature fusion map into the feature encoding module for multiple downsamplings to obtain a downsampled feature fusion map.
[0015] Input the downsampled feature fusion map into the classification module for classification to obtain the identity of the maritime target.
[0016] In a second aspect, the present application provides a multi-modal image maritime target individual recognition system based on wavelet decomposition, including the following units.
[0017] An acquisition unit for acquiring the image of the maritime target; the maritime target image includes: infrared image and visible light image.
[0018] A construction unit for constructing and training an individual recognition model; the individual recognition model includes: a dual-tree complex wavelet transform decomposition module, a noise processing module, a feature fusion module, a feature encoding module, and a classification module connected in sequence; the dual-tree complex wavelet transform decomposition module is used for wavelet decomposition of the image to obtain multiple subbands; the subbands include: low-frequency subbands and high-frequency subbands; the noise processing module is used to remove the noise in the high-frequency subbands through soft thresholding; the feature fusion module is used for feature fusion of multiple low-frequency subbands and multiple high-frequency subbands after removing noise to obtain a feature fusion map; the feature encoding module is used for multiple downsamplings of the feature fusion map; the classification module is used for classifying the downsampled feature fusion map to obtain the identity of the maritime target.
[0019] An individual recognition unit for inputting the image of the maritime target into the trained individual recognition model to obtain the identity of the maritime target, thus completing the individual recognition of the maritime target.
[0020] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the steps of the multi-modal image-based maritime target individual recognition method based on wavelet decomposition described in the first aspect above.
[0021] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the multi-modal image-based maritime target individual recognition method based on wavelet decomposition described in the first aspect above.
[0022] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of the multi-modal image-based maritime target individual recognition method based on wavelet decomposition described in the first aspect above.
[0023] According to the specific embodiments provided by the present application, the present application has the following technical effects.
[0024] The present application provides a multi-modal image-based maritime target individual recognition method, system, device, medium and product. By combining infrared and visible light images of two modalities and using wavelet transform to perform multi-scale feature decomposition on the images, it can capture the global structure information and local detail information in the images, enabling the individual recognition model to better recognize the shape, texture and other features of maritime targets in complex scenes, thereby improving the accuracy of maritime target individual recognition; through the noise processing module, the noise in the sub-bands is processed by soft thresholding, which can effectively remove high-frequency noise while retaining the important edges and detail information in the images, so as to reduce the noise level in the high-frequency sub-bands, thereby improving the accuracy of subsequent feature extraction and classification, enhancing the robustness of the algorithm, and improving the robustness of maritime target individual recognition; finally, through the feature fusion module, feature encoding module and classification module for fusion classification processing, the identity of the maritime target is finally obtained. In the whole recognition process, different frequency features are obtained by decomposing the infrared image and visible light image through the dual-tree complex wavelet transform decomposition module, then the noise is filtered and fused, and finally the robustness and accuracy of maritime target individual recognition are improved. Description of the Drawings
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0026] Figure 1Schematic flow chart of a multi-modal image-based maritime target individual recognition method provided by an embodiment of the present application.
[0027] Figure 2 Schematic framework diagram of a multi-modal image-based maritime target individual recognition method provided by an embodiment of the present application.
[0028] Figure 3 Schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0029] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0030] Based on the discussion of the background technology, using multi-modal images to achieve individual recognition has become a very promising and important development direction. The multi-modal image fusion technology can provide richer and more comprehensive target information by combining the advantages of different modalities (such as visible light and infrared images), and has good information complementarity. For example, visible light images can provide the appearance details of the target, while infrared images can capture the thermal radiation characteristics of the target. The combination of the two can significantly improve the effect of individual recognition.
[0031] Although certain progress has been made in image-based individual recognition, there are still many challenges in the refined extraction of features related to individual identities. Some existing studies have tried to extract and process image features through different algorithms. However, overall, there are few individual recognition algorithms based on multi-modal images, and most methods directly extract features from the input images, resulting in insufficiently learned features and still having great room for improvement in the effect.
[0032] As a powerful mathematical tool, wavelet decomposition has achieved remarkable success in the field of machine learning. It can obtain fine features at different scales through frequency decomposition, enabling subsequent neural networks to learn task-related features more fully. The advantage of wavelet decomposition lies in its ability to capture multi-scale information in images, especially suitable for processing complex images. In particular, in multi-modal image processing, wavelet decomposition can help better extract and fuse features of different modalities, thereby improving recognition performance. A few methods have also applied it to individual recognition based on signals. The related technology is mainly applied to the recognition of radiation source signals of unmanned aerial vehicles, and mainly focuses on analyzing the energy characteristics of radiation source signals, which is essentially different from the features learned by images. Generally speaking, there is currently no algorithm that directly applies wavelet decomposition to image tasks to achieve individual recognition.
[0033] Aiming at the problems that single-modal images are difficult to support individual recognition of maritime targets based on images and existing methods are difficult to fully extract individual identity features, this application proposes a multi-modal image target individual recognition algorithm based on wavelet decomposition. By combining infrared and visible light images of two modalities, wavelet transform is used to extract multi-scale features of the images, and a residual attention mechanism is introduced to enhance the feature representation of the target area. This method can not only give full play to the advantages of infrared and visible light images, but also effectively cope with the challenges brought by complex sea conditions, improving the robustness and accuracy of target individual recognition.
[0034] To make the above objects, features, and advantages of this application more obvious and understandable, the following further details this application in combination with the accompanying drawings and specific embodiments.
[0035] In an exemplary embodiment, as Figure 1 shown, a multi-modal image maritime target individual recognition method based on wavelet decomposition is provided. This method is executed by a computer device, which can be specifically executed by a computer device such as a terminal or a server alone, or jointly executed by a terminal and a server. In the embodiments of this application, this method is described by taking its application to a server as an example, including the following steps 1 to step 8.
[0036] Step 1: Obtain maritime target images; the maritime target images include: infrared images and visible light images.
[0037] Step 2: Construct and train an individual recognition model; the individual recognition model includes: a dual-tree complex wavelet transform decomposition module, a noise processing module, a feature fusion module, a feature encoding module, and a classification module connected in sequence. Among them, the feature encoding module includes: multiple encoding layers connected in sequence; the encoding layer includes: a convolutional layer, a batch normalization layer, and a max pooling layer connected in sequence.
[0038] Specifically, the training process of the individual recognition model specifically includes steps 21 to 23.
[0039] Step 21: Construct sample data pairs; the sample data pairs include: sample infrared images, sample visible light images, and sample identities. Specifically, constructing sample data pairs specifically includes steps 211 to 214.
[0040] Step 211: Obtain sample maritime target images; the sample maritime target images include: sample infrared images and sample visible light images.
[0041] Step 212: Crop the sample maritime target images to a unified size to obtain the cropped sample maritime target images.
[0042] Step 213: Use the manual annotation method to annotate the individual identities of the cropped sample maritime target images to obtain sample identities.
[0043] Step 214: Based on the sample infrared images, sample visible light images, and sample identities, construct sample data pairs.
[0044] Specifically, use a drone to fly forward to collect infrared and visible light images of maritime ship targets (i.e., sample infrared images and sample visible light images), use the object detection method or the manual annotation method to obtain the target positions, then crop the targets and unify the size to 128×128×3, and then annotate the individual identities (i.e., sample identities). The number of individuals is N, construct <infrared image, visible light image, identity> data pairs, thus constructing a multi-modal sample recognition task, and divide the data pairs according to a ratio of 7:3 to construct a training set and a test set.
[0045] Step 22: Input the sample infrared images and the sample visible light images into the individual recognition model to obtain sample recognition results.
[0046] Step 23: Construct a loss function according to the sample recognition results and the corresponding sample identity labels, and iteratively optimize the parameters of the individual recognition model according to the loss function until the loss function reaches the minimum value or the number of iterative optimization rounds reaches the maximum value, then stop the iterative optimization to obtain the trained individual recognition model.
[0047] Specifically, the loss function is the cross-entropy loss function, set the learning rate to 0.0002, the number of training times to 500, the batch size to e = 64, and use the data pairs constructed in step 1 to train the model 500 times. For the collected test samples, input them into the model to obtain the predicted individual identities (i.e., sample recognition results).
[0048] Step 3: Input the maritime target images into the trained individual recognition model to obtain the identities of the maritime targets, and complete the individual recognition of the maritime targets.
[0049] Among them, the image of the maritime target is input into the trained individual recognition model to obtain the identity of the maritime target, specifically including steps 31 to 35.
[0050] Step 31: Input the maritime target image into the dual-tree complex wavelet transform decomposition module for wavelet decomposition to obtain multiple subbands; the subbands include: a low-frequency subband and high-frequency subbands.
[0051] Specifically, as Figure 2 shown, construct a dual-tree complex wavelet transform (DTCWT) decomposition module to perform wavelet decomposition on the infrared image and the visible light image respectively. The input of this module is an infrared image or a visible light image with a size of 128×128×3, and 4 subbands are output, and the size of each subband is 64×64×3. This module uses the dual-tree complex wavelet transform (DTCWT) to decompose the infrared image into 4 subbands: a low-frequency subband (LL), a horizontal high-frequency subband (LH), a vertical high-frequency subband (HL), and a diagonal high-frequency subband (HH). The horizontal high-frequency subband (LH), the vertical high-frequency subband (HL), and the diagonal high-frequency subband (HH) form the high-frequency subbands.
[0052] Step 32: Input the high-frequency subbands into the noise processing module for denoising to obtain multiple denoised high-frequency subbands.
[0053] Specifically, step 32 specifically includes steps 321 to 323.
[0054] Step 321: Calculate the standard deviation of the noise of the high-frequency subband based on the median of all eigenvalues in the high-frequency subband and all eigenvalues in the high-frequency subband.
[0055] Step 322: Calculate the soft threshold of the high-frequency subband based on the standard deviation of the noise and the number of features of the high-frequency subband.
[0056] Step 323: Denoise the high-frequency subband based on the soft threshold to obtain the denoised high-frequency subband.
[0057] Specifically, the expression of the denoised high-frequency subband is as follows.
[0058] 。
[0059] 。
[0060] 。
[0061] Among them, is the j-th denoised high-frequency subband; is the sign function; are all the eigenvalues in the j-th high-frequency subband, that is , where n is the total number of wavelet coefficients; is the soft threshold of the j-th high-frequency subband; is the standard deviation of the noise in the j-th high-frequency subband; is the number of features in the j-th high-frequency subband; is the i-th eigenvalue in the j-th high-frequency subband; is the median of all the eigenvalues in the j-th high-frequency subband, that is, the value at the middle position after arranging the eigenvalue set in ascending order; 0.6745 is a constant factor of the normal distribution.
[0062] Specifically, soft threshold (ST) processing is performed on the three high-frequency subbands (LH, HL, HH) respectively. The input is the three high-frequency subbands with a size of 64×64×3, and the output is the three denoised high-frequency subbands with a size of 64×64×3. Soft threshold processing is applied to each high-frequency subband to remove noise.
[0063] Step 33: Input multiple low-frequency subbands and multiple denoised high-frequency subbands into the feature fusion module for feature fusion to obtain a feature fusion map.
[0064] Specifically, the operations in Step 31 and Step 32 are performed on the infrared image and the visible light image respectively to obtain 4 subbands after decomposition of the two modalities. Then, the subbands corresponding to the infrared image and the visible light image are concatenated together along the channel dimension to form a multi-channel feature map. Specifically, as Figure 2 shown, the 4 subbands with a size of 64×64×3 from the infrared image and the visible light image are concatenated to obtain 4 concatenated subbands LLc, LHc, HLc, HHc, and the size of each subband is 64×64×6.
[0065] Apply a 3×3 convolutional layer CONV to each concatenated subband, with the output channel number being 64, to extract richer features, and then through the batch normalization layer BN and the RELU activation function, 4 extracted subband feature maps are obtained, and the size of each feature map is 64×64×64.
[0066] Finally, as Figure 2 shown, the 4 extracted subband feature maps are concatenated together along the channel dimension to form a concatenated feature map F with a size of 64×64×256 t .
[0067] Step 34: Input the feature fusion map into the feature encoding module for multiple downsamplings to obtain a downsampled feature fusion map.
[0068] Specifically, asFigure 2 As shown, there is a construction feature encoding module, whose input is a feature map F with a size of 64×64×256 t , and the output is a feature map with a size of 4×4×1024. This module contains four consecutive encoding layers, and each encoding layer consists of a convolutional layer CONV, a batch normalization BN, and a max pooling layer MAXPOOL. Among them, the convolutional layer has a 3×3 convolutional kernel, a stride of 1, and a padding of 1, and the number of output channels is set to 128, 256, 512, and 1024 in sequence.
[0069] Through these four encoding layers, the original 64×64×256 feature map is gradually downsampled to 4×4×1024, and at the same time, the feature extraction ability is continuously enhanced, providing a rich feature representation for subsequent classification tasks.
[0070] Step 35: Input the sampled feature fusion map into the classification module for classification to obtain the identity of the maritime target.
[0071] Specifically, as Figure 2 shown, a classification module is constructed, whose input is a feature map with a size of 4×4×1024, and the output is a feature vector with a size of 1×1×2048. Specifically, a global average pooling layer GAPOOLING is used to compress the feature map into 1×1×1024, and then a fully connected layer FC with a length of 512 is used to map the 1024-dimensional feature to a 512-dimensional feature vector. Then, through a batch normalization layer BN and a RELU activation function, and finally through a fully connected layer FC with a length of N and a SOFTMAX layer, the final individual classification result is obtained.
[0072] The beneficial effects of the maritime target individual recognition method based on wavelet decomposition proposed in this application are mainly manifested in the following aspects.
[0073] (1) Through the dual-tree complex wavelet transform decomposition module in step 31, the maritime target image is decomposed into multiple sub-bands (LL, LH, HL, HH), and each sub-band represents features in different frequency ranges. Among them, the low-frequency sub-band (LL) mainly contains the structural information and large-scale features of the image, and the high-frequency sub-bands (LH, HL, HH) contain edge and detail information in the horizontal, vertical, and diagonal directions respectively. Through multi-scale decomposition, the model can simultaneously capture the global structural information (low-frequency sub-band) and local detail information (high-frequency sub-bands) in the image. This multi-scale feature extraction helps to improve the model's ability to understand complex scenes. Especially when dealing with maritime ship targets, it can better identify features such as the shape and texture of the ship. Among them, the edge and detail information in the high-frequency sub-bands are very important for target detection and classification tasks. Especially when the contrast between the target and the background is low, the high-frequency sub-bands can provide more discriminative features. Through the above method, the model can better learn the fine features related to identity.
[0074] (2) Considering that in infrared images, due to sensor limitations, noise may be more obvious. Therefore, an adaptive soft thresholding process is designed to be applied to the LH, HL, and HH subbands. This method can effectively remove high-frequency noise while retaining important edges and detailed information in the image, reduce the noise level in the high-frequency subbands, thereby improving the accuracy of subsequent feature extraction and classification, and enhancing the robustness of the algorithm.
[0075] (3) A method of first decomposing and then fusing is designed. The subbands corresponding to the infrared image and the visible light image are stitched together along the channel dimension to form a multi-channel feature map. This fusion method makes full use of the information of both modalities and enhances the representation ability of the model.
[0076] The multi-modal image sea target individual recognition method based on wavelet decomposition proposed in this application can obtain features of different frequencies by DTCWT decomposition of infrared images and visible light images collected by drones, then fuse and filter the noise, and finally improve the effectiveness of individual identity discrimination.
[0077] Based on the same inventive concept, the embodiment of this application also provides a multi-modal image sea target individual recognition system based on wavelet decomposition. The implementation solutions provided by this system to solve problems are similar to the implementation solutions recorded in the above method. Therefore, the specific limitations in one or more embodiments of the multi-modal image sea target individual recognition system based on wavelet decomposition provided below can refer to the limitations on the multi-modal image sea target individual recognition method based on wavelet decomposition in the above text, and will not be repeated here.
[0078] In an exemplary embodiment, a multi-modal image sea target individual recognition system based on wavelet decomposition is provided, which includes the following units.
[0079] An acquisition unit for acquiring sea target images; the sea target images include: infrared images and visible light images.
[0080] A construction unit for constructing and training an individual recognition model; the individual recognition model includes: a dual-tree complex wavelet transform decomposition module, a noise processing module, a feature fusion module, a feature encoding module, and a classification module connected in sequence; the dual-tree complex wavelet transform decomposition module is used to perform wavelet decomposition on the image to obtain multiple subbands; the subbands include: low-frequency subbands and high-frequency subbands; the noise processing module is used to remove the noise in the high-frequency subbands through soft thresholding; the feature fusion module is used to perform feature fusion on multiple low-frequency subbands and multiple high-frequency subbands after removing noise to obtain a feature fusion map; the feature encoding module is used to perform multiple downsamplings on the feature fusion map; the classification module is used to classify the downsampled feature fusion map to obtain the identity of the sea target.
[0081] An individual recognition unit is configured to input an image of a maritime target into a trained individual recognition model to obtain the identity of the maritime target, thereby completing the individual recognition of the maritime target.
[0082] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal, and its internal structural diagram may be as shown in Figure 3 the figure. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store individual recognition data of maritime targets. The input / output interface of the computer device is used for exchanging information between the processor and external devices. The communication interface of the computer device is used for communicating with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for individual recognition of maritime targets based on multi-modal images using wavelet decomposition.
[0083] Those skilled in the art can understand that Figure 3 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements. In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements the steps in the above method embodiments.
[0084] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, it implements the steps in the above method embodiments.
[0085] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, it implements the steps in the above method embodiments.
[0086] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0087] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0088] The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0089] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0090] In this article, specific examples are used to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A method for identifying individual targets at sea in multimodal images based on wavelet decomposition, characterized in that: The multimodal image marine target individual recognition method based on wavelet decomposition comprises: Acquire an image of a target at sea; the image of the target at sea includes: an infrared image and a visible light image; Constructing and training an individual recognition model; the individual recognition model comprises: a dual-tree complex wavelet transform decomposition module, a noise processing module, a feature fusion module, a feature encoding module and a classification module connected in sequence; Input the marine target image into the trained individual recognition model to obtain the identity of the marine target and complete the individual recognition of the marine target; The image of the maritime target is input into the trained individual recognition model to obtain the identity of the maritime target, which specifically includes: Inputting the marine target image into the dual-tree complex wavelet transform decomposition module for wavelet decomposition to obtain a plurality of sub-bands; the sub-bands include: a low-frequency sub-band and a high-frequency sub-band; Inputting the high frequency sub-band into the noise processing module for denoising to obtain a plurality of denoised high frequency sub-bands; Inputting a plurality of low-frequency sub-bands and a plurality of denoised high-frequency sub-bands into the feature fusion module for feature fusion to obtain a feature fusion graph; Inputting the feature fusion map into the feature encoding module for multiple downsampling to obtain a sampled feature fusion map; The sampled feature fusion map is input into the classification module for classification to obtain the identity of the marine target.
2. The method for identifying individual targets at sea in multimodal images based on wavelet decomposition according to claim 1 is characterized in that: The training process of the individual recognition model specifically includes: Constructing a sample data pair; the sample data pair includes: a sample infrared image, a sample visible light image and a sample identity; Inputting the sample infrared image and the sample visible light image into an individual recognition model to obtain a sample recognition result; A loss function is constructed according to the sample recognition results and the corresponding sample identity labels, and the individual recognition model parameters are iteratively optimized according to the loss function until the loss function reaches a minimum value or the iterative optimization round reaches a maximum value, and the iterative optimization is stopped to obtain the trained individual recognition model.
3. The method for identifying individual targets at sea in multimodal images based on wavelet decomposition according to claim 2 is characterized in that: Construct sample data pairs, including: Acquire a sample marine target image; the sample marine target image includes: a sample infrared image and a sample visible light image; Cropping the sample marine target image into a uniform size to obtain a cropped sample marine target image; The cropped sample marine target images are annotated with individual identities using a manual annotation method to obtain the sample identities; A sample data pair is constructed based on the sample infrared image, the sample visible light image and the sample identity.
4. The method for identifying individual targets at sea in multimodal images based on wavelet decomposition according to claim 1 is characterized in that: The high frequency sub-band is input into the noise processing module for denoising to obtain a plurality of denoised high frequency sub-bands, specifically including: Calculate the standard deviation of the noise in the high frequency sub-band based on the median of all eigenvalues in the high frequency sub-band and all eigenvalues in the high frequency sub-band; Calculate the soft threshold of the high frequency sub-band based on the standard deviation of the noise and the characteristic number of the high frequency sub-band; Based on the soft threshold, the high frequency sub-band is denoised to obtain the denoised high frequency sub-band.
5. The method for identifying individual targets at sea in multimodal images based on wavelet decomposition according to claim 4 is characterized in that: The expression of the high frequency subband after denoising is: ; ; ; in, is the jth high frequency subband after denoising; is a symbolic function; are all eigenvalues in the jth high-frequency subband; is the soft threshold of the j-th high frequency subband; is the standard deviation of the noise in the jth high frequency subband; is the characteristic number of the jth high frequency subband; is the i-th eigenvalue in the j-th high-frequency subband; is the median of all eigenvalues in the jth high-frequency subband.
6. The method for identifying individual targets at sea using multimodal images based on wavelet decomposition according to claim 1, characterized in that: The feature encoding module comprises: a plurality of encoding layers connected in sequence; The encoding layer includes: a convolutional layer, a batch normalization layer and a maximum pooling layer connected in sequence.
7. A multimodal image marine target individual recognition system based on wavelet decomposition, characterized in that: The method for identifying individual targets at sea in multimodal images based on wavelet decomposition as described in any one of claims 1 to 6, wherein the system for identifying individual targets at sea in multimodal images based on wavelet decomposition comprises: An acquisition unit is used to acquire an image of a target at sea; the image of the target at sea includes an infrared image and a visible light image; A construction unit is used to construct and train an individual recognition model; the individual recognition model includes: a dual-tree complex wavelet transform decomposition module, a noise processing module, a feature fusion module, a feature encoding module and a classification module connected in sequence; the dual-tree complex wavelet transform decomposition module is used to perform wavelet decomposition on the image to obtain multiple sub-bands; the sub-bands include: low-frequency sub-bands and high-frequency sub-bands; the noise processing module is used to remove noise in the high-frequency sub-bands through a soft threshold; the feature fusion module is used to perform feature fusion on multiple low-frequency sub-bands and multiple high-frequency sub-bands after noise removal to obtain a feature fusion map; the feature encoding module is used to perform multiple downsampling on the feature fusion map; the classification module is used to classify the downsampled feature fusion map to obtain the identity of the target at sea; The individual recognition unit is used to input the image of the maritime target into the trained individual recognition model to obtain the identity of the maritime target and complete the individual recognition of the maritime target.
8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the multimodal image marine target individual recognition method based on wavelet decomposition as described in any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for identifying individual targets at sea in multimodal images based on wavelet decomposition as described in any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for identifying individual targets at sea in multimodal images based on wavelet decomposition as described in any one of claims 1 to 6 is implemented.