Zero sample infrared target identification method and system based on cosine similarity matching

Through the method based on cosine similarity matching, the identification of weak infrared targets is solved, and the problem of insufficient recognition accuracy and robustness in the prior art is achieved, and the efficient identification of weak infrared targets and good generalization capabilities of the model are achieved.

CN119992566AActive Publication Date: 2025-05-13CHANGCHUN INST OF OPTICS FINE MECHANICS & PHYSICS CHINESE ACAD OF SCI

Patent Information

Application Number
CN202510131602.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-13
Estimated Expiration
2045-02-06

AI Technical Summary

Technical Problem

The existing infrared weak target recognition technology has shortcomings in recognition accuracy and robustness, especially the suppression of the classification domain generalization ability of the detection model by limited training data samples in diverse real domains.

Method used

A zero-sample infrared target recognition method based on cosine similarity matching is adopted to preprocess infrared images of unknown categories, and an enhanced image to be identified is generated, and a multi-scale target feature information is extracted using an infrared weak target feature extraction network with an encoder-decoder structure. Then, these feature information are input into the comparative language-image pre-trained model, and the cosine similarity is used to match, and text recognition labels, confidence and position information are generated.

Benefits of technology

It improves the recognition accuracy and robustness of weak infrared targets, enhances the generalization ability of the model in diverse real domains, and is suitable for zero-sample recognition scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992566A_ABST
    Figure CN119992566A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of infrared weak and small target recognition, and discloses a zero-sample infrared target recognition method and system based on cosine similarity matching, and the method comprises the steps: carrying out the preprocessing of an infrared image of which the category is not seen, and generating a to-be-recognized enhanced image; inputting the to-be-recognized enhanced image into an infrared weak feature extraction network with an encoder-decoder structure to extract multi-scale target features of the to-be-recognized image; and comparing a language-image pre-training model to generate image features of a plurality of types of unseen text features and multi-scale target features, and matching the text features and the image features by using cosine similarity to obtain text identification tags of the infrared images of the unseen types. According to the invention, the identification capability of the infrared weak and small target is improved through the target detection network with an encoder-decoder structure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of infrared dim small target recognition, and in particular to a zero-sample infrared target recognition method and system based on cosine similarity matching. Background Art

[0002] With the development of convolutional neural networks, the application of intelligent detection and recognition technology of targets in the recognition field has become increasingly mature. Infrared weak target recognition plays an important role in early warning, reconnaissance, interception and other fields. It has become a research hotspot and has extremely important engineering application prospects and practical value.

[0003] However, infrared weak targets are small in size, have few imaging pixels, low resolution, and little structure and texture information. In addition, the current intelligent detection model is data source driven, with a small number of samples and poor diversity. The above two difficulties restrict the accuracy of intelligent target detection and greatly limit the model's transferability and generalization capabilities in a wide range of real-world environments.

[0004] Conventional image classification models are often fully supervised and trained based on image datasets with category labels. They are highly data-dependent and are not suitable for small sample recognition application scenarios. They also limit the applicability and generalization ability of the model and are not suitable for task migration. Summary of the invention

[0005] The purpose of the present invention is to provide a zero-sample infrared target recognition method and system based on cosine similarity matching, which can solve at least one of the above-mentioned technical problems. The specific scheme is as follows: According to the specific embodiments disclosed in the present invention, the first aspect of the present invention discloses a zero-sample infrared target recognition method based on cosine similarity matching, which pre-processes infrared images of unseen categories to generate enhanced images to be recognized; Inputting the enhanced image to be identified into an infrared dim small target feature extraction network having an encoder-decoder structure, and using the infrared dim small target feature extraction network to extract multi-scale target feature information of the enhanced image to be identified; Inputting the unseen multiple categories and the multi-scale target feature information into the comparative language-image pre-training model to generate the unseen multiple categories of text features and the image feature I of the multi-scale target feature information, and the text feature I is converted into The text recognition label, confidence level and position information of the infrared image of the unseen category are matched with the image feature I to obtain the text recognition label, confidence level and position information of the infrared image of the unseen category.

[0006] Preferably, preprocessing the infrared image of the unseen category to generate an enhanced image to be identified includes: A one-stage blind denoising network is used to preprocess the infrared image of the unseen category to generate an enhanced image to be identified; wherein the one-stage blind denoising network includes: The feature extraction module includes multiple stacked convolutional layers, and a ReLU activation function is input after each convolution operation; The feature learning module based on the residual structure is composed of two ordinary convolutional layers stacked together, and the ReLU activation function is input after each convolutional layer; The image reconstruction module takes the output of the residual structure-based feature learning module as input, and includes: an average pooling layer and multiple stacked convolutional layers.

[0007] Preferably, the feature extraction module consists of a convolution , a convolutional layer with a dilation factor of 2 And a convolutional layer with a dilation factor of 4 The output of the feature extraction module is: ; in, X is the input infrared image of the unseen category, To pass through the convolutional layer The output after For the p The convolution kernel of the convolution layer, p A natural number from 1 to 8.

[0008] Preferably, in the image reconstruction module, the output of the residual structure-based feature learning module is input into an average pooling layer to obtain a pooled result. , the expression is: ; in, is a sliding window, is the number of elements in the sliding window; represents the central element of the sliding window, represents the elements of the sliding window, i , j is a natural number; Represents the sum of the values ​​of all elements in the sliding window; The pooled result is passed through the convolutional layer and convolutional layers Perform feature extraction to obtain feature output , the expression is: ; Output the features The output of the feature extraction module Perform dot multiplication to get the result of dot multiplication , the expression is: ; The result after dot multiplication Input convolutional layer , get the convolutional layer The characteristic output , the expression is: ; The convolutional layer Output Added to the input of the infrared image X of the unseen category, the feature output OUT of the enhanced image to be identified is generated, and the expression is: ; Among them, ReLU is the activation function.

[0009] Preferably, the infrared small target feature extraction network includes: Encoder layers, each of the encoder layers having different scales, for downsampling the enhanced image to be identified to generate a multi-scale feature map; A decoder layer, wherein the scale of each decoder layer corresponds to the scale of the encoder layer, and the multi-scale feature map is upsampled to extract texture structure and / or contour features; The context information extraction modules are respectively arranged between two adjacent encoder layers and two adjacent decoder layers, and are respectively used to extract features of the enhanced images to be identified at different scales and to fuse the texture structure information of the enhanced images to be identified obtained by each decoder layer.

[0010] Preferably, the infrared small target feature extraction network is used to extract multi-scale target feature information of the enhanced image to be identified, including: Extracting multi-scale features of the enhanced image to be identified by using full-scale skip connections between an encoder and a decoder based on a UNet3+ network structure; Fusing the multi-scale features of the enhanced image to be identified to generate a multi-scale fused feature map; Sending the multi-scale fusion feature map to a region proposal network to generate a proposal region for a candidate box; Pooling the proposed region with the multi-scale fusion feature map; The pooled features are passed through a fully connected layer network for classification and bounding box regression, and the category, confidence, and location information of the target are output.

[0011] Preferably, the context information extraction module comprises: Two consecutive cross-center convolution modules are used to obtain gradient features and expand the receptive field, and the dilated convolution rates of the cross-center convolution are 1 and 3 respectively; The cross attention mechanism module takes the feature map output by the cross center convolution module as input and passes through two Convolution, dimension reduction to obtain two feature maps, the two feature maps obtained by dimension reduction are subjected to affinity operation and aggregation operation to obtain features of different fine-grained levels.

[0012] According to the specific implementation mode disclosed in the present invention, the second aspect of the present invention discloses a zero-sample infrared target recognition device based on cosine similarity matching, comprising: Image processing unit: pre-processes infrared images of unseen categories and generates enhanced images to be identified; The unseen target detection unit inputs the to-be-recognized enhanced image into an infrared dim small target feature extraction network having an encoder-decoder structure, and uses the infrared dim small target feature extraction network to extract multi-scale target feature information of the to-be-recognized enhanced image; The unseen target recognition unit inputs the unseen multiple categories and the multi-scale target feature information into a comparative language-image pre-training model, generates text features of the unseen multiple categories and image features of the multi-scale target feature information, matches the text features with the image features using cosine similarity, and obtains text recognition labels, confidence levels, and position information of the infrared images of the unseen categories.

[0013] According to the specific embodiments disclosed in the present invention, the third aspect of the present invention discloses a computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method for editing the content in a document as described in any one of the above items is implemented.

[0014] According to the specific embodiments disclosed in the present invention, in a fourth aspect, the present invention discloses an electronic device, comprising: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method for editing the content in a document as described in any one of the above items.

[0015] Compared with the prior art, the above solution disclosed in the present invention has at least the following beneficial effects: The present invention improves the recognition capability of infrared dim small targets through a target detection network with an encoder-decoder structure. By comparing the language-image pre-training model combined with cosine similarity matching, the inhibition of the classification domain generalization ability of the detection model by limited training data samples in diverse real domains is reduced, and the robustness of zero-sample infrared dim small target recognition is improved, thereby achieving a coordinated improvement in the accuracy and robustness of infrared dim small target recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the disclosure of the present invention, and together with the specification, are used to explain the principles disclosed in the present invention. Obviously, the drawings described below are only some embodiments disclosed in the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings: Figure 1 It is a flow chart of a zero-sample infrared target recognition method based on cosine similarity matching provided by the present invention; Figure 2 is a schematic diagram of the network structure of an image denoising module according to an embodiment of the present invention; Figure 3 It is an architecture diagram of infrared dim small target feature extraction and network with encoder-decoder structure according to one embodiment of the present invention; Figure 4 It is a schematic diagram of a recognition algorithm flow based on cosine similarity matching according to an embodiment of the present invention; Figure 5 A schematic structural diagram of a zero-sample infrared target recognition device based on cosine similarity matching provided by the present invention; Figure 6 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0017] In order to make the purpose, technical scheme and advantages disclosed in the present invention clearer, the present invention will be further described in detail below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments disclosed in the present invention, rather than all the embodiments. Based on the embodiments disclosed in the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection disclosed in the present invention.

[0018] The terms used in the embodiments disclosed in the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the disclosure of the present invention. The singular forms "a", "said" and "the" used in the embodiments disclosed in the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings, and "multiple" generally includes at least two.

[0019] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0020] It should be understood that although the terms first, second, third, etc. may be used to describe in the disclosed embodiments of the present invention, these descriptions should not be limited to these terms. These terms are only used to distinguish the descriptions. For example, without departing from the scope of the disclosed embodiments of the present invention, the first may also be referred to as the second, and similarly, the second may also be referred to as the first.

[0021] It should also be noted that the term "includes", "comprising" or any other variation thereof is intended to cover non-exclusive inclusion, so that a commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprising a ..." do not exclude the existence of other identical elements in the commodity or device including the elements.

[0022] The following is combined with Figure 1-5 Detailed description of the alternative embodiments of the present disclosure. Example 1

[0023] Contrastive Language–Image Pretraining (CLIP) is a multimodal model developed by OpenAI that can understand the association between images and text. It converts images into feature vectors through image encoders (such as VisionTransformer or ResNet) and text encoders (such as Transformer) into feature vectors. It learns to distinguish between matching and mismatching image-text pairs by calculating the similarity of image and text feature vectors and comparing positive and negative samples, and performs reasoning and classification on unseen categories.

[0024] When using fully supervised learning to identify infrared dim small targets in related fields, the small number of samples in some fields leads to poor model diversity and is not suitable for task migration. The CILP model is trained with 400 million text-image pairs and can be applied to complex environments. It can solve the problems of low accuracy and poor robustness in identifying infrared dim small targets under domain offset conditions in wide-area backgrounds, and effectively improve the model's domain translation invariance and zero-sample detection capabilities for dim small targets.

[0025] Therefore, the first embodiment of the present invention provides a zero-sample infrared target recognition method based on cosine similarity matching, comprising the following steps: Step S102: pre-process the infrared image of the unseen category to generate an enhanced image to be identified.

[0026] This embodiment uses a one-stage blind denoising network based on a feature attention mechanism to preprocess the infrared images of unseen categories. Figure 2 As shown in Figure 1, it contains 8 convolutional layers and 1 average pooling layer. p The convolution kernel of the convolution operation is , p A natural number from 1 to 8, with a step length of s Both are 1.

[0027] Specifically, the one-stage blind denoising network includes: The feature extraction module consists of a common convolutional layer , a dilated convolutional layer with a dilation factor of 2 And a dilated convolutional layer with a dilation factor of 4 The convolution operations are stacked in sequence, and each convolution operation is output through a rectified linear unit ReLU to obtain the output of the feature extraction module. , The expression is: ; ; in, X is the input infrared image of the unseen category, To pass through the convolutional layer After the output.

[0028] The feature learning module based on the residual structure consists of two ordinary convolutional layers. , The output of the feature learning module based on the residual structure is obtained by stacking and inputting the ReLU activation function after each convolution operation: .

[0029] Image reconstruction module, including: an average pooling layer and three stacked convolutional layers , and .

[0030] Specifically, in the image reconstruction module, the output of the feature learning module is first Pooling in the average pooling layer, its expression is: ; in, is a sliding window, is the number of elements in the sliding window; represents the central element of the sliding window, represents the elements of the sliding window, i , j is a natural number; Represents the sum of the values ​​of all elements in the sliding window.

[0031] The pooled result is passed through the convolutional layer and convolutional layers Perform feature extraction to obtain feature output , the expression is: ; The output here Image reconstruction module The input is multiplied by: ; The result is input into the convolutional layer , the expression is: ; Finally, the convolutional layer The output of is added to the original input to generate the enhanced image to be identified, and the expression is: .

[0032] Step S104: input the enhanced image to be identified into an infrared dim small target feature extraction network having an encoder-decoder structure, and use the infrared dim small target feature extraction network to extract multi-scale target feature information of the enhanced image to be identified.

[0033] In this embodiment, by adopting Figure 3 The infrared dim target feature extraction network RAMF-Net shown in the figure extracts features from the enhanced image to be identified and generates the category, confidence, location information and other results of the target in the enhanced image to be identified.

[0034] like Figure 3 As shown, it includes: feature extraction network, feature fusion network, region proposal network, ROI pooling network and fully connected layer network.

[0035] The feature extraction network is a UNet3+ residual attention multi-scale skip connection network based on an encoder-decoder structure, and a feature extraction network of the context information extraction module DCCA based on differential convolution and cross attention mechanism is introduced in the encoding layer to extract multi-scale features of the enhanced image to be identified. The combination of UNet3+ network and DCCA module can adaptively refine feature extraction, deeply mine infrared weak target features, and improve the fine perception and recognition capabilities of targets.

[0036] The feature fusion network fuses the multi-scale features extracted from the enhanced image to be identified and generates a multi-scale fusion feature map; The region proposal network sends the multi-scale fusion feature map to the region proposal network RPN to generate the proposed region of the candidate box; The ROI pooling network inputs the proposed region and the multi-scale fusion feature map into the ROI pooling network for refinement; the regression network classifies and regresses the pooled features and outputs the category and precise position of the target.

[0037] Specifically, in the feature extraction network, each encoder layer has a different scale, which is used to downsample the enhanced image to be identified and generate a multi-scale feature map. The scale of each decoding layer corresponds to the scale of the encoder layer, and the multi-scale feature map is upsampled to extract the texture structure and contour features of the enhanced image to be identified; The multi-scale context feature extraction module DCCA based on differential convolution and residual attention mechanism is set between two adjacent encoder layers and two adjacent decoder layers, respectively, to extract the features of enhanced images to be identified at different scales and to fuse the texture structure information of the enhanced images to be identified obtained by each decoder layer.

[0038] The context information extraction module DCCA includes: Two consecutive cross-center convolution modules are used to obtain more gradient features and a larger receptive field. The dilated convolution rates of the cross-center convolution are 1 and 3 respectively. The cross attention mechanism module takes the feature map output by the cross center convolution module as input and passes through two Convolution,dimensionality reduction is performed to obtain two feature maps, and affinity and aggregation operations are performed on the two feature maps obtained by dimensionality reduction to obtain features of different fine-grained levels.

[0039] In this embodiment, the DCCA module effectively applies the cross-center difference convolution to the residual attention mechanism module. The cross-center difference convolution is a plug-and-play convolution module that can improve the network's ability to represent detailed features while reducing network redundancy by decoupling the two symmetrical cross sub-operators of horizontal, vertical and diagonal.

[0040] Since the cross-center difference convolution can effectively extract detail features, two consecutive cross-center convolutions with dilated convolution rates of 1 and 3 are used here to obtain more gradient features and a larger receptive field, thereby providing more powerful spatial feature information for infrared weak target recognition.

[0041] Then, the feature map output by the cross-center convolution module is sent to the repeated cross-attention mechanism module, and the input feature is recorded as , the input features are respectively passed through two Convolution and dimensionality reduction are used to obtain two feature maps Q and K. Among them, , 。

[0042] After obtaining the two matrices Q and K, the attention map is generated by applying the affinity operation and the Softmax operation. .

[0043] At every position in the space of Q u Corresponding to a vector , and can also be extracted from K and the position u Get the set of feature vectors in the same row or column .

[0044] represent The ith element of .

[0045] The calculation formula for affinity operation is as follows: ; in, yes and degree of relevance.

[0046] D is normalized by Softmax to obtain the feature map A , .

[0047] Add to the input feature map The filter is adaptive to the feature dimension. V Each position u has: vector and collection . Then the aggregation operation Aggregation can be defined as: ; In the formula, .

[0048] Further, Take the UNet3+ network as an example to illustrate how to construct a multi-scale feature map.

[0049] First, receive the feature maps of the encoder layer of the same scale The two skip connections above use different maximum pooling operations to convert the smaller scale encoder layer and After pooling and downsampling to unify the resolution of the feature map, the low-level semantic information is transmitted. Need to reduce the resolution by 4 times, The other two jump connections below use bilinear interpolation to interpolate the decoder. and Upsampling is performed to enlarge the resolution of the feature map by 4 times and 2 times respectively. After unifying the resolution, it is necessary to unify the number of feature maps through convolution operations with the same size of convolution kernel. Here, 64 3*3 filters are used for convolution operations to obtain 64 channels of feature maps. .

[0050] Feature Map The calculation formula is shown below.

[0051] ; In the formula, i Indicates the first i downsampling layers, and N represents the number of encoders.

[0052] Figure 3 The UNet3+ network uses 5 encoders. C Represents the convolution operation, function D and function U They represent downsampling and upsampling respectively, and Concat represents channel dimension splicing and fusion. H The function represents the feature aggregation mechanism (convolutional layer + batch normalization + ReLU activation function).

[0053] The feature map After being sent to the DCCA module, the convolution layer + batch normalization + ReLU activation function operation is performed through the feature aggregation mechanism to obtain .

[0054] Therefore, by The constructed multi-scale feature map expression is: .

[0055] Where H represents the feature aggregation mechanism, and DCCA represents the multi-scale context feature extraction module based on differential convolution and residual attention mechanism.

[0056] Step S106: Input the unseen multiple categories and the multi-scale target features into the comparative language-image pre-training model CLIP to generate unseen multiple categories of text features As well as the image features of the multi-scale target features I, the text features are transformed using cosine similarity Match it with the image feature I to obtain the text recognition label, confidence and location information of the infrared image of the unseen category, such as Figure 4 shown.

[0057] Due to the distribution difference between the model’s training data (source domain) and test data (target domain), the model’s performance on the target domain degrades. Therefore, in order to maintain the robust object detection invariance of the real-world “domain shift”, this embodiment uses the CLIP model and combines contrastive language-image pre-training.

[0058] First, convert N categories of unseen categories, such as Kongming lanterns, kites, hydrogen balloons, etc., into text data, and then input the converted N text data into the text encoder of the CLIP model to obtain N text features. .

[0059] Then, the multi-scale target feature information extracted by the infrared dim target feature extraction network RAMF-Net in step S104 is input into the image encoder of the CLIP model to obtain the image feature I , Calculate separately N Text features The cosine similarity with image feature I is as follows: ;

[0060] The text label corresponding to the maximum value of cosine similarity is the label corresponding to the best matching category. Finally, the text recognition label, confidence level and location information are output to complete the classification of unseen category images and achieve the coordinated improvement of the robustness and accuracy of infrared weak target recognition.

[0061] Example 2 The present invention also provides an apparatus embodiment that is consistent with the above embodiment, which is used to implement the method steps described in the above embodiment. The explanation based on the same name meaning is the same as the above embodiment, and has the same technical effect as the above embodiment, which will not be repeated here.

[0062] like Figure 5 As shown, the present invention discloses a zero-sample infrared target recognition device based on cosine similarity matching, comprising: Image processing unit 302: pre-processes the infrared image of the unseen category to generate an enhanced image to be identified; The unseen target detection unit 304 inputs the to-be-recognized enhanced image into a target detection network having an encoder-decoder structure, and uses the target detection network to extract multi-scale target features of the to-be-recognized enhanced image; Unseen class object recognition unit 306 generates unseen text features of multiple categories by comparing the language-image pre-training model and the image feature I of the multi-scale target feature, and the text feature I is converted into The text recognition label of the infrared image of the unseen category is obtained by matching the text recognition label of the infrared image of the unseen category.

[0063] Example 3 like Figure 6 As shown, this embodiment provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method steps described in the above embodiment.

[0064] Example 4 The disclosed embodiment of the present invention provides a non-volatile computer storage medium, wherein the computer storage medium stores computer executable instructions, and the computer executable instructions can execute the method steps described in the above embodiment.

[0065] Example 5 Reference below Figure 6 , which shows a schematic diagram of the structure of an electronic device suitable for implementing the disclosed embodiment of the present invention. The terminal device in the disclosed embodiment of the present invention may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments disclosed in the present invention.

[0066] like Figure 6As shown, the electronic device may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 to a random access memory (RAM) 403. In RAM 403, various programs and data required for the operation of the electronic device are also stored. The processing device 401, ROM 402, and RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0067] Typically, the following devices may be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Figure 6 An electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.

[0068] In particular, according to the embodiments disclosed in the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 409, or installed from the storage device 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above-mentioned functions defined in the method of the embodiment disclosed in the present invention are executed.

[0069] It should be noted that the computer-readable medium disclosed in the present invention may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present invention, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer readable signal media may also be any computer readable medium other than computer readable storage media, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0070] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0071] Computer program code for performing the operations disclosed in the present invention may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0072] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to the various embodiments disclosed in the present invention. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

Claims

1. A zero-sample infrared target recognition method based on cosine similarity matching, characterized in that: include: Preprocess the infrared images of unseen categories to generate enhanced images to be identified; Inputting the enhanced image to be identified into an infrared dim small target feature extraction network having an encoder-decoder structure, and using the infrared dim small target feature extraction network to extract multi-scale target feature information of the enhanced image to be identified; Inputting the unseen multiple categories and the multi-scale target feature information into the comparative language-image pre-training model to generate the unseen multiple categories of text features and the image feature I of the multi-scale target feature information, and the text feature I is converted into The text recognition label, confidence level and position information of the infrared image of the unseen category are obtained by matching the text recognition label, confidence level and position information of the infrared image of the unseen category.

2. The method according to claim 1, characterized in that The preprocessing of the infrared image of the unseen category to generate the enhanced image to be identified includes: A one-stage blind denoising network is used to preprocess the infrared image of the unseen category to generate an enhanced image to be identified; wherein the one-stage blind denoising network includes: The feature extraction module includes multiple stacked convolutional layers, and a ReLU activation function is input after each convolution operation; The feature learning module based on the residual structure is composed of two ordinary convolutional layers stacked together, and the ReLU activation function is input after each convolutional layer; The image reconstruction module takes the output of the residual structure-based feature learning module as input, and includes: an average pooling layer and multiple stacked convolutional layers.

3. The method according to claim 2, characterized in that The feature extraction module consists of a convolution , a convolutional layer with a dilation factor of 2 And a convolutional layer with a dilation factor of 4 The output of the feature extraction module is: ; in, X is the input infrared image of the unseen category, To pass through the convolutional layer The output after For the p The convolution kernel of the convolution layer, p A natural number from 1 to 8.

4. The method according to claim 2, characterized in that: In the image reconstruction module, The output of the residual structure-based feature learning module is input into the average pooling layer to obtain the pooled result , the expression is: , in, is a sliding window, is the number of elements in the sliding window; represents the central element of the sliding window, represents the elements of the sliding window, i , j is a natural number; Represents the sum of the values ​​of all elements in the sliding window; The pooled result is passed through the convolutional layer and convolutional layers Perform feature extraction to obtain feature output , the expression is: ; Output the features The output of the feature extraction module Perform dot multiplication to get the result of dot multiplication , the expression is: ; The result after dot multiplication Input convolutional layer , get the convolutional layer The characteristic output , the expression is: ; The convolutional layer Output Added to the input of the infrared image X of the unseen category, the feature output OUT of the enhanced image to be identified is generated, and the expression is: ; Among them, ReLU is the activation function.

5. The method according to claim 1, characterized in that The infrared small target feature extraction network comprises: Encoder layers, each of the encoder layers having different scales, for downsampling the enhanced image to be identified to generate a multi-scale feature map; A decoder layer, wherein the scale of each decoder layer corresponds to the scale of the encoder layer, and the multi-scale feature map is upsampled to extract texture structure and / or contour features; The context information extraction modules are respectively arranged between two adjacent encoder layers and two adjacent decoder layers, and are respectively used to extract features of the enhanced images to be identified at different scales and to fuse the texture structure information of the enhanced images to be identified obtained by each decoder layer.

6. The method according to claim 1, characterized in that The step of extracting multi-scale target feature information of the enhanced image to be identified by using the infrared small target feature extraction network includes: Extracting multi-scale features of the enhanced image to be identified by using full-scale skip connections between an encoder and a decoder based on a UNet3+ network structure; Fusing the multi-scale features of the enhanced image to be identified to generate a multi-scale fused feature map; Sending the multi-scale fusion feature map to a region proposal network to generate a proposal region for a candidate box; Pooling the proposed region with the multi-scale fusion feature map; The pooled features are passed through a fully connected layer network for classification and bounding box regression, and the category, confidence, and location information of the target are output.

7. The method according to claim 5, characterized in that The context information extraction module comprises: Two consecutive cross-center convolution modules are used to obtain gradient features and expand the receptive field, and the dilated convolution rates of the cross-center convolution are 1 and 3 respectively; The cross attention mechanism module takes the feature map output by the cross center convolution module as input and passes through two Convolution, dimension reduction to obtain two feature maps, the two feature maps obtained by dimension reduction are subjected to affinity operation and aggregation operation to obtain features of different fine-grained levels.

8. A zero-sample infrared target recognition device based on cosine similarity matching, characterized in that: include: Image processing unit: pre-processes infrared images of unseen categories and generates enhanced images to be identified; The unseen target detection unit inputs the to-be-recognized enhanced image into a target detection network having an encoder-decoder structure, and uses the target detection network to extract multi-scale target features of the to-be-recognized enhanced image; Unseen class object recognition unit, which generates unseen text features of multiple categories by comparing the language-image pre-training model and the image feature I of the multi-scale target feature, and the text feature I is converted into The text recognition label of the infrared image of the unseen category is obtained by matching the text recognition label of the infrared image of the unseen category.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the one or more processors to implement the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image blind denoising method and device and computer device

    CN110197183A

  • Infrared and visible light fusion imaging method based on deep learning

    CN113487530A

  • Method, device and equipment for detecting weak and small target in air and medium

    CN118097486A

  • Infrared small target detection method based on scene text information guidance

    CN118762364A

  • Infrared small target detection model establishment method and detection method based on multi-scale attention feature superposition

    CN118968012A

Cited By

  • Infrared image background suppression method based on zero-order learning

    CN120525742A

  • Language guidance feature decoupling infrared target detection method

    CN121330249A