Cross-modal image-text retrieval method and device based on pulse fusion

By adopting the pulse fusion method in image-text retrieval, using the pulse cross-attention fusion module and contrastive learning method, deep alignment of image and text information is achieved, the accuracy and efficiency of cross-modal retrieval are improved, and the problems of low retrieval accuracy and high energy consumption in existing technologies are solved.

CN120744149AActive Publication Date: 2025-10-03WUHAN UNIV OF TECH

Patent Information

Application Number
CN202511148012.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-16
Publication Date
2025-10-03
Estimated Expiration
2045-08-16

AI Technical Summary

Technical Problem

Existing image-text retrieval methods in the cross-modal field have problems such as low retrieval accuracy and efficiency, high retrieval energy consumption, and difficulty in aligning image-text pairs. In addition, the sparse characteristics of pulse neural networks lead to severe information loss, making it difficult to achieve effective cross-modal task learning.

Method used

By inputting the target image set and the target text set into the target detection network and word segmenter, the image region features and text word tokens are extracted offline, the pulse encoder is combined to perform pulse coding conversion of images and texts, the pulse cross attention fusion module is used to exchange and fusion inter-modal information, the weighted accumulation and average pooling are used to obtain the floating-point feature vector set, and the contrastive learning method is used to calculate the cosine similarity for image-text alignment.

Benefits of technology

It improves the accuracy and efficiency of image and text retrieval, reduces retrieval energy consumption and interference from irrelevant information, achieves deep alignment of image and text information, and enhances the feature expression capability between modalities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744149A_ABST
    Figure CN120744149A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-modal image-text retrieval method based on pulse fusion, and belongs to the technical field of image-text retrieval. The method comprises the following steps: inputting a target image set into a target detection network to obtain an image floating point code, and inputting a target text set into a word segmentation device to obtain a word floating point code; inputting the image floating point code and the word floating point code into a pulse encoder to obtain a first image pulse code and a first text pulse code; inputting the first image pulse code and the first text pulse code into a pulse cross attention fusion module to obtain a second image pulse code and a second text pulse code; respectively carrying out weighted accumulation and average pooling on the second image pulse code and the second text pulse code to obtain an image floating point feature vector set and a text floating point feature vector set, and calculating cosine similarity to obtain an image-text alignment result; and the text retrieval result of each image in the target image set is obtained based on the image-text alignment result, so that the accuracy and efficiency of image-text retrieval are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of image and text retrieval technology, and in particular relates to a cross-modal image and text retrieval method and device based on pulse fusion. Background Art

[0002] With the development of deep learning, image and text retrieval methods have also made great progress. Some methods using pre-trained networks can achieve good performance in terms of indicators. However, due to the complex scenarios and actual requirements of real-world data, these methods often ignore issues such as speed and energy consumption during model inference, which are very important factors for devices and systems. Image and text retrieval is an important task in the cross-modal field. It aims to retrieve the corresponding text in the dataset using an image as a query, and retrieve the corresponding image in the dataset using the text. With the development of the big data era, how to efficiently, quickly and accurately retrieve the corresponding samples in large-scale image and text databases has become a challenge. This task is of great significance in search engines, recommendation algorithms, user privacy and other fields. Therefore, reducing computing energy consumption while maintaining retrieval accuracy is a key challenge to achieve large-scale application of cross-modal retrieval technology.

[0003] While existing image-text retrieval methods have achieved impressive performance in single-modal tasks involving images and text, they remain largely unexplored in the multimodal domain of images and text. The application of spiking neural networks in this multimodal domain faces two major challenges: 1) Spiking neural networks with the same structure perform inferior to artificial neural networks; 2) the sparsity of spiking networks leads to significant information loss, hindering cross-modal learning. Existing spike coding schemes struggle to simultaneously preserve the complete information of both visual and textual semantics, making it difficult to align image-text features and exacerbating the cross-modal semantic gap. Existing image-text retrieval methods suffer from low accuracy and efficiency, high retrieval energy consumption, and difficulty aligning image-text pairs. Summary of the Invention

[0004] The present application aims to solve at least one of the technical problems existing in the prior art. To this end, the present application proposes a cross-modal image-text retrieval method and device based on pulse fusion, which improves the accuracy and efficiency of image-text retrieval.

[0005] In a first aspect, the present application provides a cross-modal image-text retrieval method based on pulse fusion, the method comprising: Acquire a target image set and a target text set, wherein the target image set includes n images and the target text set includes n texts; Input the target image set into the target detection network, extract the regional features of the target image set offline to obtain image floating point codes, input the target text set into the word segmenter to divide it into word tokens, and obtain word floating point codes; Inputting the image floating point code and the word floating point code into a pulse encoder and converting them into pulse codes to obtain a first image pulse code and a first text pulse code; The first image pulse code is used as a query and the first text pulse code is input into the pulse cross attention fusion module as a key-value pair to perform information exchange and pulse fusion between modalities to obtain a second image pulse code fused with text pulse information; The first text pulse code is used as a query and the first image pulse code is input into the pulse cross attention fusion module as a key-value pair to perform inter-modal information exchange and pulse fusion to obtain a second text pulse code fused with the image pulse information; Performing weighted accumulation and average pooling on the second image pulse code and the second text pulse code respectively to obtain an image floating-point feature vector set and a text floating-point feature vector set; Calculating the cosine similarity between the image floating-point feature vector set and the text floating-point feature vector set based on a contrastive learning method to obtain an image-text alignment result; Based on the image-text alignment result, a text retrieval result of each image in the target image set is obtained.

[0006] According to one embodiment of the present application, the target image set is input into a target detection network, regional features of the target image set are extracted offline to obtain image floating-point codes, and the target text set is input into a word segmenter to be divided into word tokens to obtain word floating-point codes, including: Input the target image set into the FasterRCNN network, perform target detection in a bottom-up manner, obtain floating-point features of the target area of ​​each image in the target image set and save them offline to obtain image floating-point encoding; Each text in the target text set is segmented according to a preset word length to obtain multiple text segments; Input the multiple text segments into the Tokenizer, divide them into separate tokens according to grammatical rules, and add index values; Encode each token after division to obtain the floating point encoding of the word.

[0007] According to one embodiment of the present application, the step of inputting the image floating-point code and the word floating-point code into a pulse encoder to convert them into pulse codes to obtain a first image pulse code and a first text pulse code includes: Repeat the image floating-point code and the word floating-point code T times respectively through a Repeat operation to obtain T image floating-point codes and T word floating-point codes; The T image floating-point codes are stacked and spliced ​​to obtain a spliced ​​image floating-point code, and the T word floating-point codes are stacked and spliced ​​to obtain a spliced ​​word floating-point code; The spliced ​​image floating-point code and the spliced ​​word floating-point code are respectively input into the BN layer for normalization, and then activated by neurons with learnable thresholds to obtain a first image pulse code and a first text pulse code.

[0008] According to one embodiment of the present application, the first image pulse code is used as a query, and the first text pulse code is input as a key-value pair into a pulse cross attention fusion module to perform inter-modal information exchange and pulse fusion to obtain a second image pulse code fused with text pulse information, including: Pulse-encoding the first image sequentially through a Linear layer, a BN layer, and a LIF layer to obtain an image query matrix; Passing the first text pulse code sequentially through the Linear layer, the BN layer, and the LIF layer to obtain a text key matrix and a text value matrix; The image query matrix, text key matrix and text value matrix are input into the pulse cross attention fusion module, and the information exchange and pulse fusion between modalities are carried out in turn through residual connection and pulse multilayer perceptron to obtain the second image pulse code that integrates the text pulse information.

[0009] According to one embodiment of the present application, the first text pulse code is used as a query, and the first image pulse code is input as a key-value pair into the pulse cross attention fusion module to perform inter-modal information exchange and pulse fusion to obtain the second text pulse code that fuses the image pulse information, including: Passing the first text pulse code through the Linear layer, the BN layer and the LIF layer in sequence to obtain a text query matrix; Pulse-encoding the first image sequentially through a Linear layer, a BN layer, and a LIF layer to obtain an image key matrix and an image value matrix; The text query matrix, image key matrix and image value matrix are input into the pulse cross attention fusion module, and the information exchange and pulse fusion between modalities are carried out in sequence through residual connection and pulse multilayer perceptron to obtain the second text pulse code that fused the image pulse information.

[0010] According to one embodiment of the present application, performing weighted accumulation and average pooling on the second image pulse coding and the second text pulse coding to obtain an image floating-point feature vector set and a text floating-point feature vector set includes: Multiplying the second image pulse code and the second text pulse code by a learnable weight coefficient respectively to obtain a weight value for each time step; Performing weighted accumulation on the second image pulse code and the second text pulse code in the time dimension based on the weight value of each time step to obtain a third image pulse code and a third text pulse code; Global average pooling is performed on the third image pulse code and the third text pulse code to obtain an image floating-point feature vector set and a text floating-point feature vector set.

[0011] According to one embodiment of the present application, the step of calculating the cosine similarity between the image floating-point feature vector set and the text floating-point feature vector set based on a contrastive learning method to obtain an image-text alignment result includes: The cosine similarity between the image floating-point feature vector set and the text floating-point feature vector set is calculated by the following formula to obtain the image-text alignment result:

[0012] in, represents a single image floating-point feature vector, represents a single text floating point feature vector, is the image floating point feature vector set, is the text floating point feature vector set, is the cosine similarity.

[0013] In a second aspect, the present application provides a cross-modal image-text retrieval device based on pulse fusion, the device comprising: An acquisition module, configured to acquire a target image set and a target text set, wherein the target image set includes n images and the target text set includes n texts; A floating-point encoding module is configured to input the target image set into a target detection network, extract regional features of the target image set offline, obtain image floating-point encoding, and input the target text set into a word segmenter to divide it into word tokens, thereby obtaining word floating-point encoding; a pulse coding module, configured to input the image floating point code and the word floating point code into a pulse coder and convert them into pulse codes to obtain a first image pulse code and a first text pulse code; a pulse fusion module configured to input the first image pulse code as a query and the first text pulse code as a key-value pair into a pulse cross attention fusion module to perform inter-modal information exchange and pulse fusion to obtain a second image pulse code fused with text pulse information; The first text pulse code is used as a query and the first image pulse code is input into the pulse cross attention fusion module as a key-value pair to perform inter-modal information exchange and pulse fusion to obtain a second text pulse code fused with the image pulse information; a temporal aggregation module, configured to perform weighted accumulation and average pooling on the second image pulse code and the second text pulse code, respectively, to obtain an image floating-point feature vector set and a text floating-point feature vector set; An image-text alignment module is used to calculate the cosine similarity between the image floating-point feature vector set and the text floating-point feature vector set based on a contrastive learning method to obtain an image-text alignment result; The retrieval module is used to obtain a text retrieval result for each image in the target image set based on the image-text alignment result.

[0014] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the cross-modal image and text retrieval method based on pulse fusion as described in the first aspect above is implemented.

[0015] In a fourth aspect, the present application provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the cross-modal image-text retrieval method based on pulse fusion as described in the first aspect above.

[0016] In a fifth aspect, the present application provides a chip comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the cross-modal image and text retrieval method based on pulse fusion as described in the first aspect.

[0017] In a sixth aspect, the present application provides a computer program product, comprising a computer program, which, when executed by a processor, implements the cross-modal image-text retrieval method based on pulse fusion as described in the first aspect above.

[0018] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become obvious from the description below, or will be learned through practice of the present application.

[0019] The present invention provides a cross-modal image-text retrieval method based on pulse fusion, which has the following advantages over the prior art: (1) The present invention inputs the target image set and target text set into the target detection network and word segmenter, extracts the regional features of the image offline and divides the text into word tokens, combines the pulse encoder to perform pulse coding conversion between the image and text, and uses the pulse cross-attention fusion module to perform information exchange and pulse fusion between modalities, effectively achieving deep alignment of image and text information and improving the accuracy of image-text retrieval. The floating-point feature vector set of the image and text is obtained by weighted accumulation and average pooling, and the cosine similarity is calculated using the contrastive learning method to obtain the alignment result of the image and text. It can more accurately provide the corresponding text retrieval result for each image in the target image set, improve the efficiency and accuracy of image-text retrieval, and reduce the retrieval energy consumption and the interference of irrelevant information.

[0020] (2) The present invention can achieve efficient information exchange and pulse fusion between image and text modalities by inputting the first image pulse code as a query and the first text pulse code as a key-value pair into the pulse cross-attention fusion module. By introducing the pulse cross-attention fusion module, pulse information from different modalities is fused and exchanged, so that the pulse code of a certain modality can remove its own noise information while obtaining the pulse information of another modality, maintaining its own sparse characteristics, and achieving alignment between the sparse representations of the image-text pair, thereby improving the efficiency and accuracy of image-text retrieval.

[0021] (3) The present invention inputs the first image pulse code as a query and the first text pulse code as a key-value pair into the pulse cross-attention fusion module, thereby achieving efficient information exchange and pulse fusion between the image and text modalities, effectively improving the feature expression capability between the modalities. By sequentially processing the image and text pulse codes through the Linear layer, the BN layer, and the LIF layer, the different modal data are subjected to cross-modal pulse code information exchange and fusion after the pulse-level {activation, retention, inhibition} process, thereby reducing information loss in the pulse neural network and achieving alignment of the image and text codes in the sparse pulse code space. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which: Figure 1 This is one of the flow charts of the cross-modal image-text retrieval method based on pulse fusion provided in an embodiment of the present application; Figure 2 Schematic diagram of the structure of the pulse cross attention module provided in the embodiment of the present application; Figure 3 This is a flow chart of implementing image-text retrieval based on a similarity matrix according to an embodiment of the present application; Figure 4This is the second flow chart of the cross-modal image-text retrieval method based on pulse fusion provided in an embodiment of the present application; Figure 5 Schematic diagram of the structure of a cross-modal image-text retrieval device based on pulse fusion provided in an embodiment of the present application; Figure 6 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0023] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.

[0024] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.

[0025] In the following, in combination with the accompanying drawings, the cross-modal image-text retrieval method based on pulse fusion, the cross-modal image-text retrieval device based on pulse fusion, the electronic device and the readable storage medium provided in the embodiments of the present application are described in detail through specific embodiments and their application scenarios.

[0026] Among them, the cross-modal image-text retrieval method based on pulse fusion can be applied to a terminal, and can be specifically executed by hardware or software in the terminal.

[0027] The terminal includes, but is not limited to, a portable communication device such as a mobile phone or tablet computer having a touch-sensitive surface (e.g., a touch screen display and / or a touch pad). It should also be understood that in some embodiments, the terminal may not be a portable communication device, but rather a desktop computer having a touch-sensitive surface (e.g., a touch screen display and / or a touch pad).

[0028] In the following embodiments, a terminal including a display and a touch-sensitive surface is described. However, it should be understood that the terminal may include one or more other physical user interface devices such as a physical keyboard, a mouse, and a joystick.

[0029] The embodiment of the present application provides a cross-modal image and text retrieval method based on pulse fusion. The execution subject of the cross-modal image and text retrieval method based on pulse fusion can be an electronic device or a functional module or functional entity in the electronic device that can implement the cross-modal image and text retrieval method based on pulse fusion. The electronic devices mentioned in the embodiment of the present application include but are not limited to mobile phones, tablets, computers, cameras and wearable devices, etc. The cross-modal image and text retrieval method based on pulse fusion provided in the embodiment of the present application is explained below using an electronic device as an example of the execution subject.

[0030] Figure 1 This is one of the flow charts of the cross-modal image-text retrieval method based on pulse fusion provided in the embodiment of the present application, such as Figure 1 As shown, the cross-modal image-text retrieval method based on pulse fusion includes: step 110, step 120, step 130, step 140, step 150, step 160, step 170 and step 180.

[0031] Step 110: Acquire a target image set and a target text set, wherein the target image set includes n images and the target text set includes n texts; It is easy to understand that a target image set and a target text set are obtained, the target image set includes n images, the target text set includes n texts, and each image corresponds to only one text.

[0032] Step 120: Input the target image set into a target detection network, extract regional features of the target image set offline to obtain image floating-point codes, input the target text set into a word segmenter to divide it into word tokens, and obtain word floating-point codes; In some embodiments, the step of inputting the target image set into a target detection network, extracting regional features of the target image set offline to obtain image floating-point codes, and inputting the target text set into a word segmenter to divide the text into word tokens to obtain word floating-point codes comprises: Input the target image set into the FasterRCNN network, perform target detection in a bottom-up manner, obtain floating-point features of the target area of ​​each image in the target image set and save them offline to obtain image floating-point encoding; Each text in the target text set is segmented according to a preset word length to obtain multiple text segments; Input the multiple text segments into the Tokenizer, divide them into separate tokens according to grammatical rules, and add index values; Encode each token after division to obtain the floating point encoding of the word.

[0033] It is easy to understand that the region-based retrieval method is used to align and match graphic regions and text words. First, the target image set is subjected to regional feature extraction. FasterRCNN is used as the backbone network to extract the regional features of the image offline in the form of floating-point encoding, which is expressed as For example, the pre-trained FasterRCNN network is used in a bottom-up manner on the experimental datasets Flickr30K and MSCOCO to perform the target detection task and obtain the floating-point feature encoding of the target area in the image. 36 regions are extracted from each image, and the dimension of each region encoding is 2048. The data is saved locally offline as direct input for subsequent networks.

[0034] The target text set uses Tokenizer to divide the complete sentence into word tokens and encode them into word encodings in floating point format, represented as , word floating point encoding includes the following steps: (1) The text sentences are truncated to 36 words in length, and [CLA] and [SEP] are added at the beginning and end to indicate the start and end of the sentence. If the length is less than 36, the [PAD] mark is added to indicate a blank.

[0035] (2) Use the Tokenizer provided by the BERT network to divide the text into separate tokens and corresponding indexes according to the syntax.

[0036] (3) Use torch.nn.embedding to directly encode the divided tokens to obtain the corresponding floating-point format encoding vector with an encoding dimension of 1024.

[0037] In this embodiment, the target image set is input into the FasterRCNN network and target detection is performed, the floating-point features of the target area are extracted offline, and the floating-point encoding of the image is obtained. The target text set is input into the word segmenter and divided into word tokens and encoded to obtain the floating-point encoding of the words. This facilitates subsequent image and text fusion processing, can improve the efficiency and accuracy of image and text data processing, reduce the consumption of computing resources, and provide higher quality feature representation for subsequent cross-modal fusion.

[0038] Step 130: Input the image floating point code and the word floating point code into a pulse encoder and convert them into pulse codes to obtain a first image pulse code and a first text pulse code; In some embodiments, the step of inputting the image floating point code and the word floating point code into a pulse encoder to convert them into pulse codes to obtain a first image pulse code and a first text pulse code comprises: Repeat the image floating-point code and the word floating-point code T times respectively through a Repeat operation to obtain T image floating-point codes and T word floating-point codes; The T image floating-point codes are stacked and spliced ​​to obtain a spliced ​​image floating-point code, and the T word floating-point codes are stacked and spliced ​​to obtain a spliced ​​word floating-point code; The spliced ​​image floating-point code and the spliced ​​word floating-point code are respectively input into the BN layer for normalization, and then activated by neurons with learnable thresholds to obtain a first image pulse code and a first text pulse code.

[0039] In order to meet the data requirements of the spiking neural network, that is, the binary pulse format, it is necessary to convert it into a pulse data format before inputting the network. The image floating point code and the word floating point code are input into the pulse encoder to increase the time dimension and convert the floating point code into a pulse code. The output pulse code is represented as and , specifically including the following steps: (1) Use the Repeat operation to encode the image into floating point and word floating point encoding Repeat T times, then stack and splice together to increase the time dimension. The calculation formula is as follows:

[0040]

[0041] in, Floating point encoding for the stacked image, Floating point encoding for the stacked words, Floating point encoding for images, Floating point encoding for words.

[0042] (2) The stacked image floating-point code and word floating-point code are passed through the BN layer to obtain normalized values. After the neurons are activated with a learnable threshold, a binary format pulse code is obtained. The calculation formula is as follows:

[0043]

[0044] in, For the first image pulse code, is the first text pulse coding, BN is batch normalization, and LIF is leaky integral firing neuron.

[0045] In this embodiment, by repeating, stacking, and concatenating image and word floating-point codes, the expressive power of information is enhanced while preserving the characteristics of the original data. Applying a batch normalization layer to the concatenated image and text data, combined with neuron activation using a learnable threshold, further improves the accuracy and robustness of the spike coding.

[0046] Step 140: Input the first image pulse code as a query and the first text pulse code as a key-value pair into a pulse cross attention fusion module to perform inter-modal information exchange and pulse fusion to obtain a second image pulse code fused with text pulse information; Figure 2 This is a schematic diagram of the structure of the pulse cross attention module provided in the embodiment of the present application. Figure 2 As shown in the figure, the image features extracted from FasterRCNN and the text features obtained directly by Tokenizer, after passing through the linear layer and batch normalization layer, are then passed through the LIF neurons with learnable thresholds to obtain pulse coding features, which are respectively used as Q, K, and V matrices. The cross-modal pulse coding is then fused through matrix dot product, and finally after passing through a linear layer and activation, the pulse coding that integrates the information of another modality is obtained.

[0047] It is worth noting that, unlike artificial neural networks, in pulse neural networks, the binary representation of pulse information flow faces a huge information loss problem. If pulse fusion is not adopted, the information obtained only from the sparse coding of its own modality is seriously insufficient. The lack of information makes it difficult to align cross-modal samples. The pulse fusion operation is indispensable. Therefore, the first image pulse code is used as the query Query and the first text pulse code is used as the key-value Key-Value input into the pulse cross-attention fusion module to perform information exchange and pulse fusion between modalities to obtain the second image pulse code that fused the text pulse information.

[0048] Step 150: Input the first text pulse code as a query and the first image pulse code as a key-value pair into a pulse cross attention fusion module to perform inter-modal information exchange and pulse fusion to obtain a second text pulse code fused with the image pulse information; Step 160: Perform weighted accumulation and average pooling on the second image pulse code and the second text pulse code respectively to obtain an image floating-point feature vector set and a text floating-point feature vector set; In some embodiments, performing weighted accumulation and average pooling on the second image pulse code and the second text pulse code to obtain an image floating-point feature vector set and a text floating-point feature vector set includes: Multiplying the second image pulse code and the second text pulse code by a learnable weight coefficient respectively to obtain a weight value for each time step; Performing weighted accumulation on the second image pulse code and the second text pulse code in the time dimension based on the weight value of each time step to obtain a third image pulse code and a third text pulse code; Global average pooling is performed on the third image pulse code and the third text pulse code to obtain an image floating-point feature vector set and a text floating-point feature vector set.

[0049] It is worth noting that the time dimension introduced before the network input needs to be eliminated after the network output in order to facilitate the calculation of similarity. Compared with the uncertainty brought by directly taking a certain time step, weighted addition of each time step can better reflect the global information. The second image pulse coding and the second text pulse coding need to be weighted accumulation in the time dimension to eliminate the time dimension, as well as average pooling in the region and token dimensions. The image floating point feature vector set and text floating point feature vector set , specifically including the following steps: (1) The second image pulse coding and the second text pulse coding with a time dimension need to be multiplied by a learnable weight coefficient to obtain the weight of each time step, and then accumulated in the time dimension to eliminate the time dimension. The calculation formula is as follows:

[0050]

[0051] in, For the The second image pulse code of time steps, For the The second text pulse code of time steps, Pulse code for the third image, For the third text pulse code, , is the learnable weight parameter corresponding to each time step.

[0052] (2) By performing global average pooling on the third image pulse code and the third text pulse code in N dimensions, an image floating-point feature vector set and a text floating-point feature vector set are obtained. The calculation formula is as follows:

[0053]

[0054] in, is the image floating point feature vector set, is a set of text floating-point feature vectors, and AveragePool represents the average pooling operation.

[0055] In this embodiment, by performing weighted accumulation and global average pooling on the second image pulse code and the second text pulse code, deep features of the image and text can be effectively extracted and their representation capabilities enhanced. By introducing a learnable weight coefficient for weighted accumulation at each time step, the correlation between images and text can be more accurately captured. Global average pooling further reduces the feature dimensionality and improves the comprehensiveness and accuracy of the feature information.

[0056] Step 170: Calculate the cosine similarity between the image floating-point feature vector set and the text floating-point feature vector set based on a contrastive learning method to obtain an image-text alignment result; Furthermore, the cosine similarity between the image and text feature vectors is calculated using the contrastive learning method for the image floating-point feature vector set and the text floating-point feature vector set to achieve alignment between the two. The contrastive learning method is the most commonly used method for cross-modal alignment, with the advantages of robustness and versatility. By calculating the similarity, the distance between the positive samples of the two modalities in the encoding space can be effectively shortened to achieve the purpose of matching alignment.

[0057] Step 180: Based on the image-text alignment result, obtain a text retrieval result for each image in the target image set.

[0058] Finally, based on the image-text alignment results, the text corresponding to each image in the target image set is obtained. Figure 3 This is a flow chart of implementing image-text retrieval based on a similarity matrix provided in an embodiment of the present application. Figure 3 As shown in the figure, the final output is a cosine similarity matrix, with the horizontal axis being the image number and the vertical axis being the text number. The closer the value on the diagonal is to 1, the higher the matching degree of the correct image-text pair is, and the closer the value on the other positions is to 0, the lower the matching degree of the incorrect image-text pair is. The final text retrieval result is: Figure 1 Matches with text 1, and picture n matches with text n; Figure 1 It does not match text 2, 3, 4, or n. The ideal state of the similarity matrix is ​​that all values ​​except those on the diagonal are 1 and all other values ​​are 0.

[0059] According to the cross-modal image-text retrieval method based on pulse fusion provided by the embodiment of the present application, the target image set and the target text set are input into the target detection network and the word segmenter, the regional features of the image are extracted offline and the text is divided into word tokens, the pulse encoder is combined to perform pulse coding conversion of the image and text, and the pulse cross-attention fusion module is used to perform information exchange and pulse fusion between modalities, effectively achieving deep alignment of image and text information and improving the accuracy of image-text retrieval. The floating-point feature vector set of the image and text is obtained by weighted accumulation and average pooling, and the cosine similarity is calculated using the contrastive learning method to obtain the alignment result of the image and text. It can more accurately provide corresponding text retrieval results for each image in the target image set, improve the efficiency and accuracy of image-text retrieval, and reduce retrieval energy consumption and interference from irrelevant information.

[0060] In some embodiments, the first image pulse code is used as a query and the first text pulse code is used as a key-value pair to be input into a pulse cross attention fusion module to perform inter-modal information exchange and pulse fusion to obtain a second image pulse code fused with text pulse information, including: Pulse-encoding the first image sequentially through a Linear layer, a BN layer, and a LIF layer to obtain an image query matrix; Passing the first text pulse code sequentially through the Linear layer, the BN layer, and the LIF layer to obtain a text key matrix and a text value matrix; The image query matrix, text key matrix and text value matrix are input into the pulse cross attention fusion module, and the information exchange and pulse fusion between modalities are carried out in turn through residual connection and pulse multilayer perceptron to obtain the second image pulse code that integrates the text pulse information.

[0061] It is easy to understand that the Q, K, and V matrices obtained after the first image pulse coding and the first text pulse coding pass through the {Linear, BN, LIF} layer are all in pulse form. The matrix dot product calculation can be reduced to logical and sum operations. The calculation formula is as follows:

[0062]

[0063]

[0064] in, is the image query matrix, is the text key matrix, is a matrix of text values.

[0065] The first image pulse code and the first text pulse code undergo a matrix multiplication in the pulse cross-attention process, undergoing a process of {excitation, retention, and inhibition}, which determines the pulse state after fusion. Important information from the text is activated and reflected in the image code; the common information between the image and the text is retained unchanged; the membrane potential information of the image itself that is not conducive to modal alignment is treated as noise suppression. The image query matrix, text key matrix, and text value matrix are input into the pulse cross-attention fusion module. The inter-modal information exchange and pulse fusion are carried out through residual connections and pulse multilayer perceptrons in turn, resulting in the second image pulse code that incorporates the text pulse information. The use of residual connections and pulse multilayer perceptrons (SMLPs) can ensure training stability. The calculation formula is as follows:

[0066]

[0067]

[0068] in, For pulse cross attention, For query, is the key, For value, Pulse code for the second image, The intermediate value of the second image pulse code.

[0069] It should be noted that since the pulse cross attention does not involve Softmax and scale operations, by exchanging the calculation order of Q, K and V matrix multiplication, that is, , the time complexity can be expressed as down to , to speed up the inference.

[0070] In this embodiment, by inputting the first image pulse code as a query and the first text pulse code as a key-value pair into the pulse cross-attention fusion module, efficient information exchange and pulse fusion between the image and text modalities can be achieved. By introducing the pulse cross-attention fusion module, pulse information from different modalities can be fused and exchanged, allowing the pulse code of one modality to remove its own noise information while obtaining the pulse information of the other modality, thereby maintaining its own sparse characteristics. This improves the alignment between the sparse representations of the image-text pair.

[0071] In some embodiments, the first text pulse code is used as a query and the first image pulse code is used as a key-value pair to input into a pulse cross attention fusion module to perform inter-modal information exchange and pulse fusion to obtain a second text pulse code that fuses the image pulse information, including: Passing the first text pulse code through the Linear layer, the BN layer and the LIF layer in sequence to obtain a text query matrix; Pulse-encoding the first image sequentially through a Linear layer, a BN layer, and a LIF layer to obtain an image key matrix and an image value matrix; The text query matrix, image key matrix and image value matrix are input into the pulse cross attention fusion module, and the information exchange and pulse fusion between modalities are carried out in sequence through residual connection and pulse multilayer perceptron to obtain the second text pulse code that fused the image pulse information.

[0072] It is easy to understand that the Q, K and V matrices obtained after the first image pulse coding and the first text pulse coding pass through the {Linear, BN, LIF} layer are all in pulse form. The calculation formula is as follows:

[0073]

[0074]

[0075] in, is the text query matrix, is the image key matrix, is the image value matrix.

[0076] The text query matrix, image key matrix, and image value matrix are input into the pulse cross attention fusion module. The information exchange and pulse fusion between modalities are carried out through residual connection and pulse multilayer perceptron in turn to obtain the second text pulse code that fuses the image pulse information. The calculation formula is as follows:

[0077]

[0078]

[0079] in, For the second text pulse code, The middle value of the second text pulse code.

[0080] In this embodiment, by inputting the first image pulse code as a query and the first text pulse code as a key-value pair into the pulse cross-attention fusion module, efficient information exchange and pulse fusion between the image and text modalities can be achieved, effectively improving the feature expression capability between the modalities. By sequentially processing the image and text pulse codes through the Linear layer, the BN layer, and the LIF layer, the different modal data undergo cross-modal pulse code information exchange and fusion after the pulse-level {activation, retention, inhibition} process, reducing information loss in the pulse neural network and achieving alignment of the image and text codes in the sparse pulse code space.

[0081] In some embodiments, the calculating the cosine similarity between the image floating-point feature vector set and the text floating-point feature vector set based on the contrastive learning method to obtain the image-text alignment result includes: The cosine similarity between the image floating-point feature vector set and the text floating-point feature vector set is calculated by the following formula to obtain the image-text alignment result:

[0082] in, represents a single image floating-point feature vector, represents a single text floating point feature vector, is the image floating point feature vector set, is the text floating point feature vector set, is the cosine similarity.

[0083] It is easy to understand that after obtaining the cosine similarity between the image floating point feature vector set and the text floating point feature vector set, we calculate The positive sample pairs that match each other are represented by and The similarity between and The similarity between the unmatched negative sample pairs represented.

[0084] In some embodiments, the present application also provides a cross-modal image-text retrieval model based on pulse fusion, including the following modules: The image region feature extraction module uses FasterRCNN in a bottom-up manner to extract regional features from images, covering objects appearing in the image. It extracts 36 regional features from each image and outputs a regional feature code with a dimension of 36×2048, in the format of a floating-point value vector.

[0085] The text token segmentation and encoding module uses BERT's Tokenizer to segment text sentences into separate tokens. The text data is truncated at a length of 36 and the [CLS] and [SEP] tags are added to the beginning and end. If the length is less than 36, the [PAD] tag is added. The tokens are directly encoded into floating-point vectors using torch.nn.embedding, with an encoding dimension of 1024.

[0086] The pulse coding conversion module uses a method of repeating T times and stacking to add the time dimension T to the original data dimension. After normalization, it is sent to the neuron activation with a learnable threshold to obtain the feature encoding in pulse format, that is, binary format.

[0087] The pulse cross attention fusion module interacts and fuses the pulse codes of different modalities, and the pulse signals undergo {excitation, maintenance, inhibition} to be aligned. The pulse codes of the input pulse cross attention fusion module are respectively subjected to {Linear, BN, LIF} to obtain Q, K, V matrices, which do not include Softmax and Relu activation functions and are replaced by LIF neurons. The matrix multiplication operations are all 0-1 matrix multiplications, which can be converted to logical AND and addition operations. By exchanging the matrix calculation order, we have Linear complexity.

[0088] The time dimension aggregation and pooling module weights and accumulates the time dimension using T learnable weight coefficients to eliminate the time dimension T. It then uses a global average pooling operation to pool the token dimension, ultimately obtaining a feature vector of dimension 1×D. The cosine similarity calculation and alignment module calculates the cosine similarity between feature vectors from different modalities, subtracts the similarity between positive samples from the similarity between negative samples, and keeps the loss decreasing in the downward direction.

[0089] The loss function is set to train the pulse fusion cross-modal image-text retrieval model. The calculation formula of the loss function is as follows:

[0090] in, is the temperature coefficient, used to adjust the loss function curve, is the boundary value of the hyperparameter control loss, B is the batch size, is the loss function, is the first cosine similarity between unmatched negative sample pairs, is the second cosine similarity between unmatched negative sample pairs, is the similarity between the positive sample pairs that match each other.

[0091] Figure 4This is the second flow chart of the cross-modal image-text retrieval method based on pulse fusion provided in the embodiment of the present application, such as Figure 4 As shown, FasterRCNN is used offline to extract regional feature codes of the image on the image side, and Tokenizer is used on the text side to convert the text into tokens and directly encode the word feature codes. The image and text floating-point feature codes are sent to the pulse code converter, and the floating-point code is increased in time dimension and converted into a binary pulse code. The image pulse code is used as the query Query, and the text pulse code is used as the key value Key-Value, which is input into the pulse cross-attention fusion module to obtain the image pulse code that is fused with the text pulse information. The text pulse code is used as the query Query, and the image pulse code is used as the key value Key-Value, which is input into the pulse cross-attention fusion module to obtain the text pulse code that is fused with the image pulse information. The pulse codes are weightedly accumulated in the time dimension and averagely pooled in the token number dimension to obtain the final feature vector. The contrastive learning method is used to calculate the similarity between the image feature vector and the text feature vector to achieve cross-modal image-text alignment.

[0092] In this embodiment, by calculating the cosine similarity between the image floating-point feature vector set and the text floating-point feature vector set, image and text alignment can be achieved efficiently. By comparing the image floating-point feature vector set with the text floating-point feature vector set to obtain the image-text alignment result, the efficiency and accuracy of image-text retrieval are improved, and retrieval energy consumption and interference from irrelevant information are reduced.

[0093] The cross-modal image-text retrieval method based on pulse fusion provided in the embodiment of the present application can be executed by a cross-modal image-text retrieval device based on pulse fusion. In the embodiment of the present application, the cross-modal image-text retrieval method based on pulse fusion is executed by a cross-modal image-text retrieval device based on pulse fusion as an example to illustrate the cross-modal image-text retrieval device based on pulse fusion provided in the embodiment of the present application.

[0094] The present application also provides a cross-modal image and text retrieval device based on pulse fusion, such as Figure 5 As shown, the cross-modal image-text retrieval device based on pulse fusion includes: an acquisition module 510, a floating-point encoding module 520, a pulse encoding module 530, a pulse fusion module 540, a time aggregation module 550, an image-text alignment module 560 and a retrieval module 570.

[0095] An acquisition module 510 is configured to acquire a target image set and a target text set, wherein the target image set includes n images and the target text set includes n texts; A floating-point encoding module 520 is configured to input the target image set into a target detection network, extract regional features of the target image set offline, obtain image floating-point encodings, and input the target text set into a word segmenter to divide the text into word tokens to obtain word floating-point encodings. A pulse coding module 530 is configured to input the image floating point code and the word floating point code into a pulse coder to convert them into pulse codes, thereby obtaining a first image pulse code and a first text pulse code; a pulse fusion module 540 configured to input the first image pulse code as a query and the first text pulse code as a key-value pair into a pulse cross attention fusion module to perform inter-modal information exchange and pulse fusion to obtain a second image pulse code fused with text pulse information; The first text pulse code is used as a query and the first image pulse code is input into the pulse cross attention fusion module as a key-value pair to perform inter-modal information exchange and pulse fusion to obtain a second text pulse code fused with the image pulse information; A temporal aggregation module 550 is configured to perform weighted accumulation and average pooling on the second image pulse code and the second text pulse code, respectively, to obtain an image floating-point feature vector set and a text floating-point feature vector set; An image-text alignment module 560 is configured to calculate the cosine similarity between the image floating-point feature vector set and the text floating-point feature vector set based on a contrastive learning method to obtain an image-text alignment result; The retrieval module 570 is configured to obtain a text retrieval result for each image in the target image set based on the image-text alignment result.

[0096] According to the cross-modal image-text retrieval method based on pulse fusion provided by the embodiment of the present application, the target image set and the target text set are input into the target detection network and the word segmenter, the regional features of the image are extracted offline and the text is divided into word tokens, the pulse encoder is combined to perform pulse coding conversion of the image and text, and the pulse cross-attention fusion module is used to perform information exchange and pulse fusion between modalities, effectively achieving deep alignment of image and text information and improving the accuracy of image-text retrieval. The floating-point feature vector set of the image and text is obtained by weighted accumulation and average pooling, and the cosine similarity is calculated using the contrastive learning method to obtain the alignment result of the image and text. It can more accurately provide corresponding text retrieval results for each image in the target image set, improve the efficiency and accuracy of image-text retrieval, and reduce retrieval energy consumption and interference from irrelevant information.

[0097] The cross-modal image and text retrieval device based on pulse fusion provided in the embodiment of the present application can achieve Figures 1 to 4 To avoid repetition, the various processes implemented in the embodiment of the cross-modal image-text retrieval method based on pulse fusion are not described here.

[0098] In some embodiments, as Figure 6 As shown, an embodiment of the present application also provides an electronic device 600, including a processor 601, a memory 602, and a computer program stored in the memory 602 and executable on the processor 601. When the program is executed by the processor 601, each process of the embodiment of the above-mentioned cross-modal image and text retrieval method based on pulse fusion is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.

[0099] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.

[0100] An embodiment of the present application also provides a non-transitory computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the various processes of the above-mentioned cross-modal image and text retrieval method embodiment based on pulse fusion, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0101] The processor is the processor in the electronic device in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk.

[0102] An embodiment of the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned cross-modal image-text retrieval method based on pulse fusion.

[0103] The processor is the processor in the electronic device in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.

[0104] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, which is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-mentioned cross-modal image and text retrieval method embodiment based on pulse fusion, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0105] It should be understood that the chip mentioned in the embodiments of the present application can also be called a device-level chip, a device chip, a chip device, or an on-chip device chip, etc.

[0106] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0107] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the cross-modal image-text retrieval method based on pulse fusion of each embodiment of the present application.

[0108] In the description of this application, "first feature" and "second feature" may include one or more such features.

[0109] In the description of this application, “plurality” means two or more.

[0110] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

[0111] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "illustrative embodiments," "examples," "specific examples," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0112] Although the embodiments of the present application have been shown and described, those skilled in the art will appreciate that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and intent of the present application, and that the scope of the present application is defined by the claims and their equivalents.

Claims

1. A cross-modal image-text retrieval method based on pulse fusion, characterized in that: The method comprises: Acquire a target image set and a target text set, wherein the target image set includes n images and the target text set includes n texts; Input the target image set into the target detection network, extract the regional features of the target image set offline to obtain image floating point codes, input the target text set into the word segmenter to divide it into word tokens, and obtain word floating point codes; Inputting the image floating point code and the word floating point code into a pulse encoder and converting them into pulse codes to obtain a first image pulse code and a first text pulse code; The first image pulse code is used as a query and the first text pulse code is input into the pulse cross attention fusion module as a key-value pair to perform information exchange and pulse fusion between modalities to obtain a second image pulse code fused with text pulse information; The first text pulse code is used as a query and the first image pulse code is input into the pulse cross attention fusion module as a key-value pair to perform inter-modal information exchange and pulse fusion to obtain a second text pulse code fused with the image pulse information; Performing weighted accumulation and average pooling on the second image pulse code and the second text pulse code respectively to obtain an image floating-point feature vector set and a text floating-point feature vector set; Calculating the cosine similarity between the image floating-point feature vector set and the text floating-point feature vector set based on a contrastive learning method to obtain an image-text alignment result; Based on the image-text alignment result, a text retrieval result of each image in the target image set is obtained.

2. The cross-modal image-text retrieval method based on pulse fusion according to claim 1 is characterized in that: The target image set is input into a target detection network, regional features of the target image set are extracted offline to obtain image floating point codes, and the target text set is input into a word segmenter to be divided into word tokens to obtain word floating point codes, including: Input the target image set into the FasterRCNN network, perform target detection in a bottom-up manner, obtain floating-point features of the target area of ​​each image in the target image set and save them offline to obtain image floating-point encoding; Each text in the target text set is segmented according to a preset word length to obtain multiple text segments; Input the multiple text segments into the Tokenizer, divide them into separate tokens according to grammatical rules, and add index values; Encode each token after division to obtain the floating point encoding of the word.

3. The cross-modal image-text retrieval method based on pulse fusion according to claim 1, characterized in that: The step of inputting the image floating point code and the word floating point code into a pulse encoder and converting them into pulse codes to obtain a first image pulse code and a first text pulse code comprises: Repeat the image floating-point code and the word floating-point code T times respectively through a Repeat operation to obtain T image floating-point codes and T word floating-point codes; The T image floating-point codes are stacked and spliced ​​to obtain a spliced ​​image floating-point code, and the T word floating-point codes are stacked and spliced ​​to obtain a spliced ​​word floating-point code; The spliced ​​image floating-point code and the spliced ​​word floating-point code are respectively input into the BN layer for normalization, and then activated by neurons with learnable thresholds to obtain a first image pulse code and a first text pulse code.

4. The cross-modal image-text retrieval method based on pulse fusion according to claim 1, characterized in that: The first image pulse code is used as a query, and the first text pulse code is input into the pulse cross attention fusion module as a key-value pair to perform inter-modal information exchange and pulse fusion to obtain a second image pulse code fused with text pulse information, including: Pulse-encoding the first image sequentially through a Linear layer, a BN layer, and a LIF layer to obtain an image query matrix; Passing the first text pulse code sequentially through the Linear layer, the BN layer, and the LIF layer to obtain a text key matrix and a text value matrix; The image query matrix, text key matrix and text value matrix are input into the pulse cross attention fusion module, and the information exchange and pulse fusion between modalities are carried out in turn through residual connection and pulse multilayer perceptron to obtain the second image pulse code that integrates the text pulse information.

5. The cross-modal image-text retrieval method based on pulse fusion according to claim 1 is characterized in that: The first text pulse code is used as a query, and the first image pulse code is input as a key-value pair into a pulse cross attention fusion module to perform inter-modal information exchange and pulse fusion to obtain a second text pulse code that fuses the image pulse information, including: Passing the first text pulse code through the Linear layer, the BN layer and the LIF layer in sequence to obtain a text query matrix; Pulse-encoding the first image sequentially through a Linear layer, a BN layer, and a LIF layer to obtain an image key matrix and an image value matrix; The text query matrix, image key matrix and image value matrix are input into the pulse cross attention fusion module, and the information exchange and pulse fusion between modalities are carried out in sequence through residual connection and pulse multilayer perceptron to obtain the second text pulse code that fused the image pulse information.

6. The cross-modal image-text retrieval method based on pulse fusion according to claim 1, characterized in that: The step of performing weighted accumulation and average pooling on the second image pulse code and the second text pulse code to obtain an image floating-point feature vector set and a text floating-point feature vector set includes: Multiplying the second image pulse code and the second text pulse code by a learnable weight coefficient respectively to obtain a weight value for each time step; Performing weighted accumulation on the second image pulse code and the second text pulse code in the time dimension based on the weight value of each time step to obtain a third image pulse code and a third text pulse code; Global average pooling is performed on the third image pulse code and the third text pulse code to obtain an image floating-point feature vector set and a text floating-point feature vector set.

7. The cross-modal image-text retrieval method based on pulse fusion according to claim 1 is characterized in that: The method of calculating the cosine similarity between the image floating point feature vector set and the text floating point feature vector set based on the contrastive learning method to obtain the image-text alignment result includes: The cosine similarity between the image floating-point feature vector set and the text floating-point feature vector set is calculated by the following formula to obtain the image-text alignment result: ; in, represents a single image floating-point feature vector, represents a single text floating point feature vector, is the image floating point feature vector set, is the text floating point feature vector set, is the cosine similarity.

8. A cross-modal image-text retrieval device based on pulse fusion, implemented using the cross-modal image-text retrieval method based on pulse fusion according to any one of claims 1 to 7, characterized in that: The device comprises: An acquisition module, configured to acquire a target image set and a target text set, wherein the target image set includes n images and the target text set includes n texts; A floating-point encoding module is configured to input the target image set into a target detection network, extract regional features of the target image set offline, obtain image floating-point encoding, and input the target text set into a word segmenter to divide it into word tokens, thereby obtaining word floating-point encoding; a pulse coding module, configured to input the image floating point code and the word floating point code into a pulse coder and convert them into pulse codes to obtain a first image pulse code and a first text pulse code; a pulse fusion module configured to input the first image pulse code as a query and the first text pulse code as a key-value pair into a pulse cross attention fusion module to perform inter-modal information exchange and pulse fusion to obtain a second image pulse code fused with text pulse information; The first text pulse code is used as a query and the first image pulse code is input into the pulse cross attention fusion module as a key-value pair to perform inter-modal information exchange and pulse fusion to obtain a second text pulse code fused with the image pulse information; a temporal aggregation module, configured to perform weighted accumulation and average pooling on the second image pulse code and the second text pulse code, respectively, to obtain an image floating-point feature vector set and a text floating-point feature vector set; An image-text alignment module is used to calculate the cosine similarity between the image floating-point feature vector set and the text floating-point feature vector set based on a contrastive learning method to obtain an image-text alignment result; The retrieval module is used to obtain a text retrieval result for each image in the target image set based on the image-text alignment result.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the pulse fusion-based cross-modal image-text retrieval method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the cross-modal image-text retrieval method based on pulse fusion as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Pulse neural network multi-mode lip reading method and system based on attention mechanism

    CN115482582A

  • Multi-modal emotion recognition method based on spiking neural network and attention mechanism

    CN118861805A

  • Multi-scale information dynamic fusion remote sensing cross-modal image-text retrieval method

    CN118939821A

Cited By

  • Bearing fault diagnosis method and fault self-monitoring device based on pulse neural network

    CN121257618A

  • Bearing fault diagnosis method and fault self-monitoring device based on pulse neural network

    CN121257618B

  • Channel buoy detection method based on fusion of multi-mode pulse neural network and visual Transform

    CN121616952A