An image retrieval system and method based on complementary semantic alignment and symmetric retrieval

CN121434432BActive Publication Date: 2026-09-29NINGBO UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511281533.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2026-09-29
Estimated Expiration
2045-09-09

AI Technical Summary

Technical Problem

[0005]本发明要解决的技术问题是如何克服现有利用卷积神经网络提取图像特征的技术方案存在特征表达割裂、对齐机制僵化和无法根据图像内容复杂度动态调整权重的技术缺陷,为克服以上现有技术的缺陷,本发明提供一种基于互补语义对齐和对称检索的图像检索系统和方法,具体包含一种基于互补语义对齐和对称检索的图像检索系统和一种基于互补语义对齐和对称检索的图像检索方法

Benefits of technology

[0024]本发明所公开的方法,在构建图像数据库之后,利用跨模态约束损失函数对层级特征提取装置进行参数优化,以获得经过参数优化后的层级特征提取装置,从而结合跨模态语义约束引入文本监督信号,在损失函数中联合优化视觉与文本特征空间,达到增强语义对齐能力的技术效果。随后通过参数优化后的层级特征提取装置以采用全局平均池化方式获得查询图像和各数据库图像的全局特征向量,并通过计算归一化空间权重矩阵与掩码矩阵的哈达玛积的方式获得查询图像和各数据库图像的加权局部特征集合,从而不仅实现特征提取,而且在特征提取阶段实现全局语义与局部细节的深度耦合,克服传统方案中特征割裂问题。随后通过条件式对称对齐装置以利用条件式路径切换策略,根据全局特征置信度动态选择对称对齐(即全局-局部与局部-全局)或容错对齐(即全局-局部与局部-局部)路径,显著提升遮挡及复杂场景的鲁棒性,避免对齐机制僵化。最后通过结果生成装置,基于双向分数差异度与全局置信度生成自适应权重,实现针对不同内容复杂度的动态决策优化,从而在获得查询图像的检索结果的同时,达到根据图像内容复杂度动态调整权重的技术效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121434432B_ABST
    Figure CN121434432B_ABST
Patent Text Reader

Abstract

The application relates to an image retrieval system and method based on complementary semantic alignment and symmetric retrieval, wherein a hierarchical feature extraction device is arranged to obtain global feature vectors of a query image and each database image in a global average pooling manner, and a weighted local feature set of the query image and each database image is obtained in a manner of calculating a Hadamard product of a normalized space weight matrix and a mask matrix, so as to overcome the feature fragmentation problem in a traditional scheme. A conditional symmetric alignment device is arranged to utilize a conditional path switching strategy, to dynamically select a symmetric alignment path or a fault-tolerant alignment path according to a global feature confidence, to significantly improve the robustness of a shielding and complex scene, and to avoid the alignment mechanism from being rigid. A result generation device is arranged, an adaptive weight is generated based on a bidirectional score difference degree and a global confidence to obtain a search result, and dynamic decision optimization for different content complexities is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer modeling and systems technology, and more specifically, to an image retrieval system and method based on complementary semantic alignment and symmetric retrieval. Background Technology

[0002] Content-based image retrieval (CBIR) is a core requirement for managing massive amounts of visual data, especially in mobile internet scenarios where users' demand for real-time visual searches by photographing objects has exploded. Traditional solutions for content-based image retrieval typically rely on text-based tagging, but this approach suffers from high annotation costs, strong subjectivity, and difficulty in capturing visual details, thus failing to meet the need for precise retrieval.

[0003] To overcome these shortcomings, existing technologies disclose technical solutions for extracting image features using convolutional neural networks (CNNs). These solutions utilize encoder and decoder structures formed by combining CNNs to extract global or local features, and employ fusion strategies and cross-modal alignment methods to achieve feature fusion and alignment, thereby achieving content-based image retrieval.

[0004] However, this technical solution faces significant challenges in complex scenarios: On the one hand, it requires local feature aggregation during feature extraction or fusion, but the local feature aggregation scheme used in this solution independently uses global or local features, making it difficult to establish semantic associations across images, thus leading to fragmented feature representation. On the other hand, in the cross-modal alignment process, the cross-modal alignment scheme (such as image-text alignment in VSE++) adopts a one-way fixed matching strategy (only global to local or local to global), ignoring the synergistic effect of bidirectional semantic interaction. This not only leads to a rigid alignment mechanism but also causes a lack of dynamic adaptability, making it impossible to dynamically adjust weights according to the complexity of image content, resulting in excessive fluctuations in retrieval accuracy for simple objects (such as single products) and complex scenes (such as pedestrians in street view). Summary of the Invention

[0005] The technical problem to be solved by this invention is how to overcome the technical defects of existing technologies that use convolutional neural networks to extract image features, such as fragmented feature representation, rigid alignment mechanism, and inability to dynamically adjust weights according to the complexity of image content. In order to overcome the above defects of the prior art, this invention provides an image retrieval system and method based on complementary semantic alignment and symmetric retrieval, specifically including an image retrieval system based on complementary semantic alignment and symmetric retrieval and an image retrieval method based on complementary semantic alignment and symmetric retrieval.

[0006] This invention provides an image retrieval system based on complementary semantic alignment and symmetric retrieval, comprising: The data module is configured to collect images in real time and use the collected images as database images to build and update the image database; monitor; The hierarchical feature extraction device, which communicates with both the display and the data module, is configured to obtain global feature vectors of the query image and each of the database images through global average pooling, and to obtain weighted local feature sets of the query image and each of the database images by calculating the Hadamard product of the normalized spatial weight matrix and the mask matrix. The conditional symmetric alignment device communicates with the hierarchical feature extraction device and is configured to dynamically select a symmetric alignment or fault-tolerant alignment path based on the global feature confidence level, so as to obtain bidirectional scores and their differences using all global feature vectors and weighted local feature sets obtained by the hierarchical feature extraction device. The result generation device, which communicates with both the conditional symmetric alignment device and the display, is configured to use a fully connected layer model algorithm to convert the global feature confidence of the database image, the global feature complexity of the query image, and the bidirectional score difference into adaptive weights. Then, it uses a weighted sum of the bidirectional scores and the adaptive weights to obtain the retrieval result of the query image and calls the display to show the retrieval result.

[0007] The image retrieval system based on complementary semantic alignment and symmetric retrieval disclosed in this invention addresses the aforementioned technical shortcomings by employing a hierarchical feature extraction device to obtain global feature vectors for the query image and each database image using global average pooling. Furthermore, it calculates the Hadamard product of the normalized spatial weight matrix and the mask matrix to obtain a weighted set of local features for both the query image and each database image. This not only achieves feature extraction but also enables deep coupling between global semantics and local details during the feature extraction stage, overcoming the feature fragmentation problem in traditional solutions. Moreover, by setting up a conditional symmetric alignment device, a conditional path switching strategy is employed to dynamically select either a symmetric alignment path (i.e., global-local and local-global) or a fault-tolerant alignment path (i.e., global-local and local-local) based on global feature confidence, significantly improving robustness to occlusion and complex scenes and avoiding rigid alignment mechanisms. Finally, a result generation device generates adaptive weights based on bidirectional score difference and global confidence, enabling dynamic decision optimization for different content complexities. This achieves the technical effect of dynamically adjusting weights according to image content complexity while obtaining retrieval results for the query image.

[0008] In one possible implementation, the hierarchical feature extraction device includes: The backbone module, which communicates with both the display and the data module, is configured to extract initial feature maps of the query image and each of the database images using a convolutional neural network model algorithm. The global feature module, which communicates with both the backbone module and the conditional symmetric alignment device, is configured to perform global average pooling on each of the initial feature maps to obtain global feature vectors for the query image and each of the database images. The spatial weight module, which communicates with the backbone module, is configured to perform channel attention mechanism mapping processing on each of the initial feature maps to obtain the normalized spatial weight matrix of the query image and each of the database images. The local detail module, which communicates with the main module, is configured to use a deformable convolution and spatial attention mechanism fusion algorithm to map each of the initial feature maps into a mask matrix to obtain the mask matrix of the query image and each of the database images. The feature modulation module, which communicates with both the spatial weight module and the local detail module, is configured to calculate the Hadamard product of the normalized spatial weight matrix and the mask matrix of the query image, and to form a set of all column vectors of the obtained Hadamard product to obtain a weighted local feature set of the query image; it is also configured to calculate the Hadamard product of the normalized spatial weight matrix and the mask matrix of each of the database images, and to form a set of all column vectors of the obtained Hadamard product to obtain a weighted local feature set of each of the database images.

[0009] The hierarchical feature extraction device with the above structure first obtains an initial feature map using a backbone module. Then, it can extract global feature vectors in parallel using a global feature module, extract a normalized spatial weight matrix using a spatial weight module, and extract a mask matrix using a local detail module. Finally, it uses a feature modulation module to calculate the Hadamard product of the normalized spatial weight matrix and the mask matrix to obtain a weighted set of local features for the query image and each database image. Furthermore, a hierarchical guidance mechanism is proposed, which generates a spatial-channel joint weight matrix by calculating the Hadamard product of the normalized spatial weight matrix and the mask matrix. This further achieves deep coupling between global semantics and local details in the feature extraction stage, overcoming the feature fragmentation problem in traditional schemes.

[0010] In one possible implementation, the backbone module is a network structure consisting of multiple convolutional units, at least one residual unit, and at least five pooling units connected in series, with one of the pooling units located at the end of the network structure, to extract the initial feature maps of the query image and each database image respectively, thereby achieving controllable extraction accuracy and efficiency.

[0011] In one possible implementation, the global feature module includes: The quantization unit, which communicates with the backbone module, is configured to perform INT8 quantization on each of the initial feature maps to obtain the quantization results of the query image and each of the database images. The global average pooling unit, which communicates with both the quantization unit and the conditional symmetric alignment device, is configured to perform global average pooling on each of the quantization results to obtain global feature vectors for the query image and each of the database images.

[0012] The global feature module with the above structure and functions performs global average pooling after INT8 quantization, which can compress the vector size, reduce computational complexity, save energy, and improve the extraction efficiency of global feature vectors.

[0013] In one possible implementation, the spatial weighting module includes: The channel feature unit, which communicates with the backbone module, is configured to convert each of the initial feature maps into channel attention vectors using the SE-block model algorithm, so as to obtain the channel attention vectors of the query image and each of the database images; The channel multiplication unit, which communicates with both the channel feature unit and the backbone module, is configured to perform channel multiplication on the initial feature map of the query image and the channel attention vector to obtain the channel attention feature map of the query image; it is also configured to perform channel multiplication on the initial feature map and the channel attention vector of each database image to obtain the channel attention feature map of each database image. The weight matrix generation unit communicates with the channel multiplication unit and is configured to perform convolution processing with a kernel size of 1×1 on each of the channel attention feature maps to obtain the spatial weight matrix of the query image and each of the database images. The normalization unit, which communicates with both the weight matrix generation unit and the feature modulation module, is configured to perform bilinear interpolation and global normalization operations on each of the spatial weight matrices in turn to obtain the normalized spatial weight matrices of the query image and each of the database images.

[0014] The spatial weight module, which has the above structure and functions, first calculates the channel attention vector using the SE-block model (which contains two fully connected layers and has a good compression ratio) and performs channel multiplication with the initial feature map; then it reduces the dimensionality to a spatial matrix through 1×1 convolution, adjusts the size through bilinear interpolation, and performs global normalization; finally, it outputs a normalized spatial weight matrix to enhance channel attention and achieve global semantic modulation of local features.

[0015] In one possible implementation, the local detail module includes: The offset generation device, which communicates with the backbone module, is configured to perform mapping processing on each of the initial feature maps respectively through a composite operation of hyperbolic tangent function and convolution with a kernel size of 3×3, so as to obtain the kernel coordinate offset of the query image and each of the database images. The variable convolution unit, communicating with the offset generation device, is configured to perform variable convolution processing on each of the initial feature maps to obtain salient region feature maps of the query image and each of the database images; wherein, the variable convolution processing is performed by dynamically adjusting the convolution kernel geometry according to the convolution kernel coordinate offset corresponding to each of the initial feature maps, so as to use the dynamically adjusted convolution kernel to perform convolution processing on the corresponding initial feature map; The spatial attention device, which communicates with both the variable convolutional unit and the feature modulation module, is configured to perform spatial attention mechanism mapping processing on each of the salient region feature maps to obtain the mask matrix of the query image and each of the database images.

[0016] The local detail module with the above structure and functions dynamically adjusts the geometry of the convolution kernel to adapt to the deformation of the object by configuring programmable kernel coordinate offsets for deformable convolution units. The offsets are confirmed by activation of the convolution and hyperbolic tangent functions, thereby improving the efficiency and practicality of mask matrix generation.

[0017] In one possible implementation, the spatial attention device includes: The max pooling unit, which communicates with the variable convolutional unit, is configured to perform max pooling operations on each of the salient region feature maps to obtain spatial salient maps of the query image and each of the database images. The average pooling unit, which communicates with the variable convolutional unit, is configured to perform average pooling operations on each of the salient region feature maps to obtain background context information maps of the query image and each of the database images. The activation unit, which communicates with the max pooling unit, the average pooling unit, and the feature modulation module, is configured to perform convolution processing on the spatial saliency map and the background context information map of the query image with a kernel size of 1×1, sum the results, and then perform sigmoid activation function mapping processing on the summed results to obtain the mask matrix of the query image; it is also configured to perform convolution processing on the spatial saliency map and the background context information map of each of the database images with a kernel size of 1×1, sum the results, and then perform sigmoid activation function mapping processing on the summed results to obtain the mask matrix of each of the database images.

[0018] The spatial attention device with the above structure and function adopts a dual-gating mechanism: the first path extracts salient regions through max pooling, and the second path captures the context through average pooling. The outputs of the two paths are added after being compressed by 1×1 convolution, and a spatial mask (i.e., mask matrix) is generated through a sigmoid activation function, thereby achieving the technical effect of region enhancement.

[0019] In one possible implementation, the conditional symmetry alignment device is configured to perform the following steps: A1: Based on the global feature vector of the query image and the weighted local feature set of the database images, the global-local alignment scores of the query image and each of the database images are obtained using the maximum cosine similarity calculation formula; A2: Use the confidence calculation formula to obtain the global feature confidence of the query image and several database images respectively, and calculate the average global feature confidence of these several database images. Then, take 0.7 times the obtained average value as the confidence threshold. A3: Determine whether the global feature confidence score of the query image is less than the confidence threshold. If so, proceed to the next step; If not, proceed to step A6; A4: Based on the weighted local feature set of the query image and the weighted local feature set of the database image, the local-local alignment score between the query image and each of the database images is obtained using the maximum cosine similarity calculation formula, and then the next step is executed; A5: The global-local alignment score and local-local alignment score of the query image with each of the database images are used as the bidirectional score. The difference between the global-local alignment score and local-local alignment score of the query image with each of the database images is used as the bidirectional score difference degree. The obtained bidirectional score and bidirectional score difference degree are transmitted to the result generation device. After receiving the new global feature vector and weighted local feature set, the process returns to step A1. A6: Based on the weighted local feature set of the query image and the global feature vector of the database image, the local-global alignment score between the query image and each of the database images is obtained using the maximum cosine similarity calculation formula, and then the next step is executed; A7: The global-local alignment score and local-global alignment score of the query image with each of the database images are used as the bidirectional score. The difference between the global-local alignment score and local-global alignment score of the query image with each of the database images is used as the bidirectional score difference degree. The obtained bidirectional score and bidirectional score difference degree are transmitted to the result generation device. After receiving the new global feature vector and weighted local feature set, the execution of step A1 is reversed.

[0020] The conditional symmetric alignment device, which operates according to the above process, dynamically selects global-local and local-global or global-local and local-local paths based on the global feature confidence level, thereby further improving the robustness of occlusion and complex scenes.

[0021] In one possible implementation, the result generation device includes: The computing unit, which communicates with the conditional symmetric alignment device, is configured to calculate the complexity of the global feature vector of the query image using a vector complexity calculation formula, so as to obtain the global feature complexity of the query image. The fully connected unit, communicating with the computing unit, is configured to convert the global feature confidence of the database image, the global feature complexity of the query image, and the bidirectional score difference into adaptive weights by calling the fully connected layer model it contains. The output unit, which communicates with both the fully connected unit and the display, is configured to obtain the retrieval result of the query image by weighted summation of the bidirectional score and the adaptive weight, and to display the retrieval result on the display.

[0022] The result generation device with the above structure and functions forms an uncertainty-aware fusion model based on bidirectional score difference degree and global confidence degree, and generates adaptive weights through a fully connected layer model to achieve dynamic decision optimization for different content complexities.

[0023] Another technical solution of the present invention is to provide an image retrieval method based on complementary semantic alignment and symmetric retrieval, the method comprising the following steps: S1: The collected images are used as database images through the data module to build and update the image database; S2: Optimize the parameters of the hierarchical feature extraction device through the cross-modal constraint loss function to obtain the parameter-optimized hierarchical feature extraction device; S3: Using a parameter-optimized hierarchical feature extraction device, global feature vectors of the query image and each of the database images are obtained by global average pooling, and weighted local feature sets of the query image and each of the database images are obtained by calculating the Hadamard product of the normalized spatial weight matrix and the mask matrix. S4: The conditional symmetric alignment device dynamically selects the symmetric alignment path or the fault-tolerant alignment path based on the global feature confidence, so as to obtain the bidirectional score and its difference degree using all global feature vectors and weighted local feature sets obtained by the hierarchical feature extraction device. S5: The result generation device executes a fully connected layer model algorithm to convert the global feature confidence, the global feature complexity of the query image, and the bidirectional score difference into adaptive weights. Then, the retrieval result of the query image is obtained by weighted summation of the bidirectional score and the adaptive weights, and the retrieval result is displayed on the display.

[0024] The method disclosed in this invention, after constructing an image database, optimizes the parameters of a hierarchical feature extraction device using a cross-modal constraint loss function to obtain a parameter-optimized hierarchical feature extraction device. This, combined with cross-modal semantic constraints, introduces text supervision signals and jointly optimizes the visual and text feature spaces in the loss function, achieving enhanced semantic alignment capabilities. Subsequently, the parameter-optimized hierarchical feature extraction device uses global average pooling to obtain global feature vectors for the query image and each database image. It then calculates the Hadamard product of the normalized spatial weight matrix and the mask matrix to obtain a weighted local feature set for the query image and each database image. This not only achieves feature extraction but also realizes deep coupling between global semantics and local details during the feature extraction stage, overcoming the feature fragmentation problem in traditional schemes. Finally, a conditional symmetric alignment device utilizes a conditional path switching strategy to dynamically select symmetric alignment (i.e., global-local and local-global) or fault-tolerant alignment (i.e., global-local and local-local) paths based on global feature confidence, significantly improving robustness to occlusion and complex scenes and avoiding rigid alignment mechanisms. Finally, through the result generation device, adaptive weights are generated based on bidirectional score difference and global confidence, realizing dynamic decision optimization for different content complexities. Thus, while obtaining the retrieval results of the query image, the technical effect of dynamically adjusting the weights according to the complexity of the image content is achieved. Attached Figure Description

[0025] Figure 1 This is a schematic diagram of the structure of an image retrieval system based on complementary semantic alignment and symmetric retrieval disclosed in the embodiments of this application; Figure 2 This is a schematic diagram of the hierarchical feature extraction device disclosed in the embodiments of this application; Figure 3 This is a schematic diagram of the main module structure disclosed in the embodiments of this application; Figure 4 This is a schematic diagram of the global feature module structure disclosed in the embodiments of this application; Figure 5This is a schematic diagram of the spatial weighting module structure disclosed in the embodiments of this application; Figure 6 This is a schematic diagram of the partial detail module structure disclosed in the embodiments of this application; Figure 7 This is a schematic diagram of the spatial attention device structure disclosed in the embodiments of this application; Figure 8 This is a flowchart illustrating the operation of the conditional symmetry alignment device disclosed in the embodiments of this application. Figure 9 This is a schematic diagram of the result generation device disclosed in the embodiments of this application; Figure 10 This is a schematic diagram of the fully connected layer model structure disclosed in the embodiments of this application; Figure 11 This is a flowchart of the method disclosed in the embodiments of this application. Detailed Implementation

[0026] First, those skilled in the art should understand that these embodiments are merely used to explain the technical principles of the embodiments of this application and are not intended to limit the scope of protection of the embodiments of this application. Those skilled in the art can make adjustments as needed to adapt to specific application scenarios.

[0027] In the embodiments of this application, unless otherwise explicitly specified and limited, communication or communication connection between the first feature and the second feature means that there is information transmission between the first feature and the second feature. This information transmission can be unidirectional or bidirectional, and the communication connection can be achieved through electrical connection of wires, wireless communication, communication through electromagnetic media (such as semiconductors), communication through channels, etc. Furthermore, unless otherwise specified, the base of the logarithmic function used is 2.

[0028] In the embodiments of this application, unless otherwise explicitly specified and limited, "above," "below," "in front of," or "behind" the second feature can mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediate medium. Furthermore, "above," "on top of," and "over" the second feature can mean that the first feature is directly above or diagonally above the second feature, or simply indicates that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature can mean that the first feature is directly below or diagonally below the second feature, or simply indicates that the first feature is at a lower horizontal level than the second feature. "Before," "in front of," and "in front of" the second feature can mean that the first feature is directly in front of or diagonally in front of the second feature, or simply indicates that the first feature precedes the second feature in sequence. "After," "behind," and "behind" the second feature can mean that the first feature is directly behind or diagonally behind the second feature, or simply indicates that the first feature is after the second feature in sequence.

[0029] The technical solution of this application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0030] See Figures 1 to 11 This application discloses an image retrieval system based on complementary semantic alignment and symmetric retrieval. Figure 1 This is a schematic diagram of the image retrieval system structure, as shown below. Figure 1 As shown, the image retrieval system includes a data module, a display, a hierarchical feature extraction device, a conditional symmetry alignment device, and a result generation device. The hierarchical feature extraction device communicates with both the display and the data module; the conditional symmetry alignment device communicates with the hierarchical feature extraction device; and the result generation device communicates with both the conditional symmetry alignment device and the display. The data module is configured to collect images in real time, using the collected images as database images to build and update the image database. The display is configured to show the query image and the retrieval results.

[0031] See Figure 1 and Figure 2 In this image retrieval system, the hierarchical feature extraction device is configured to obtain global feature vectors of the query image and each database image through global average pooling, and to obtain weighted local feature sets of the query image and each database image by calculating the Hadamard product of the normalized spatial weight matrix and the mask matrix. See also Figure 2 In this embodiment, the hierarchical feature extraction device includes a backbone module, a global feature module, a spatial weight module, a local detail module, and a feature modulation module. The backbone module communicates with both the display and the data module. The global feature module communicates with both the backbone module and the conditional symmetric alignment device. The spatial weight module communicates with the backbone module. The local detail module communicates with the backbone module. The feature modulation module communicates with both the spatial weight module and the local detail module.

[0032] See Figure 3 In the hierarchical feature extraction device, the backbone module is configured to extract the initial feature maps of the query image and each database image respectively through a convolutional neural network model algorithm. Figure 3 This is a schematic diagram of the backbone module structure in this embodiment. The backbone module is a network structure composed of multiple convolutional units, at least one residual unit, and at least five pooling units connected in series, with one pooling unit located at the end of the network structure. In this embodiment, the backbone module uses a residual network as the residual unit and has five pooling units.

[0033] See Figure 4In the hierarchical feature extraction device, the global feature module is configured to perform global average pooling on each initial feature map to obtain global feature vectors for the query image and each database image. In this embodiment, the global feature vector is a 2048-dimensional vector. Figure 4 As shown, in this embodiment, the global feature module includes a quantization unit and a global average pooling unit. The quantization unit communicates with the backbone module, and the global average pooling unit communicates with both the quantization unit and the conditional symmetric alignment device. Within the global feature module, the quantization unit is configured to perform INT8 quantization on each initial feature map to obtain the quantization results for the query image and each database image. The global average pooling unit is configured to perform global average pooling on each quantization result to obtain the global feature vectors for the query image and each database image.

[0034] See Figure 5 In the hierarchical feature extraction device, the spatial weight module is configured to perform channel attention mechanism mapping processing on each initial feature map to obtain the normalized spatial weight matrix of the query image and each database image. For example... Figure 5 As shown, in this embodiment, the spatial weighting module includes a channel feature unit, a channel multiplication unit, a weight matrix generation unit, and a normalization unit. The channel feature unit communicates with the backbone module, the channel multiplication unit communicates with both the channel feature unit and the backbone module, the weight matrix generation unit communicates with the channel multiplication unit, and the normalization unit communicates with both the weight matrix generation unit and the feature modulation module.

[0035] In the spatial weighting module, the channel feature unit is configured to convert each initial feature map into a channel attention vector using the SE-block model algorithm to obtain the channel attention vectors of the query image and each database image. The channel multiplication unit is configured to perform channel multiplication on the initial feature map and the channel attention vector of the query image to obtain the channel attention feature map of the query image; additionally, the channel multiplication unit is also configured to perform channel multiplication on the initial feature map and the channel attention vector of each database image to obtain the channel attention feature map of each database image. The weight matrix generation unit is configured to perform convolution processing with a kernel size of 1×1 on each channel attention feature map to obtain the spatial weight matrix of the query image and each database image. The normalization unit is configured to perform bilinear interpolation and global normalization on each spatial weight matrix sequentially to obtain the normalized spatial weight matrix of the query image and each of the database images; the purpose of resizing is to ensure that the normalized spatial weight matrix and the mask matrix have the same number of rows and columns, so that they can perform Hadamard product operations.

[0036] See Figure 6In the hierarchical feature extraction device, the local detail module is configured to use a fusion algorithm of deformable convolution and spatial attention to map each of the initial feature maps into a mask matrix, thereby obtaining the mask matrices of the query image and each database image. For example... Figure 6 As shown, in this embodiment, the local detail module includes an offset generation device, a variable convolutional unit, and a spatial attention device. The offset generation device communicates with the backbone module, the variable convolutional unit communicates with the offset generation device, and the spatial attention device communicates with both the variable convolutional unit and the feature modulation module.

[0037] In the local detail module, the offset generation device is configured to perform a composite operation of hyperbolic tangent function and 3×3 convolution kernel on each initial feature map to obtain the convolution kernel coordinate offsets of the query image and each database image. Specifically, the initial feature map is first convolved, and then activated with the hyperbolic tangent function to obtain the convolution kernel coordinate offsets.

[0038] In the local detail module, the variable convolution unit is configured to perform variable convolution processing on each initial feature map to obtain salient region feature maps of the query image and each database image. The variable convolution processing is performed by dynamically adjusting the convolution kernel geometry based on the kernel coordinate offset corresponding to each initial feature map, so as to use the dynamically adjusted convolution kernel to perform convolution processing on the corresponding initial feature map. In this embodiment, the variable convolution processing supports dynamic kernel size adjustment of 3×3, 5×5, and 7×7.

[0039] See Figure 7 In the local detail module, the spatial attention mechanism is configured to perform spatial attention mapping on the feature maps of each salient region to obtain the mask matrices of the query image and each database image. For example... Figure 7 As shown, in this embodiment, the spatial attention device includes a max pooling unit, an average pooling unit, and an activation unit. The max pooling unit communicates with the variable convolution unit, the average pooling unit communicates with the variable convolution unit, and the activation unit communicates with the max pooling unit, the average pooling unit, and the feature modulation module simultaneously.

[0040] In the spatial attention device, the max pooling unit is configured to perform max pooling operations on each salient region feature map to obtain spatial saliency maps of the query image and each database image. The average pooling unit is configured to perform average pooling operations on each salient region feature map to obtain background context information maps of the query image and each database image. The activation unit is configured to perform convolution processing on the spatial saliency map and the background context information map of the query image with a kernel size of 1×1, sum the results, and apply a sigmoid activation function to obtain the mask matrix of the query image; in addition, the activation unit is also configured to perform convolution processing on the spatial saliency map and the background context information map of each database image with a kernel size of 1×1, sum the results, and apply a sigmoid activation function to obtain the mask matrix of each database image.

[0041] In this embodiment, the sigmoid activation function used by the activation unit is the sigmoid activation function. The spatial attention device employs dual-gating in obtaining the mask matrix: the first path extracts the spatial saliency map through max pooling, and the second path extracts the background context information through an average pooling layer. The outputs of both paths are activated by the sigmoid function to generate the spatial mask matrix. The spatial mask matrix is ​​then multiplied element-wise (i.e., Hadamard product) with the normalized spatial weight matrix to achieve region enhancement. The activation unit first performs 1×1 convolution processing on the spatial saliency map and the background context information map, then sums the results of the two convolutions, and finally applies the summation result to the sigmoid activation function to obtain the mask matrix.

[0042] In the hierarchical feature extraction device, the feature modulation module is configured to calculate the Hadamard product of the normalized spatial weight matrix and the mask matrix of the query image, use the column vectors of the obtained Hadamard product as weighted local feature vectors, and form a set of all column vectors of the obtained Hadamard product to obtain a weighted local feature set of the query image. Furthermore, the feature modulation module is also configured to calculate the Hadamard product of the normalized spatial weight matrix and the mask matrix of each database image, use the column vectors of the obtained Hadamard product as weighted local feature vectors, and form a set of all column vectors of the obtained Hadamard product to obtain a weighted local feature set of each database image. In this embodiment, the local feature vector is a 128-dimensional vector, and all weighted local feature sets contain the same number of local feature vectors, with the number controlled between 20 and 50.

[0043] See Figure 8 In this image retrieval system, the conditional symmetric alignment device is configured to dynamically select a symmetric alignment path or a fault-tolerant alignment path based on the global feature confidence level, so as to obtain bidirectional scores and their differences by utilizing all global feature vectors and weighted local feature sets obtained by the hierarchical feature extraction device. Figure 8To illustrate the operation flow of the conditional symmetry alignment device, in this embodiment, the conditional symmetry alignment device is configured to perform the following steps: A1: Based on the global feature vector of the query image and the weighted local feature set of the database images, the global-local alignment score between the query image and each database image is obtained using the maximum cosine similarity formula; the specific calculation formula is as follows: , In the formula, Represents the query image and the first Global-local alignment score of images in a database; This represents the number of weighted local feature vectors contained in any weighted local feature set; Represents the global feature vector of the queried image; Representing the The weighted local feature set of the i-th database image contains the th A weighted local feature vector; This represents the dot product operation of vectors.

[0044] A2: Obtain the global feature confidence scores for the query image and several database images using the confidence score calculation formula, and calculate the average global feature confidence scores of these database images. Then, use 0.7 times the obtained average value as the confidence threshold. The confidence score calculation formula is: , in The global feature confidence (scalar, range [0, 1]). Global feature vector Information entropy The feature dimension is fixed at 2048. This is the normalization factor. Specifically, in this embodiment, the global feature confidence of the query image (similar to that of database images) is: , In the formula, Represents the global feature confidence of the queried image; The global feature vector information entropy represents the query image.

[0045] A3: Determine whether the global feature confidence of the query image is less than the confidence threshold. If so, proceed to the next step to select a fault-tolerant alignment path; If not, proceed to step A6 to select a symmetrical alignment path.

[0046] A4: Based on the weighted local feature set of the query image and the weighted local feature set of the database images, the local-local alignment score between the query image and each database image is obtained using the maximum cosine similarity formula, and then the next step is executed; the specific calculation formula is as follows: , In the formula, Represents the query image and the first Local-local alignment score of images in a database; The weighted set of local features representing the query image contains the first... A weighted local feature vector.

[0047] A5: Take the global-local alignment score and local-local alignment score of the query image with each database image as bidirectional scores, and take the difference between the global-local alignment score and local-local alignment score of the query image with each database image as the bidirectional score difference degree. Transmit the obtained bidirectional scores and bidirectional score difference degree to the result generation device. After receiving the new global feature vector and weighted local feature set, return to step A1.

[0048] A6: Based on the weighted local feature set of the query image and the global feature vector of the database images, the local-global alignment score between the query image and each database image is obtained using the maximum cosine similarity formula, and then the next step is executed; the specific calculation formula is as follows: , Represents the query image and the first Local-global alignment score of images in a database; Representing the Global feature vectors of images in a database.

[0049] A7: The global-local alignment score and local-global alignment score of the query image with each database image are used as bidirectional scores. The difference between the global-local alignment score and local-global alignment score of the query image with each database image is used as the bidirectional score difference degree. The obtained bidirectional scores and bidirectional score difference degree are transmitted to the result generation device. After receiving the new global feature vector and weighted local feature set, the execution of step A1 is reversed.

[0050] See Figure 9In this image retrieval system, the result generation device is configured to use a fully connected layer model algorithm to transform the global feature confidence of the database image, the global feature complexity of the query image, and the bidirectional score difference into adaptive weights. Then, the retrieval result of the query image is obtained by weighted summation of the bidirectional score and the adaptive weights, and the retrieval result is displayed on a screen. Figure 9 As shown, in this embodiment, the result generation device includes a computing unit, a fully connected unit, and an output unit. The computing unit communicates with the conditional symmetric alignment device, the fully connected unit communicates with the computing unit, and the output unit communicates with both the fully connected unit and the display.

[0051] In the result generation device, the computing unit is configured to calculate the complexity of the global feature vector of the query image using a vector complexity calculation formula, thereby obtaining the global feature complexity of the query image. In this embodiment, the vector complexity calculation formula used is: , In the formula, Represents vector complexity. The representative vector; according to the above formula, the complexity of querying the global features of an image is: .

[0052] See Figure 10 In the result generation device, the fully connected unit is configured to convert the global feature confidence of the database image, the global feature complexity of the query image, and the bidirectional score dissimilarity into adaptive weights by calling its contained fully connected layer model. For example... Figure 10 As shown, in this embodiment, the fully connected layer model consists of an input layer, a hidden layer, an activation layer, and an output layer. The input layer contains three neurons to respectively input the global feature confidence of the database image, the global feature complexity of the query image, and the bidirectional score difference. The hidden layer has one neuron and contains 128 neurons. The activation layer uses the ReLU activation function, and the output layer is a single neuron. This embodiment uses the standard triplet loss function to optimize the parameters of the fully connected layer model. Let's assume that the first... The global feature confidence of each database image is Query image and the first The two-way score dissimilarity of the images in the database is Furthermore, in the process of transforming into adaptive weights, the fully connected layer model in the 1st... The input vector for this input is The adaptive weights obtained after transformation through the fully connected layer model are: .

[0053] In the result generation device, the output unit is configured to obtain the retrieval result of the query image by weighted summation of bidirectional scores and adaptive weights, and then display the retrieval result on the display. Specifically, in this embodiment, if the global feature confidence is less than the confidence threshold, that is, if the query image output by the conditional symmetric alignment device is less than the confidence threshold, the retrieval result is not found to be true. The bidirectional score of each database image is and At this point, adaptive weights are used. After weighted summation, the similarity score is obtained as follows: ; If the global feature confidence level is not less than the confidence threshold, that is, the query image output by the conditional symmetric alignment device is similar to the first... The bidirectional score of each database image is and At this point, adaptive weights are used. After weighted summation, the similarity score is obtained as follows: ; in, To query the image and the first The similarity score of each database image is calculated. Images with similarity scores less than a set value (e.g., 60) can then be filtered out, and the images are sorted from highest to lowest score to obtain the search results, which are then displayed on the screen.

[0054] See Figure 11 The image retrieval method corresponding to the image retrieval system based on complementary semantic alignment and symmetric retrieval in this embodiment will be further disclosed below. Figure 11 Here is a flowchart of the method, which includes the following steps: S1: The collected images are used as database images through the data module to build and update the image database.

[0055] S2: The parameters of the hierarchical feature extraction device are optimized using a cross-modal constraint loss function to obtain the optimized hierarchical feature extraction device. During parameter optimization, the associated text descriptions of the images in the database are extracted, and a 128-dimensional text feature vector T is generated within 3ms using a low-latency BERT model; subsequently, a cross-modal constraint loss function is constructed. ,in, This is the global feature vector (2048 dimensions). To match the text feature vector (128 dimensions). These are the negative sample text feature vectors (M=256 in total). The dot product method is... (2048 dimensions) and composed of 16 The dot product operation is performed on the column vectors formed by (128 dimensions × 16 = 2048 dimensions). It should be noted that... Generally refers to global feature vectors, while This only represents the global feature vector of the queried image.

[0056] Furthermore, this embodiment also optimizes the parameters of the network structure formed by the hierarchical feature extraction device, the conditional symmetric alignment device, and the result generation device by constructing a total loss function. The total loss function is: ,in , , For the standard triplet loss function, The square of the two-way score difference is calculated as follows: Title / tag text of images is retrieved from the ElasticSearch database; a cropped BERT model (retaining the first 6 Transformer layers) is used to extract a 128-dimensional text feature vector T of the [CLS] tag; the negative sample set is clustered using k-means (k=256) to select the cluster center closest to the current text from the database text feature vector; when calculating the cross-modal loss function, the... After L2 normalization, the dot product similarity is calculated with the text vector; in the total loss function Control path consistency, Adjust the cross-modal constraint strength, decaying it every epoch during training. 2%.

[0057] S3: Input the query image into the parameter-optimized hierarchical feature extraction device, use the parameter-optimized hierarchical feature extraction device to obtain the global feature vectors of the query image and each database image in the form of global average pooling, and obtain the weighted local feature set of the query image and each database image by calculating the Hadamard product of the normalized spatial weight matrix and the mask matrix. S4: The conditional symmetric alignment device dynamically selects the symmetric alignment path or the fault-tolerant alignment path based on the global feature confidence, so as to obtain the bidirectional score and its difference degree by using all global feature vectors and weighted local feature sets obtained by the hierarchical feature extraction device. S5: The fully connected layer model algorithm is executed using the result generation device to convert the global feature confidence of the database image, the global feature complexity of the query image, and the bidirectional score difference into adaptive weights. Then, the retrieval result of the query image is obtained by weighted summation of the bidirectional score and the adaptive weights, and the retrieval result is displayed on the screen.

[0058] The image retrieval system based on complementary semantic alignment and symmetric retrieval disclosed in this embodiment obtains global feature vectors of the query image and each database image by setting up a hierarchical feature extraction device and using global average pooling. It then obtains weighted local feature sets of the query image and each database image by calculating the Hadamard product of the normalized spatial weight matrix and the mask matrix. This not only achieves feature extraction but also achieves deep coupling between global semantics and local details during the feature extraction stage, overcoming the feature fragmentation problem in traditional schemes. Furthermore, by setting up a conditional symmetric alignment device, a conditional path switching strategy is used to dynamically select symmetric alignment paths (i.e., global-local and local-global) or fault-tolerant alignment paths (i.e., global-local and local-local) based on global feature confidence, significantly improving robustness to occlusion and complex scenes and avoiding rigid alignment mechanisms. Finally, by setting up a result generation device, adaptive weights are generated based on bidirectional score difference and global confidence, achieving dynamic decision optimization for different content complexities. This achieves the technical effect of dynamically adjusting weights according to image content complexity while obtaining retrieval results for the query image.

[0059] In the description of the embodiments of this application, it should be noted that the terms "inner" and "outer" and other terms indicating direction or positional relationship are based on the direction or positional relationship shown in the drawings. This is only for the convenience of description and does not indicate or imply that the device or component must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this application.

[0060] In the description of this application, the references to terms such as "an embodiment," "some embodiments," "in this embodiment," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0061] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An image retrieval system based on complementary semantic alignment and symmetric retrieval, characterized in that, include: The data module is configured to collect images in real time and use the collected images as database images to build and update the image database; monitor; The hierarchical feature extraction device, which communicates with both the display and the data module, is configured to obtain global feature vectors of the query image and each of the database images through global average pooling, and to obtain weighted local feature sets of the query image and each of the database images by calculating the Hadamard product of the normalized spatial weight matrix and the mask matrix. The conditional symmetric alignment device, which communicates with the hierarchical feature extraction device, is configured as follows: The global feature confidence of the query image and several database images are obtained by using the confidence calculation formula, and the average global feature confidence of the several database images is calculated. 0.7 times the obtained average value is used as the confidence threshold. Determine whether the global feature confidence score of the query image is less than the confidence threshold. If so, then the fault-tolerant alignment path is selected, that is, a global-local alignment score is obtained based on the global feature vector of the query image and the weighted local feature set of the database image, and a local-local alignment score is obtained based on the weighted local feature set of the query image and the weighted local feature set of the database image, as a bidirectional score; If not, then the symmetrical alignment path is selected, that is, a global-local alignment score is obtained based on the global feature vector of the query image and the weighted local feature set of the database image, and a local-global alignment score is obtained based on the weighted local feature set of the query image and the global feature vector of the database image, as a bidirectional score; Furthermore, the difference between the two scores in the two-way score is taken as the two-way score difference degree, and the obtained two-way score and the two-way score difference degree are transmitted to the result generation device; The result generation device, which communicates with both the conditional symmetric alignment device and the display, is configured to use a fully connected layer model algorithm to convert the global feature confidence of the database image, the global feature complexity of the query image, and the bidirectional score difference into adaptive weights. Then, it uses a weighted sum of the bidirectional scores and the adaptive weights to obtain the retrieval result of the query image and calls the display to show the retrieval result.

2. The image retrieval system based on complementary semantic alignment and symmetric retrieval according to claim 1, characterized in that, The hierarchical feature extraction device includes: The backbone module, which communicates with both the display and the data module, is configured to extract initial feature maps of the query image and each of the database images using a convolutional neural network model algorithm. The global feature module, which communicates with both the backbone module and the conditional symmetric alignment device, is configured to perform global average pooling on each of the initial feature maps to obtain global feature vectors for the query image and each of the database images. The spatial weight module, which communicates with the backbone module, is configured to perform channel attention mechanism mapping processing on each of the initial feature maps to obtain the normalized spatial weight matrix of the query image and each of the database images. The local detail module, which communicates with the main module, is configured to use a deformable convolution and spatial attention mechanism fusion algorithm to map each of the initial feature maps into a mask matrix to obtain the mask matrix of the query image and each of the database images. The feature modulation module, which communicates with both the spatial weight module and the local detail module, is configured to calculate the Hadamard product of the normalized spatial weight matrix and the mask matrix of the query image, and to form a set of all column vectors of the obtained Hadamard product to obtain a weighted local feature set of the query image; it is also configured to calculate the Hadamard product of the normalized spatial weight matrix and the mask matrix of each of the database images, and to form a set of all column vectors of the obtained Hadamard product to obtain a weighted local feature set of each of the database images.

3. The image retrieval system based on complementary semantic alignment and symmetric retrieval according to claim 2, characterized in that, The backbone module is a network structure consisting of multiple convolutional units, at least one residual unit, and at least five pooling units connected in series, with one of the pooling units located at the end of the network structure.

4. The image retrieval system based on complementary semantic alignment and symmetric retrieval according to claim 2, characterized in that, The global feature module includes: The quantization unit, which communicates with the backbone module, is configured to perform INT8 quantization on each of the initial feature maps to obtain the quantization results of the query image and each of the database images. The global average pooling unit, which communicates with both the quantization unit and the conditional symmetric alignment device, is configured to perform global average pooling on each of the quantization results to obtain global feature vectors for the query image and each of the database images.

5. The image retrieval system based on complementary semantic alignment and symmetric retrieval according to claim 2, characterized in that, The spatial weighting module includes: The channel feature unit, which communicates with the backbone module, is configured to convert each of the initial feature maps into channel attention vectors using the SE-block model algorithm, so as to obtain the channel attention vectors of the query image and each of the database images; The channel multiplication unit, which communicates with both the channel feature unit and the backbone module, is configured to perform channel multiplication on the initial feature map of the query image and the channel attention vector to obtain the channel attention feature map of the query image; it is also configured to perform channel multiplication on the initial feature map and the channel attention vector of each database image to obtain the channel attention feature map of each database image. The weight matrix generation unit communicates with the channel multiplication unit and is configured to perform convolution processing with a kernel size of 1×1 on each of the channel attention feature maps to obtain the spatial weight matrix of the query image and each of the database images. The normalization unit, which communicates with both the weight matrix generation unit and the feature modulation module, is configured to perform bilinear interpolation and global normalization operations on each of the spatial weight matrices in turn to obtain the normalized spatial weight matrices of the query image and each of the database images.

6. The image retrieval system based on complementary semantic alignment and symmetric retrieval according to claim 2, characterized in that, The local detail module includes: The offset generation device, which communicates with the backbone module, is configured to perform mapping processing on each of the initial feature maps respectively through a composite operation of hyperbolic tangent function and convolution with a kernel size of 3×3, so as to obtain the kernel coordinate offset of the query image and each of the database images. The variable convolution unit, communicating with the offset generation device, is configured to perform variable convolution processing on each of the initial feature maps to obtain salient region feature maps of the query image and each of the database images; wherein, the variable convolution processing is performed by dynamically adjusting the convolution kernel geometry according to the convolution kernel coordinate offset corresponding to each of the initial feature maps, so as to use the dynamically adjusted convolution kernel to perform convolution processing on the corresponding initial feature map; The spatial attention device, which communicates with both the variable convolutional unit and the feature modulation module, is configured to perform spatial attention mechanism mapping processing on each of the salient region feature maps to obtain the mask matrix of the query image and each of the database images.

7. The image retrieval system based on complementary semantic alignment and symmetric retrieval according to claim 6, characterized in that, The spatial attention device includes: The max pooling unit, which communicates with the variable convolutional unit, is configured to perform max pooling operations on each of the salient region feature maps to obtain spatial salient maps of the query image and each of the database images. The average pooling unit, which communicates with the variable convolutional unit, is configured to perform average pooling operations on each of the salient region feature maps to obtain background context information maps of the query image and each of the database images. The activation unit, which communicates with the max pooling unit, the average pooling unit, and the feature modulation module, is configured to perform convolution processing on the spatial saliency map and the background context information map of the query image with a kernel size of 1×1, sum the results, and then perform sigmoid activation function mapping processing on the summed results to obtain the mask matrix of the query image; it is also configured to perform convolution processing on the spatial saliency map and the background context information map of each of the database images with a kernel size of 1×1, sum the results, and then perform sigmoid activation function mapping processing on the summed results to obtain the mask matrix of each of the database images.

8. The image retrieval system based on complementary semantic alignment and symmetric retrieval according to claim 1, characterized in that, The result generation device includes: The computing unit, which communicates with the conditional symmetric alignment device, is configured to calculate the complexity of the global feature vector of the query image using a vector complexity calculation formula, so as to obtain the global feature complexity of the query image. The fully connected unit, communicating with the computing unit, is configured to convert the global feature confidence of the database image, the global feature complexity of the query image, and the bidirectional score difference into adaptive weights by calling the fully connected layer model it contains. The output unit, which communicates with both the fully connected unit and the display, is configured to obtain the retrieval result of the query image by weighted summation of the bidirectional score and the adaptive weight, and to display the retrieval result on the display.

9. An image retrieval method based on complementary semantic alignment and symmetric retrieval, characterized in that, The image retrieval system based on complementary semantic alignment and symmetric retrieval according to any one of claims 1-8 includes the following steps: S1: The collected images are used as database images through the data module to build and update the image database; S2: Optimize the parameters of the hierarchical feature extraction device through the cross-modal constraint loss function to obtain the optimized hierarchical feature extraction device; S3: Using a parameter-optimized hierarchical feature extraction device, global feature vectors of the query image and each of the database images are obtained by global average pooling, and weighted local feature sets of the query image and each of the database images are obtained by calculating the Hadamard product of the normalized spatial weight matrix and the mask matrix. S4: The conditional symmetric alignment device dynamically selects the symmetric alignment path or the fault-tolerant alignment path based on the global feature confidence, so as to obtain the bidirectional score and its difference degree using all global feature vectors and weighted local feature sets obtained by the hierarchical feature extraction device. S5: The result generation device executes a fully connected layer model algorithm to convert the global feature confidence, the global feature complexity of the query image, and the bidirectional score difference into adaptive weights. Then, the retrieval result of the query image is obtained by weighted summation of the bidirectional score and the adaptive weights, and the retrieval result is displayed on the display.

Citation Information

Patent Citations

  • Cross-modal image text retrieval method based on credibility self-adaptive matching network

    CN111026894A

  • Searching system and method for image semantic enhancement and symmetric semantic completion

    CN119964162A