Searching system and method for image semantic enhancement and symmetric semantic completion

Through the search system of image semantic enhancement and symmetric semantic completion, the problem of high computational cost of text-to-image search in the prior art is solved, and efficient application in large-scale scenarios and target search results with low computational cost are achieved.

CN119964162AInactive Publication Date: 2025-05-09NINGBO UNIV

Patent Information

Application Number
CN202510428896.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-05-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, text-to-image target search has technical defects that are computationally expensive, limiting its application in large-scale scenarios.

Method used

It provides a search system for image semantic enhancement and symmetric semantic completion. The input image is superpixel segmented and semantic encoding through the image semantic enhancement device, and local-local and global-local cross-modal alignment is performed through the symmetric semantic completion device to realize semantic completion of images and text.

Benefits of technology

It reduces the calculation pressure for establishing the correspondence between text and image, reduces the calculation cost, realizes efficient application in large-scale scenarios, and avoids the design of attention modules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964162A_ABST
    Figure CN119964162A_ABST
Patent Text Reader

Abstract

The invention relates to a search system and method for image semantic enhancement and symmetric semantic complementation, an image semantic enhancement device is arranged to divide an image region by using a superpixel segmentation algorithm, so that low-level semantic information in the region is more consistent, and image data enhancement carried out before processing by the superpixel segmentation algorithm is combined, so that the image semantic complementation efficiency is improved. The high-level semantics called from the global image can be ensured to be transmitted to the local image blocks. Moreover, local-local cross-modal alignment and global-local cross-modal alignment can be carried out through an arranged symmetric semantic completion device, so that semantic completion in local and global directions of the image and the text is realized, and global and local features are recovered through cross-modal interaction. Finally, the purpose of global alignment of the input image and the description text thereof is achieved, semantic completion is completed, and the calculation cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer model and system technology, and in particular to a search system and method for image semantic enhancement and symmetric semantic completion. Background Art

[0002] Text-to-image target search is a technical means to retrieve the target image from a large-scale image library using the text description of the target image. This technical means needs to achieve the purpose of aligning image and text features to achieve fast and accurate retrieval. This technical means faces two major challenges. First, as a cross-modal retrieval technical means, it has to deal with the inherent modal heterogeneity between vision and language; second, this technical means is essentially a fine-grained visual recognition task, which has the problem of large intra-class differences and small inter-class differences.

[0003] In response to these two major challenges, a local matching method is disclosed in the prior art. This method attempts to establish a correspondence between the search target part in the image and the noun phrase in the text, and integrates multi-level, multi-granularity matching strategies and specific attention modules to focus on different local areas, thereby achieving the technical effect of text-to-image target search.

[0004] However, this method requires a carefully designed attention module when acquiring image semantics, and has technical defects such as high computational cost in establishing the correspondence between text and image and integrating multi-level and multi-granularity matching strategies, which limits the application of this method in large-scale scenarios. Summary of the invention

[0005] The technical problem to be solved by the present invention is how to overcome the technical defect of high computational cost in the target search from text to image in the prior art. In order to overcome the above defects of the prior art, the present invention provides a search system and method for image semantic enhancement and symmetric semantic completion, including a search system for image semantic enhancement and symmetric semantic completion and a search method for image semantic enhancement and symmetric semantic completion.

[0006] The present invention provides a search system for image semantic enhancement and symmetric semantic completion, comprising: The image semantic enhancement device is configured to randomly enhance the image data of the input image to obtain a visual Figure 1 And Vision Figure 2 Then, the visual Figure 1 and the visual Figure 2 The superpixel segmentation is performed to obtain the corresponding superpixel block sets, and finally the video image is obtained based on the obtained superpixel block sets through the encoding algorithm. Figure 1 The semantic coding sequence and the visual Figure 2 Semantic coding sequence of A symmetric semantic completion device, in communication with the image semantic enhancement device, is configured to transform the image semantics into the image semantics by masking and cross-attention extraction. Figure 1 The semantic coding sequence or the visual Figure 2 The semantic coding sequence is locally aligned with the description text corresponding to the input image and is semantically completed by locally-local cross-modal alignment and globally-local cross-modal alignment to obtain local semantic completion results and global semantic completion results.

[0007] The search system for image semantic enhancement and symmetrical semantic completion disclosed in the present invention aims at the technical defects listed above by setting up an image semantic enhancement device and a symmetrical semantic completion device, so as to utilize the set image semantic enhancement device to divide the image area using a superpixel segmentation algorithm, so as to make the low-level semantic information in the area more consistent, and combined with the random enhancement of the image data before the superpixel segmentation algorithm is processed, it can ensure that the high-level semantics retrieved from the global image is transmitted to the local image block. In addition, the symmetrical semantic completion device can perform local-local cross-modal alignment and global-local cross-modal alignment, thereby realizing semantic completion of the image and text in both local and global directions, realizing the restoration of global and local features through cross-modal interaction, and finally achieving the purpose of global alignment of the input image and its caption, completing the semantic completion, and thus making the search system in the present invention have three output results (i.e., visual Figure 1 Semantic coding sequence or visual Figure 2 The invention also provides a method for realizing text-to-image target search by arranging the search system, and the method comprises the following steps: a) a search engine architecture is formed by combining the semantic coding sequence, the local semantic completion result and the global semantic completion result, and arranging the search system in the invention can realize the search of target description text through image or the search of target image through description text, so as to achieve the purpose of target search from text to image. In addition, since the algorithm based on the search system in the invention has low operation complexity and strict logic, the design of attention module is avoided, and the search system in large-scale scenarios can be arranged relatively easily, and finally the calculation pressure is relieved and the calculation cost is reduced in the process of establishing the correspondence between text and image.

[0008] In a possible implementation, the image semantic enhancement device includes an enhancement device, a superpixel segmentation algorithm module, an original encoder and a momentum encoder, the superpixel segmentation algorithm module communicates with the enhancement device, the original encoder communicates with the superpixel segmentation algorithm module, the momentum encoder communicates with the superpixel segmentation algorithm module, and the symmetric semantic completion device communicates with the momentum encoder; The original encoder has the same structure as the momentum encoder, and the structural parameters of the momentum encoder are obtained by performing exponential moving average on the structural parameters of the original encoder; The combination of the superpixel segmentation algorithm module, the original encoder and the momentum encoder of this scheme can realize self-supervised consistency search and semantic enhancement of image regions in the shared coding space, thereby further acquiring high-level semantics from the global image and obtaining a semantic coding sequence.

[0009] In a possible implementation, in the image semantic enhancement device, the enhancement device is configured to perform random enhancement on the input image to obtain the visual Figure 1 and the visual Figure 2 ; The superpixel segmentation algorithm module is configured to segment the image by a superpixel segmentation algorithm. Figure 1 Perform superpixel segmentation to obtain the visual Figure 1 The corresponding super pixel block set, and also the visual Figure 2 Perform superpixel segmentation to obtain the visual Figure 2 The corresponding super pixel block set, and the visual Figure 1 The corresponding super pixel block set is transmitted to the original encoder, and the video Figure 2 The corresponding superpixel block set is transmitted to the momentum encoder; The original encoder is configured to Figure 1 The corresponding super pixel block set performs semantically enhanced pixel encoding to obtain the visual Figure 1 Semantic coding sequence of The momentum encoder is configured to Figure 2 The corresponding super pixel block set performs semantically enhanced pixel encoding to obtain the visual Figure 2 The semantic coding sequence of Figure 2 The semantic coding sequence is transmitted to the symmetric semantic completion device; This scheme randomly enhances the image data of the input image to obtain the visual Figure 1 And Vision Figure 2 , use the superpixel segmentation algorithm (SLIC) for segmentation, and input the obtained superpixel block set as the subsequent local image block into the encoder to ensure that the information in the image block is as consistent as possible. After processing by the momentum encoder and the original encoder, the local image block features can be converted into visual encoding to obtain the corresponding encoding series. Since the original encoder and the momentum encoder share the same encoding space, it can ensure that the acquired high-level semantics are naturally migrated to the local image block encoding, and finally achieve local semantic enhancement of the image.

[0010] In a possible implementation manner, the image semantic enhancement device further includes two mapping heads that communicate with each other, wherein one of the mapping heads communicates with the original encoder, and the other mapping head communicates with the momentum encoder; The mapping head is configured to convert the semantic coding sequence it receives into a category distribution in a coding space of a specified dimension; By adding two mapping heads to the output of the momentum encoder and the original encoder, the extracted image features can be converted into category distributions in the specified dimensional encoding space, achieving feature alignment in self-supervised learning. In addition, for the mask image learning task, the mapping head converts the features of the mask image modeling task into more semantic information, rather than just restoring at the pixel level, which pays more attention to the acquisition of semantic features.

[0011] In a possible implementation, the symmetric semantic completion device includes a local semantic completion module and a global semantic completion module, and the global semantic completion module communicates with the local semantic completion module; The symmetric semantics completion device is configured to perform the following steps: A1: The local semantic completion module is used to complete the visual Figure 2 The semantic coding sequence of is randomly masked to obtain a first mask; A2: fusing the self-attention mechanism mapping result of the first mask and the first mask through the local semantic completion module to obtain a first fusion result; A3: performing text encoding and random masking on the caption text corresponding to the input image in sequence through the global semantic completion module to obtain a second mask; A4: fusing the self-attention mechanism mapping result of the second mask and the second mask through the global semantic completion module to obtain a second fusion result; A5: mapping the first fusion result and the second fusion result by the local semantic completion module using a cross attention mechanism to obtain a first mapping result, and mapping the first fusion result and the second fusion result by the global semantic completion module using a cross attention mechanism to obtain a second mapping result; A6: fusing the first mapping result and the first fusion result through the local semantic completion module to obtain a third fusion result, and fusing the second mapping result and the second fusion result through the global semantic completion module to obtain a fourth fusion result; A7: fusing the mapping result of the feedforward network mapping of the third fusion result with the third fusion result through the local semantic completion module to obtain a local semantic completion result, and fusing the mapping result of the feedforward network mapping of the fourth fusion result with the fourth fusion result through the global semantic completion module to obtain a global semantic completion result; This solution performs semantic tagging in both image and text directions by setting the operation of the global semantic completion module and the local semantic completion module to fully obtain the semantic information between different modalities. In addition, by setting the global semantic completion module and the local semantic completion module, the symmetric semantic completion device can fully obtain the semantic information between different modalities and achieve local-local and global-local cross-modal alignment.

[0012] In a possible implementation, the local semantic completion module includes a first random mask unit, a first self-attention unit, a first summation unit, a first cross-attention unit, a second summation unit, a first feedforward network unit and a third summation unit, wherein the first random mask unit communicates with the momentum encoder, the first self-attention unit communicates with the first random mask unit, the first summation unit communicates with the first self-attention unit and the first random mask unit at the same time, the first cross-attention unit communicates with the first summation unit and the global semantic completion module at the same time, the second summation unit communicates with the first cross-attention unit and the first summation unit at the same time, the first feedforward network unit communicates with the second summation unit, and the third summation unit communicates with the first feedforward network unit and the second summation unit at the same time; this scheme can ensure the smooth operation of the local semantic completion module and better reveal the relationship between words in the text and local areas in the image.

[0013] In a possible implementation, the global semantic completion module includes a text encoder, a second random mask unit, a second self-attention unit, a fourth summation unit, a second cross-attention unit, a fifth summation unit, a second feedforward network unit, and a sixth summation unit, wherein the second random mask unit communicates with the text encoder, the second self-attention unit communicates with the second random mask unit, the fourth summation unit communicates with the second self-attention unit and the second random mask unit at the same time, the second cross-attention unit communicates with the fourth summation unit and the first summation unit at the same time, the fifth summation unit communicates with the second cross-attention unit and the fourth summation unit at the same time, the second feedforward network unit communicates with the fifth summation unit, and the sixth summation unit communicates with the second feedforward network unit and the fifth summation unit at the same time; The first cross-attention unit is in communication with the fourth summation unit; This solution can ensure the smooth operation of the global semantic completion module and better reveal the relationship between words in the text and the global area in the image.

[0014] In a possible implementation, the first feedforward network unit and the second feedforward network unit are each composed of two fully connected layer networks connected in series, and the activation function of each of the fully connected layer networks is a nonlinear activation function; after the fusion features are mapped by self-attention and cross-attention, they are further processed by the feedforward network unit to further transform the features after the attention mechanism is processed, thereby improving the expressive power of the model.

[0015] Another technical solution of the present invention is to provide a search method for image semantic enhancement and symmetric semantic completion, the method comprising the following steps: S1: using a loss function to optimize model parameters of a network formed by an image semantic enhancement device to be optimized and a symmetric semantic completion device to be optimized to obtain an image semantic enhancement device and a symmetric semantic completion device; S2: The image semantic enhancement device randomly enhances the image data of the input image to obtain a visual Figure 1 And Vision Figure 2 Then, the visual Figure 1 and the visual Figure 2 The superpixel segmentation is performed to obtain the corresponding superpixel block sets, and then the video image is obtained based on the obtained superpixel block sets through the encoding algorithm. Figure 1 The semantic coding sequence and the visual Figure 2 Semantic coding sequence of S3: The visual Figure 1 The semantic coding sequence or the visual Figure 2 The semantic coding sequence is locally aligned with the description text corresponding to the input image and is semantically completed by locally-local cross-modal alignment and globally-local cross-modal alignment to obtain local semantic completion results and global semantic completion results.

[0016] The method disclosed in the present invention first uses a loss function to optimize model parameters, and then uses an image semantic enhancement device to call a superpixel segmentation algorithm to divide the image area, so that the low-level semantic information in the area is more consistent. Combined with the previously performed image data enhancement, it can ensure that the high-level semantics retrieved from the global image are transferred to the local image block. In addition, a symmetrical semantic completion device is used to perform local-local and global-local cross-modal alignment, thereby achieving semantic completion in both image and text directions, and realizing cross-modal interactive restoration of global and local features, and finally achieving the purpose of global alignment of the input image and its caption, completing semantic completion, and then being able to search for the corresponding caption through the image or search for the corresponding image through the caption, achieving the purpose of text-to-image target search, and avoiding the design of the attention module, relieving the computational pressure in the process of establishing the correspondence between text and image, and reducing the computational cost.

[0017] In a possible implementation, the loss function is calculated as follows: , , , , In the formula, A function value representing the loss function; Represents the view Figure 1 The semantic coding sequence of the visual Figure 2 The cross entropy loss function value of the semantic encoding sequence; Representatives based on the view Figure 1 The corresponding super pixel block set and the visual Figure 2 The mask image modeling loss function value of the corresponding superpixel block set; Representatives based on the view Figure 2 A local semantic completion loss function value constructed by locally-locally cross-modally aligning the semantic coding sequence of the image with the caption text corresponding to the input image; Representatives based on the view Figure 2 The masked language modeling loss function value of the semantic encoding sequence of ; Representatives based on the view Figure 2 A global semantic completion loss function value constructed by globally-locally cross-modal alignment of the semantic coding sequence of the input image with the caption text corresponding to the input image; represents the identity loss function value; Represents the contrast loss function value from text to image; Represents the contrast loss function value from image to text; Represents the alignment loss function value from image to text; Represents the alignment loss function value from text to image; This scheme can optimize the model parameters of the network formed by the image semantic enhancement device to be optimized and the symmetric semantic completion device to be optimized, thereby making the obtained image semantic enhancement device and symmetric semantic completion device have the characteristics of efficient operation and reasonable output. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 A schematic diagram of the structure of a search system for image semantic enhancement and symmetric semantic completion disclosed in an embodiment of the present application; Figure 2 This is a flow chart of the operation of the symmetric semantic completion device disclosed in the embodiments of the present application; Figure 3 It is a flow chart of the method disclosed in the embodiments of this application; Figure 4 This is a schematic diagram of the training data flow disclosed in the embodiments of this application. DETAILED DESCRIPTION

[0019] First, those skilled in the art should understand that these implementations are only used to explain the technical principles of the embodiments of the present application, and are not intended to limit the protection scope of the embodiments of the present application. Those skilled in the art can make adjustments to them as needed to adapt to specific application scenarios.

[0020] In the embodiments of the present application, unless otherwise clearly specified and limited, the communication or communication connection between the first feature and the second feature refers to the transmission of information between the first feature and the second feature, and this information transmission can be either unidirectional or bidirectional, and the communication connection can be achieved by wire electrical connection, radio connection, electrical connection of electromagnetic media (such as semiconductors), communication achieved by channels, etc. In addition, unless otherwise specified, the base of the logarithmic function used is 2. At the same time, unless otherwise specified, the symbol "=" represents equal to.

[0021] In the embodiments of the present application, unless otherwise clearly specified and limited, the first feature being "above", "below", "in front of" or "behind" the second feature may mean that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. Moreover, the first feature being "above", "above" and "above" the second feature may mean that the first feature is directly above or obliquely above the second feature, or simply means that the first feature is higher in level than the second feature. The first feature being "below", "below" and "below" the second feature may mean that the first feature is directly below or obliquely below the second feature, or simply means that the first feature is lower in level than the second feature. The first feature being "before", "in front of" and "in front of" the second feature may mean that the first feature is directly in front of or obliquely in front of the second feature, or simply means that the first feature is before the second feature in order. The first feature being "after", "behind" and "behind" the second feature may mean that the first feature is directly behind or obliquely behind the second feature, or simply means that the first feature is after the second feature in order.

[0022] The technical solution of the present application is further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0023] See also Figure 1 to Figure 4 , the embodiment of the present application discloses a search system for image semantic enhancement and symmetric semantic completion, Figure 1 is a schematic diagram of the search system structure, such as Figure 1 As shown, the search system includes an image semantic enhancement device and a symmetric semantic completion device, and the symmetric semantic completion device communicates with the image semantic enhancement device.

[0024] In the search system, the image semantic enhancement device is configured to randomly enhance the image data of the input image to obtain a visual Figure 1 And Vision Figure 2 , and then the superpixel segmentation algorithm is used to segment the Figure 1 And Vision Figure 2 The superpixel segmentation is performed to obtain the corresponding superpixel block sets, and then the video is obtained based on the superpixel block set through the encoding algorithm. Figure 1 Semantic coding sequence and visual Figure 2 Semantic encoding sequence of . Figure 1 In this embodiment, the image semantic enhancement device includes an enhancement device, a superpixel segmentation algorithm module, an original encoder, a momentum encoder and two mapping heads. The superpixel segmentation algorithm module communicates with the enhancement device, the original encoder communicates with the superpixel segmentation algorithm module, the momentum encoder communicates with the superpixel segmentation algorithm module, and the symmetric semantic completion device communicates with the momentum encoder. One of the mapping heads communicates with the original encoder, and the other mapping head communicates with the momentum encoder. The original encoder and the momentum encoder have the same structure, and the structural parameters of the momentum encoder are obtained by performing exponential moving average on the structural parameters of the original encoder.

[0025] In this embodiment, the structural parameters of the momentum encoder are obtained by the following formula during the iterative update process of the model: , In the formula, = Represents assignment to; represents the structural parameters of the momentum encoder; Represents the structural parameters of the original encoder; m represents the momentum coefficient.

[0026] In the image semantic enhancement device, the enhancement device is configured to randomly enhance the image data of the input image to obtain a visual Figure 1 And Vision Figure 2 ; The superpixel segmentation algorithm module is configured to segment the image by a superpixel segmentation algorithm. Figure 1 Perform superpixel segmentation to obtain visual Figure 1 The corresponding super pixel block set, and also the view Figure 2 Perform superpixel segmentation to obtain visual Figure 2 The corresponding super pixel block set, and the visual Figure 1 The corresponding superpixel block set is passed to the original encoder, Figure 2 The corresponding set of superpixel blocks is passed to the momentum encoder; the original encoder is set to view Figure 1 The corresponding superpixel block set performs semantically enhanced pixel encoding to obtain the visual Figure 1 The momentum encoder is set to look at Figure 2 The corresponding superpixel block set performs semantically enhanced pixel encoding to obtain the visual Figure 2 The semantic coding sequence of Figure 2 The semantic coding sequence is transmitted to the symmetric semantic completion device. The mapping head is configured to convert the received semantic coding sequence into a category distribution in a coding space of a specified dimension, and the specified dimension is the same as the number of superpixel blocks output by the superpixel segmentation algorithm module.

[0027] In the search system, the symmetric semantic completion device is configured to extract the visual Figure 1 Semantic coding sequence or visual Figure 2 The semantic coding sequence and the caption text corresponding to the input image (in this embodiment, the visual Figure 2 The semantic coding sequence of the image and the caption text corresponding to the input image are used for local-local cross-modal alignment and global-local cross-modal alignment semantic completion to obtain local semantic completion results and global semantic completion results.

[0028] See also Figure 1 In this embodiment, the symmetric semantic completion device includes a local semantic completion module and a global semantic completion module, and the global semantic completion module communicates with the local semantic completion module.

[0029] In a symmetric semantic completion device, a local semantic completion module includes a first random mask unit, a first self-attention unit, a first summation unit, a first cross-attention unit, a second summation unit, a first feedforward network unit and a third summation unit. The first random mask unit communicates with the momentum encoder, the first self-attention unit communicates with the first random mask unit, the first summation unit communicates with the first self-attention unit and the first random mask unit at the same time, the first cross-attention unit communicates with the first summation unit and the global semantic completion module at the same time, the second summation unit communicates with the first cross-attention unit and the first summation unit at the same time, the first feedforward network unit communicates with the second summation unit, and the third summation unit communicates with the first feedforward network unit and the second summation unit at the same time.

[0030] In the symmetric semantic completion device, the global semantic completion module includes a text encoder, a second random mask unit, a second self-attention unit, a fourth summation unit, a second cross-attention unit, a fifth summation unit, a second feedforward network unit and a sixth summation unit, the second random mask unit communicates with the text encoder, the second self-attention unit communicates with the second random mask unit, the fourth summation unit communicates with the second self-attention unit and the second random mask unit at the same time, the second cross-attention unit communicates with the fourth summation unit and the first summation unit at the same time, the fifth summation unit communicates with the second cross-attention unit and the fourth summation unit at the same time, the second feedforward network unit communicates with the fifth summation unit, and the sixth summation unit communicates with the second feedforward network unit and the fifth summation unit at the same time. The first cross-attention unit communicates with the fourth summation unit.

[0031] In the symmetric semantic completion device, the first feedforward network unit and the second feedforward network unit are both composed of two fully connected layer networks connected in series, and the activation function of each fully connected layer network is a nonlinear activation function.

[0032] See also Figure 2 In this embodiment, the symmetric semantics completion device is configured to perform the following steps: A1: Via the first random mask unit in the local semantic completion module Figure 2 The semantic coding sequence of is randomly masked to obtain a first mask; A2: obtaining a self-attention mechanism mapping result of the first mask by using a self-attention mechanism through the first self-attention unit in the local semantic completion module, and then fusing the self-attention mechanism mapping result of the first mask with the first mask through the first summation unit in the local semantic completion module to obtain a first fusion result; A3: performing text encoding on the caption text corresponding to the input image by the text encoder in the global semantic completion module to obtain a text encoding result, and then performing random masking on the text encoding result by the second random mask unit in the global semantic completion module to obtain a second mask; A4: performing self-attention mechanism mapping on the second mask through the second self-attention unit in the global semantic completion module to obtain a self-attention mechanism mapping result of the second mask, and then fusing the self-attention mechanism mapping result of the second mask with the second mask through the fourth summation unit in the global semantic completion module to obtain a second fusion result; A5: The first cross attention unit in the local semantic completion module uses a cross attention mechanism to map the first fusion result and the second fusion result (the Cartesian product) to obtain a first mapping result, and the second cross attention unit in the global semantic completion module uses a cross attention mechanism to map the first fusion result and the second fusion result (the Cartesian product) to obtain a second mapping result; A6: fusing the first mapping result and the first fusion result through the second summing unit in the local semantic completion module to obtain a third fusion result, and fusing the second mapping result and the second fusion result through the fifth summing unit in the global semantic completion module to obtain a fourth fusion result; A7: The following two processes are carried out in parallel: (1) performing feedforward network mapping on the third fusion result through the first feedforward network unit in the local semantic completion module to obtain a mapping result of the feedforward network mapping of the third fusion result, and then fusing the mapping result of the feedforward network mapping of the third fusion result with the third fusion result through the third summation unit in the local semantic completion module to obtain a local semantic completion result; (2) performing feedforward network mapping on the fourth fusion result through the second feedforward network unit in the global semantic completion module to obtain a mapping result of the feedforward network mapping of the fourth fusion result, and then fusing the mapping result of the feedforward network mapping of the fourth fusion result with the fourth fusion result through the sixth summation unit in the global semantic completion module to obtain a global semantic completion result.

[0033] See also Figure 3 and Figure 4 The following further discloses the usage method corresponding to the search system for image semantic enhancement and symmetric semantic completion in this embodiment. Figure 3 is a flow chart of the method, the method comprising the following steps: S1: using a loss function to optimize model parameters of a network formed by an image semantic enhancement device to be optimized and a symmetric semantic completion device to be optimized to obtain an image semantic enhancement device and a symmetric semantic completion device; S2: The image semantic enhancement device randomly enhances the input image data to obtain visual Figure 1 And Vision Figure 2 , and then the superpixel segmentation algorithm is used to segment the Figure 1 And Vision Figure 2 The superpixel segmentation is performed to obtain the corresponding superpixel block sets, and then the visual image is obtained based on the obtained superpixel block sets through the encoding algorithm. Figure 1 Semantic coding sequence and visual Figure 2 Semantic coding sequence of S3: The symmetrical semantic completion device uses masking and cross-attention extraction to convert the visual Figure 2 The semantic coding sequence is locally aligned with the caption text corresponding to the input image and the semantic completion is performed locally and globally to obtain local and global semantic completion results.

[0034] In step S1, the loss function is calculated as follows: , , , , In the formula, Represents the function value of the loss function; Representative video Figure 1 Semantic coding sequence and visual Figure 2 The cross entropy loss function value of the semantic encoding sequence; Represents based on video Figure 1 The corresponding superpixel block set and view Figure 2 The mask image modeling loss function value of the corresponding superpixel block set; Represents based on video Figure 2 The local semantic completion loss function value constructed by local-local cross-modal alignment of the semantic coding sequence with the caption text corresponding to the input image; Represents based on video Figure 2 The masked language modeling loss function value of the semantic encoding sequence of ; Represents based on video Figure 2 The global semantic completion loss function value constructed by globally-locally cross-modal alignment of the semantic encoding sequence of the input image with the caption text corresponding to the input image; Represents the identity loss (ID loss for short) function value; Represents the contrast loss function value from text to image; Represents the contrast loss function value from image to text; Represents the alignment loss function value from image to text; Represents the alignment loss function value of text to image.

[0035] The construction of the loss function is further described in detail below. For the image semantic enhancement device, the following self-supervised learning constraint loss function is first used to force the visual Figure 1 And Vision Figure 2 Semantic consistency between: , In the formula, Representative video Figure 2 Semantic coding sequence of Represents the view Figure 1 Semantic coding sequence of represents the cross entropy loss.

[0036] In addition, in order to ensure that the image semantic enhancement device (specifically the momentum encoder therein) transfers the high-level semantics of the superpixel block set to the coding sequence of the local block, this embodiment introduces a mask image modeling (MIM) learning mechanism to Figure 1 And Vision Figure 2 The local blocks of the superpixel block set are randomly masked, and the mask image modeling loss function is used to constrain the semantic consistency between local blocks. The formula of the mask image modeling loss function is as follows: , In the formula, M Represents the number of superpixel blocks obtained by superpixel segmentation; Representatives look at each other Figure 1 Perform superpixel segmentation to obtain j Superpixel blocks; Representatives look at each other Figure 2 Perform superpixel segmentation to obtain j Superpixel blocks; Then, the above two loss functions are summed to obtain the loss function of the image semantic enhancement device as follows: .

[0037] For the symmetrical semantic completion device, since the symmetrical semantic completion device has two branches, namely the local semantic completion (LSC for short) module and the global semantic completion (GSC for short) module, these two modules focus on the completed local semantic information and global semantic information respectively. The local semantic completion module aims to use unmasked data to complete the semantic completion of masked local markers, and try to reveal the relationship between words in the text and local areas in the image. In this embodiment, the InfoNCE loss function is used as the local semantic completion loss function to ensure that the features after mask completion (i.e. semantic completion) are closely similar to the corresponding pre-mask features (input image and its corresponding caption), and the corresponding formula is as follows: , , , , , In the formula, Represents the number of training samples; represents a similarity function, and this embodiment adopts a cosine similarity function; Model output function representing the symmetric semantic completion device; Represents the masked features of the image used for training; Represents the caption text features corresponding to the images used during training; Represents the encoded features of the input image used during training; Represents the features of the captions corresponding to the input images used during training after random masking; Represents the scale factor.

[0038] In addition, the mask language modeling loss function is introduced in the training of the local semantic completion module. First, based on the visual Figure 2 In the semantic coding sequence of , 15% of the length of the coded text of the semantic coding sequence is used to randomly mark the coded text with probabilities of 80%, 10% and 10%, or keep it unchanged to obtain the marked coded text. Then a vocabulary-based classification task is performed to predict the masked words, and finally the masked language modeling loss function value is obtained by the following calculation formula: , In the formula, Represents the classification results; The original tags representing the tagged encoded text; This method of using cross-modal interaction to reconstruct the semantic features of masked local labels can effectively obtain the local-local correspondence between images and texts.

[0039] The loss function for contrastive learning of global tags is calculated as follows: , , In the formula, and They represent the complete global features of image tags and text tags output by the symmetric semantic completion device, respectively. represents the number of complete global features of the image label, while and Represents the corresponding global mark of the recovery. Negative samples are global features of other complete images or texts in a batch of data. It is worth noting that during the gradient back propagation, and This strategy can enhance the model’s focus on global feature recovery.

[0040] The global semantic completion loss function is defined as: , Minimizing the above formula helps to make the global features extracted from the mask image Compared with the one obtained from the complete image Similar (for and Similarly). This minimization process helps to recover the semantic information of the mask data. It enables the global representation to acquire complementary knowledge from the corresponding tags of the other modality, thus achieving accurate global-local alignment.

[0041] The loss function of the symmetric semantic completion device is the sum of the above three loss functions, that is, its formula is as follows: .

[0042] In addition to the loss function of the symmetric semantic completion device, this embodiment further includes an identity loss function and a contrast loss function applied to global features. The contrast loss function includes a text-to-image (T2I) contrast loss function and an image-to-text (I2T) contrast loss function. In a batch of image-text pairs, the text-to-image (T2I) contrast loss function and the image-to-text (I2T) contrast loss function are as follows: , , and represents a positive sample pair, and Represents a negative sample pair. The formula of the global feature contrast loss function is: .

[0043] In addition, this embodiment further increases the image-to-text alignment loss function and the text-to-image alignment loss function values. Taking the text-to-image (T2I) alignment loss function as an example, for the query text , the image library is defined as First, the query text With Image Library Each adjacent image in Perform cross-modal positive sampling. Specifically, for a given sample, its positive match can be searched in the opposite modality based on the paired true information. When there are multiple positive matches, one of them can be uniformly sampled. This process produces and The corresponding relationship of represents the sampling function. Then, around and Calculate two separate similarity distributions as follows: , .

[0044] Finally, based on KL divergence The alignment loss function for text to image (T2I) is defined as: .

[0045] The calculation process of mutual alignment loss in image to text (I2T) is similar to that of text to image, so we will not elaborate on it here. The final total loss function is as follows: .

[0046] Figure 4 The data trend of a training run of the search system is shown in the figure. The images used are pedestrian images and are accompanied by explanatory text. In order to distinguish different encoders and mapping heads, symbols are added after the relevant text. In addition, in the figure, the first summing unit, the second summing unit, the third summing unit, the fourth summing unit, the fifth summing unit, and the sixth summing unit are all marked with symbols. In order to achieve a visual expression, the “contrast / mutual learning loss function” in the figure represents the sum of the global feature contrast loss function, the image-to-text alignment loss function, and the text-to-image alignment loss function.

[0047] The search system for image semantic enhancement and symmetrical semantic completion disclosed in this embodiment is provided with an image semantic enhancement device and a symmetrical semantic completion device. The image semantic enhancement device uses a superpixel segmentation algorithm to divide the image area, so that the low-level semantic information in the area is more consistent. Combined with the image enhancement performed before the superpixel segmentation algorithm is processed, it can ensure that the high-level semantics retrieved from the global image is transmitted to the local image block. In addition, the symmetrical semantic completion device can perform local-local cross-modal alignment and global-local cross-modal alignment, thereby realizing semantic completion in both local and global directions of the image and text, and realizing the restoration of global and local features through cross-modal interaction, and finally achieving the purpose of global alignment of the input image and its caption, completing the semantic completion, and thus making the search system in the present invention have three output results (i.e., visual Figure 1 Semantic coding sequence or visual Figure 2 The invention also provides a method for realizing text-to-image target search by arranging the search system, which can realize the search of corresponding captions through images or the search of corresponding images through captions, and achieve the purpose of text-to-image target search. In addition, since the algorithm operation complexity given by the search system in the present invention is relatively small and the logic is rigorous, the design of the attention module is avoided, and it can be relatively easily arranged in the search system in large-scale scenarios, and finally the calculation pressure is relieved and the calculation cost is reduced in the process of establishing the correspondence between text and image.

[0048] In the description of the embodiments of the present application, it should be noted that in the description of the present application, terms such as "inside" and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the drawings. This is only for the convenience of description, and does not indicate or imply that the device or component must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present application.

[0049] In the description of the present application, the description with reference to the terms "one embodiment", "some embodiments", "in the present embodiment", "specific example", or "some examples" etc. means that the specific features, mechanisms, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, mechanisms, materials or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0050] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.

Claims

1. A search system for image semantic enhancement and symmetric semantic completion, characterized in that: include: An image semantic enhancement device is configured to randomly enhance image data of an input image to obtain a view 1 and a view 2, then perform superpixel segmentation on the view 1 and the view 2 respectively through a superpixel segmentation algorithm to obtain respective corresponding superpixel block sets, and finally obtain a semantic coding sequence of the view 1 and a semantic coding sequence of the view 2 based on the obtained superpixel block sets through a coding algorithm; The symmetric semantic completion device communicates with the image semantic enhancement device and is configured to perform semantic completion of the semantic coding sequence of view one or the semantic coding sequence of view two with the caption text corresponding to the input image by local-local cross-modal alignment and global-local cross-modal alignment through masking and cross-attention extraction to obtain local semantic completion results and global semantic completion results.

2. The image semantic enhancement and symmetric semantic completion search system according to claim 1, characterized in that: The image semantic enhancement device comprises an enhancement device, a superpixel segmentation algorithm module, an original encoder and a momentum encoder, the superpixel segmentation algorithm module communicates with the enhancement device, the original encoder communicates with the superpixel segmentation algorithm module, the momentum encoder communicates with the superpixel segmentation algorithm module, and the symmetric semantics completion device communicates with the momentum encoder; The original encoder and the momentum encoder have the same structure, and the model structure parameters of the momentum encoder are obtained by performing exponential moving average on the model structure parameters of the original encoder.

3. The image semantic enhancement and symmetric semantic completion search system according to claim 2, characterized in that: In the image semantic enhancement device, the enhancement device is configured to perform random enhancement of image data on the input image to obtain the view 1 and the view 2; The superpixel segmentation algorithm module is configured to perform superpixel segmentation on the view 1 by using a superpixel segmentation algorithm to obtain a superpixel block set corresponding to the view 1, and also perform superpixel segmentation on the view 2 to obtain a superpixel block set corresponding to the view 2, and transmit the superpixel block set corresponding to the view 1 to the original encoder, and transmit the superpixel block set corresponding to the view 2 to the momentum encoder; The original encoder is configured to perform semantically enhanced pixel encoding on a superpixel block set corresponding to the view one to obtain a semantically encoded sequence of the view one; The momentum encoder is configured to perform semantically enhanced pixel encoding on the superpixel block set corresponding to the view two to obtain a semantic encoding sequence of the view two, and transmit the semantic encoding sequence of the view two to the symmetric semantic completion device.

4. The image semantic enhancement and symmetric semantic completion search system according to claim 3, characterized in that: The image semantic enhancement device further comprises two mapping heads communicating with each other, wherein one of the mapping heads communicates with the original encoder, and the other mapping head communicates with the momentum encoder; The mapping head is configured to convert the semantic coding sequence received therefrom into a category distribution in a coding space of a specified dimension.

5. The image semantic enhancement and symmetric semantic completion search system according to claim 3 or 4, characterized in that: The symmetric semantic completion device includes a local semantic completion module and a global semantic completion module, and the global semantic completion module communicates with the local semantic completion module; The symmetric semantics completion device is configured to perform the following steps: A1: randomly masking the semantic coding sequence of the view 2 by the local semantic completion module to obtain a first mask; A2: fusing the self-attention mechanism mapping result of the first mask and the first mask through the local semantic completion module to obtain a first fusion result; A3: performing text encoding and random masking on the caption text corresponding to the input image in sequence through the global semantic completion module to obtain a second mask; A4: fusing the self-attention mechanism mapping result of the second mask and the second mask through the global semantic completion module to obtain a second fusion result; A5: mapping the first fusion result and the second fusion result by the local semantic completion module using a cross attention mechanism to obtain a first mapping result, and mapping the first fusion result and the second fusion result by the global semantic completion module using a cross attention mechanism to obtain a second mapping result; A6: fusing the first mapping result and the first fusion result through the local semantic completion module to obtain a third fusion result, and fusing the second mapping result and the second fusion result through the global semantic completion module to obtain a fourth fusion result; A7: The mapping result of the feedforward network mapping of the third fusion result is fused with the third fusion result through the local semantic completion module to obtain a local semantic completion result, and the mapping result of the feedforward network mapping of the fourth fusion result is fused with the fourth fusion result through the global semantic completion module to obtain a global semantic completion result.

6. The image semantic enhancement and symmetric semantic completion search system according to claim 5, characterized in that: The local semantic completion module includes a first random mask unit, a first self-attention unit, a first summation unit, a first cross-attention unit, a second summation unit, a first feedforward network unit and a third summation unit. The first random mask unit communicates with the momentum encoder, the first self-attention unit communicates with the first random mask unit, the first summation unit communicates with the first self-attention unit and the first random mask unit at the same time, the first cross-attention unit communicates with the first summation unit and the global semantic completion module at the same time, the second summation unit communicates with the first cross-attention unit and the first summation unit at the same time, the first feedforward network unit communicates with the second summation unit, and the third summation unit communicates with the first feedforward network unit and the second summation unit at the same time.

7. The image semantic enhancement and symmetric semantic completion search system according to claim 6, characterized in that: The global semantic completion module includes a text encoder, a second random mask unit, a second self-attention unit, a fourth summation unit, a second cross-attention unit, a fifth summation unit, a second feedforward network unit and a sixth summation unit, wherein the second random mask unit communicates with the text encoder, the second self-attention unit communicates with the second random mask unit, the fourth summation unit communicates with the second self-attention unit and the second random mask unit at the same time, the second cross-attention unit communicates with the fourth summation unit and the first summation unit at the same time, the fifth summation unit communicates with the second cross-attention unit and the fourth summation unit at the same time, the second feedforward network unit communicates with the fifth summation unit, and the sixth summation unit communicates with the second feedforward network unit and the fifth summation unit at the same time; The first cross-attention unit communicates with the fourth summation unit.

8. The image semantic enhancement and symmetric semantic completion search system according to claim 7, characterized in that: The first feedforward network unit and the second feedforward network unit are both composed of two fully connected layer networks connected in series, and the activation function of each of the fully connected layer networks is a nonlinear activation function.

9. A search method for image semantic enhancement and symmetric semantic completion, characterized in that: A search system for image semantic enhancement and symmetric semantic completion applicable to any one of claims 1 to 8, comprising the following steps: S1: using a loss function to optimize model parameters of a network formed by an image semantic enhancement device to be optimized and a symmetric semantic completion device to be optimized to obtain an image semantic enhancement device and a symmetric semantic completion device; S2: performing random image data enhancement on the input image by the image semantic enhancement device to obtain view 1 and view 2, then performing superpixel segmentation on the view 1 and view 2 by a superpixel segmentation algorithm to obtain respective corresponding superpixel block sets, then obtaining a semantic coding sequence of the view 1 and a semantic coding sequence of the view 2 based on the obtained superpixel block sets by a coding algorithm; S3: The symmetrical semantic completion device uses masking and cross-attention extraction methods to perform local-local cross-modal alignment and global-local cross-modal alignment of the semantic coding sequence of view one or the semantic coding sequence of view two with the description text corresponding to the input image to obtain local semantic completion results and global semantic completion results.

10. The image semantic enhancement and symmetric semantic completion search method according to claim 9, characterized in that: The loss function is calculated as follows: , , , , In the formula, A function value representing the loss function; A cross entropy loss function value representing the semantic coding sequence of the view one and the semantic coding sequence of the view two; represents a loss function value of mask image modeling based on the super pixel block set corresponding to the view one and the super pixel block set corresponding to the view two; represents a local semantic completion loss function value constructed by performing local-local cross-modal alignment between the semantic coding sequence of the view 2 and the caption text corresponding to the input image; represents a masked language modeling loss function value based on the semantic coding sequence of the view 2; represents a global semantic completion loss function value constructed by globally-locally cross-modal alignment of the semantic coding sequence of the view 2 with the caption text corresponding to the input image; represents the identity loss function value; Represents the contrast loss function value from text to image; Represents the contrast loss function value from image to text; Represents the alignment loss function value from image to text; Represents the alignment loss function value of text to image.

Citation Information

Patent Citations

  • Farmland crop identification method based on fusion of semantic segmentation and superpixel segmentation

    CN114067219A

  • Remote sensing image self-supervision semantic segmentation method based on position coding

    CN114913412A

  • Three-dimensional semantic scene completion method based on semantic segmentation prior

    CN119540456A

Cited By

  • Image retrieval system and method based on complementary semantic alignment and symmetric retrieval

    CN121434432A

  • An image retrieval system and method based on complementary semantic alignment and symmetric retrieval

    CN121434432B