Image description method, apparatus, and program product

CN122799129APending Publication Date: 2026-09-22GUANGDONG INST OF INTELLIGENT SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611281985.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-24
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

然而,现有的注意力机制往往采用连续加权混合的方式聚合全局特征,这导致前景特征和背景特征在聚合过程中发生语义混叠,使得图像描述子受到背景噪声的干扰,进而致使图像描述子的精度低下

Benefits of technology

[0014]According to the image description method, apparatus, and program product provided in the embodiments of this application, a target image is divided into multiple image blocks. For each image block, its single-channel features and multi-channel features are extracted first, and a unique unipolar space mask is matched based on the single-channel features. Then, feature enhancement is performed on the multi-channel features based on the unipolar space mask. The aim is to use 0/1 discrete symbols to force the background energy to zero at the physical level, fundamentally blocking the propagation path of background noise in the feature aggregation process, while effectively perceiving and enhancing the image foreground structure, thereby obtaining sparse image features. Finally, the sparse image features of all image blocks are fused to generate an image descriptor with high purity and strong discriminative power. This can effectively suppress background noise interference and semantic aliasing interference between foreground and background, thereby effectively improving the accuracy of image description.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122799129A_ABST
    Figure CN122799129A_ABST
Patent Text Reader

Abstract

This application discloses an image description method, apparatus, and program product, belonging to the field of image processing technology. The method includes: acquiring a target image; performing feature extraction and block adaptation processing on the target image to obtain single-channel features and multi-channel features corresponding to multiple image blocks respectively; matching corresponding unipolar space masks for multiple image blocks based on the single-channel features corresponding to the multiple image blocks respectively; performing feature enhancement processing on the multi-channel features corresponding to the multiple image blocks respectively based on the unipolar space masks corresponding to the multiple image blocks respectively to obtain sparse image features corresponding to the multiple image blocks respectively; and fusing the sparse image features corresponding to the multiple image blocks respectively to obtain an image descriptor. This application can effectively suppress background noise interference and semantic aliasing interference between foreground and background, thereby effectively improving the accuracy of image description.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, in particular to an image description method, device and program product. BACKGROUND

[0002] Image description aims to enable a machine to understand the visual content of an image and automatically generate a description sentence that is semantically accurate and grammatically correct. In the related art, an initial feature map of an image is processed by means of an attention mechanism to obtain a refined feature map of the image. The refined feature map is aggregated to generate a one-dimensional continuous floating-point feature vector, which is an image description sub-vector, so as to realize image description. However, the existing attention mechanism often aggregates global features in a continuous weighted mixing manner, which causes semantic aliasing of foreground features and background features in the aggregation process, so that the image description sub-vector is interfered by background noise, and the accuracy of the image description sub-vector is low. SUMMARY

[0003] The main purpose of the embodiments of the present application is to propose an image description method, device and program product, which aims to effectively suppress the interference of background noise and the semantic aliasing of foreground and background, so as to effectively improve the accuracy of image description.

[0004] To achieve the above purpose, one aspect of the embodiments of the present application proposes an image description method, which comprises: obtaining a target image; performing feature extraction and block adaptation processing on the target image to obtain single-channel features and multi-channel features corresponding to a plurality of image blocks respectively; matching corresponding monopole spatial masks for the plurality of image blocks according to the single-channel features corresponding to the plurality of image blocks respectively; performing feature enhancement processing on the multi-channel features corresponding to the plurality of image blocks respectively according to the monopole spatial masks corresponding to the plurality of image blocks respectively, to obtain sparse image features corresponding to the plurality of image blocks respectively; performing fusion processing on the sparse image features corresponding to the plurality of image blocks respectively to obtain an image description sub-vector.

[0005] In some embodiments, the feature extraction and block adaptation processing on the target image to obtain single-channel features and multi-channel features corresponding to a plurality of image blocks respectively comprises: performing feature extraction processing on the target image to obtain an image feature map; dividing the target image into a plurality of image blocks, and extracting block features corresponding to the plurality of image blocks respectively from the image feature map; Channel adaptation processing is performed on the block features corresponding to the multiple image blocks respectively to obtain single-channel features and multi-channel features corresponding to the multiple image blocks respectively.

[0006] In some embodiments, the step of performing feature extraction processing on the target image to obtain an image feature map includes: The target image is subjected to feature extraction processing to obtain an initial feature map; The initial feature map is scaled to bipolar space to obtain the image feature map.

[0007] In some embodiments, performing channel adaptation processing on the block features corresponding to the plurality of image blocks respectively to obtain single-channel features and multi-channel features corresponding to the plurality of image blocks includes: The block features of the target image block are determined as the multi-channel features of the target image block; The block features of the target image block are subjected to channel fusion processing to obtain the single-channel features of the target image block; Wherein, the target image block refers to any one of the image blocks.

[0008] In some embodiments, matching corresponding single-pole spatial masks for the plurality of image patches based on the single-channel features corresponding to the plurality of image patches includes: Dictionary retrieval processing is performed based on the single-channel features of the target image block to obtain a binary exhaustive dictionary of the target image block; The binary exhaustive dictionary of the target image patch is mapped to obtain the unipolar space mask of the target image patch; Wherein, the target image block refers to any one of the image blocks.

[0009] In some embodiments, the step of performing dictionary lookup processing based on the single-channel features of the target image patch to obtain a binary exhaustive dictionary of the target image patch includes: The routing probability of the target image patch is obtained by performing routing operations based on the target dictionary and the single-channel features of the target image patch. Based on the routing probability of the target image patch, the target dictionary is searched to obtain a binary exhaustive dictionary of the target image patch.

[0010] In some embodiments, the step of performing feature enhancement processing on the multi-channel features corresponding to the plurality of image patches respectively, based on the single-pole spatial masks corresponding to the plurality of image patches respectively, to obtain sparse image features corresponding to the plurality of image patches respectively, includes: Based on the unipolar space mask of the target image block, energy feature extraction processing is performed on the multi-channel features of the target image block to obtain the channel energy features of the target image block; The channel energy features of the target image block are broadcast to the unipolar spatial mask of the target image block to obtain the sparse image features of the target image block; Wherein, the target image block refers to any one of the image blocks.

[0011] In some embodiments, fusing the sparse image features corresponding to the plurality of image blocks to obtain an image descriptor includes: Dimensionality reduction is performed on the sparse image features corresponding to the multiple image blocks respectively to obtain the semantic features corresponding to the multiple image blocks respectively; The semantic features corresponding to the multiple image blocks are fused according to the spatial slot order to obtain the image descriptor.

[0012] To achieve the above objectives, another aspect of this application provides an image description apparatus, the apparatus comprising: The acquisition module is used to acquire the target image; The first processing module is used to perform feature extraction and block adaptation processing on the target image to obtain single-channel features and multi-channel features corresponding to multiple image blocks respectively. The second processing module is used to match corresponding unipolar space masks for the multiple image blocks according to the single-channel features corresponding to the multiple image blocks respectively; The third processing module is used to perform feature enhancement processing on the multi-channel features corresponding to the multiple image blocks respectively according to the single-pole space mask corresponding to the multiple image blocks respectively, so as to obtain the sparse image features corresponding to the multiple image blocks respectively. The fourth processing module is used to fuse the sparse image features corresponding to the multiple image blocks to obtain image descriptors.

[0013] To achieve the above objectives, another aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the above-described image description method.

[0014] According to the image description method, apparatus, and program product provided in the embodiments of this application, a target image is divided into multiple image blocks. For each image block, its single-channel features and multi-channel features are extracted first, and a unique unipolar space mask is matched based on the single-channel features. Then, feature enhancement is performed on the multi-channel features based on the unipolar space mask. The aim is to use 0 / 1 discrete symbols to force the background energy to zero at the physical level, fundamentally blocking the propagation path of background noise in the feature aggregation process, while effectively perceiving and enhancing the image foreground structure, thereby obtaining sparse image features. Finally, the sparse image features of all image blocks are fused to generate an image descriptor with high purity and strong discriminative power. This can effectively suppress background noise interference and semantic aliasing interference between foreground and background, thereby effectively improving the accuracy of image description.

[0015] Other features and advantages of this application will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the description and the accompanying drawings. Attached Figure Description

[0016] Figure 1 This is a flowchart of an image description method provided in this application; Figure 2 yes Figure 1 Flowchart of step S102; Figure 3 yes Figure 1 Flowchart of step S103; Figure 4 yes Figure 1 Flowchart of step S104; Figure 5 yes Figure 1 Flowchart of step S105; Figure 6 This is a schematic diagram illustrating the implementation process of an image description method provided in this application; Figure 7 This is a structural diagram of an image description device provided in this application. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.

[0018] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”

[0019] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0021] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as target images, is required, the user's permission or consent will be obtained first. Furthermore, the collection, use, and processing of this data will comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user will be obtained through pop-ups or redirects to confirmation pages. Only after obtaining the user's separate permission or consent will the necessary user-related data for the normal operation of the embodiments of this application be obtained.

[0022] First, the terms and concepts involved in one or more embodiments of this application will be explained.

[0023] 1) Image Captioning is a multimodal artificial intelligence task that aims to bridge the semantic gap between vision and language. It aims to analyze the visual content of an input image, understand the entities, attributes, spatial locations and interactive relationships within it, and generate a grammatically sound, semantically coherent natural language sentence that conforms to human cognitive habits, thereby achieving a high-level abstract generalization of the overall scene of the image.

[0024] 2) Unipolar space refers to a data representation in which the range of continuous features is constrained to the non-negative interval [0, 1], aiming to measure the presence and strength of features. In unipolar space, the magnitude of the value represents the absolute strength or probability of occurrence of the feature, with 0 indicating that the feature is completely absent and 1 indicating that the feature is fully activated.

[0025] 3) Bipolar space refers to a data representation that constrains the range of continuous features to an interval [-1, 1] symmetrical about the origin. It aims to measure the direction and intensity of a feature's deviation from a baseline. In bipolar space, the positive or negative sign of a value represents the direction or polarity of the feature. A positive sign indicates that the feature is activated or above the baseline, while a negative sign indicates that the feature is suppressed or below the baseline. 0 represents the neutral equilibrium point.

[0026] Currently, image captioning is one of the core tasks in the intersection of Computer Vision (CV) and Natural Language Processing (NLP), aiming to enable machines to understand the visual content of images and automatically generate semantically accurate and grammatically correct descriptive sentences. With the rapid development of deep learning (DL) technology, image captioning has gradually evolved from the early extensive generation paradigm based on template filling or retrieval-based splicing to an end-to-end generation paradigm represented by encoder-decoder architecture and Transformer self-attention mechanism. This paradigm can not only effectively capture salient objects in images and their spatial semantic relationships, but also learn cross-modal semantic alignment representations on large-scale multimodal data with the help of pre-trained vision-language models, thereby generating more fluent, accurate and diverse text descriptions, namely image descriptors.

[0027] In related technologies, attention mechanisms are typically used to process the initial feature map of an image to enhance the foreground region and suppress irrelevant background responses, thereby obtaining a refined feature map. This refined feature map is then aggregated to generate a one-dimensional continuous floating-point feature vector, which is the image descriptor. This descriptor describes the visual content of the image as semantic textual information, thus achieving image description. However, existing attention mechanisms often aggregate global features using a continuous weighted mixing method. Their attention is smoothly diffused across multiple objects rather than focusing on the core foreground. This leads to semantic aliasing between foreground and background features during the aggregation process, causing the image descriptor to be interfered with by background noise, resulting in low accuracy of the image descriptor.

[0028] In view of this, this application provides an image description method, apparatus, and program product. This scheme divides a target image into multiple image blocks. For each image block, its single-channel and multi-channel features are first extracted. A unique unipolar space mask is then matched based on the single-channel features. Subsequently, feature enhancement is performed on the multi-channel features based on the unipolar space mask. This aims to force the background energy to zero at the physical level using 0 / 1 discrete symbols, fundamentally blocking the propagation path of background noise during feature aggregation. Simultaneously, it effectively perceives and enhances the foreground structure of the image, thereby obtaining sparse image features. Finally, the sparse image features of all image blocks are fused to generate an image descriptor with high purity and strong discriminative power. This effectively suppresses background noise interference and semantic aliasing interference between the foreground and background, thereby effectively improving the accuracy of image description.

[0029] The image description methods, apparatuses, and program products provided in the embodiments of this application mainly relate to various application scenarios such as image retrieval, visual positioning, and target recognition. Those skilled in the art will understand that the image description methods, apparatuses, and program products provided in the embodiments of this application can be executed in various application scenarios. Specifically, taking the image description method in the embodiments of this application as an example: For example, the image description method provided in this application embodiment can be applied to image retrieval scenarios. For instance, a user uploads an image through a front-end such as a mobile phone or laptop, hoping to retrieve images of similar or identical scenes. In this case, a back-end such as a cloud server can receive the image to be queried uploaded by the user and use the image description method of this application embodiment to convert the image to be queried into a corresponding image descriptor. By quickly matching the image descriptor to be queried with the descriptors pre-stored in a preset image database, at least one image with high similarity can be returned to the front-end, thereby achieving high-precision image retrieval.

[0030] For example, the image description method provided in this application embodiment can be applied to visual positioning scenarios. For instance, when a robot is performing a specific task, it usually needs to determine its own position in real time. At this time, the robot can acquire the environmental image of the current frame and use the image description method of this application embodiment to convert the environmental image of the current frame into a corresponding image descriptor. By quickly matching the image descriptor of the current frame with the descriptors pre-stored in the preset location database, its own position can be determined, thereby achieving high-precision visual positioning.

[0031] For example, the image description method provided in this application embodiment can be applied to target recognition scenarios. For instance, in industrial production, industrial cameras need to classify and identify parts on a conveyor belt in real time. At this time, the industrial processor can receive the parts images captured in real time by the industrial camera and use the image description method of this application embodiment to convert the parts images into corresponding image descriptors. By quickly matching the parts image descriptors with the descriptors pre-stored in the preset parts database, the model and category of the parts can be identified, thereby achieving high-precision target recognition.

[0032] It is understood that the above application scenarios are merely illustrative and do not imply any limitation on the actual application of the image description method provided in the embodiments of this application. Those skilled in the art will understand that the image description method provided in the embodiments of this application can be used to perform specified tasks in different application scenarios.

[0033] The implementation steps of an image description method provided in this application will be described in detail below with reference to the accompanying drawings.

[0034] This application provides an image description method that can be applied to a terminal, a server, or software running on either a terminal or a server. The terminal can be a tablet, laptop, desktop computer, etc., but is not limited to these. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Furthermore, the server can be a node server in a blockchain network, but is not limited to these. Blockchain is a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms.

[0035] Reference Figure 1 , Figure 1This is a flowchart of an image description method provided in this application, which may include the following steps S101-S105.

[0036] S101, acquire the target image.

[0037] In this step, the image to be described, i.e., the target image, is acquired. The type of target image is not specifically limited. For example, in an image retrieval scenario, the target image can be an image uploaded by a target object; the target object can be a real user or an account, but is not limited to these. As another example, in a visual positioning scenario, the target image can be an environmental image captured in real time by a robot. Yet another example, in a target recognition scenario, the target image can be an image of components captured by an industrial camera.

[0038] S102, perform feature extraction and block adaptation processing on the target image to obtain single-channel and multi-channel features corresponding to multiple image blocks respectively.

[0039] In this step, after obtaining the target image, feature extraction and patch adaptation are performed. This aims to divide the target image into multiple image patches and extract the single-channel and multi-channel features corresponding to each patch. Single-channel features refer to features with one channel, corresponding to the structural branch, describing the spatial response intensity of the target image, but not containing complex semantic and texture information. Multi-channel features, on the other hand, refer to features with more than one channel, corresponding to the energy branch, describing the complex semantic and texture information of the target image.

[0040] S103, based on the single-channel features corresponding to multiple image blocks, match the corresponding single-pole space mask for multiple image blocks.

[0041] In this step, a unique spatial mask is matched for each image patch based on single-channel features. The spatial mask for each image patch can be obtained by traversing the image patches. It is worth noting that in the spatial mask, pixels with a value of 0 represent background pixels, while pixels with a value of 1 represent foreground structure pixels.

[0042] S104. Based on the single-pole space mask corresponding to each of the multiple image blocks, feature enhancement processing is performed on the multi-channel features corresponding to each of the multiple image blocks to obtain the sparse image features corresponding to each of the multiple image blocks.

[0043] In this step, for each image patch, feature enhancement is performed on multi-channel features based on a unipolar space mask to obtain sparse image features with physical strength. The sparse image features corresponding to each image patch can be obtained by traversing each patch. It is worth noting that since pixels with a value of 0 in the unipolar space mask represent background pixels and pixels with a value of 1 represent foreground structure pixels, using this mask to enhance multi-channel features can force the background energy to zero at the physical level using the discrete symbols of 0 / 1. This fundamentally blocks the propagation path of background noise during feature aggregation, while effectively perceiving and enhancing the foreground structure of the image, thereby improving the accuracy of image features.

[0044] S105, the sparse image features corresponding to multiple image blocks are fused to obtain image descriptors.

[0045] In this step, the sparse image features corresponding to each image block can be obtained through the processing of the aforementioned steps. These features are then fused into a feature vector with standard semantic dimensions. This feature vector is the image descriptor, which accurately describes the visual content of the image as semantic text information, thereby achieving high-precision image description.

[0046] Therefore, the embodiments of this application divide the target image into multiple image blocks. For each image block, its single-channel and multi-channel features are extracted first, and a unique unipolar space mask is matched based on the single-channel features. Then, feature enhancement is performed on the multi-channel features based on the unipolar space mask. The aim is to use 0 / 1 discrete symbols to force the background energy to zero at the physical level, fundamentally blocking the propagation path of background noise in the feature aggregation process, while effectively perceiving and enhancing the image foreground structure, thereby obtaining sparse image features. Finally, the sparse image features of all image blocks are fused to generate an image descriptor with high purity and strong discriminative power. This can effectively suppress background noise interference and semantic aliasing interference between foreground and background, thereby effectively improving the accuracy of image description.

[0047] The steps described above will be explained in further detail below.

[0048] In some implementations, refer to Figure 2 In step S102 above, feature extraction and block adaptation processing are performed on the target image to obtain single-channel and multi-channel features corresponding to multiple image blocks, which may include: Feature extraction is performed on the target image to obtain an image feature map; The target image is divided into multiple image blocks, and the block features corresponding to each image block are extracted from the image feature map. Channel adaptation processing is performed on the block features corresponding to multiple image blocks respectively to obtain single-channel features and multi-channel features corresponding to multiple image blocks respectively.

[0049] In this embodiment, firstly, feature extraction processing is performed on the target image to extract key feature information and obtain an image feature map. Then, the target image is spatially segmented into multiple image blocks of specific sizes, discarding fragments at the boundaries that do not meet specific sizes. All image blocks are independent and do not overlap. The size of the image blocks is not specifically limited; for example, an image block can be 3×3, but is not limited to this. Next, each image block is flattened into a feature vector corresponding to its own size, i.e., block features, which represent the key feature information of the image block. Finally, for each image block, channel adaptation processing is performed on the block features to further extract structural contour information and spatial distribution information to obtain single-channel features, and to further extract semantic and texture information to obtain multi-channel features. By traversing each image block, the single-channel and multi-channel features corresponding to each image block can be obtained. In this way, the visual content of the image can be converted into a standard format suitable for local concept matching, thereby providing accurate data for subsequent processing.

[0050] In some implementations, the above-described feature extraction process on the target image to obtain an image feature map may include: Feature extraction is performed on the target image to obtain an initial feature map; The initial feature map is scaled to bipolar space to obtain the image feature map.

[0051] In this embodiment, firstly, feature extraction processing is performed on the target image to initially extract key feature information and obtain an initial feature map. The method of feature extraction is not specifically limited. For example, a backbone network such as a Convolutional Neural Network (CNN) or Transformer can be used to extract features from the target image, but it is not limited to this. Optionally, after completing the feature extraction processing, a lightweight adaptation structure is introduced. This structure may include, but is not limited to, sequentially connected upsampling layers, channel integration layers, and normalization layers. The upsampling layer is used to restore the spatial resolution of the initial feature map, the channel integration layer (such as a 1×1 convolutional layer) is used to unify the number of channels in the initial feature map, and the normalization layer is used to normalize the initial feature map. Further processing of the initial feature map using this lightweight adaptation structure can effectively improve the accuracy of the initial feature map; it should be understood that the processed initial feature map is the new initial feature map. Then, the initial feature map is scaled and mapped to the [-1, 1] bipolar space to obtain the mapped initial feature map as the image feature map. Thus, by simulating the bipolar characteristics of activation potentials in biological neurons, the feature map can have a positive and negative symmetrical response distribution because its own feature values ​​all satisfy the range of [-1, 1]. This amplifies the difference between foreground and background features and alleviates the feature saturation problem caused by the mean shift of the unipolar space during subsequent mask matching, thereby effectively improving the image feature extraction accuracy and providing accurate data for subsequent processing. The scaling method is not specifically limited; for example, activation functions such as Sigmoid can be used to scale the initial feature map to the bipolar space, as shown in the following formula (1): (1); In equation (1), Represents image feature maps; Represents the initial feature map; This represents the Sigmoid activation function.

[0052] In some implementations, the above-mentioned channel adaptation processing of the block features corresponding to multiple image blocks to obtain single-channel features and multi-channel features corresponding to multiple image blocks may include: The block features of the target image block are determined as the multi-channel features of the target image block; Channel fusion processing is performed on the block features of the target image block to obtain the single-channel features of the target image block; Here, the target image block represents any image block.

[0053] In this embodiment, for ease of understanding, any image block is defined as a target image block. After obtaining the block features of the target image block, the block features are further divided into two feature branches: a structure branch and an energy branch. In the energy branch, the block features of the target image block can be directly determined as the multi-channel features of the target image block. These features are the multi-channel original local features used for energy calculation, which describe complex semantic and texture information. In the structure branch, the block features of the target image block need to be fused to obtain the single-channel features of the target image block. These features are the single-channel saliency features used for structure matching, which describe the spatial response intensity and do not contain complex semantic and texture information. The channel fusion method can be flexibly set according to the actual situation. For example, the block features can be compressed by a 1×1 convolutional layer to fuse the multi-channel features into a single channel, but it is not limited to this. It should be noted that both the single-channel features and the multi-channel features are located within the bipolar space [-1, 1], that is, the feature values ​​of both satisfy the value range of [-1, 1], but the number of channels of the two are different. Thus, by dividing the block features into two feature branches, structural and energy, we can effectively capture the structural contours and spatial response intensity in the target image and generate corresponding single-channel features. At the same time, we can deeply mine the complex semantic and texture information in the target image and generate corresponding multi-channel features, thereby providing accurate data for subsequent processing.

[0054] In some implementations, refer to Figure 3 In step S103 above, matching corresponding single-pole spatial masks for multiple image patches based on their respective single-channel features may include: Dictionary retrieval processing is performed based on the single-channel features of the target image patch to obtain a binary exhaustive dictionary of the target image patch; The binary exhaustive dictionary of the target image patch is mapped to obtain the unipolar space mask of the target image patch; Here, the target image block represents any image block.

[0055] In this embodiment, for ease of understanding, any image patch is defined as a target image patch. After obtaining the single-channel features of the target image patch, a dictionary lookup process is first performed based on the single-channel features. This aims to retrieve a binary exhaustive dictionary matching the target image patch from a pre-built target dictionary, using the single-channel features as a benchmark. It is worth noting that the target dictionary is a pre-built combined concept memory that exhaustively enumerates all {-1, 1} bipolar binary patterns that satisfy the image patch size (e.g., 3×3). The {-1, 1} bipolar binary pattern is a standard binary pattern, representing that each pixel in the image patch can only be either a foreground with a value of 1 or a background with a value of -1. The binary exhaustive dictionary represents the {-1, 1} bipolar binary patterns that match the target image patch. Simply put, the target dictionary covers any possible spatial structure within the size of the target image patch. Regardless of the spatial distribution of the single-channel features, the corresponding {-1, 1} bipolar binary pattern can be retrieved from the target dictionary. Then, the binary exhaustive dictionary is further mapped to a {0, 1} unipolar space to obtain a unipolar space mask. In this mask, pixels with a value of 0 represent background pixels, and pixels with a value of 1 represent foreground structure pixels. In this way, continuous single-channel features can be hard-bound into a discrete binary exhaustive dictionary and mapped to a {0, 1} unipolar space mask, thereby achieving Hardmax hard binding from continuous sub-symbol signals to discrete visual symbols, that is, converting feature matching from continuous floating-point calculations to discrete symbol determination. In subsequent processing, feature enhancement is performed based on this template, which can force image patches to clearly belong to the foreground or background, thereby cutting off the propagation link of background noise at the physical level, while effectively perceiving and enhancing the foreground structure of the image.

[0056] In some implementations, the dictionary lookup process based on the single-channel features of the target image patch to obtain a binary exhaustive dictionary of the target image patch may include: The routing probability of the target image patch is obtained by performing routing operations based on the target dictionary and the single-channel features of the target image patch. Based on the routing probability of the target image patch, the target dictionary is searched to obtain a binary exhaustive dictionary of the target image patch.

[0057] In this embodiment, for ease of understanding, any image patch is defined as the target image patch. In the dictionary retrieval process, routing operations are performed based on the target dictionary and single-channel features, aiming to calculate the corresponding routing probability based on the single-channel features. This probability determines which bipolar binary mode is selected in the target dictionary. In specific implementation, the target dictionary and single-channel features are first processed using the concept attention mechanism, that is, the single-channel features are used as the query vector and the dot product of the single-channel features and the transposed target dictionary is calculated as the score vector, as shown in the following formula (2): (2); In equation (2), Represents the score vector; Indicates single-channel characteristics; Indicates the target dictionary.

[0058] Subsequently, using the dual-state routing mechanism, the score vector is further converted into the corresponding probability distribution, i.e., the routing probability, thereby achieving decisive concept clustering, as shown in the following formula (3): (3); In equation (3), Indicates the probability of routing; This indicates an operation that returns the input parameter that results in the maximum value. This indicates one-hot encoding.

[0059] Optionally, during training, the score vector is backpropagated using the Gumbel Softmax activation function to update the parameters of the target dictionary. , This represents the parameters for backpropagation differentiation. This represents the Gumbel Softmax activation function.

[0060] Finally, based on the routing probability, the optimal bipolar binary pattern that matches the target image patch is adaptively selected from the target dictionary as the binary exhaustive dictionary. Specifically, the dot product of the routing probability and the target dictionary is calculated as the binary exhaustive dictionary, as shown in the following formula (4): (4); In equation (4), This represents a binary exhaustive dictionary.

[0061] After obtaining the binary exhaustive dictionary, it is necessary to map the binary exhaustive dictionary to the {0, 1} unipolar space. Specifically, this involves calculating the sum of the binary exhaustive dictionary and one, and then calculating the ratio of this sum to two as a unipolar space mask, as shown in the following formula (5): (5); In equation (5), This represents a single-space mask.

[0062] Therefore, this implementation method, through a dual-state routing mechanism, converts the matching degree between single-channel features and each bipolar binary pattern in the target dictionary into a routing probability distribution. Based on this, it adaptively selects the optimal bipolar binary pattern from the target dictionary as a binary exhaustive dictionary, thereby hard-binding continuous single-channel features into a discrete binary exhaustive dictionary and mapping it to a {0,1} unipolar space mask. This method can transform the matching problem of high-dimensional continuous features into a probability retrieval in discrete space, without the need for complex high-dimensional convolution or attention calculations. At the same time, since the target dictionary is a predefined discrete combination, the retrieval process is not affected by continuous background noise disturbances. Thus, it can improve the mask matching accuracy while reducing the mask matching overhead, thereby effectively improving the accuracy and efficiency of Hardmax hard binding.

[0063] In some implementations, refer to Figure 4 In step S104 above, feature enhancement processing is performed on the multi-channel features corresponding to the multiple image patches based on the single-pole space masks corresponding to the multiple image patches respectively, to obtain the sparse image features corresponding to the multiple image patches, which may include: Based on the unipolar space mask of the target image patch, energy feature extraction processing is performed on the multi-channel features of the target image patch to obtain the channel energy features of the target image patch. The channel energy features of the target image patch are broadcast to the unipolar space mask of the target image patch to obtain the sparse image features of the target image patch. Here, the target image block represents any image block.

[0064] In this embodiment, for ease of understanding, any image block is defined as the target image block. After obtaining the unipolar spatial mask of the target image block, a channel-wise energy modulation mechanism is introduced, which uses the unipolar spatial mask to extract the average physical intensity of the structural region from the multi-channel features to obtain the channel energy feature. This feature represents the channel-wise energy vector, as shown in the following formula (6): (6); In equation (6), Indicates channel energy characteristics; Indicates multi-channel characteristics; This is a very small constant used to prevent the denominator from being zero; it can be flexibly set according to the actual situation. This indicates element-wise multiplication.

[0065] It is worth noting that since pixels with a value of 0 in a single-pole space mask represent background pixels and pixels with a value of 1 represent foreground structure pixels, using this mask to enhance multi-channel features can force image blocks to make clear foreground or background classification. At the same time, by using 0 / 1 discrete symbols to force the background energy to zero at the physical level, the propagation path of background noise in the feature aggregation process is fundamentally blocked, thereby improving the accuracy of image features.

[0066] Subsequently, a concept encoding mechanism is introduced, which further broadcasts the channel energy features to the {0,1} unipolar space mask, aiming to inject physical visual sensory intensity into the channel energy features, thereby effectively perceiving and enhancing the foreground structure of the image. In specific implementation, the element-wise product of the channel energy features and the unipolar space mask is calculated as the sparse image features, as shown in the following formula (7): (7); In equation (7), This represents sparse image features.

[0067] In some implementations, refer to Figure 5 In step S105 above, the sparse image features corresponding to multiple image patches are fused to obtain an image descriptor, which may include: Dimensionality reduction is performed on the sparse image features corresponding to multiple image patches to obtain the semantic features corresponding to each image patch. The semantic features corresponding to multiple image patches are fused according to the spatial slot order to obtain image descriptors.

[0068] In this embodiment, after obtaining the refined feature map, related technologies typically use Global Average Pooling (GAP) to process the refined feature map, thereby generating a one-dimensional continuous floating-point feature vector. However, the traditional GAP method destroys the two-dimensional spatial topology of the feature map, which causes the model performing downstream tasks to lose its local region retrieval ability and fail to accurately locate key foreground structures. To address this, this embodiment introduces a Memory Stitching & Global Representation mechanism. This mechanism aims to reduce the dimensionality of the sparse image features corresponding to all image patches and stitch them together according to the spatial slot order (original two-dimensional spatial coordinate order) into a complete image descriptor containing accurate spatial topology. This aggregates local discrete symbols into a global image descriptor that retains the absolute spatial topology, thus fully preserving the two-dimensional spatial topology of the feature map and effectively improving the accuracy of the image descriptor.

[0069] Specifically, after obtaining the sparse image features corresponding to multiple image patches, the sparse image features corresponding to multiple image patches are first reduced in dimensionality based on Local Linear Projection. This aims to promote the interaction and fusion of visual cues in the spatial dimension and visual cues in the channel dimension, and reduce the dimensionality to the standard semantic dimension to obtain the semantic features corresponding to multiple image patches. In the specific implementation, for ease of understanding, any image patch is defined as the target image patch. The sparse image features of the target image patch are flattened and input into the fully connected layer (Linear). The semantic features of the target image patch are output by the fully connected layer, as shown in the following formula (8): (8); In equation (8), This represents the semantic slot vector after local aggregation, i.e., semantic features; Indicates the flattening operation; This indicates a fully connected layer.

[0070] After dimensionality reduction, a fusion operation is performed. In practice, the semantic features corresponding to multiple image blocks are first flattened and stitched together in strict accordance with the original two-dimensional spatial coordinate order to construct a spatial slot sequence, as shown in the following formula (9): (9); In equation (9), Indicates the spatial slot sequence; Indicates the first Semantic features of each image patch; Indicates the number of image patches; This indicates a flattening and sewing process.

[0071] Subsequently, the spatial slot sequence is subjected to global L2 normalization to eliminate scale differences, thereby generating an image descriptor. This descriptor fully preserves the absolute spatial location information, which enables the downstream task model to have the ability to retrieve local regions based on specific spatial locations, as shown in the following formula (10): (10); In equation (10), Represents an image descriptor; This indicates global L2 normalization.

[0072] To facilitate understanding of the image description method described above in this application, an example of a practical application scenario of the image description method described above in this application is provided below.

[0073] This application scenario is an image retrieval scenario. Users upload the image to be queried through the front end, and this image becomes the target image. (See reference...)Figure 6 The process of converting the target image into an image descriptor and realizing the downstream image retrieval task in this application scenario is shown in the following steps (1)-(3).

[0074] (1) Visual Perception & Spatial Adaptation aims to convert the heterogeneous features output by different backbone networks into a standard bipolar spatial format suitable for local concept matching. The specific operation is as follows: (1.1) Feature Extraction and Adaptation: CNN is used to extract features from the target image to obtain the feature extraction results. A lightweight adaptation structure with upsampling layers, 1×1 convolutional layers (channel integration layers), and normalization layers is used to process the feature extraction results to obtain the original dense feature map, denoted as the initial feature map. ,in Indicates the number of channels. Indicates high, It indicates width.

[0075] (1.2) Bipolar Mapping: Mapping the initial feature map The image feature map is obtained by scaling and mapping to the [-1, 1] bipolar space as shown in formula (1) above. .

[0076] (1.3) Patchwise Unfold and Two-Branch Preparation: In the spatial dimension, from the image feature map Multiple 3×3 local image patches are extracted, and each patch is flattened into a feature vector of length 9, denoted as the patch feature. For each image patch, the patch feature is directly determined as a multi-channel feature. Simultaneously, a 1×1 convolutional layer is used to compress the block features into channels, resulting in single-channel features. .

[0077] (2) Concept Binding & Foveal Encoding aims to achieve hard binding from continuous sub-symbol signals to discrete visual symbols and inject physical sensory intensity. Specifically, for each image block, the following operations are performed: (2.1) Concept Attention: using single-channel features Used as a query vector, and single-channel features are calculated. With the transposed target dictionary The dot product is used as the score vector. As shown in formula (2) above.

[0078] (2.2) Dual-State Routing: The score vector is... Transform into probability distribution As shown in formula (3) above, decisive concept clustering is achieved. Based on this, the best-matching binary exhaustive dictionary is further retrieved. As shown in formula (4) above.

[0079] (2.3) Spatial Mask Generation: Enumerating a binary dictionary Mapped to a single-pole space mask In this mask, pixels with a value of 0 represent background pixels, and pixels with a value of 1 represent foreground structure pixels, as shown in the above formula (5).

[0080] (2.4) Channel-wise Energy Modulation: Utilizing a unipolar spatial mask From multi-channel characteristics The average physical intensity of the structural region is extracted to obtain the channel energy characteristics. As shown in formula (6) above.

[0081] (2.5) Concept Encoding: Encoding the channel energy features Broadcast to unipolar space mask Generate sparse image features with physical strength. As shown in formula (7) above.

[0082] (3) Memory Stitching & Global Representation aims to aggregate local discrete symbols into global descriptors that preserve the absolute spatial topology. The specific operations are as follows: (3.1) Local Linear Projection: For each image patch, sparse image features are projected onto the local linear projection plotter. After flattening, visual cues from spatial and channel dimensions are fused through a fully connected layer and reduced to the standard semantic dimension. As shown in formula (8) above, the semantic features are obtained. .

[0083] (3.2) Memory Stitching: Flatten and stitch the semantic features corresponding to all image patches strictly according to the original two-dimensional spatial coordinate order to construct a spatial slot sequence. As shown in formula (9) above.

[0084] (3.3) Global L2 Normalization: For spatial slot sequences Global L2 normalization is performed to eliminate scale differences, as shown in formula (10) above, to obtain the image descriptor. .

[0085] (4) Image retrieval: A preset image database is invoked, which stores descriptors corresponding to several image samples. The similarity between the image descriptors obtained in step (3) above and the descriptors of each image sample is calculated by means such as Euclidean distance and cosine similarity. Image samples with similarity greater than a preset similarity threshold are returned to the front end. The front end receives the image samples and renders and displays them, thereby realizing image retrieval.

[0086] In addition, refer to Figure 7 This application also provides an image description device, which may include: Acquisition module 201 is used to acquire the target image; The first processing module 202 is used to perform feature extraction and block adaptation processing on the target image to obtain single-channel features and multi-channel features corresponding to multiple image blocks respectively. The second processing module 203 is used to match corresponding unipolar space masks for multiple image blocks based on the single-channel features corresponding to each of the multiple image blocks. The third processing module 204 is used to perform feature enhancement processing on the multi-channel features corresponding to the multiple image blocks according to the single-pole space mask corresponding to the multiple image blocks respectively, so as to obtain the sparse image features corresponding to the multiple image blocks respectively. The fourth processing module 205 is used to fuse the sparse image features corresponding to multiple image blocks to obtain image descriptors.

[0087] The content of the above method embodiments is applicable to the device embodiments. The specific functions implemented by the device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0088] Finally, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described image description method.

[0089] The content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0090] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0091] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0092] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0093] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0094] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0095] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0096] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0097] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0098] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0099] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0100] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. An image description method, characterized in that, The method includes: Acquire the target image; The target image is subjected to feature extraction and block adaptation processing to obtain single-channel features and multi-channel features corresponding to multiple image blocks respectively; Based on the single-channel features corresponding to the multiple image blocks, a corresponding unipolar space mask is matched for the multiple image blocks; Based on the single-pole space mask corresponding to each of the multiple image blocks, feature enhancement processing is performed on the multi-channel features corresponding to each of the multiple image blocks to obtain the sparse image features corresponding to each of the multiple image blocks. The sparse image features corresponding to the multiple image blocks are fused to obtain image descriptors.

2. The method according to claim 1, characterized in that, The step of performing feature extraction and block-based adaptation on the target image to obtain single-channel and multi-channel features corresponding to multiple image blocks includes: The target image is subjected to feature extraction processing to obtain an image feature map; The target image is divided into multiple image blocks, and block features corresponding to each of the multiple image blocks are extracted from the image feature map; Channel adaptation processing is performed on the block features corresponding to the multiple image blocks respectively to obtain single-channel features and multi-channel features corresponding to the multiple image blocks respectively.

3. The method according to claim 2, characterized in that, The step of performing feature extraction processing on the target image to obtain an image feature map includes: The target image is subjected to feature extraction processing to obtain an initial feature map; The initial feature map is scaled to bipolar space to obtain the image feature map.

4. The method according to claim 2, characterized in that, The step of performing channel adaptation processing on the block features corresponding to the multiple image blocks respectively to obtain single-channel features and multi-channel features corresponding to the multiple image blocks includes: The block features of the target image block are determined as the multi-channel features of the target image block; The block features of the target image block are subjected to channel fusion processing to obtain the single-channel features of the target image block; Wherein, the target image block refers to any one of the image blocks.

5. The method according to claim 1, characterized in that, The step of matching corresponding single-pole spatial masks for the multiple image blocks based on their respective single-channel features includes: Dictionary retrieval processing is performed based on the single-channel features of the target image block to obtain a binary exhaustive dictionary of the target image block; The binary exhaustive dictionary of the target image block is mapped to obtain the unipolar space mask of the target image block; Wherein, the target image block refers to any one of the image blocks.

6. The method according to claim 5, characterized in that, The step of performing dictionary lookup processing based on the single-channel features of the target image patch to obtain a binary exhaustive dictionary for the target image patch includes: The routing probability of the target image patch is obtained by performing routing operations based on the target dictionary and the single-channel features of the target image patch. Based on the routing probability of the target image patch, the target dictionary is searched to obtain a binary exhaustive dictionary of the target image patch.

7. The method according to claim 1, characterized in that, The step of performing feature enhancement processing on the multi-channel features corresponding to the multiple image patches respectively, based on the single-pole spatial masks corresponding to the multiple image patches respectively, to obtain the sparse image features corresponding to the multiple image patches respectively, includes: Based on the unipolar space mask of the target image block, energy feature extraction processing is performed on the multi-channel features of the target image block to obtain the channel energy features of the target image block; The channel energy features of the target image block are broadcast to the unipolar spatial mask of the target image block to obtain the sparse image features of the target image block; Wherein, the target image block refers to any one of the image blocks.

8. The method according to claim 1, characterized in that, The step of fusing the sparse image features corresponding to the multiple image blocks to obtain an image descriptor includes: Dimensionality reduction is performed on the sparse image features corresponding to the multiple image blocks respectively to obtain the semantic features corresponding to the multiple image blocks respectively; The semantic features corresponding to the multiple image blocks are fused according to the spatial slot order to obtain the image descriptor.

9. An image description device, characterized in that, The device includes: The acquisition module is used to acquire the target image; The first processing module is used to perform feature extraction and block adaptation processing on the target image to obtain single-channel features and multi-channel features corresponding to multiple image blocks respectively. The second processing module is used to match corresponding unipolar space masks for the multiple image blocks according to the single-channel features corresponding to the multiple image blocks respectively; The third processing module is used to perform feature enhancement processing on the multi-channel features corresponding to the multiple image blocks respectively according to the single-pole space mask corresponding to the multiple image blocks respectively, so as to obtain the sparse image features corresponding to the multiple image blocks respectively. The fourth processing module is used to fuse the sparse image features corresponding to the multiple image blocks to obtain image descriptors.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the image description method according to any one of claims 1 to 8.