Image information structuring method and device, equipment and storage medium
By extending the CLIP model's Transformer architecture and cross-modal cosine similarity calculation, multi-scale visual features are extracted and matched with text attributes, solving the problem of low matching degree between image and text details and achieving high-precision recognition of fine-grained attributes.
Patent Information
- Application Number
- CN202511111282.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies have low matching accuracy between image and text details and are difficult to effectively identify fine-grained attributes such as "frameless glasses", "long curly hair" and "wear on the left front tire of a vehicle".
By extending the Transformer architecture of the native CLIP model to 24 layers, multi-scale visual feature vectors are extracted. Combined with a pre-trained cross-modal cosine similarity calculation mechanism and an attribute-aware text encoder, high-precision matching of visual features and text attributes is achieved.
It improves the matching accuracy between image and text details, enabling more accurate identification of fine-grained attributes and solving the problem that traditional single-level features are difficult to match complex text attributes.
Smart Images

Figure CN120953994A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image information processing, and in particular to a method, apparatus, device, and storage medium for structuring image information. Background Technology
[0002] Cross-modal models based on Contrastive Language-Image Pre-training (CLIP) have achieved efficient recognition of open-word attributes such as "wearing glasses," "backpack," and "white vehicle" in human and vehicle feature labeling tasks through text-image semantic alignment technology, providing core technical support for fields such as intelligent security, autonomous driving, and smart retail. Existing improvement schemes (such as ViLD extending open-word detection capabilities through knowledge distillation, GLIP achieving language-guided target localization, and SAFE enhancing few-sample attribute recognition through semantic perception fine-tuning) have improved labeling performance to some extent, but with the increasing demand for fine-grained attribute recognition in real-world scenarios (such as "frameless glasses," "long curly hair," and "wear on the left front tire of a vehicle"), the problem of low image-text detail matching in existing technologies is becoming increasingly prominent. Summary of the Invention
[0003] This invention provides a method, apparatus, computer device, and storage medium for structuring image information to solve the problem of low matching degree between image and text details in the prior art.
[0004] Firstly, an image information structuring method is provided, including: Input the image to be recognized; For the image, a multi-scale visual feature vector is extracted using a visual encoder; The multi-scale visual feature vectors are matched with the text attribute semantic space to obtain the corresponding fine-grained attribute text.
[0005] Optionally, the step of extracting multi-scale visual feature vectors from the image using a visual encoder specifically includes: Based on the native CLIP model, by expanding the number of layers in its Transformer architecture, corresponding visual features are output from different Transformer architecture layers, and then the multi-layer visual features are fused to obtain a multi-scale visual feature vector.
[0006] Optionally, based on the native CLIP model, by extending the number of layers in its Transformer architecture, corresponding visual features are output from different Transformer architecture layers, including: The number of layers in the Transformer architecture of the original CLIP model is expanded to 24, and feature output interfaces are set in layers 6, 12 and 24 to output shallow visual features, mid-level visual features and deep visual features, respectively.
[0007] Optionally, based on the native CLIP model, by extending the number of layers in its Transformer architecture, corresponding visual features are output from different Transformer architecture layers, including: The image is divided into multiple small blocks and each block is converted into a corresponding vector. The vectors corresponding to multiple small blocks are arranged according to the position order of each small block in the image to obtain a vector sequence; A learnable classification-specific identifier is added to the beginning position of the vector sequence to obtain the input sequence; The input sequence is processed by a Transformer architecture with an extended number of layers to obtain multi-scale visual feature vectors.
[0008] Optionally, matching the multi-scale visual feature vector with the text attribute semantic space to obtain the corresponding fine-grained attribute text includes: By using a pre-trained cross-modal cosine similarity calculation mechanism, the multi-scale visual feature vectors are compared with the predefined fine-grained attribute text vectors in the text attribute semantic space, and the fine-grained attribute text corresponding to the fine-grained attribute text vector with the highest similarity is selected.
[0009] Optionally, the text attribute semantic space includes: Based on the native CLIP model, the text encoder is extended to an attribute-aware text encoder to segment predefined fine-grained attribute text into basic words and attribute words. The basic words are processed through a basic word embedding matrix to generate basic word vectors, and the attribute words are processed through an attribute word embedding matrix to generate attribute word vectors. The basic word vectors and attribute word vectors are then fused to obtain a composite embedding vector. The composite embedding vector is encoded to obtain fine-grained attribute text vectors, and all fine-grained attribute text vectors constitute the text attribute semantic space.
[0010] Optionally, the step of encoding the composite embedding vector to obtain fine-grained attribute text vectors, wherein all fine-grained attribute text vectors constitute a text attribute semantic space, including: The composite embedding vector is constrained to interact with the attribute words through an attribute mask attention mechanism, and then encoded to obtain a fine-grained attribute text vector.
[0011] Secondly, an image information structuring device is provided, comprising: The input module is used to input the image to be recognized; The processing module is used to extract multi-scale visual feature vectors from the image using a visual encoder. The output module is used to match the multi-scale visual feature vector with the text attribute semantic space to obtain the corresponding fine-grained attribute text.
[0012] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the steps of the above-described image information structuring method.
[0013] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described image information structuring method.
[0014] In the above-mentioned image information structuring method, device, computer equipment and storage medium, the multi-scale visual feature vector covers visual information at different levels. The matching mechanism with the semantic space of text attributes can establish a direct association between visual features and fine-grained text descriptions, which solves the problem that traditional single-level features are difficult to match complex text attributes, making the matching degree between images and text details higher and more accurate. Attached Figure Description
[0015] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a flowchart illustrating an image information structuring method according to an embodiment of the present invention; Figure 2 This is a schematic diagram of an image information structuring device according to an embodiment of the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0018] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0019] It should also be understood that the term “and / or” as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0020] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," or "in response to determination." Similarly, the phrase "if determined" or "if matched to [described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once matched to [described condition or event]," or "in response to matched to [described condition or event]."
[0021] Furthermore, in the description of this invention and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0022] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of the invention include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0023] It should be understood that the sequence number of each step in the following embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0024] To illustrate the technical solution of the present invention, specific embodiments are described below.
[0025] Please see Figure 1 As shown, Figure 1 A flowchart illustrating the image information structuring method provided in this embodiment of the invention includes the following steps: S11: Input the image to be recognized.
[0026] This step is the starting point of the image information structuring method. It involves inputting the raw image data containing the target object (such as a person or vehicle) into the trained model. Its core function is to provide raw visual information for the entire process and is a prerequisite for achieving structured image parsing. This step does not involve complex feature processing; it only needs to ensure that the image format meets the model's input requirements (such as common RGB format, preset resolution, etc.).
[0027] The image to be identified must contain a processable target (such as a "person" or "vehicle"), which forms the basis for the subsequent "fine-grained attribute text" output. If the image does not contain a target object (such as a pure landscape image), the model may output an empty result or no valid attributes.
[0028] In one example, such as a traffic management scenario, the input is an image to be identified captured by a road surveillance camera (containing a "white sedan with a damaged left rear tire"); or in a person recognition scenario, the input is a captured image of a person (such as "a woman with long hair wearing black-rimmed glasses").
[0029] S12: For the image, extract multi-scale visual feature vectors using a visual encoder.
[0030] The multi-scale visual feature vector is a vector containing visual features at different levels, generally referring to a vector that includes visual features from local to global. This step is the core processing step of the image information structuring method. It refers to extracting visual features at different levels (such as shallow details, mid-level components, deep global features, etc.) from the input image through a visual encoder (such as the improved CLIP model) and fusing them into a unified multi-scale visual feature vector representation, providing a visual feature foundation for subsequent matching with the semantic space of text attributes.
[0031] In one embodiment, extracting multi-scale visual feature vectors from the image using a visual encoder specifically includes: Based on the native CLIP model, by expanding the number of layers in its Transformer architecture, corresponding visual features are output from different Transformer architecture layers, and then the multi-layer visual features are fused to obtain a multi-scale visual feature vector.
[0032] This step clarifies the extraction method for multi-scale visual feature vectors: based on the native CLIP model, visual features are extracted from Transformer layers of different depths by expanding the number of layers in its Transformer architecture, and then these multi-layer features are fused to obtain the final multi-scale visual feature vectors. The core lies in three steps: "expanding the number of layers," "layered extraction," and "feature fusion," which achieves a complete representation of visual information from details to the global picture through hierarchical processing.
[0033] The native CLIP model's visual encoder uses a Transformer architecture (typically 12 layers). This step enhances the ability to extract features from different levels by increasing the number of layers. The extended Transformer architecture retains the cross-modal learning advantages of the native model while achieving finer feature differentiation through deeper layers.
[0034] Taking an image containing "a woman with long hair wearing black-rimmed glasses" as an example, the native CLIP 12-layer Transformer is extended to enhance the extraction capability of fine-grained features such as "long hair" and "black-rimmed glasses." Features are output from shallow layers (such as the earlier Transformer layers): including the frame lines of the black-rimmed glasses, the reflective texture of the lenses, and local features such as the curl details and edges of the long hair strands. Features are output from middle layers (such as the middle Transformer layers): including component-related features such as "black-rimmed glasses covering the eyes" and "long hair wrapping around both sides of the head." Features are output from deep layers (such as the later Transformer layers): including the global semantic feature "a woman with long hair wearing glasses." These three layers of features are then concatenated and compressed into a vector of a unified dimension to obtain a multi-scale visual feature vector.
[0035] In one embodiment, the step of outputting corresponding visual features from different Transformer architecture layers based on the native CLIP model by expanding the number of Transformer architecture layers includes: The number of layers in the Transformer architecture of the original CLIP model is expanded to 24, and feature output interfaces are set in layers 6, 12 and 24 to output shallow visual features, mid-level visual features and deep visual features, respectively.
[0036] The native CLIP model's visual encoder typically employs a 12-layer Transformer architecture. This step doubles the number of layers to 24, enhancing the model's ability to differentiate features at different levels. The extended architecture retains the Transformer architecture's self-attention mechanism and feedforward network core structure, improving its ability to resolve fine-grained features by increasing the number of layers. To address the feature requirements at different semantic levels, output interfaces are set in three key layers: Layer 6 (shallow layer): Positioned in the middle of the extended Transformer architecture, its output features focus on the low-level visual details of the image (such as edge contours, texture, and local pixel distribution); Layer 12 (middle layer): Located in the middle of the extended Transformer architecture, its output features include the structured relationships of target components (such as the spatial position of glasses and face, and the connection between hair and shoulders); Layer 24 (deep layer): As the final layer of the extended Transformer architecture, its output features reflect the overall category and core content of the image (such as major categories like "person" and "vehicle" and their associated high-level attributes).
[0037] Taking an image containing vehicle type, color, and components as an example, based on the native CLIP model, the ability to analyze fine-grained vehicle features is improved by extending it to a 24-layer Transformer architecture. Layer 6 (shallow layer): Outputs features focusing on underlying visual details, including: the distribution of red pixels on the vehicle surface (color brightness and saturation variations), the irregular edge contour of the damaged left headlight, the reflective texture of the window glass, and the rubber texture of the tires. These features directly correspond to visually observable local elements in the image. Layer 12 (middle layer): Outputs features containing structured relationships between components, including: "the positional relationship between the left headlight and the front of the vehicle" (the left headlight is located on the left side of the front of the vehicle), "the boundary relationship between the damaged area and the headlight body" (the damaged area is located in the upper right corner of the headlight body), and "the regional division of the vehicle body color and the window glass" (red paint covers the main body of the vehicle, and the glass area is transparent), integrating shallow details into component-level structural information. Layer 24 (Deep Layer): Outputs features that aggregate global semantics, clearly defining the core content of the image as "red sedan with a broken left front light." This includes high-level semantics such as vehicle type (sedan), color (red), and the state of the key component (broken left front light), providing a global anchor point for subsequent matching with text attributes (e.g., "red sedan with a broken left front light"). Through these three layers of feature output, complete feature capture of the vehicle, from local details to its overall state, is achieved.
[0038] In one embodiment, the step of outputting corresponding visual features from different Transformer architecture layers based on the native CLIP model by extending the number of Transformer architecture layers includes: segmenting the image into multiple small blocks and converting each small block into a corresponding vector; arranging the vectors corresponding to the multiple small blocks according to the position order of each small block in the image to obtain a vector sequence; adding a learnable classification-specific identifier at the beginning of the vector sequence to obtain an input sequence; and processing the input sequence through a Transformer architecture with extended layers to obtain multi-scale visual feature vectors.
[0039] This step is the core computational process for obtaining multi-scale visual feature vectors. Through four stages—image segmentation, vector transformation, sequence construction, and Transformer processing—the image is transformed into a visual feature vector containing spatial location information and multi-scale semantics. The key lies in preserving the spatial structure of the image (through positional order) and introducing global semantic anchors (classification-specific identifiers) to provide structured input for the extended-layer Transformer architecture, ultimately obtaining multi-scale visual feature vectors.
[0040] Taking the image of a "red car with a broken left headlight" as an example, the first step is to divide the image into 14×14=196 small blocks of 16×16 pixels. For example, the small block of the broken left headlight area is converted into a 768-dimensional vector containing features of "damaged edge" and "dark crack" using linear mapping; the small block of the red car body area is converted into a 768-dimensional vector containing features of "red pixel distribution" and "smooth paint texture" using linear mapping; and the small block of the window area is converted into a 768-dimensional vector containing features of "transparent texture" and "glass reflection" using linear mapping. The second step is to arrange the 196 vectors corresponding to the 196 small blocks according to the position of each block in the original image. For example, the 196 small blocks can be numbered in the order of "from left to right and from top to bottom": the first block corresponds to the first small block in the upper left corner of the image, the second block is its adjacent block to the right, and so on. The 14th block is the rightmost small block in the first row; the 15th block is the leftmost small block in the second row, the 16th block is its adjacent block to the right, and so on, until the 196th block corresponds to the last small block in the lower right corner of the image. This numbering method is essentially to record the two-dimensional spatial position with a linear sequence number to ensure that the "nth vector" in the vector sequence strictly corresponds to the spatial position of the "nth small block" in the image. The third step involves inserting a learnable classification-specific identifier at the beginning of the vector sequence. This learnable classification-specific identifier is essentially a vector with the same dimension as the image patch vectors (e.g., 768-dimensional), ultimately forming a 197×768 input sequence (196 patch vectors + 1 learnable classification-specific identifier). This learnable classification-specific identifier, through the self-attention mechanism of the Transformer architecture, continuously interacts with all image patch vectors (i.e., learns to focus on the features of the patches) and continuously adjusts its own value during backpropagation. For example, in an image of a "red car with a broken left headlight," it gradually learns to focus on associating features of patches such as "red car body" and "broken left headlight area." The fourth step involves processing the input sequence through a 24-layer Transformer architecture: the lower-layer output retains details such as "crack texture at the damaged left front light" and "local red color blocks on the car body"; the middle-layer output integrates structural information such as "positional relationship between the left front light and the front of the car" and "red paint covering the main body of the car body"; and the higher-layer output is aggregated into a multi-scale visual feature vector containing the global semantics of "red sedan, damaged left front light" through a classification-specific identifier.
[0041] S13: Match the multi-scale visual feature vector with the text attribute semantic space to obtain the corresponding fine-grained attribute text.
[0042] This step is the final output of image information structuring. By performing cross-modal matching between the multi-scale visual feature vectors extracted in the previous steps (which contain comprehensive information from local details to global semantics of the image) and the pre-constructed text attribute semantic space (composed of a large number of vectors of predefined fine-grained attribute texts), the text description with the closest semantics is found, thereby transforming the image information into structured fine-grained attribute text.
[0043] The text attribute semantic space consists of vectors generated by encoding a large number of predefined fine-grained attribute texts (such as "red sedan, broken left headlight", "black SUV, missing right rearview mirror", "long-haired woman, wearing black-rimmed glasses", etc.). These text vectors have the same dimension (e.g., 768-dimensional) as multi-scale visual feature vectors.
[0044] In one embodiment, matching the multi-scale visual feature vector with the text attribute semantic space to obtain the corresponding fine-grained attribute text includes: using a pre-trained cross-modal cosine similarity calculation mechanism to compare the similarity between the multi-scale visual feature vector and the predefined fine-grained attribute text vector in the text attribute semantic space, and selecting the fine-grained attribute text corresponding to the fine-grained attribute text vector with the highest similarity.
[0045] This step clarifies the specific implementation of the matching: by using a pre-trained cross-modal cosine similarity calculation mechanism, the semantic correlation between multi-scale visual feature vectors and predefined text vectors in the text attribute semantic space is quantified, and finally the text with the highest similarity is selected as the output.
[0046] For example, the multi-scale visual feature vector contains visual information such as "red car body," "broken left headlight," and "car model," with 768 dimensions. Predefined text vectors (partial) in the text attribute semantic space: A: "Red car, broken left headlight" (768 dimensions); B: "Red car, right headlight intact" (768 dimensions); C: "Blue car, broken left headlight" (768 dimensions); D: "Red SUV, broken left headlight" (768 dimensions). After calculation using a pre-trained cosine similarity mechanism, the results are as follows: Visual vector similarity with A: 0.92 (highest) – because they match across all dimensions: "color (red)," "car model," "part (left headlight)," and "state (broken)"; Similarity with B: 0.75 – the difference lies in the mismatch in part position and state; Similarity with C: 0.68 – the difference lies in the mismatch in color attribute; Similarity with D: 0.71 – the difference lies in the mismatch in car model attribute. The text corresponding to A with the highest similarity, "red sedan, broken left front light", is selected as the fine-grained attribute text of the final output, thus completing the conversion of image information into structured text.
[0047] In one embodiment, the text attribute semantic space includes: based on the native CLIP model, extending the text encoder therein to an attribute-aware text encoder to segment predefined fine-grained attribute text to obtain base words and attribute words; generating base word vectors from the base words through a base word embedding matrix, generating attribute word vectors from the attribute words through an attribute word embedding matrix, fusing the base word vectors and attribute word vectors to obtain a composite embedding vector; encoding the composite embedding vector to obtain fine-grained attribute text vectors, and all fine-grained attribute text vectors constitute the text attribute semantic space.
[0048] Taking the text "red sedan, left front light damaged" as an example, the first step is to segment the text "red sedan, left front light damaged" into the following words: basic words: "sedan" and "damaged"; attribute words: "red" (color attribute) and "left front light" (component attribute). The second step involves generating a 768-dimensional vector for the basic word vector: "sedan" is embedded using a basic word embedding matrix; a corresponding vector is generated for "damaged". Attribute word vectors are also generated: "red" is embedded using an attribute word embedding matrix; a corresponding vector is generated for "left front light". A composite embedding vector is formed by fusing these four vectors to create a composite embedding vector containing "car model + color + component + status". The third step involves encoding the composite embedding vector using an attribute-aware text encoder to generate a 768-dimensional fine-grained attribute text vector. The encoding process uses an attribute-aware mechanism to enhance the attribute specificity of "red" and "left front light" (e.g., the "red" vector component is bound to the color perception dimension, and the "left front light" component is bound to the spatial location dimension). The resulting fine-grained attribute text vector retains the complete semantics of "red sedan, left front light damaged" while also aligning with multi-scale visual feature vectors in terms of dimensions. This fine-grained attribute text vector, together with other fine-grained attribute text vectors (such as "the right front light of the blue sedan is intact" and "the left front light of the red SUV is damaged"), constitutes the text attribute semantic space.
[0049] In one embodiment, encoding the composite embedding vector to obtain fine-grained attribute text vectors, where all fine-grained attribute text vectors constitute a text attribute semantic space, includes: constraining the interaction of the composite embedding vectors with attribute words through an attribute mask attention mechanism, and then encoding them to obtain fine-grained attribute text vectors.
[0050] This step clarifies the generation method of fine-grained attribute text vectors: in the encoding process of composite embedding vectors (integrating basic word and attribute word vectors), the interaction between words with the same attribute is first constrained by the attribute mask attention mechanism (e.g., color attribute words such as "red" and "blue" interact first, and component attribute words such as "left front light" and "right rearview mirror" interact first), and then fine-grained attribute text vectors are generated through encoding. Finally, all such vectors constitute the text attribute semantic space.
[0051] Attribute masking attention is an attention mechanism that provides targeted constraints on the interactions of attribute words in text. Its core is to construct a mask matrix to control the intensity of attention interactions between words of different attribute categories. A detailed explanation follows: When processing composite embedding vectors, this mechanism first classifies attribute words in the text according to attribute categories (such as color, component, state, etc.), and then constructs a corresponding mask matrix. In the matrix, the position mask value corresponding to words of the same attribute category (such as "red" and "blue" both belonging to the color attribute) is 1, allowing them to fully interact in self-attention calculation and strengthening the semantic association within the same attribute; the position mask value corresponding to words of different attribute categories (such as "red" belonging to the color attribute and "left front light" belonging to the component attribute) is 0, restricting their attention interaction and avoiding mutual interference between the semantic features of different attributes.
[0052] Taking "red car, broken left headlight" as an example, the mask matrix will ensure that "red" (color attribute) only receives effective attention from other color attribute words (if there are no other color words in the text, it mainly strengthens its own semantics), "left headlight" (component attribute) only receives effective attention from other component attribute words, and the attention interaction between "red" and "left headlight" is suppressed by the mask because of their different attribute categories.
[0053] The aforementioned image information structuring method involves inputting an image to be recognized; extracting multi-scale visual feature vectors from the image using a visual encoder; and matching the multi-scale visual feature vectors with the text attribute semantic space to obtain the corresponding fine-grained attribute text. In this invention, the multi-scale visual feature vectors encompass visual information at different levels, and the matching mechanism with the text attribute semantic space establishes a direct association between visual features and fine-grained text descriptions. This solves the problem that traditional single-level features are difficult to match complex text attributes, resulting in a higher and more accurate matching degree between image and text details.
[0054] In addition, the present invention also provides an image information structuring model, which mainly includes five core parts: a visual feature extractor (including a visual encoder module and a multi-layer feature fusion module), a text feature extractor (including a word segmentation and embedding module, an attribute-aware Transformer layer, and a global average pooling layer), and a loss function.
[0055] The visual encoder module is a key component of the model. Its core function is to extract multi-scale features from the input image while effectively preserving both global semantics and local details, thus solving the problem of small object information loss caused by the traditional CLIP model relying solely on final layer features. This module is based on the native ViT-B / 16 architecture of the CLIP model, extending the Transformer architecture from 12 layers to 24 layers. The input image is first segmented into 14×14=196 16×16 pixel blocks. Each block is transformed into a 768-dimensional vector through linear mapping. The vectors corresponding to all blocks are combined into a vector sequence, and a learnable classification-specific identifier, also a 768-dimensional vector, is added at the beginning of the vector sequence, forming an input sequence with dimensions of 197×768. The classification-specific identifier acts as a proxy for global semantics, interacting with the vectors corresponding to all blocks through a self-attention mechanism, while the remaining 196 vectors correspond to the visual features of local image regions, preserving the spatial location information of each block within the image. During the feature extraction process, the visual encoder module sets intermediate feature output interfaces at layers 6, 12, and 24 of the Transformer architecture to form three feature representations with different semantic levels.
[0056] Shallow features (layer 6): This layer outputs low-level visual information such as edges and textures. It has a high response to small local details such as eyeglass frames and backpack straps, providing basic feature representations for subsequent fine detection.
[0057] Mid-level features (12th layer): Through the interaction of multiple self-attention mechanisms, this layer can capture structured information of components such as faces and wheels, and has both local detail features and preliminary semantic aggregation capabilities, realizing the transition from low-level features to mid-level semantics.
[0058] Deep features (layer 24): After processing by the complete Transformer architecture, the output of this layer forms a semantic representation of the overall category of the image (such as "pedestrian" or "vehicle"). The feature dimension of all three layers is maintained at 197×768.
[0059] The multi-layer feature fusion module concatenates the features output from layers 6, 12, and 24 along the channel dimension to form a high-dimensional feature of 197×2304 (768×3). The concatenated 2304-dimensional feature is compressed to 768-dimensional feature through linear transformation to obtain a 197×768 fused feature. The classification-specific identifier (768-dimensional) is extracted from the fused feature and normalized by LayerNorm to serve as the final multi-scale visual feature vector.
[0060] The text feature extractor is an attribute-aware hierarchical text encoder that, for manually annotated fine-grained attribute text (such as "long-haired woman, wearing black-rimmed glasses"), achieves hierarchical encoding from natural language text to structured attribute semantic vectors through attribute type embedding and attention constraint mechanisms. It consists of three core units: a word segmentation and embedding layer, an attribute-aware Transformer layer, and a global average pooling layer, ultimately outputting a 768-dimensional fine-grained attribute text vector.
[0061] Word segmentation and embedding layers convert natural language text into basic semantic vectors suitable for deep neural network processing.
[0062] The word segmentation uses the CLIP native byte-pair encoding word segmenter, which supports sub-word segmentation of mixed Chinese / English text, generating a token sequence containing special markers with a maximum length of 77. The token sequence is then divided into basic words and attribute words, and word vector processing is performed on these two types of words.
[0063] The embedding layer includes word embeddings and attribute embeddings. Word embeddings are used to generate basic word vectors, which are learned embedding matrices. V=49152 is the size of the CLIP pre-trained vocabulary. Each token is mapped to a 768-dimensional vector to form the basic word vector. (L is the length of the token sequence). Mathematical expression: , where token_indices is the token index sequence after word segmentation.
[0064] The method of attribute embedding is the same as that of word embedding. The only difference is that, based on the previous word embedding, the words are labeled with attribute tags as attribute words. Then, the attribute words are processed in the same way. Finally, the 64-dimensional word vector of the attribute words and the 768-dimensional word vector of the base words are concatenated to obtain the word vector of 832 dimensions.
[0065] The attribute-aware Transformer layer achieves fine-grained explicit modeling and hierarchical aggregation of attributes through attribute semantic enhancement and attention constraints, including: Attribute type definition: A predefined set of domain-specific attributes (character scenes include 8 categories: gender, hairstyle, glasses type, clothing color, carried items, age, expression, and posture), and assigns a unique attribute type label to each token (e.g., "long hair" → "hairstyle", "black frame" → "glasses type").
[0066] Composite embedding vector generation: via attribute embedding matrix (A is the number of attribute types, and A=8 for characters and scenes) Generate a 64-dimensional attribute type code, concatenate it with the basic word vector to form an 832-dimensional composite embedding vector, and then transform the composite embedding vector into a 768-dimensional composite embedding vector through linear projection.
[0067] Attribute-masked multi-head attention mechanism: Semantic interaction constraint: Introducing an attribute mask matrix into self-attention computation. Defined as: By using masking operations, the attention weights between tokens of different attribute types (such as "hairstyle" and "glasses type") are forced to approach 0, allowing only tokens of the same attribute type to semantically interact. Masked attention calculation: Where Q, K, and V are the Query, Key, and Value vectors (single-head dimension). This involves using an 8-head attention mechanism to achieve semantic aggregation within attributes. Masked attention computation is a well-known technique and will not be elaborated upon here.
[0068] Hierarchical feature extraction network: By stacking 24 Transformer encoder layers, each Transformer encoder layer contains an attribute mask attention layer and a feedforward neural network.
[0069] The mask attention layer performs mask attention calculations, which have already been introduced in the previous section on attribute mask multi-head attention mechanisms, and will not be repeated here.
[0070] Feedforward Neural Network: This module is a two-layer linear transformation + activation function structure. Its core is a non-linear feature transformation. The first linear transformation maps the input 768-dimensional vector to 3072 dimensions through a 768×3072 weight matrix, enabling the model to capture more complex and finer-grained attribute relationships. The second linear transformation "compresses" the activated 3072-dimensional features back to 768 dimensions, ensuring input and output dimension alignment. This facilitates subsequent residual connections with the original input (if any) and stabilizes the final output at 768 dimensions. Here, a Gaussian Error Linear Unit (GELU) is used as the activation function. GELU is a smooth non-linear activation function with advantages over traditional activation gradients, offering greater stability and smoothness. The formula is expressed as: Output features: After 24 layers of Transformer encoding, a feature sequence containing hierarchical attribute information is generated. .
[0071] Global average pooling layer: Calculates the average value of all local features of the text across all dimensions. After 24 layers of Transformer encoding, the resulting feature sequence containing hierarchical attribute information is represented as a tensor of shape [N, D], where N is the number of local features and D is the feature dimension (which will be converted to 768 dimensions later, so D is adjusted after processing). Global average pooling calculates the average value dimension-wise, resulting in a tensor of shape [1, 768], ultimately yielding a 768-dimensional fine-grained attribute text vector. The formula can be simply expressed as: in It is the vector after global pooling. Dimension value, It is the first The local feature is the first Dimension value, It represents the total number of local features. This approach preserves the overall distribution trend of text features and encodes the average information of the global semantics of the text. In this scenario, it can smoothly aggregate the semantics after the interaction of multiple attributes.
[0072] Loss function: This solution adopts CLIP's native comparison loss function. This achieves global semantic alignment of visual and textual features. The loss function is a variant of the InfoNCE loss function, which constructs a discriminative feature space for cross-modal retrieval by maximizing the cosine similarity of positive sample pairs while minimizing the similarity with negative sample pairs. The specific calculation formula is as follows: ,in This represents the contrastive loss, used to quantify the difference between the model's predicted results and the actual results in a contrastive learning task. The smaller the value, the better the model performs in distinguishing between positive and negative sample pairs. The logarithmic function (usually the natural logarithm with base e) serves to compress the numerical range, transform multiplication (after exponential operation) into addition (logarithmic operation), facilitate gradient calculation for backpropagation optimization, and convert the probabilistic ratio into a loss value that can be directly used for optimization. This represents an exponential function used to calculate powers with the natural constant e as the base, and to convert similarity values ( or Convert the values to non-negative values to construct a relative relationship similar to probability. This typically represents the similarity score between positive sample pairs, that is, the similarity between sample pairs that the model considers to belong to the same semantics and whose feature distance needs to be reduced (such as fine-grained attribute text and its corresponding matching content). It is a temperature coefficient used to adjust the "sharpness" of similarity, controlling the ease with which positive and negative sample pairs can be distinguished. The smaller the value, the more concentrated the similarity distribution, and the more "strict" the model is in distinguishing between positive and negative sample pairs. It can typically be equal to 0.07. Through this loss function, the model can learn the shared semantic space across modal data, making the feature vectors of the same semantic concept highly similar in visual and textual modalities, thus laying the foundation for fine-grained attribute alignment.
[0073] After the model is built, dataset collection and model training are required. The dataset consists of two parts: first, data selected from classic open-source libraries, including COCO and VOC datasets for object detection, and ImageNet and Caltech101 datasets for image recognition, retaining approximately 200,000 images after deduplication and quality screening; second, a proprietary dataset of 500,000 images collected from multiple devices, covering targets such as bicycles, pedestrians, and vehicles, and various indoor and outdoor scenes. A large model is used to predict attribute labels on all data. After manual review and correction, blurred and conflicting samples are removed to complete data cleaning. The processed data is then divided into training, validation, and test sets in a 6:3:1 ratio. The model is built based on the PyTorch framework, and training is performed after parameter initialization and hyperparameter tuning.
[0074] In one embodiment, an image information structuring apparatus is provided. For example... Figure 2 As shown, it includes an input module 21, a processing module 22, and an output module 23. Detailed descriptions of each functional module are as follows: Input module 21 is used to input the image to be recognized; Processing module 22 is used to extract multi-scale visual feature vectors from the image using a visual encoder; Output module 23 is used to match the multi-scale visual feature vector with the text attribute semantic space to obtain the corresponding fine-grained attribute text.
[0075] This invention also provides an electronic device, including a memory, a processor, and a program stored in the memory and executable on the processor. The processor, when executing the program, implements the aforementioned image information structuring method; to avoid repetition, this will not be described further. Alternatively, the electronic device can implement the functions of each module in this embodiment of the image information structuring device; this will not be described further.
[0076] This invention also provides a readable storage medium storing a program, characterized in that, when executed by a processor, the program implements the aforementioned image information structuring method; to avoid repetition, this will not be described again here. Alternatively, when executed by a processor, the program implements the functions of each module in this embodiment of the image information structuring device; this will not be described again here.
[0077] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
[0078] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0079] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
Claims
1. A method for structuring image information, characterized in that, include: Input the image to be recognized; For the image, a multi-scale visual feature vector is extracted using a visual encoder; The multi-scale visual feature vectors are matched with the text attribute semantic space to obtain the corresponding fine-grained attribute text.
2. The image information structuring method according to claim 1, characterized in that, The step of extracting multi-scale visual feature vectors from the image using a visual encoder specifically includes: Based on the native CLIP model, by expanding the number of layers in its Transformer architecture, corresponding visual features are output from different Transformer architecture layers, and then the multi-layer visual features are fused to obtain a multi-scale visual feature vector.
3. The image information structuring method according to claim 2, characterized in that, The native CLIP model, by extending the number of layers in its Transformer architecture, outputs corresponding visual features from different Transformer architecture layers, including: The number of layers in the Transformer architecture of the original CLIP model is expanded to 24, and feature output interfaces are set in layers 6, 12 and 24 to output shallow visual features, mid-level visual features and deep visual features, respectively.
4. The image information structuring method according to claim 1, characterized in that, The native CLIP model, by extending the number of layers in its Transformer architecture, outputs corresponding visual features from different Transformer architecture layers, including: The image is divided into multiple small blocks and each block is converted into a corresponding vector. The vectors corresponding to multiple small blocks are arranged according to the position order of each small block in the image to obtain a vector sequence; A learnable classification-specific identifier is added to the beginning position of the vector sequence to obtain the input sequence; The input sequence is processed by a Transformer architecture with an extended number of layers to obtain multi-scale visual feature vectors.
5. The image information structuring method according to claim 4, characterized in that, The step of matching the multi-scale visual feature vector with the text attribute semantic space to obtain the corresponding fine-grained attribute text includes: By using a pre-trained cross-modal cosine similarity calculation mechanism, the multi-scale visual feature vectors are compared with the predefined fine-grained attribute text vectors in the text attribute semantic space, and the fine-grained attribute text corresponding to the fine-grained attribute text vector with the highest similarity is selected.
6. The image information structuring method according to claim 1, characterized in that, The text attribute semantic space includes: Based on the native CLIP model, the text encoder is extended to an attribute-aware text encoder to segment predefined fine-grained attribute text into basic words and attribute words. The basic words are processed through a basic word embedding matrix to generate basic word vectors, and the attribute words are processed through an attribute word embedding matrix to generate attribute word vectors. The basic word vectors and attribute word vectors are then fused to obtain a composite embedding vector. The composite embedding vector is encoded to obtain fine-grained attribute text vectors, and all fine-grained attribute text vectors constitute the text attribute semantic space.
7. The image information structuring method according to claim 6, characterized in that, The process of encoding the composite embedding vector to obtain fine-grained attribute text vectors, whereby all fine-grained attribute text vectors constitute a text attribute semantic space, including: The composite embedding vector is constrained to interact with the attribute words through an attribute mask attention mechanism, and then encoded to obtain a fine-grained attribute text vector.
8. An image information structuring device, characterized in that, include: The input module is used to input the image to be recognized; The processing module is used to extract multi-scale visual feature vectors from the image using a visual encoder. The output module is used to match the multi-scale visual feature vector with the text attribute semantic space to obtain the corresponding fine-grained attribute text.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the image information structuring method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the image information structuring method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Zero sample image anomaly detection method and device
CN119130931A
Combined zero sample image classification method based on hierarchical feature fusion
CN119478551A