Intelligent Image Recognition System and Method Based on Deep Learning

Through an intelligent image recognition system based on deep learning, pre-trained feature extraction network and extended network map image and text features to high-dimensional space, and optimize feature similarity through alignment layers and contrast learning strategies, solving the problem of image and text information fusion and achieving efficient image recognition.

CN119963950BActive Publication Date: 2025-06-20TAIYUAN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510451037.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-06-20
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

The prior art is difficult to effectively integrate the information of images and text, resulting in separation of processing between images and text in image recognition tasks, and it is difficult to directly compare or work together in a unified space.

Method used

Using a deep learning-based intelligent image recognition system, high-level semantic features are extracted through pre-trained image and text feature extraction networks, extended networks are used to map features to high-dimensional spaces, and similarity between images and text features is optimized through alignment layers and contrast learning strategies, ultimately achieving efficient matching of images and text in a unified embedded space.

Benefits of technology

It realizes the close alignment of images and text in the unified embedding space, enhances the deep fusion of multimodal information, and improves the accuracy of the image recognition system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963950B_ABST
    Figure CN119963950B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent image recognition system and method based on deep learning, including: extracting high-level semantic features from images and texts respectively through pre-trained image and text feature extraction networks; using an extended network to map image and text features from a low-dimensional space to a higher dimension; introducing local and global contrast learning strategies to optimize the similarity of image and text features through dynamic resampling and memory enhancement mechanisms; generating auxiliary text information from image features through a generative adversarial network or a variational autoencoder and performing contrast learning with real text labels; in a high-dimensional unified embedding space, through fine regression optimization and semantic matching scoring mechanisms, enabling images and texts to be efficiently matched and complete the image recognition task during the inference stage. The present invention can enable the semantic similarity between images and texts to be compared through the same distance metric, thereby achieving efficient image-text matching during the inference stage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and particularly to an intelligent image recognition system and method based on deep learning. Background Art

[0002] In image recognition tasks, images and texts are two main modal information, which respectively express the understanding of objects, scenes, and events from different perspectives. Images present visual information through the combination of pixels and colors, while texts capture the related context and meaning through language and semantics. Although images and texts differ in the way of transmitting information, they are essentially interrelated. For example, the objects shown in an image can be further defined and explained through text descriptions, and the descriptions in the text can also provide a clearer context for understanding the image. However, how to effectively fuse the information of these two modalities remains a major challenge in the field of image recognition. In current technical methods, the processing of images and texts is often separated, and image features and text features are respectively extracted and processed through independent models, making it difficult for them to be directly compared or work together in a unified space. Summary of the Invention

[0003] To solve the above problems, the present invention provides an intelligent image recognition system and method based on deep learning.

[0004] To achieve the above object, the technical solutions adopted by the present invention are as follows:

[0005] On the one hand, the present invention discloses an intelligent image recognition method based on deep learning, including:

[0006] Step 1: Extract high-level semantic features from images and texts respectively through pre-trained image and text feature extraction networks;

[0007] Step 2: Use an extension network to map image and text features from a low-dimensional space to a higher dimension, enhance semantic information, and map the two to a unified high-dimensional semantic space through an alignment layer;

[0008] Step 3: Introduce local and global contrast learning strategies, optimize the similarity of image and text features through dynamic resampling and memory enhancement mechanisms, and strengthen the semantic association between multi-modalities;

[0009] Step 4: Generate auxiliary text information from image features through a generative adversarial network or variational autoencoder, and perform contrast learning with real text labels to further refine the semantic mapping between images and texts;

[0010] Step 5: In the high-dimensional unified embedding space, through a refined regression optimization and semantic matching scoring mechanism, the image and text can be efficiently matched during the inference stage to complete the image recognition task.

[0011] Further: The said step 1 includes:

[0012] For image feature extraction, the Vision Transformer method is adopted. The image is segmented into blocks of a fixed size, and the relationships between the blocks are modeled through the self-attention mechanism to capture global context information. Finally, a high-dimensional embedding representation of the image is obtained from the penultimate layer of the ViT model. For the extraction of corresponding text features, the BERT model is selected. First, the input text is tokenized, and then it is fed into the BERT encoder. Using its bidirectional Transformer architecture, the semantic representation of each word is obtained, and the global representation vector of the text is aggregated through the [CLS] token.

[0013] Further: The said step 2 includes:

[0014] First, perform feature dimension expansion. For image feature expansion, a multi-layer perceptron network is used. Through a linear transformation layer and a non-linear activation layer, the initial embedding of the image is mapped to a higher dimension; text feature expansion also adopts a similar MLP network to map the text embedding to the same high-dimensional space as the image features. Then, perform spatial alignment. Through the bidirectional cross-attention mechanism and similarity metric learning of the alignment layer, the semantic relationship between the image and text features in the high-dimensional space is strengthened to achieve precise alignment.

[0015] Further: The said step 3 includes:

[0016] By introducing local and global contrast learning strategies, the similarity of image and text features is optimized, and the semantic association between multi-modalities is strengthened. Local contrast learning focuses on object-level semantic contrast. A local contrast loss function is designed to maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs, guiding the network to learn effective local features; global contrast learning focuses on the overall scene or context contrast. A global contrast loss function is designed to optimize the model's understanding of the overall scene; a dynamic resampling mechanism is also introduced. According to the current learning situation of the model, more difficult samples are dynamically selected to increase the probability of difficult samples appearing and avoid overfitting to simple categories. At the same time, with the help of memory-augmented contrast learning, the memory network is used to store historical contrast samples, and the replay mechanism helps the model remember the semantic differences between categories.

[0017] Further: The said step 4 includes:

[0018] Select a generative adversarial network as the generative model. Input the image features into the generator, combine with random noise to generate text descriptions, and then use the discriminator to judge the similarity between the generated text and the real text. Design the discriminator loss function to optimize the generated text. Compare the generated text description with the original real text for contrast learning. Adopt a new contrast learning strategy to maximize the similarity of similar text pairs and minimize the distance of dissimilar text pairs, further refining the semantic mapping. Introduce a diversity penalty mechanism to encourage the generator to explore text generation paths. At the same time, conduct joint training of the generated text and image features to ensure the consistency of the generated text and image features in the embedding space, and obtain a more accurate image-text semantic mapping.

[0019] Further: Step 5 includes:

[0020] Through the regression optimization mechanism, adopt a bidirectional distance loss function to minimize the distance of matching samples and maximize the distance of non-matching samples, realizing the optimization of the image and text embedding spaces. Introduce a balance factor to adjust the optimization intensity, maintaining the semantic consistency and appropriate diversity of the image and text. In the inference stage, for new images and text descriptions, first extract high-level features through the feature extraction network, then map them to the high-dimensional embedding space for alignment through the extension network, and measure the similarity between the image and text through the semantic matching score to judge the matching degree. For the object detection task, combine the image region features and text features, use the region convolutional neural network for object localization, and combine the matching score to confirm the object category, ultimately achieving high-precision image recognition.

[0021] On the other hand, the present invention discloses an intelligent image recognition system based on deep learning, including: An image and text preliminary feature extraction module: Extract high-level semantic features from images and texts respectively through pre-trained image and text feature extraction networks, laying a foundation for subsequent mapping and contrast learning;

[0022] A progressive feature extension module: Use the extension network to map the image and text features from the low-dimensional space to a higher dimension, enhance the semantic information, and map the two to a unified high-dimensional semantic space through the alignment layer;

[0023] A dynamic contrast learning module: Introduce local and global contrast learning strategies, optimize the similarity of image and text features through dynamic resampling and memory enhancement mechanisms, and strengthen the semantic association between multi-modalities;

[0024] A self-supervised text generation module: Generate auxiliary text information from image features through a generative adversarial network or variational autoencoder, and conduct contrast learning with real text labels to further refine the semantic mapping of images and texts;

[0025] Optimization and Inference Module for Unified Embedding Space: In the high-dimensional unified embedding space, through fine-tuned regression optimization and semantic matching scoring mechanism, images and texts can be efficiently matched during the inference stage to complete the image recognition task.

[0026] Compared with the prior art, the technical progress achieved by the present invention lies in:

[0027] The present invention can not only extract meaningful high-dimensional features from images and texts respectively from their respective modal spaces, but also align these two types of features with each other in a common embedding space, thereby providing a more compact and accurate semantic mapping for subsequent inference and task execution. By mapping images and texts to a unified embedding space, the present invention can achieve deep fusion of multimodal information. This fusion can not only enhance the understanding of the relationship between images and texts, but also effectively improve the accuracy of the image recognition system. The design of the unified embedding space enables the semantic similarity between images and texts to be compared through the same distance metric, thus achieving efficient image-text matching during the inference stage. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention.

[0029] In the drawings:

[0030] Figure 1 is a flowchart of the present invention;

[0031] Figure 2 is a system structure diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will be described below with reference to the drawings.

[0033] Embodiment 1

[0034] As Figure 1 shown, this embodiment discloses an intelligent image recognition method based on deep learning, including:

[0035] Step 1: Extract high-level semantic features from images and texts respectively through pre-trained image and text feature extraction networks;

[0036] Step 2: Use an extended network to map image and text features from a low-dimensional space to a higher dimension to enhance semantic information, and map the two to a unified high-dimensional semantic space through an alignment layer;

[0037] Step 3: Introduce local and global contrastive learning strategies, optimize the similarity between image and text features through dynamic resampling and memory enhancement mechanisms, and strengthen the semantic association between multimodalities;

[0038] Step 4: Generate auxiliary text information from image features through a generative adversarial network or variational autoencoder, and perform contrastive learning with the true text labels to further refine the semantic mapping between images and texts;

[0039] Step 5: In the high-dimensional unified embedding space, through fine-grained regression optimization and semantic matching scoring mechanisms, enable images and texts to be efficiently matched during the inference stage and complete the image recognition task.

[0040] The purpose of Step 1 is to extract preliminary features from the original images and texts, laying a foundation for the subsequent mapping and contrastive learning stages.

[0041] First, perform image feature extraction.

[0042] To extract features with rich semantic information from images, this embodiment adopts a method based on VisionTransformer ViT. In this step, the goal of this embodiment is not only to extract traditional low-level visual features, but also to enable the model to capture more detailed semantic information in images through global feature integration and local information encoding.

[0043] 1. Select the base model:

[0044] Use ViT as the base network for feature extraction. Different from traditional convolutional neural networks, ViT divides an image into fixed-size patches and models the relationships between patches through self-attention mechanisms to capture global context information. Therefore, ViT can preserve long-range dependencies in images in high-dimensional spaces and provide a feature representation that fuses global and local information.

[0045] 2. Image patching and positional encoding:

[0046] First, divide the input image (with size ) into several patches of a fixed size, where the size of each patch is , and is the scale parameter of patching. Each image patch is flattened and mapped to a vector of a fixed dimension , where is the feature dimension of each patch. To preserve the spatial position information of patches in the image, this embodiment introduces positional encoding. By adding the positional encoding of each patch ( ) to its corresponding image patch vector, an image patch representation containing spatial position information is obtained:

[0047]

[0048] 3. Transformer Encoder:

[0049] The encodings of all image patches are input into the Transformer Encoder, and the self-attention mechanism and the feed-forward neural network are used for global interaction of information. In this process, the model encodes the local information of the image through the self-attention mechanism and captures the cross-patch correlations through the multi-head attention layer. The output sequence of the Transformer contains rich context information, and the finally generated feature vector is used to represent the high-dimensional feature representation of the image.

[0050] 4. Output layer:

[0051] In this embodiment, the final image feature representation is obtained from the penultimate layer of the ViT model (i.e., the embedding vector output by the last multi-head attention layer) , which will be the high-dimensional embedding representation of the image and is used for subsequent contrast learning and modality alignment tasks.

[0052] Then text feature extraction is performed.

[0053] In terms of text feature extraction, this embodiment selects BERT (Bidirectional Encoder Representations from Transformers) as the base model. BERT models language features through its bidirectional context learning ability and can well capture the semantic relationships and nuances in the text.

[0054] 1. Text preprocessing and encoding:

[0055] For the input text description (e.g., "A dog is running on the grass"), first, it is processed through Tokenization (word segmentation) to break the sentence into a series of words or lexical units. Then, these token sequences are passed into the BERT model for encoding, and BERT will obtain the semantic representation of each word from the context using its bidirectional Transformer architecture, generating the vector representation of each vocabulary (wherein, is the dimension of the text feature).

[0056] 2. Word vector synthesis and global representation:

[0057] To obtain the global feature representation of the entire sentence, the vectors of each word output by BERT are aggregated through the [CLS] token (a special classification token). This [CLS] token is designed to represent the semantic information of the entire sentence. Therefore, by extracting the output of the [CLS] token, this embodiment can obtain the global representation vector of the text , which is the embedding representation of the text.

[0058] 3. Feature fusion:

[0059] After extracting the image and text features, the next task is to align the embedding representations of the image and text and map them to a unified feature space. This process is usually carried out in subsequent contrast learning and mapping steps. However, in this step, this embodiment has already obtained the preliminary features of the two modalities: image features and text features , which respectively retain the basic semantic information of the image and text, laying a foundation for subsequent collaborative learning and alignment.

[0060] In step 1, this embodiment extracts the high-level features of the image and text by selecting powerful pre-trained models such as ViT and BERT. In image processing, ViT uses block-based decomposition and self-attention mechanisms, enabling the image to be effectively encoded as a high-dimensional feature vector , while the text features generate the global representation of the text through the BERT model , which provides a reliable foundation for subsequent contrast learning and cross-modal feature alignment. The results of these preliminary feature extractions will support the subsequent feature mapping and contrast learning stages and ultimately promote the deep fusion and understanding of image and text semantics.

[0061] The purpose of step 2 is to ensure that these two modalities can be mapped to a unified high-dimensional embedding space by gradually expanding the dimensions of the image and text features and capture more complex semantic information between them. The core idea of expanding the features is to enable the image and text features to effectively fuse and align with each other through non-linear transformation and multi-layer interaction, thereby providing a richer representation for subsequent cross-modal learning.

[0062] First, perform feature dimension expansion.

[0063] 1. Image feature expansion:

[0064] The key to the image feature expansion process lies in how to expand the original image features from space to a higher dimension , enabling the image features to have stronger expressive power and establish effective semantic connections with the text features in a unified space. Here, in this embodiment, an extended network is adopted, which consists of a series of non-linear transformation layers.

[0065] 1.1 Network Structure Design:

[0066] In this embodiment, a multi-layer perceptron (MLP) network is used as the extended module. This MLP network consists of two main parts:

[0067] The first part is a linear transformation layer that maps the initial embedding of the image to an intermediate space , where is the dimension of the intermediate space, usually taken as .

[0068] The second part is a non-linear activation layer that introduces non-linear transformations through a series of ReLU activation functions to further improve the expressive power of the feature space and finally outputs a high-dimensional feature representation , and this dimension can be designed as the maximum dimension when the image and text features are aligned.

[0069] The mathematical expression is as follows:

[0070]

[0071]

[0072] where and are weight matrices, and are bias terms, is the activation function.

[0073] 2. Text Feature Extension:

[0074] During the text feature extension process, this embodiment follows the same idea as the image feature extension and also uses an extended network to map the embedding of the text from space to space, enabling the image and text features to be mapped to the same embedding space, thus ensuring that the semantics between the two can be better aligned and compared.

[0075] 2.1 Network Structure Design:

[0076] Similar to the extension of image features, the text feature extension is also implemented through a multi-layer perceptron (MLP) network. First, the text feature will be mapped to the intermediate space through a linear transformation layer In it, usually take . Then, further improve the expression ability of features through a non-linear activation function, and finally output a high-dimensional text embedding .

[0077] The mathematical expression is as follows:

[0078]

[0079]

[0080] Through this process, this embodiment maps the preliminary feature embedding of the text to the same high-dimensional space as the image features , ensuring the effective alignment of image and text features in a unified semantic space.

[0081] Then perform spatial alignment.

[0082] After completing the expansion of image and text features, this embodiment obtains high-dimensional feature representations of the two modalities and , and next this embodiment will achieve the precise alignment of the two through the alignment layer to ensure their semantic consistency in the high-dimensional space.

[0083] 1. Alignment layer design:

[0084] The core goal of the alignment layer is to strengthen the semantic relationship between images and texts in the high-dimensional feature spaces of images and texts through a cross-fusion mechanism. Specifically, the alignment layer will ensure the correct alignment of image and text features in the high-dimensional space through a bidirectional cross-attention mechanism and similarity metric learning.

[0085] 1.1 Cross-attention mechanism:

[0086] This embodiment adopts a cross-self-attention mechanism, enabling the image features to be adjusted under the guidance of the text features, and at the same time the text features to be optimized under the influence of the image features. The calculation of the cross-attention mechanism can be completed through the following steps:

[0087]

[0088]

[0089]

[0090] Among them, and are the query and key matrices of images and texts respectively, and It is a value matrix. Through this cross-fusion process, this embodiment can mutually reinforce the high-dimensional features of images and texts, thereby making the two more closely related semantically.

[0091] 2. Similarity measurement:

[0092] After cross-fusion, this embodiment further introduces similarity measurement learning to optimize the alignment of images and texts in the unified space by calculating the similarity between image and text features. Commonly used similarity measurement methods include cosine similarity or Euclidean distance, and the specific forms are as follows:

[0093]

[0094] By optimizing this similarity measurement, this embodiment can further refine the relationship between images and texts in the unified embedding space, thereby enhancing their semantic consistency.

[0095] In step 2, this embodiment adopts the expansion networks of images and texts through a progressive feature expansion mechanism to expand the features of the two modalities from and spaces to a unified high-dimensional space , and further aligns their semantic relationships through cross-fusion and similarity measurement. Through this process, this embodiment not only effectively expands the expression ability of image and text features, but also ensures their semantic alignment in the high-dimensional space, laying a solid foundation for subsequent cross-modal learning and reasoning tasks.

[0096] The purpose of step 3 is to strengthen the semantic correlation between images and texts in the high-dimensional embedding space by introducing a dynamic contrast learning mechanism, optimize the distance relationship between the two, and ensure that the model can maintain the learning effect when facing complex and difficult semantic contrasts through dynamic resampling and memory-enhanced contrast learning strategies, avoiding overfitting on simple samples.

[0097] First, perform local and global contrast learning.

[0098] 1. Local contrast learning:

[0099] In the contrast learning of images and texts, local contrast learning focuses on semantic contrast at the object level, that is, contrasting the semantic similarity at the local level (such as a single object or the details of an object). For example, in an image, perform object-level contrast such as "cat" and "dog"; in text, perform semantic contrast through corresponding descriptions (such as the text descriptions of "cat" and "dog").

[0100] Here, in this embodiment, a local contrastive loss function is designed. Its core idea is to guide the network to learn effective local features by maximizing the similarity of positive sample pairs and minimizing the similarity of negative sample pairs.

[0101] The mathematical expression is:

[0102]

[0103] Where and are the high-dimensional feature representations of the -th image and text samples respectively. is the temperature parameter that controls the scale of the similarity between samples. Through this loss function, the model can gradually learn effective contrastive features at the object level.

[0104] 2. Global contrastive learning:

[0105] In global contrastive learning, the focus shifts to more macroscopic semantic comparisons, involving the comparison of the entire scene or overall context. For example, the comparison of the entire scene "dining table" and "sofa" in an image, or the scene comparison in text descriptions (such as text describing a dining table and a sofa). The goal of global contrastive learning is to optimize the model's understanding of the overall scene and capture the semantic relationships between images and texts on a larger scale.

[0106] For global contrast, a global contrastive loss function (Global Contrastive Loss) is designed. This loss function aims to maximize the distance from positive samples (similar scenes) and minimize the distance from negative samples (dissimilar scenes) at the same time.

[0107] The mathematical expression is:

[0108]

[0109] By comparing the global semantics of images and texts, the model can not only capture local information in detail but also understand the overall semantic structure of images and texts, thereby enhancing the generalization ability for diverse scenes.

[0110] Then dynamic resampling is carried out.

[0111] To avoid overfitting to certain simple categories in contrastive learning, this embodiment introduces a dynamic resampling mechanism, aiming to automatically increase the proportion of difficult samples by gradually adjusting the sampled contrastive samples in each training batch and introduce challenges during the training process.

[0112] 1. Design of the resampling mechanism:

[0113] Within each training cycle, the system dynamically selects samples with greater sampling difficulty based on the learning situation of the current comparison samples. For example, in the initial stage, the model may be more likely to distinguish common objects (such as "cat" and "dog"), while in the later stage, more complex or ambiguous semantic pairs (such as "cat" and "rabbit") will be sampled for learning.

[0114] The key to the resampling strategy is to use the output of the current model to predict which samples are more difficult to distinguish and increase the probability of these difficult samples appearing in the training. In this embodiment, a Hard Sample Mining module is used to dynamically adjust the selection of training samples:

[0115]

[0116] where is the current loss of sample i, is the temperature parameter that controls the weight of difficult samples. In this way, the probability of difficult samples will gradually increase as the training progresses, thereby helping the model learn more complex semantic relationships.

[0117] Finally, memory-augmented contrastive learning is performed.

[0118] To further enhance the stability of contrastive learning, this embodiment introduces a memory network, which can store historical comparison samples during the training process and helps the current model better remember and distinguish the semantic differences between different categories through a replay mechanism.

[0119] 1. Memory network design:

[0120] The memory network here serves as an auxiliary module to store and replay previous comparison samples. In this embodiment, the effective comparison pairs during the training process are stored in the memory bank through a soft storage mechanism, and then these samples are input into the model together with the current training samples through a replay mechanism, thereby enhancing the model's understanding of the similarities and differences between different categories.

[0121] Mathematically, the output of the memory network can be expressed as:

[0122]

[0123] where represents the memory bank that stores historical effective comparison samples; the " " operation retrieves the most relevant memories from the memory bank and helps optimize the current comparison samples through these memories. The introduction of the memory replay mechanism can prevent the model from relying solely on the current batch of samples, ensure the stability of training, and enhance the generalization ability.

[0124] In step 3, in this embodiment, through a dynamic contrast learning mechanism, local contrast and global contrast learning are combined, not only paying attention to the semantic differences of individual objects, but also extending to the contrast of the entire scene. Through dynamic resampling, the learning challenges of more complex samples are introduced to prevent the model from falling into overfitting to a single category. Finally, with the help of memory-augmented contrast learning and using historical valid samples and replay mechanisms, the stability and generalization ability of the model during training are further improved. Through these designs, the model can not only perform effective contrast learning in high-dimensional space, but also continuously adapt to and optimize the semantic relationships between different categories, laying a solid foundation for subsequent cross-modal learning.

[0125] The purpose of step 4 is to generate auxiliary text information through self-supervised learning to enhance the semantic consistency between image and text features. Specifically, through a generative model (such as a generative adversarial network GAN or a variational autoencoder VAE), corresponding text descriptions are generated from image features as a self-supervised training signal to help the semantic alignment of image and text features. This process not only improves the ability of image understanding, but also further strengthens the accuracy of text descriptions for image semantics, ultimately making the joint embedding of images and texts more accurate.

[0126] First, self-supervised text generation is performed.

[0127] 1. Selection and design of the generative model:

[0128] In this embodiment, the generative adversarial network (GAN) is selected as the main generative model, and combined with the structure of the conditional generative adversarial network, text descriptions are generated from image features. During this process, the image features in this embodiment are input into the generator network through the high-dimensional space representation extended in the previous step to generate corresponding text descriptions. The generator combines the image features with random noise, performs feature mapping through multiple convolutional layers and fully connected layers, and finally generates text descriptions with semantic consistency.

[0129] Mathematically expressed as:

[0130]

[0131] where represents the generator, is the extended image feature, is the random noise, is the generated text embedding representation.

[0132] 2. Discriminator design:

[0133] Similar to the traditional GAN structure, in this embodiment, a discriminator is introduced to judge the similarity between the generated text and the real text. The goal of the discriminator is to maximize the difference between the real and generated texts. The discriminator encodes the input text through a convolutional neural network (CNN) and compares it with the features of the image.

[0134] Loss function of the discriminator It is designed to maximize the matching degree between the real text and the image features, while minimizing the matching degree between the generated text and the image features, and is specifically expressed as:

[0135]

[0136] Among them, represents the features of the real text, is the discriminator, and the goal is to make the generated text as distinguishable from the real text as possible.

[0137] 3. Contrastive learning of generated text:

[0138] The generated text description is subjected to contrastive learning with the original real text description, so as to further refine the semantic mapping between the image and the text. In this embodiment, a new contrastive learning strategy is adopted, and the generated text is embedded is compared with the real text embedding to maximize the similarity of similar text pairs and minimize the distance of dissimilar text pairs.

[0139] The loss function of contrastive learning can be expressed as:

[0140]

[0141] Among them, and respectively represent the embedding representations of the th generated text and real text, is the temperature parameter, which controls the scale of contrast.

[0142] Through this process, the generated text not only serves as an auxiliary supervision signal to help image understanding, but also further enhances the semantic consistency between the image and the text through contrastive learning, enabling the two to be more precisely aligned in the high-dimensional embedding space.

[0143] Then, the optimization of self-supervised text generation is carried out.

[0144] 1. Enhancement of text generation diversity:

[0145] To further enhance the diversity and accuracy of the generated text, this embodiment introduces a diversity penalty mechanism, which can prevent the generated text from being too monotonous. By optimizing through comparing multiple potential text descriptions during the text generation process, this embodiment can encourage the generator to explore more text generation paths, thereby enhancing the diversity and semantic richness of the text.

[0146] The goal of diversity penalty is to increase the diversity of the generated text in the following ways:

[0147]

[0148] Among them, represents the cosine similarity. The goal is to minimize the similarity between the generated texts, so as to make the generated texts more diverse.

[0149] 2. Joint training of generated text and image features:

[0150] While generating text, this embodiment jointly optimizes the training of the generator and the discriminator to ensure that the generated text can be consistent with the image features in the embedding space. Under this joint training strategy, the generator not only needs to generate semantically reasonable text descriptions, but also ensure that these text descriptions can be closely aligned with the image features.

[0151] The joint optimization of the loss function can be expressed as:

[0152]

[0153] Among them, is the balancing term, which controls the intensity of the diversity penalty.

[0154] In step 4, this embodiment uses a self-supervised text generation mechanism, using a generative adversarial network or a conditional generative adversarial network, to generate text descriptions from image features as self-supervised signals to help strengthen the semantic consistency between images and texts. Through the contrastive learning of the generated text, the generated text is compared with the real text to further refine the semantic alignment between the two in the high-dimensional embedding space. At the same time, this embodiment introduces a diversity penalty mechanism and a joint training strategy to improve the diversity and quality of the generated text, and ensure the consistency between the generated text and the image features, so as to obtain a more accurate image-text semantic mapping in the whole cross-modal learning.

[0155] The purpose of Step 5 is to enable efficient and accurate matching between images and texts during the inference stage through the optimization of the unified embedding space, thereby effectively completing image recognition tasks (such as image classification, object detection, etc.). Specifically, the optimization process of the unified embedding space will strive to reduce the distance between similar images and texts, while maximizing the distance between heterogeneous features (images and texts), and through an accurate matching scoring mechanism, perform efficient image-text comparison during the inference stage, ultimately achieving high-precision completion of the inference task.

[0156] First, perform the optimization of the unified embedding space.

[0157] 1. Optimization objective design:

[0158] After the image and text features are co-embedded into the same high-dimensional space, the next step is to ensure that the matching between images and texts is more accurate in this space. This embodiment designs a refined regression optimization mechanism to minimize the distance between similar image and text pairs in the embedding space and maximize the distance between heterogeneous feature pairs. This optimization objective is achieved by designing a bidirectional distance loss function, which takes into account both the alignment of image-text pairs and the reverse text-image pairs.

[0159] The core form of the loss function is as follows:

[0160]

[0161] Where, and respectively represent the high-dimensional embedding representations of the image and text of the th sample, is the weight coefficient that controls the negative distance term between the image and text, represents a negative sample (non-matching sample). This loss function optimizes the image and text embedding space by minimizing the distance of matching samples and maximizing the distance between non-matching samples.

[0162] 2. Balance in the optimization process:

[0163] During the actual training process, the embedding space of images and texts is dynamically changing. Therefore, this embodiment introduces a balance factor to adjust the optimization intensity between the two. This embodiment needs to ensure that the representations of images and texts are highly consistent semantically, but at the same time avoid them becoming overly dependent on a certain modal feature during the optimization process. Therefore, the distance calculation between images and texts adopts a weighted method:

[0164]

[0165] Where, It is a loss function used to enhance the diversity of image and text features. It is the weight hyperparameter that balances the loss terms between the image and the text. This balancing strategy enables the image and the text to maintain semantic consistency during the optimization process while keeping an appropriate level of diversity, further improving the matching accuracy in the inference stage.

[0166] Then, the matching in the inference stage is carried out.

[0167] 1. Image Feature Extraction and Mapping:

[0168] In the inference stage, given a new image and text description, first, high-level features are extracted from the image through a feature extraction network (such as using an optimized ResNet or ViT). Then, they are mapped to a high-dimensional embedding space through an expansion network to align them with the high-dimensional embedding space of the text.

[0169] Mathematically, the high-dimensional mapping of the image feature can be expressed as:

[0170]

[0171] where is the image expansion network.

[0172] 2. Text Feature Extraction and Enhancement:

[0173] Meanwhile, the given text description is input into a pre-trained language model (such as BERT or GPT) to generate preliminary text features . Then, through a generation network and a contrastive learning mechanism, the semantic information of this text feature is enhanced to map it to the same high-dimensional space as the image feature.

[0174] The high-dimensional mapping of the text feature is expressed as:

[0175]

[0176] where is the text expansion network, is the original text description.

[0177] 3. Semantic Matching Scoring:

[0178] After the image and the text are mapped in the high-dimensional space, next, the similarity between the two needs to be measured through semantic matching scoring. The semantic matching score S_{match} can be achieved by calculating the cosine similarity between the image feature and the text feature:

[0179]

[0180] Among them, represents the dot product of vectors, represents the norm of a vector.

[0181] Based on this matching score, the system can judge the matching degree between the image and the text, and complete subsequent tasks (such as classification or object detection) according to the threshold.

[0182] Then perform object detection in the inference stage.

[0183] In the object detection task, this embodiment not only requires the matching score of the image and the text, but also further enhances the matching accuracy through the object position in the image. Specifically, by combining the regional features in the image with the text features, a region convolutional neural network (R-CNN) is used to perform object localization on the image, and the object category is confirmed by combining the matching score.

[0184] Finally, the loss function of object detection combines the position deviation and the matching degree, and the score S_{det} of object detection is calculated by combining the position and semantic similarity.

[0185] In step 5, after the features of the image and the text are optimized in the unified embedding space, semantic alignment between the image and text features is achieved through a fine-grained regression optimization mechanism, ensuring that they have the minimum matching distance in the high-dimensional space while maximizing the distance between heterogeneous feature pairs. In the inference stage, the features of the image and the text are mapped to the unified embedding space through an extended network and a contrast learning mechanism, and then the matching degree between the two is judged through the semantic matching score to complete tasks such as classification or object detection in the image recognition task.

[0186] Embodiment 2

[0187] As Figure 2 described, this embodiment discloses an intelligent image recognition system based on deep learning, including: an image and text preliminary feature extraction module: extracting high-level semantic features from the image and text respectively through pre-trained image and text feature extraction networks, laying a foundation for subsequent mapping and contrast learning;

[0188] A progressive feature expansion module: using an extended network to map the image and text features from a low-dimensional space to a higher dimension, enhancing semantic information, and mapping the two to a unified high-dimensional semantic space through an alignment layer;

[0189] A dynamic contrast learning module: introducing local and global contrast learning strategies, optimizing the similarity of the image and text features through dynamic resampling and memory enhancement mechanisms, and strengthening the semantic association between multi-modalities;

[0190] Self-supervised text generation module: Generate auxiliary text information from image features through a generative adversarial network or variational autoencoder, and perform contrastive learning with real text labels to further refine the semantic mapping between images and text;

[0191] Optimization and inference module for unified embedding space: In the high-dimensional unified embedding space, through fine-grained regression optimization and semantic matching scoring mechanism, images and text can be efficiently matched and the image recognition task can be completed during the inference stage.

[0192] The module in Embodiment 2 is used to implement the functions in Embodiment 1. This embodiment can be implemented by a system, which includes a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, the intelligent image recognition system and method based on deep learning described in Embodiment 1 of the present application are implemented. The system also includes other components well-known to those skilled in the art such as a communication bus and a communication interface, and their settings and functions are known in the art, so they will not be described in detail here.

[0193] In the present application, the aforementioned memory can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, system, or device. For example, a computer-readable storage medium can be any suitable magnetic storage medium or magneto-optical storage medium, such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high-bandwidth memory (HBM), hybrid memory cube (HMC), etc., or any other medium that can be used to store the required information and can be accessed by an application, module, or both. Any such computer storage medium can be part of a device or accessible or connectable to the device. Any application or module described in the present application can be implemented using computer-readable / executable instructions that can be stored or otherwise maintained by such a computer-readable medium.

[0194] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the scope of protection of the claims of the present invention.

Claims

1. An intelligent image recognition method based on deep learning, characterized in that: include: Step 1: Extract high-level semantic features from images and texts respectively through pre-trained image and text feature extraction networks; Step 2: Use an extended network to map image and text features from a low-dimensional space to a higher dimension, enhance semantic information, and map both to a unified high-dimensional semantic space through an alignment layer; Step 3: Introduce local and global contrastive learning strategies to optimize the similarity of image and text features through dynamic resampling and memory enhancement mechanisms, and strengthen the semantic association between multimodal features, including: By introducing local and global contrast learning strategies, the similarity of image and text features is optimized and the semantic association between multimodalities is strengthened. Local contrast learning focuses on object-level semantic contrast, designs a local contrast loss function, maximizes the similarity of positive sample pairs, minimizes the similarity of negative sample pairs, and guides the network to learn effective local features; global contrast learning focuses on the overall scene or context comparison, designs a global contrast loss function, and optimizes the model's understanding of the overall scene; a dynamic resampling mechanism is also introduced to dynamically select difficult samples according to the current learning status of the model, increase the probability of difficult samples, and avoid overfitting simple categories. At the same time, with the help of memory-enhanced contrast learning, the memory network is used to store historical comparison samples, and the playback mechanism is used to help the model remember the semantic differences between categories; Step 4: Generate auxiliary text information from image features through generative adversarial networks or variational autoencoders, and compare and learn with real text labels to further refine the semantic mapping between image and text; Step 5: In the high-dimensional unified embedding space, through regression optimization and semantic matching scoring mechanism, images and texts can be efficiently matched in the inference stage and the image recognition task can be completed, including: Through the regression optimization mechanism, a bidirectional distance loss function is adopted to minimize the distance of matching samples and maximize the distance of non-matching samples, so as to achieve image and text embedding space optimization. A balance factor is introduced to adjust the optimization intensity to maintain the semantic consistency and moderate diversity of images and texts. In the reasoning stage, for new images and text descriptions, high-level features are first extracted through the feature extraction network, and then mapped to the high-dimensional embedding space for alignment through the extended network. The similarity of images and texts is measured by the semantic matching score to judge the degree of matching. For target detection tasks, the regional convolutional neural network is used to locate the target by combining the image region features and text features, and the target category is confirmed by combining the matching score, finally achieving high-precision image recognition.

2. The deep learning-based intelligent image recognition method according to claim 1, characterized in that: The step 1 comprises: For image feature extraction, the Vision Transformer method is used to divide the image into blocks of fixed size. The relationship between blocks is modeled through the self-attention mechanism to capture global context information. Finally, the high-dimensional embedding representation of the image is obtained from the penultimate layer of the ViT model. The BERT model is used to extract the corresponding text features. The input text is first segmented and then passed into the BERT encoding. The bidirectional Transformer architecture is used to obtain the semantic representation of each word, and the global representation vector of the text is obtained through [CLS] token aggregation.

3. The deep learning-based intelligent image recognition method according to claim 2, characterized in that: The step 2 comprises: First, feature dimension expansion is performed. Image feature expansion uses a multi-layer perceptron network, which uses linear transformation layers and nonlinear activation layers to initially embed and map the image to a higher dimension. Text feature expansion also uses a similar MLP network to map text embedding to the same high-dimensional space as the image features, followed by spatial alignment. Through the bidirectional cross-attention mechanism and similarity metric learning of the alignment layer, the semantic relationship between image and text features in the high-dimensional space is strengthened to achieve precise alignment.

4. The deep learning-based intelligent image recognition method according to claim 3, characterized in that: The step 4 comprises: A generative adversarial network is selected as the generative model. Image features are input into the generator and combined with random noise to generate text descriptions. The discriminator is then used to judge the similarity between the generated text and the real text. The discriminator loss function is designed to optimize the generated text. The generated text description is compared with the original real text. A contrastive learning strategy is adopted to maximize the similarity of similar text pairs and minimize the distance between dissimilar text pairs to further refine the semantic mapping. A diversity penalty mechanism is introduced to encourage the generator to explore the text generation path. At the same time, the generated text and image features are jointly trained to ensure that the generated text and image features remain consistent in the embedding space, so as to obtain a more accurate image-text semantic mapping.

5. Intelligent image recognition system based on deep learning, characterized in that: include: Image and text preliminary feature extraction module: extracts high-level semantic features from images and texts respectively through pre-trained image and text feature extraction networks, laying the foundation for subsequent mapping and contrastive learning; Progressive feature expansion module: Uses an expansion network to map image and text features from a low-dimensional space to a higher dimension, enhances semantic information, and maps both to a unified high-dimensional semantic space through an alignment layer; Dynamic contrastive learning module: introduces local and global contrastive learning strategies, optimizes the similarity of image and text features through dynamic resampling and memory enhancement mechanisms, and strengthens the semantic association between multimodal features, including: By introducing local and global contrast learning strategies, the similarity of image and text features is optimized and the semantic association between multimodalities is strengthened. Local contrast learning focuses on object-level semantic contrast, designs a local contrast loss function, maximizes the similarity of positive sample pairs, minimizes the similarity of negative sample pairs, and guides the network to learn effective local features; global contrast learning focuses on the overall scene or context comparison, designs a global contrast loss function, and optimizes the model's understanding of the overall scene; a dynamic resampling mechanism is also introduced to dynamically select difficult samples according to the current learning status of the model, increase the probability of difficult samples, and avoid overfitting simple categories. At the same time, with the help of memory-enhanced contrast learning, the memory network is used to store historical comparison samples, and the playback mechanism is used to help the model remember the semantic differences between categories; Self-supervised text generation module: Generates auxiliary text information from image features through generative adversarial networks or variational autoencoders, and compares and learns with real text labels to further refine the semantic mapping between image and text; Unified embedding space optimization and reasoning module: In the high-dimensional unified embedding space, through regression optimization and semantic matching scoring mechanism, images and texts can be efficiently matched in the reasoning stage and image recognition tasks can be completed, including: Through the regression optimization mechanism, a bidirectional distance loss function is adopted to minimize the distance of matching samples and maximize the distance of non-matching samples, so as to achieve image and text embedding space optimization. A balance factor is introduced to adjust the optimization intensity to maintain the semantic consistency and moderate diversity of images and texts. In the reasoning stage, for new images and text descriptions, high-level features are first extracted through the feature extraction network, and then mapped to the high-dimensional embedding space for alignment through the extended network. The similarity of images and texts is measured by the semantic matching score to judge the degree of matching. For target detection tasks, the regional convolutional neural network is used to locate the target by combining the image region features and text features, and the target category is confirmed by combining the matching score, finally achieving high-precision image recognition.

Citation Information

Patent Citations

  • Multi-task learning model combining image-text matching and visual reasoning, visual common sense reasoning method and computer equipment

    CN114996502A

  • Cross-modal image text retrieval method based on deep learning

    CN119311911A