Intelligent image recognition system and method based on deep learning
Through the intelligent image recognition method based on deep learning, high-level semantic features of images and text are extracted and mapped, and efficient matching of images and text in a unified embedding space through alignment and contrast learning strategies, the problem of image and text information fusion is solved and the accuracy of the image recognition system is improved.
Patent Information
- Application Number
- CN202510451037.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-04-11
AI Technical Summary
The prior art is difficult to effectively integrate the information of images and text, resulting in separation of processing between images and text in image recognition tasks, and it is difficult to directly compare or work together in a unified space.
Using a deep learning-based intelligent image recognition method, high-level semantic features are extracted through pre-trained image and text feature extraction networks, extended networks are used to map features to high-dimensional spaces, and similarity between image and text features is optimized through alignment layers and contrast learning strategies, ultimately achieving efficient matching of images and text in a unified embedded space.
The deep fusion of images and text is achieved, the understanding of the relationship between images and text is enhanced, the accuracy of the image recognition system is improved, and the semantic similarity between images and text can be compared with the same distance metric.
Smart Images

Figure CN119963950A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition technology, and in particular to an intelligent image recognition system and method based on deep learning. Background Art
[0002] In image recognition tasks, images and text are two main modal information, which express the understanding of objects, scenes and events from different perspectives. Images present visual information through a combination of pixels and colors, while text captures the relevant context and meaning through language and semantic descriptions. Although images and text differ in the way they convey information, they are essentially interrelated. For example, the objects shown in the image can be further defined and explained through text descriptions, and the descriptions in the text can also provide a clearer context for the understanding of the image. However, how to effectively integrate the information of these two modalities remains a major challenge in the field of image recognition. In current technical methods, the processing of images and texts is often separate, and image features and text features are extracted and processed through independent models, which makes it difficult for them to be directly compared or work together in a unified space. Summary of the invention
[0003] In order to solve the above problems, the present invention provides an intelligent image recognition system and method based on deep learning.
[0004] To achieve the above purpose, the technical solution adopted by the present invention is as follows: On the one hand, the present invention discloses an intelligent image recognition method based on deep learning, comprising: Step 1: Extract high-level semantic features from images and texts respectively through pre-trained image and text feature extraction networks; Step 2: Use an extended network to map image and text features from a low-dimensional space to a higher dimension, enhance semantic information, and map both to a unified high-dimensional semantic space through an alignment layer; Step 3: Introduce local and global contrastive learning strategies to optimize the similarity of image and text features through dynamic resampling and memory enhancement mechanisms, and strengthen the semantic association between multimodalities; Step 4: Generate auxiliary text information from image features through generative adversarial networks or variational autoencoders, and compare and learn with real text labels to further refine the semantic mapping between image and text; Step 5: In the high-dimensional unified embedding space, through sophisticated regression optimization and semantic matching scoring mechanism, images and texts can be efficiently matched in the inference stage to complete the image recognition task.
[0005] Further: Step 1 comprises: For image feature extraction, the Vision Transformer method is used to divide the image into blocks of fixed size. The relationship between blocks is modeled through the self-attention mechanism to capture global context information. Finally, the high-dimensional embedding representation of the image is obtained from the penultimate layer of the ViT model. The BERT model is used to extract the corresponding text features. The input text is first segmented and then passed into the BERT encoding. The bidirectional Transformer architecture is used to obtain the semantic representation of each word, and the global representation vector of the text is obtained through [CLS] token aggregation.
[0006] Further: Step 2 includes: First, feature dimension expansion is performed. Image feature expansion uses a multi-layer perceptron network, which uses linear transformation layers and nonlinear activation layers to initially embed and map the image to a higher dimension. Text feature expansion also uses a similar MLP network to map text embedding to the same high-dimensional space as the image features, followed by spatial alignment. Through the bidirectional cross-attention mechanism and similarity metric learning of the alignment layer, the semantic relationship between image and text features in the high-dimensional space is strengthened to achieve precise alignment.
[0007] Further: Step 3 comprises: By introducing local and global contrastive learning strategies, the similarity of image and text features is optimized and the semantic association between multimodalities is strengthened. Local contrastive learning focuses on object-level semantic contrast and designs a local contrast loss function to maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs, thereby guiding the network to learn effective local features. Global contrastive learning focuses on the overall scene or context contrast and designs a global contrast loss function to optimize the model's understanding of the overall scene. A dynamic resampling mechanism is also introduced to dynamically select difficult samples based on the current learning status of the model, increase the probability of difficult samples, and avoid overfitting simple categories. At the same time, with the help of memory-enhanced contrastive learning, a memory network is used to store historical contrast samples, and a playback mechanism is used to help the model remember the semantic differences between categories.
[0008] Further: Step 4 includes: A generative adversarial network is selected as the generative model. Image features are input into the generator and combined with random noise to generate text descriptions. The discriminator is then used to judge the similarity between the generated text and the real text. The discriminator loss function is designed to optimize the generated text. The generated text description is compared with the original real text. A new comparative learning strategy is adopted to maximize the similarity of similar text pairs, minimize the distance between dissimilar text pairs, and further refine the semantic mapping. A diversity penalty mechanism is introduced to encourage the generator to explore the text generation path. At the same time, the generated text and image features are jointly trained to ensure that the generated text and image features remain consistent in the embedding space, so as to obtain a more accurate image-text semantic mapping.
[0009] Further: Step 5 comprises: Through the regression optimization mechanism, a bidirectional distance loss function is adopted to minimize the distance of matching samples and maximize the distance of non-matching samples, so as to achieve image and text embedding space optimization. A balance factor is introduced to adjust the optimization intensity to maintain the semantic consistency and moderate diversity of images and texts. In the reasoning stage, for new images and text descriptions, high-level features are first extracted through the feature extraction network, and then mapped to the high-dimensional embedding space for alignment through the extended network. The similarity of images and texts is measured by the semantic matching score to judge the degree of matching. For target detection tasks, the regional convolutional neural network is used to locate the target by combining the image region features and text features, and the target category is confirmed by combining the matching score, finally achieving high-precision image recognition.
[0010] On the other hand, the present invention discloses an intelligent image recognition system based on deep learning, comprising: an image and text preliminary feature extraction module: extracting high-level semantic features from images and texts respectively through pre-trained image and text feature extraction networks, laying a foundation for subsequent mapping and contrastive learning; Progressive feature expansion module: Uses an expansion network to map image and text features from a low-dimensional space to a higher dimension, enhances semantic information, and maps both to a unified high-dimensional semantic space through an alignment layer; Dynamic contrastive learning module: introduces local and global contrastive learning strategies, optimizes the similarity of image and text features through dynamic resampling and memory enhancement mechanisms, and strengthens the semantic association between multimodalities; Self-supervised text generation module: Generates auxiliary text information from image features through generative adversarial networks or variational autoencoders, and compares and learns with real text labels to further refine the semantic mapping between image and text; Unified embedding space optimization and reasoning module: In the high-dimensional unified embedding space, through sophisticated regression optimization and semantic matching scoring mechanism, images and texts can be efficiently matched in the reasoning stage to complete the image recognition task.
[0011] Compared with the prior art, the present invention has the following technical advances: The present invention can not only extract meaningful high-dimensional features from images and texts from their respective modal spaces, but also align these two features in a common embedding space, thereby providing a tighter and more accurate semantic mapping for subsequent reasoning and task execution. By mapping images and texts into a unified embedding space, the present invention can achieve deep fusion of multimodal information, which can not only enhance the understanding of the relationship between images and texts, but also effectively improve the accuracy of image recognition systems. The design of the unified embedding space enables the semantic similarity of images and texts to be compared through the same distance metric, thereby achieving efficient image and text matching in the reasoning stage. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.
[0013] In the attached picture: Figure 1 is a flow chart of the present invention; Figure 2 It is a system structure diagram of the present invention. DETAILED DESCRIPTION
[0014] The following specific embodiments may be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. The embodiments of the present invention will be described below in conjunction with the accompanying drawings.
[0015] Example 1 like Figure 1 As shown, this embodiment discloses an intelligent image recognition method based on deep learning, including: Step 1: Extract high-level semantic features from images and texts respectively through pre-trained image and text feature extraction networks; Step 2: Use an extended network to map image and text features from a low-dimensional space to a higher dimension, enhance semantic information, and map both to a unified high-dimensional semantic space through an alignment layer; Step 3: Introduce local and global contrastive learning strategies to optimize the similarity of image and text features through dynamic resampling and memory enhancement mechanisms, and strengthen the semantic association between multimodalities; Step 4: Generate auxiliary text information from image features through generative adversarial networks or variational autoencoders, and compare and learn with real text labels to further refine the semantic mapping between image and text; Step 5: In the high-dimensional unified embedding space, through sophisticated regression optimization and semantic matching scoring mechanism, images and texts can be efficiently matched in the inference stage to complete the image recognition task.
[0016] The purpose of step 1 is to extract preliminary features from the original image and text, laying the foundation for the subsequent mapping and contrastive learning stages.
[0017] First, image feature extraction is performed.
[0018] In order to extract features with rich semantic information from the image, this embodiment adopts a method based on VisionTransformer ViT. In this step, the goal of this embodiment is not only to extract traditional low-level visual features, but also to enable the model to capture more detailed semantic information in the image through global feature integration and local information encoding.
[0019] 1. Select the base model: Using ViT as the basic network for feature extraction, ViT is different from traditional convolutional neural networks in that it divides images into fixed-size patches and models the relationship between patches through a self-attention mechanism to capture global context information. Therefore, ViT can preserve long-distance dependencies in images in high-dimensional space and provide a feature representation that fuses global and local information.
[0020] 2. Image segmentation and position encoding: First, the input image (size is ) is divided into several fixed-size patches, each of which is ,in, is the scale parameter of the block, each image block is flattened and mapped to a vector of fixed dimension ,in, is the feature dimension of each block. In order to preserve the spatial position information of the blocks in the image, this embodiment introduces position coding, by encoding the position of each block ( ) is added to its corresponding image block vector to obtain an image block representation containing spatial position information:
[0021] 3. Transformer Encoder: The encoding of all image blocks is input into the Transformer Encoder, and the self-attention mechanism and feedforward neural network are used to perform global information interaction. In this process, the model encodes the local information of the image through the self-attention mechanism and captures the correlation across blocks through the multi-head attention layer. The output sequence of the Transformer Containing rich contextual information, the final generated feature vector is used to represent the high-dimensional feature representation of the image.
[0022] 4. Output layer: This example will obtain the final image feature representation from the penultimate layer of the ViT model (i.e., the embedding vector output by the last multi-head attention layer). , which will be the high-dimensional embedding representation of the image for subsequent contrastive learning and modality alignment tasks.
[0023] Then text feature extraction is performed.
[0024] In terms of text feature extraction, this embodiment selects BERT (Bidirectional Encoder Representations from Transformers) as the basic model. BERT models language features through bidirectional context learning capabilities and can well capture the semantic relationships and subtle differences in the text.
[0025] 1. Text preprocessing and encoding: For the input text description (for example, "a dog is running on the grass"), the sentence is first decomposed into a series of words or vocabulary units through Tokenization. Then, these word segmentation sequences are passed to the BERT model for encoding. BERT uses its bidirectional Transformer architecture to obtain the semantic representation of each word from the context and generate a vector representation of each word. (in, is the dimension of text features).
[0026] 2. Word vector synthesis and global representation: In order to obtain the global feature representation of the entire sentence, the vector of each word output by BERT is aggregated through the [CLS] token (special classification marker). This [CLS] token is designed to represent the semantic information of the entire sentence. Therefore, by extracting the output of the [CLS] token, this embodiment can obtain the global representation vector of the text. , which is the embedded representation of the text.
[0027] 3. Feature Fusion: After extracting image and text features, the next task is to align the embedded representations of the image and text and map them to a unified feature space. This process is usually performed in the subsequent contrastive learning and mapping steps, but in this step, this embodiment has obtained preliminary features of the two modalities: image features and text features , which retain the basic semantic information of images and texts respectively, laying the foundation for subsequent collaborative learning and alignment.
[0028] In step 1, this embodiment uses powerful pre-trained models such as ViT and BERT to extract high-level features of images and texts respectively. In terms of image processing, ViT uses block-based decomposition and self-attention mechanisms, so that images can be effectively encoded into high-dimensional feature vectors. , while the text features generate a global representation of the text through the BERT model This provides a reliable foundation for subsequent contrastive learning and cross-modal feature alignment. The results of these preliminary feature extractions will support the subsequent feature mapping and contrastive learning stages, and ultimately promote the deep fusion and understanding of image and text semantics.
[0029] The purpose of step 2 is to ensure that the two modalities can be mapped to a unified high-dimensional embedding space and capture more complex semantic information between the two by gradually expanding the dimensions of image and text features. The core idea of extending features is to effectively fuse and align image and text features with each other through nonlinear transformations and multi-layer interactions, thereby providing richer representations for subsequent cross-modal learning.
[0030] First, feature dimension expansion is performed.
[0031] 1. Image feature expansion: The key to the image feature expansion process is how to transform the original image features from Space expands to higher dimensions , so that the image features can have stronger expressive power and establish effective semantic connections with text features in a unified space. Here, this embodiment adopts an extended network, which is composed of a series of nonlinear transformation layers.
[0032] 1.1 Network structure design: This embodiment uses a multi-layer perceptron (MLP) network as an expansion module. This MLP network consists of two main parts: The first part is a linear transformation layer, which embeds the image Mapping to an intermediate space ,in, is the dimension of the intermediate space, usually taken as .
[0033] The second part is a nonlinear activation layer, which introduces nonlinear transformations through a series of ReLU activation functions to further improve the expressiveness of the feature space and finally output a high-dimensional feature representation. , the dimension It can be designed as the maximum dimension when the image and text features are aligned.
[0034] The mathematical expression is as follows:
[0035]
[0036] in, and is the weight matrix, and is the bias term, is the activation function.
[0037] 2. Text feature expansion: In the text feature expansion process, this embodiment follows the same idea as image feature expansion, and also uses an expansion network to embed the text from Space mapping to space, so that image and text features can be mapped into the same embedding space, thus ensuring that the semantics between the two can be better aligned and compared.
[0038] 2.1 Network structure design: Similar to the expansion of image features, text feature expansion is also implemented through a multi-layer perceptron (MLP) network. First, the text features It will be mapped to the intermediate space through a linear transformation layer In, usually take Then, a nonlinear activation function is used to further improve the expressiveness of the features, and finally a high-dimensional text embedding is output. .
[0039] The mathematical expression is as follows:
[0040]
[0041] Through this process, this embodiment embeds the preliminary features of the text into the same high-dimensional space as the image features. , ensuring that image and text features are effectively aligned in a unified semantic space.
[0042] Then perform spatial alignment.
[0043] After completing the expansion of image and text features, this embodiment obtains high-dimensional feature representations of the two modalities. and ,Next, this embodiment will achieve precise alignment of the two through the alignment layer to ensure their semantic consistency in the high dimensional space.
[0044] 1. Alignment layer design: The core goal of the alignment layer is to strengthen the semantic relationship between the image and text in the high-dimensional feature space through a cross-fusion mechanism. Specifically, the alignment layer will ensure that the image and text features can be correctly aligned in the high-dimensional space through a bidirectional cross-attention mechanism and similarity metric learning.
[0045] 1.1 Cross-Attention Mechanism: This embodiment adopts a cross self-attention mechanism so that the image features are adjusted under the guidance of the text features, and the text features are optimized under the influence of the image features. The calculation of the cross attention mechanism can be completed by the following steps:
[0046]
[0047]
[0048] in, and are the query and key matrices for images and texts respectively, and is a value matrix. Through this cross-fusion process, this embodiment can mutually enhance the high-dimensional features of the image and text, thereby making the two more closely linked in semantics.
[0049] 2. Similarity measurement: After cross-fusion, this embodiment further introduces similarity metric learning to optimize the alignment of the two in a unified space by calculating the similarity between image and text features. Commonly used similarity metrics include cosine similarity or Euclidean distance, which are in the following form:
[0050] By optimizing the similarity metric, this embodiment can further refine the relationship between the image and the text in the unified embedding space, thereby enhancing the semantic consistency between the two.
[0051] In step 2, this embodiment uses a progressive feature expansion mechanism to adopt an extended network of images and texts to expand the features of the two modalities from and The space is expanded to a unified high-dimensional space , and further aligns the semantic relationship between the two by means of cross-fusion and similarity measurement. Through this process, this embodiment not only effectively expands the expressive power of image and text features, but also ensures their semantic alignment in high-dimensional space, laying a solid foundation for subsequent cross-modal learning and reasoning tasks.
[0052] The purpose of step 3 is to strengthen the semantic relevance between images and texts in the high-dimensional embedding space by introducing a dynamic contrastive learning mechanism, optimize the distance relationship between the two, and ensure that the model can maintain its learning effect in the face of complex and difficult semantic contrasts through dynamic resampling and memory-enhanced contrastive learning strategies, thereby avoiding overfitting on simple samples.
[0053] First, local and global comparative learning is performed.
[0054] 1. Local contrastive learning: In contrastive learning of images and texts, local contrastive learning focuses on semantic contrast at the object level, that is, contrasting the semantic similarity at the local level (such as a single object or the details of an object). For example, in images, object-level contrasts such as "cat" and "dog" are performed; in texts, semantic contrasts are performed through corresponding descriptions (such as the text descriptions of "cat" and "dog").
[0055] Here, this embodiment designs a local contrast loss function, the core idea of which is to guide the network to learn effective local features by maximizing the similarity of positive sample pairs and minimizing the similarity of negative sample pairs.
[0056] The mathematical expression is:
[0057] in, and They are High-dimensional feature representation of image and text samples, is a temperature parameter that controls the scale of similarity between samples. Through this loss function, the model can gradually learn effective contrast features at the object level.
[0058] 2. Global contrastive learning: In global contrastive learning, the focus shifts to more macroscopic semantic contrasts, involving comparisons of entire scenes or overall contexts. For example, the contrast between the entire scene "dining table" and "sofa" in an image, or the contrast of scenes in a text description (such as a text describing a dining table and a sofa). The goal of global contrastive learning is to optimize the model's understanding of the overall scene and capture the semantic relationship between images and text in a larger scope.
[0059] For global comparison, a global contrast loss function is designed, which aims to maximize the distance with positive samples (similar scenes) while minimizing the distance with negative samples (dissimilar scenes).
[0060] The mathematical expression is:
[0061] By comparing the global semantics of images and texts, the model can not only capture local information in detail, but also understand the overall semantic structure of images and texts, thereby enhancing the generalization ability for diverse scenarios.
[0062] Then dynamic resampling is performed.
[0063] In order to avoid overfitting to certain simple categories in contrastive learning, this embodiment introduces a dynamic resampling mechanism, which aims to automatically increase the proportion of difficult samples and introduce challenges in the training process by gradually adjusting the contrast samples sampled in each training batch.
[0064] 1. Resampling mechanism design: In each training cycle, the system dynamically selects samples with higher sampling difficulty based on the learning status of the current comparison samples. For example, in the early stage, the model may find it easier to distinguish common objects (such as "cat" and "dog"), while in the later stage, it will sample more complex or ambiguous semantic pairs (such as "cat" and "rabbit") for learning.
[0065] The key to the resampling strategy is to use the output of the current model to predict which samples are more difficult to distinguish and increase the probability of these difficult samples appearing in training. This embodiment uses a hard sample mining module to dynamically adjust the selection of training samples:
[0066] in, is the current loss of sample i, is the temperature parameter, which controls the weight of difficult samples. In this way, the probability of difficult samples will gradually increase as training progresses, thus helping the model learn more complex semantic relationships.
[0067] Finally, memory-enhanced contrastive learning is performed.
[0068] In order to further enhance the stability of contrastive learning, this embodiment introduces a memory network, which can store historical contrast samples during the training process and help the current model better remember and distinguish the semantic differences between different categories through a playback mechanism.
[0069] 1. Memory Network Design: The memory network is used as an auxiliary module to store and replay previous comparison samples. This embodiment stores the effective comparison pairs in the training process into the memory bank through a soft storage mechanism, and then inputs these samples into the model together with the current training samples through a replay mechanism, thereby enhancing the model's understanding of the similarities and differences between different categories.
[0070] Mathematically, the output of the memory network can be expressed as:
[0071] in, Represents the memory bank, which stores historically valid comparison samples; The "operation" retrieves the most relevant memories from the memory bank and uses these memories to help optimize the current comparison samples. The introduction of the memory replay mechanism can prevent the model from relying solely on the samples of the current batch, ensure the stability of training, and enhance the generalization ability.
[0072] In step 3, this embodiment combines local contrast and global contrast learning through a dynamic contrast learning mechanism, not only focusing on the semantic differences of a single object, but also extending to the contrast of the entire scene. Through dynamic resampling, the learning challenge of more complex samples is introduced to avoid the model from falling into overfitting of a single category. Finally, with the help of memory-enhanced contrast learning, historical valid samples and playback mechanisms are used to further improve the stability and generalization ability of the model during training. Through these designs, the model can not only perform effective contrast learning in high-dimensional space, but also continuously adapt to and optimize the semantic relationships between different categories, laying a solid foundation for subsequent cross-modal learning.
[0073] The purpose of step 4 is to generate auxiliary text information through self-supervised learning to enhance the semantic consistency between image and text features. Specifically, by generating models (such as generative adversarial networks GAN or variational autoencoders VAE), corresponding text descriptions are generated from image features as a self-supervised training signal to help the semantic alignment of image and text features. This process not only improves the ability to understand images, but also further strengthens the accuracy of text descriptions on image semantics, ultimately making the joint embedding of images and texts more accurate.
[0074] First, self-supervised text generation is performed.
[0075] 1. Generative model selection and design: This embodiment selects the generative adversarial network (GAN) as the main generative model, and combines the structure of the conditional generative adversarial network to generate text descriptions from image features. In this process, this embodiment represents the image features through the high-dimensional space expanded in the previous step. The input is sent to the generator network to generate the corresponding text description. The generator combines the image features with random noise, performs feature mapping through multiple convolutional layers and fully connected layers, and finally generates a text description with semantic consistency.
[0076] Mathematically expressed as:
[0077] in, represents a generator, is the expanded image feature, is random noise, is the generated text embedding representation.
[0078] 2. Discriminator design: Similar to the traditional GAN structure, this embodiment introduces a discriminator to judge the similarity between the generated text and the real text. The goal of the discriminator is to maximize the difference between the real and generated text. The discriminator encodes the input text through a convolutional neural network (CNN) and compares it with the features of the image.
[0079] Loss function of the discriminator The design is to maximize the matching degree between real text and image features, while minimizing the matching degree between generated text and image features. Specifically, it is expressed as:
[0080] in, Represents the characteristics of real text, is the discriminator, whose goal is to make the generated text as distinguishable as possible from the real text.
[0081] 3. Contrastive Learning of Generated Text: The generated text description is compared with the original real text description to further refine the semantic mapping between image and text. This embodiment adopts a new contrastive learning strategy to embed the generated text into With real text embedding The comparison is performed to maximize the similarity of similar text pairs and minimize the distance of dissimilar text pairs.
[0082] The loss function of contrastive learning can be expressed as:
[0083] in, and Respectively represent The embedding representations of generated text and real text, is the temperature parameter that controls the scale of contrast.
[0084] Through this process, the generated text not only serves as an auxiliary supervisory signal to help image understanding, but also further enhances the semantic consistency between image and text through contrastive learning, so that the two can be more accurately aligned in the high-dimensional embedding space.
[0085] Then optimization for self-supervised text generation is performed.
[0086] 1. Enhanced text generation diversity: In order to further improve the diversity and accuracy of the generated text, this embodiment introduces a diversity penalty mechanism, which can prevent the generated text from being too monotonous. By optimizing multiple potential text descriptions during the text generation process, this embodiment can encourage the generator to explore more text generation paths, thereby enhancing the diversity and semantic richness of the text.
[0087] The goal of the diversity penalty is to increase the diversity of generated text by:
[0088] in, Represents cosine similarity, and the goal is to minimize the similarity between the generated texts, so that the generated texts are more diverse.
[0089] 2. Joint training of generated text and image features: While generating text, this embodiment jointly optimizes the training of the generator and the discriminator to ensure that the generated text can be consistent with the image features in the embedding space. Under this joint training strategy, the generator not only needs to generate semantically reasonable text descriptions, but also ensure that these text descriptions can be closely aligned with the image features.
[0090] The joint optimization of the loss function can be expressed as:
[0091] in, is a balancing term that controls the intensity of the diversity penalty.
[0092] In step 4, this embodiment uses a self-supervised text generation mechanism and a generative adversarial network or a conditional generative adversarial network to generate text descriptions from image features as self-supervisory signals to help strengthen the semantic consistency between image and text. Through comparative learning of generated text, the generated text is compared with the real text to further refine the semantic alignment of the two in the high-dimensional embedding space. At the same time, this embodiment introduces a diversity penalty mechanism and a joint training strategy to improve the diversity and quality of generated text and ensure consistency between the generated text and image features, thereby obtaining a more accurate image-text semantic mapping throughout the cross-modal learning.
[0093] The purpose of step 5 is to optimize the unified embedding space so that images and texts can be matched efficiently and accurately in the inference phase, thereby effectively completing image recognition tasks (such as image classification, object detection, etc.). Specifically, the optimization process of the unified embedding space will focus on reducing the distance between similar images and texts, while maximizing the distance between heterogeneous features (images and texts), and through an accurate matching scoring mechanism, perform efficient image and text comparison in the inference phase, and ultimately achieve high-precision completion of the inference task.
[0094] First, unified embedding space optimization is performed.
[0095] 1. Optimize target design: After the image and text features are embedded in the same high-dimensional space, the next step is to ensure that the match between the image and the text is more accurate in this space. This embodiment designs a sophisticated regression optimization mechanism to minimize the distance between similar image and text pairs in the embedding space and maximize the distance between heterogeneous feature pairs. This optimization goal is achieved by designing a bidirectional distance loss function that takes into account both the alignment of image-text pairs and the reverse alignment of text-image pairs.
[0096] The core form of the loss function is as follows:
[0097] in, and Respectively represent High-dimensional embedding representation of images and texts of samples, is the weight coefficient controlling the negative distance term between image and text, Represents a negative sample (non-matching sample). This loss function optimizes the image and text embedding space by minimizing the distance of matching samples and maximizing the distance between non-matching samples.
[0098] 2. Balance during optimization: In the actual training process, the embedding space of images and texts changes dynamically, so this embodiment introduces a balance factor to adjust the optimization strength between the two. This embodiment needs to ensure that the representation of images and texts is highly consistent in semantics, but at the same time avoid them becoming overly dependent on a certain modal feature during the optimization process. Therefore, the distance calculation between images and texts adopts a weighted method:
[0099] in, is a loss function used to enhance the diversity of image and text features. is a weight hyperparameter that balances the loss term between images and text. This balancing strategy enables images and text to maintain semantic consistency during the optimization process while maintaining moderate diversity, further improving the matching accuracy in the reasoning phase.
[0100] Then the matching in the inference phase is performed.
[0101] 1. Image feature extraction and mapping: In the inference phase, given a new image and text description, we first extract high-level features from the image through a feature extraction network (e.g. using an optimized ResNet or ViT) , and then mapped to a high-dimensional embedding space by expanding the network , so that it is aligned with the high-dimensional embedding space of text.
[0102] Mathematically, image features The high-dimensional mapping can be expressed as:
[0103] in, It is the image expansion network.
[0104] 2. Text feature extraction and enhancement: At the same time, the given text description will be input into a pre-trained language model (such as BERT or GPT) to generate preliminary text features. Then, the semantic information of the text feature is enhanced through the generative network and contrastive learning mechanism, so that it is mapped to the same high-dimensional space as the image feature.
[0105] The high-dimensional mapping of text features is expressed as:
[0106] in, is a text expansion network, is the original text description.
[0107] 3. Semantic matching score: After the image and text are mapped in the high-dimensional space, the next step is to measure the similarity between the two through semantic matching score. The semantic matching score S_{match} can be achieved by calculating the cosine similarity between the image features and the text features:
[0108] in, represents the dot product of vectors, Represents the magnitude of a vector.
[0109] Based on this matching score, the system can determine the degree of match between the image and the text, and complete subsequent tasks (such as classification or object detection) based on the threshold.
[0110] Then object detection is performed in the inference phase.
[0111] In the target detection task, this embodiment not only requires the matching score of the image and the text, but also further enhances the matching accuracy through the location of the object in the image. Specifically, by combining the regional features in the image with the text features, the regional convolutional neural network (R-CNN) is used to locate the target in the image, and the target category is confirmed by combining the matching score.
[0112] Finally, the loss function of object detection combines the position deviation and the matching degree, and the score S_{det} of object detection is calculated by combining the position and semantic similarity.
[0113] In step 5, after the features of the image and text are optimized in a unified embedding space, a sophisticated regression optimization mechanism is used to achieve semantic alignment between the image and text features, ensuring that they have the minimum matching distance in the high-dimensional space while maximizing the distance between heterogeneous feature pairs. In the inference stage, the features of the image and text are mapped to a unified embedding space through an extended network and contrastive learning mechanism, and then the matching degree between the two is judged by a semantic matching score to complete tasks such as image recognition classification or object detection.
[0114] Example 2 like Figure 2 As described, this embodiment discloses an intelligent image recognition system based on deep learning, including: image and text preliminary feature extraction module: extracting high-level semantic features from images and texts respectively through pre-trained image and text feature extraction networks, laying a foundation for subsequent mapping and contrastive learning; Progressive feature expansion module: Uses an expansion network to map image and text features from a low-dimensional space to a higher dimension, enhances semantic information, and maps both to a unified high-dimensional semantic space through an alignment layer; Dynamic contrastive learning module: introduces local and global contrastive learning strategies, optimizes the similarity of image and text features through dynamic resampling and memory enhancement mechanisms, and strengthens the semantic association between multimodalities; Self-supervised text generation module: Generates auxiliary text information from image features through generative adversarial networks or variational autoencoders, and compares and learns with real text labels to further refine the semantic mapping between image and text; Unified embedding space optimization and reasoning module: In the high-dimensional unified embedding space, through sophisticated regression optimization and semantic matching scoring mechanism, images and texts can be efficiently matched in the reasoning stage to complete the image recognition task.
[0115] The module in the embodiment 2 is used to implement the function in the embodiment 1. This embodiment can be implemented by a system, the system includes a processor and a memory, the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the intelligent image recognition system and method based on deep learning according to the embodiment 1 of the present application are implemented. The system also includes other components well known to those skilled in the art, such as a communication bus and a communication interface, and their settings and functions are known in the art, so they are not repeated here.
[0116] In the present application, the aforementioned memory may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, system or device. For example, a computer-readable storage medium may be any suitable magnetic storage medium or magneto-optical storage medium, such as a resistive random access memory RRAM (Resistive Random Access Memory), a dynamic random access memory DRAM (Dynamic Random Access Memory), a static random access memory SRAM (Static Random-Access Memory), an enhanced dynamic random access memory EDRAM (Enhanced Dynamic Random Access Memory), a high-bandwidth memory HBM (High-Bandwidth Memory), a hybrid memory cube HMC (Hybrid Memory Cube), etc., or any other medium that can be used to store the required information and can be accessed by an application, a module or both. Any such computer storage medium may be part of a device or accessible or connectable to a device. Any application or module described in this application may be implemented using computer-readable / executable instructions that may be stored or otherwise maintained by such a computer-readable medium.
[0117] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein with equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the scope of protection of the claims of the present invention.
Claims
1. An intelligent image recognition method based on deep learning, characterized in that: include: Step 1: Extract high-level semantic features from images and texts respectively through pre-trained image and text feature extraction networks; Step 2: Use an extended network to map image and text features from a low-dimensional space to a higher dimension, enhance semantic information, and map both to a unified high-dimensional semantic space through an alignment layer; Step 3: Introduce local and global contrastive learning strategies to optimize the similarity of image and text features through dynamic resampling and memory enhancement mechanisms, and strengthen the semantic association between multimodalities; Step 4: Generate auxiliary text information from image features through generative adversarial networks or variational autoencoders, and compare and learn with real text labels to further refine the semantic mapping between image and text; Step 5: In the high-dimensional unified embedding space, through sophisticated regression optimization and semantic matching scoring mechanism, images and texts can be efficiently matched in the inference stage to complete the image recognition task.
2. The deep learning-based intelligent image recognition method according to claim 1, characterized in that: The step 1 comprises: For image feature extraction, the Vision Transformer method is used to divide the image into blocks of fixed size. The relationship between blocks is modeled through the self-attention mechanism to capture global context information. Finally, the high-dimensional embedding representation of the image is obtained from the penultimate layer of the ViT model. The BERT model is used to extract the corresponding text features. The input text is first segmented and then passed into the BERT encoding. The bidirectional Transformer architecture is used to obtain the semantic representation of each word, and the global representation vector of the text is obtained through [CLS] token aggregation.
3. The deep learning-based intelligent image recognition method according to claim 2, characterized in that: The step 2 comprises: First, feature dimension expansion is performed. Image feature expansion uses a multi-layer perceptron network, which uses linear transformation layers and nonlinear activation layers to initially embed and map the image to a higher dimension. Text feature expansion also uses a similar MLP network to map text embedding to the same high-dimensional space as the image features, followed by spatial alignment. Through the bidirectional cross-attention mechanism and similarity metric learning of the alignment layer, the semantic relationship between image and text features in the high-dimensional space is strengthened to achieve precise alignment.
4. The deep learning-based intelligent image recognition method according to claim 3, characterized in that: The step 3 comprises: By introducing local and global contrastive learning strategies, the similarity of image and text features is optimized and the semantic association between multimodalities is strengthened. Local contrastive learning focuses on object-level semantic contrast and designs a local contrast loss function to maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs, thereby guiding the network to learn effective local features. Global contrastive learning focuses on the overall scene or context contrast and designs a global contrast loss function to optimize the model's understanding of the overall scene. A dynamic resampling mechanism is also introduced to dynamically select difficult samples based on the current learning status of the model, increase the probability of difficult samples, and avoid overfitting simple categories. At the same time, with the help of memory-enhanced contrastive learning, a memory network is used to store historical contrast samples, and a playback mechanism is used to help the model remember the semantic differences between categories.
5. The deep learning-based intelligent image recognition method according to claim 4, characterized in that: The step 4 comprises: A generative adversarial network is selected as the generative model. Image features are input into the generator and combined with random noise to generate text descriptions. The discriminator is then used to judge the similarity between the generated text and the real text. The discriminator loss function is designed to optimize the generated text. The generated text description is compared with the original real text. A new comparative learning strategy is adopted to maximize the similarity of similar text pairs, minimize the distance between dissimilar text pairs, and further refine the semantic mapping. A diversity penalty mechanism is introduced to encourage the generator to explore the text generation path. At the same time, the generated text and image features are jointly trained to ensure that the generated text and image features remain consistent in the embedding space, so as to obtain a more accurate image-text semantic mapping.
6. The deep learning-based intelligent image recognition method according to claim 5, characterized in that: The step 5 comprises: Through the regression optimization mechanism, a bidirectional distance loss function is adopted to minimize the distance of matching samples and maximize the distance of non-matching samples, so as to achieve image and text embedding space optimization. A balance factor is introduced to adjust the optimization intensity to maintain the semantic consistency and moderate diversity of images and texts. In the reasoning stage, for new images and text descriptions, high-level features are first extracted through the feature extraction network, and then mapped to the high-dimensional embedding space for alignment through the extended network. The similarity of images and texts is measured by the semantic matching score to judge the degree of matching. For target detection tasks, the regional convolutional neural network is used to locate the target by combining the image region features and text features, and the target category is confirmed by combining the matching score, finally achieving high-precision image recognition.
7. The deep learning-based intelligent image recognition system according to claim 6, characterized in that: include: Image and text preliminary feature extraction module: extracts high-level semantic features from images and texts respectively through pre-trained image and text feature extraction networks, laying the foundation for subsequent mapping and contrastive learning; Progressive feature expansion module: Uses an expansion network to map image and text features from a low-dimensional space to a higher dimension, enhances semantic information, and maps both to a unified high-dimensional semantic space through an alignment layer; Dynamic contrastive learning module: introduces local and global contrastive learning strategies, optimizes the similarity of image and text features through dynamic resampling and memory enhancement mechanisms, and strengthens the semantic association between multimodalities; Self-supervised text generation module: Generates auxiliary text information from image features through generative adversarial networks or variational autoencoders, and compares and learns with real text labels to further refine the semantic mapping between image and text; Unified embedding space optimization and reasoning module: In the high-dimensional unified embedding space, through sophisticated regression optimization and semantic matching scoring mechanism, images and texts can be efficiently matched in the reasoning stage to complete the image recognition task.
Citation Information
Patent Citations
Multi-task learning model combining image-text matching and visual reasoning, visual common sense reasoning method and computer equipment
CN114996502A
Cross-modal image text retrieval method based on deep learning
CN119311911A
News event search method and system based on multi-level image-text semantic alignment model
WO2023093574A1
Training multimodal machine learning models using cross-modality contrastive learning
WO2024237962A1
Cited By
Intelligent video analysis method and system adaptive to environment change
CN120126061A
An intelligent video analysis method and system that adapts to environmental changes
CN120126061B
Large-model medical image report generation method based on closed-loop feedback
CN122224399A