Image label generation method, system, equipment and medium

Through multiple feature extraction and Gaussian weight fusion, combined with similarity calculation, the problem of insufficient feature summary and reasoning capabilities in the prior art is solved, and the accuracy and adaptability of image processing are improved.

CN120220149APending Publication Date: 2025-06-27SHANGHAI XIDING IND CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510208801.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art relies on shallow feature extraction and complex feature fusion in image processing, but lacks the ability to summarize and reason as efficiently as humans, resulting in high cost of model training and insufficient focus on key areas.

Method used

Global features are generated by multiple feature extractions and weighted fusion is used to simulate human attention to significant areas. Compute the similarity of the fused features with the features in the feature library to ensure accurate matching and improve the accuracy of label generation.

Benefits of technology

It improves the feature expression effect and adaptability of images, improves classification accuracy, reduces the probability of misclassification, and improves the feature summary and reasoning ability in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220149A_ABST
    Figure CN120220149A_ABST
Patent Text Reader

Abstract

The invention relates to an image label generation method, system and device and a medium. The method comprises the following steps: acquiring a to-be-analyzed image; inputting the image into a feature extraction network of an image recognition model, and performing multiple times of feature extraction on the image to obtain a plurality of global features; inputting the plurality of global features into a fusion network of an image recognition model, determining the Gaussian weight of each global feature, and carrying out Gaussian weighted fusion on the corresponding global features based on the Gaussian weights to obtain fusion features; inputting the fusion features into a similarity calculation network of an image recognition model, calculating the similarity between the fusion features and each feature in a feature library, and taking the tags of the features with the similarity greater than a preset feature threshold as the tags of the image; wherein each feature in the feature library corresponds to a preset tag. The method has a human-like neuron learning function, and can significantly improve the learning ability of the current neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly relates to a method, a system, a device, and a medium for generating image tags. Background Art

[0002] At present, deep learning technology has made remarkable progress in the fields of image classification, object detection, etc., but still faces some key challenges in image processing: on the one hand, it highly depends on large-scale labeled data and high-computing-power devices, resulting in high model training costs. On the other hand, the learning mechanism of existing neural networks usually depends on repeatedly adjusting shallow features and lacks the ability to quickly extract high-level semantic information, which is essentially different from the human learning method. In addition, the same computing resources are allocated to each input sample without discrimination, ignoring the ability to focus on key regions. And existing feature fusion methods mainly rely on data-driven weight optimization and ignore the biologically inspired selective attention mechanism, resulting in insufficient ability of the model to capture significant regions. Therefore, a method, a system, a device, and a medium for generating image tags are needed. Summary of the Invention

[0003] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a method, a system, a device, and a medium for generating image tags, which improve the problem that the prior art often relies on shallow feature extraction and complex feature fusion but lacks the ability of efficient feature summarization and reasoning like humans.

[0004] To achieve the above purpose and other related purposes, the present invention provides a method for generating image tags, including: obtaining an image to be analyzed; inputting the image into the feature extraction network of an image recognition model to perform multiple feature extractions on the image to obtain multiple global features; inputting the multiple global features into the fusion network of the image recognition model to determine the Gaussian weights of each global feature, and performing Gaussian weighted fusion on the corresponding global features based on the Gaussian weights to obtain a fusion feature; inputting the fusion feature into the similarity calculation network of the image recognition model to calculate the similarity between the fusion feature and each feature in the feature library, and using the label of the feature whose similarity is greater than a preset feature threshold as the label of the image; wherein, each feature in the feature library corresponds to a preset label.

[0005] In an embodiment of the present invention, inputting the image into the feature extraction network of the image recognition model to perform multiple feature extractions on the image to obtain multiple global features, including: taking the image as the initial region to be analyzed; performing feature extraction on the region to be analyzed based on the filters of the first scale of the feature extraction network, obtaining and saving the global features of the region to be analyzed; performing feature extraction on the region to be analyzed based on the filters of the second scale of the feature extraction network, obtaining and saving the local features of the region to be analyzed; detecting whether there is a feature value greater than a preset feature threshold in the local features, and when there is, taking the region in the image corresponding to the feature value as the new region to be analyzed, and extracting the global features and local features of the region to be analyzed again until there is no feature value greater than the feature threshold in the obtained local features.

[0006] In an embodiment of the present invention, inputting the multiple global features into the fusion network of the image recognition model to determine the Gaussian weights of the respective global features, and performing Gaussian weighted fusion on the corresponding global features based on the Gaussian weights to obtain a fusion feature, including: for each global feature: dynamically identifying the center point of the Gaussian function in the global feature based on each feature carried in the global feature; determining the Gaussian weights corresponding to the respective global features based on the center point of the Gaussian function and the position of the corresponding global feature in the image; performing weighted fusion on the corresponding global features based on the Gaussian weights of the respective global features to obtain a fusion feature.

[0007] In an embodiment of the present invention, inputting the multiple global features into the fusion network of the image recognition model to determine the Gaussian weights of the respective global features, and performing Gaussian weighted fusion on the corresponding global features based on the Gaussian weights to obtain a fusion feature, including: for each global feature: determining the center point of the Gaussian function corresponding to the position based on the position of the global feature in the image; wherein, the center point of the Gaussian function is preset; determining the Gaussian weights corresponding to the respective global features based on the center point of the Gaussian function and the position of the corresponding global feature in the image; performing weighted fusion on the corresponding global features based on the Gaussian weights of the respective global features to obtain a fusion feature.

[0008] In an embodiment of the present invention, inputting the fusion feature into the similarity calculation network of the image recognition model to calculate the similarity between the fusion feature and each feature in the feature library, and taking the label of the feature whose similarity is greater than a preset feature threshold as the label of the image, including: calculating the similarity between the fusion feature and each feature in the feature library; for each similarity: determining whether the similarity is greater than a preset similarity threshold: if so, taking the label of the corresponding feature as the label of the image; if not, continuing to compare the next similarity until all similarities are compared to generate update information.

[0009] In one embodiment of the present invention, after generating the update information, it further includes: saving the true label of the image and the fusion feature of the image to the feature library, and updating the feature library.

[0010] In one embodiment of the present invention, the image recognition model is trained. When training the image recognition model, the parameters of the feature extraction network remain unchanged.

[0011] In one embodiment of the present invention, there is also provided a system for generating image labels. The system includes: an image acquisition module for acquiring an image to be analyzed; a feature extraction module for inputting the image into the feature extraction network of the image recognition model, performing multiple feature extractions on the image to obtain multiple global features; a feature fusion module for inputting the multiple global features into the fusion network of the image recognition model, performing Gaussian weighted fusion on the multiple global features to obtain a fusion feature; a label generation module for inputting the fusion feature into the similarity calculation network of the image recognition model, calculating the similarity between the fusion feature and each feature in the feature library, and using the label of the feature whose similarity is greater than a preset feature threshold as the label of the image; wherein, each feature in the feature library corresponds to a preset label.

[0012] In one embodiment of the present invention, there is also provided an electronic device, including: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, enabling the electronic device to implement the method for generating image labels described in any one of the above.

[0013] In one embodiment of the present invention, there is also provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor of a computer, the computer is enabled to execute the method for generating image labels described in any one of the above.

[0014] As described above, a method, system, device, and medium for generating image tags according to the present invention have the following beneficial effects: By performing multiple feature extractions on an image, the generated global features contain the overall semantic information of the image, making the information expression of the image richer. Since the Gaussian weights are calculated based on the feature intensity or the spatial position of the feature in the image, the Gaussian weights can emphasize the more significant or critical regions in the image, thus simulating the human attention to the significant regions. Through weighted fusion, not only the feature expression effect of the image is improved, but also the fused features have stronger adaptability in diverse tasks. Through similarity calculation, it is ensured that the fused features are precisely matched with the known features in the feature library, simulating the ability of humans to quickly generalize new information based on existing knowledge during the reasoning process, improving the accuracy of tag generation, and reducing the probability of generating incorrect tags. It improves the problem that the prior art often relies on shallow feature extraction and complex feature fusion, but lacks the efficient feature summarization and reasoning ability like humans. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a schematic flowchart of a method for generating image tags provided by an embodiment of the present invention;

[0016] Figure 2 It is a block diagram showing the structure of a system for generating image tags provided by an embodiment of the present invention;

[0017] Figure 3 It is a schematic structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] The following specific examples illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0019] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0020] In the following description, numerous details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.

[0021] Currently, deep learning relies on a computational model called Artificial Neural Networks (ANN), which consists of a large number of nodes (or "neurons"). These nodes are organized into different layers: the input layer is used to receive input data; there are usually multiple hidden layers for data processing; the output layer is used to output the final result. Each neuron is connected to multiple other neurons, and each connection has a weight representing the strength of the connection. Deep learning learns patterns in data by adjusting these weights. The key technologies involved usually include forward propagation, backpropagation, activation functions, and optimization algorithms. Its three main elements are data, computing power, and the network. Common network modules include convolutional neural networks, recurrent neural networks, long short-term memory networks, and Transformers, etc. Convolutional neural networks are a type of neural network particularly suitable for processing data with network structures such as images. It usually includes: a convolutional layer for extracting image features through convolutional operations; a pooling layer for reducing the dimensionality of data and retaining important information; a fully connected layer, which is the last few layers of the network, for mapping features to categories. A Transformer is a deep learning model based on the self-attention mechanism. It usually includes: a self-attention module for allowing the model to consider the relationships between different positions in the input sequence; a multi-head attention module for splitting the input into multiple heads and computing attention in parallel to improve the model's expressive ability; a positional encoding module for introducing the positions of elements in the sequence.

[0022] However, the inventors have found that existing various artificial neural networks mainly adopt fully supervised, semi-supervised, unsupervised, and self-supervised methods, and specifically, none of these methods can achieve online learning and summarization similar to humans; the fundamental reason is that during the learning of these methods, the shallow feature extraction is constantly changed; this is fundamentally different from the learning methods of humans and animals, resulting in the problem that a large amount of labeled data must be relied on, and it is impossible to effectively and quickly judge the feature value size of the input data. Any input requires the same amount of computing power.

[0023] The present invention provides a method for generating image tags. By performing multiple feature extractions on an image, the generated global features contain the overall semantic information of the image, making the information expression of the image richer. Since the Gaussian weights are calculated based on the feature intensity or the spatial position of the feature in the image, the Gaussian weights can emphasize the more significant or critical regions in the image, thus simulating the human attention to the significant regions. Through weighted fusion, not only the feature expression effect of the image is improved, but also the fused features have stronger adaptability in diverse tasks. Through similarity calculation, the exact matching between the fused features and the known features in the feature library is ensured, simulating the ability of humans to quickly generalize new information based on existing knowledge during the reasoning process, improving the classification accuracy and reducing the probability of misclassification. It improves the problem that the prior art often relies on shallow feature extraction and complex feature fusion, but lacks the efficient feature summarization and reasoning ability like humans. In the feature extraction network of the present invention, the human eye texture feature-based human-like learning method is fixed in advance, greatly accelerating the learning speed of the network. Based on the human online learning ability, the proposed similarity calculation network structure has the online learning ability and the function of human-like neuron learning, which will significantly improve the learning ability of the current neural network.

[0024] Please refer to Figure 1 , the method for generating image tags includes the following steps:

[0025] S1. Obtain the image to be analyzed.

[0026] When the image needs to be analyzed, first obtain the image to be analyzed. Among them, the image can be obtained by real-time acquisition through a camera, reading from local, or reading through a network interface, etc. After obtaining the image, in order to make the image meet the subsequent analysis standards, it can be judged whether the resolution of the image is less than the preset resolution threshold. If the resolution is less than the resolution threshold, the resolution can be increased to the resolution threshold or above the resolution threshold by interpolation; otherwise, if the resolution is greater than or equal to the resolution threshold, the image remains unchanged. Further, in order to adapt to the input requirements of the subsequent image recognition model, the image also needs to be normalized to map the pixel values to a preset range.

[0027] S2. Input the image into the feature extraction network of the image recognition model, perform multiple feature extractions on the image, and obtain multiple global features.

[0028] The image recognition model can be any model capable of extracting image features and performing analysis, including but not limited to convolutional neural networks, Transformer series models, deep neural networks, etc. For example, the image recognition model can be a YOLO series model. The feature extraction network is used to extract the features of the image. The feature extraction network in the present invention includes a plurality of filter banks, and each filter bank includes a preset number of filters with different sizes and different directions, which are used to capture local features and global features of different scales and different directions in the image, and obtain global features of multiple different regions in the image through layer-by-layer transmission.

[0029] In an embodiment of the present invention, the step of inputting the image into the feature extraction network of the image recognition model and performing multiple feature extractions on the image to obtain a plurality of global features includes:

[0030] Taking the image as the initial region to be analyzed;

[0031] Performing feature extraction on the region to be analyzed based on the filters of the first scale of the feature extraction network, and obtaining and saving the global features of the region to be analyzed;

[0032] Performing feature extraction on the region to be analyzed based on the filters of the second scale of the feature extraction network, and obtaining and saving the local features of the region to be analyzed;

[0033] Detecting whether there is a feature value greater than a preset feature threshold in the local features, and when there is, taking the region in the image corresponding to the feature value as a new region to be analyzed, and extracting the global features and local features of the region to be analyzed again until there is no feature value greater than the feature threshold in the obtained local features.

[0034] When extracting the features of a certain image, the image is input into the feature extraction model. First, the entire image is used as an initial region R to be analyzed. Based on a convolutional neural network or a filter, feature extraction is performed on this region R to generate the feature matrix of this region R for subsequent analysis. When performing feature extraction, a filter of the first scale is used to process the region to be analyzed to extract global features, which represent the macroscopic structure of the image, such as large textures or edge information. Further, a filter of the second scale is also used to perform local feature extraction on the same region to be analyzed to focus on the detailed features in the image, such as smaller textures or edges, where the first scale is greater than the second scale. After extraction, it is detected whether there are eigenvalues greater than the local feature threshold in the local feature matrix. If there are eigenvalues greater than the local feature threshold, the image regions corresponding to these eigenvalues are used as new regions to be analyzed, and feature extraction is continued on this region according to the first scale and the second scale. The above process is repeated until there are no eigenvalues greater than the feature threshold in the finally obtained local features. Through multiple iterative extractions, features of regions with high information content can be obtained to ensure that both detailed and global information are fully extracted. It can be understood that the feature extraction model performs feature extraction by designing a well-designed convolutional filter and combining the method of dilated convolution. Dilated convolution expands the receptive field of the filter without increasing the number of parameters through a preset dilation rate, so as to be able to capture a wider range of image features. At the same time, a sparse computing method is adopted in the calculation process, and only the regions with high information content in the image are calculated intensively, while the regions with low information content are skipped or simplified. This not only reduces the ineffective calculation but also significantly reduces the computational complexity. It should be noted that when extracting features for a new region to be analyzed each time, the dilation rate of the dilated convolution can be dynamically changed to achieve flexible adjustment of the receptive field and make it more adaptable to the characteristics of the current region to be analyzed. Further, considering that the Gabor filter is very close to the human visual cortex perception model and can well simulate the perception of the human eye for textures in different directions and frequencies, in an embodiment of the present invention, the filter is a Gabor filter.

[0035] S3. Input the multiple global features into the fusion network of the image recognition model, determine the Gaussian weights of each global feature, and perform Gaussian weighted fusion on the corresponding global features based on the Gaussian weights to obtain a fusion feature.

[0036] After the feature extraction network extracts multiple global features of the image, these global features are concatenated and input into the fusion network. A weight is assigned to each global feature according to the Gaussian function, and these global features are weighted and fused according to the corresponding weights to obtain the fusion feature of the entire image. Among them, the fusion feature contains the overall semantic features and detailed features of the image.

[0037] In an embodiment of the present invention, multiple global features are input into the fusion network of the image recognition model to determine the Gaussian weights of each global feature, and the corresponding global features are Gaussian weighted and fused based on the Gaussian weights to obtain a fused feature, including:

[0038] For each global feature: based on each feature carried in the global feature, dynamically identify the center point of the Gaussian function in the global feature;

[0039] Based on the center point of the Gaussian function and the position of the corresponding global feature in the image, determine the Gaussian weights corresponding to each global feature;

[0040] Based on the Gaussian weights of each global feature, perform weighted fusion on the corresponding global features to obtain a fused feature.

[0041] In this embodiment, the fusion network is obtained through training, so the center point of the corresponding Gaussian function can be dynamically generated based on the information of the input image. Since different global features correspond to different regions to be analyzed, for each global feature: based on the feature intensity of each feature carried in the global feature, the center point μ of the Gaussian function can be generated through formulas (1) and (2) x and μ y :

[0042]

[0043]

[0044] where M is the number of features included in the global feature, x j and y j are respectively the abscissa and ordinate of the j-th feature in the global feature in the image, strength(j) is the feature intensity of the j-th feature, and the feature intensity can be directly represented by the value of the feature. According to the above formula, the center point (μ x , μ y ) of the Gaussian function corresponding to the global feature can be calculated, where the center point refers to the region with the richest feature information. The standard deviation σ of the Gaussian function can be adaptively generated by the trained fusion network according to the scale of the global feature, which reflects the influence range of the Gaussian function. For each global feature: according to the position of each feature in the global feature in the image and the center point of the Gaussian function corresponding to the global feature calculated above, determine the Gaussian weights of each feature in the global feature, as shown in formula (3):

[0045]

[0046] where G(x,y) is the Gaussian weight of the feature (x,y), μ x and μy is the center point of the Gaussian function, and σ is the standard deviation of the Gaussian function. Therefore, for each global feature: the weights of all the features it carries can be obtained according to the position (x, y) of the feature in the image and the center point (μ x , μ y ) of the corresponding Gaussian function according to formula (3). And the feature weights closer to the center point are higher, while the feature weights farther from the center point are lower. In the present invention, the features in the significant region are naturally highlighted by the Gaussian function, and at the same time, the regions far from the center point are suppressed. By calculating the Gaussian weights of each global feature and performing weighted processing on it and the original feature, the weighted global feature is generated, and the sum of all the weighted global features can obtain the fusion feature of the image. Through the above processing, the fusion feature not only retains the global information in the image, but also highlights the feature expression in the significant region, making it have higher semantic feature expression ability to improve the accuracy of subsequent image label generation.

[0047] In another embodiment of the present invention, multiple global features are input into the fusion network of the image recognition model to determine the Gaussian weights of each global feature, and based on the Gaussian weights, the corresponding global features are Gaussian weighted and fused to obtain the fusion feature, including:

[0048] For each global feature: based on the position of the global feature in the image, determine the center point of the Gaussian function corresponding to this position; wherein, the center point of the Gaussian function is preset;

[0049] Based on the center point of the Gaussian function and the position of the corresponding global feature in the image, determine the Gaussian weights corresponding to each global feature;

[0050] Based on the Gaussian weights of each global feature, perform weighted fusion on the corresponding global features to obtain the fusion feature.

[0051] In this embodiment, the fusion network does not need to be trained. The center point of the Gaussian function can be determined in advance based on the specific task of image recognition. Exemplarily, if face recognition is performed, the center points of the eyes and mouth can be divided into the center points of the Gaussian function to highlight the feature expression of these key regions. In addition, the input image can also be segmented into multiple fixed grids, and the center of each grid can be used as the center point of the Gaussian function. By presetting multiple center points, while ensuring the accuracy and robustness of the fusion feature, it also greatly simplifies the network design of the fusion network, avoids complex training processes, and can efficiently achieve feature fusion and adapt to multiple recognition tasks. It should be noted that after determining the center point of the Gaussian function, the subsequent process of obtaining the fusion feature is the same as the relevant process of the center point obtained through training above, and will not be elaborated here.

[0052] S4. Input the fusion feature into the similarity calculation network of the image recognition model, calculate the similarity between the fusion feature and each feature in the feature library, and use the label of the feature whose similarity is greater than a preset feature threshold as the label of the image; wherein, each feature in the feature library corresponds to a preset label.

[0053] The feature library pre-stores multiple trained known features, and each feature corresponds to a preset label one by one. The label is used to characterize the category information of the corresponding feature. Input the fusion feature into the similarity calculation network, and determine the category label of the image by calculating the similarity between the fusion feature and each feature in the feature library. Among them, the calculation methods of similarity include but are not limited to cosine similarity, Euclidean distance, etc., as long as they can measure the similarity between two features, and there is no limitation here. According to the similarity calculation result, select the features greater than the feature threshold from the feature library, and use the labels corresponding to these features as the category labels of the current input image.

[0054] In an embodiment of the present invention, the step of inputting the fusion feature into the similarity calculation network of the image recognition model, calculating the similarity between the fusion feature and each feature in the feature library, and using the label of the feature whose similarity is greater than a preset feature threshold as the label of the image includes:

[0055] Calculate the similarity between the fusion feature and each feature in the feature library;

[0056] For each similarity:

[0057] Judge whether the similarity is greater than a preset similarity threshold:

[0058] If so, use the label of the corresponding feature as the label of the image;

[0059] If not, continue to compare the next similarity until all similarities are compared, and generate update information.

[0060] Input the fusion feature into the similarity calculation network, calculate the similarity between the fusion feature and all the known features stored in the feature library one by one, and after the calculation, obtain a similarity set {S1, S2, …, S N}, where S i is the similarity between the fusion feature and the i-th feature in the feature library. For each similarity S i:Compare this similarity with a preset similarity threshold to determine whether the similarity is greater than the similarity threshold. If it is greater, it means that the similarity between the fused feature and the i-th feature in the feature library is high. At this time, the label of this feature can be used as the class label of the input image for output. On the contrary, if the similarity is less than or equal to the similarity threshold, it means that the current feature does not match, and the next similarity needs to be compared. After all similarities have been compared, all qualified labels are output. If all similarities are less than or equal to the similarity threshold, it means that the current image cannot match any known class in the feature library. At this time, update information needs to be generated to prompt the operator to update the feature library. Thus, through the above method, the labels that are semantically consistent with the fused feature can be efficiently and accurately screened out to achieve precise classification. In addition, when the fused feature is similar to the features of multiple labels, multiple relevant labels can be output as the classification result, and structured label information including spatial relationships can be generated based on the relative positions of the respective labels in the image. For example, if there is a picture of a dog with cow horns, it is related to the labels of both "dog" and "cow". The historical learning knowledge based on these labels will be fused to output a more creative result. This method simulates the ability of humans to make inferences based on existing knowledge, making the classification result more flexible and intelligent.

[0061] In an embodiment of the present invention, after generating the update information, it further includes: saving the true label of the image and the fused feature of the image to the feature library to update the feature library. After generating the update information, it is necessary to manually or automatically label the true label of the current image, and save the fused feature and the true label to the feature library together. Among them, the fused feature is used as the feature vector of the new category in the feature library, and the true label and the fused feature are in one-to-one correspondence to ensure the complete mapping relationship between the feature and the category information.

[0062] In an embodiment of the present invention, the image recognition model is obtained through training. When training the image recognition model, the parameters of the feature extraction network remain unchanged. The entire image recognition model needs to be obtained through training. During the training process, the Gabor filter can be used in the feature extraction network, and its parameters are fixed and do not need to be learned. Since the parameters of each layer in the feature extraction network are fixed, the same or similar input images will go through the same processing process and finally obtain similar fusion features. Because the parameters of the feature extraction network are fixed, the focus of the learning process is concentrated on the adjustment of the similarity calculation network and the fusion network. To speed up the process, only the similarity calculation network can be trained. Therefore, the entire training process is relatively fast, and the model parameters can be quickly adjusted and adapted to new inputs on a small dataset. When processing new images, the model can gradually optimize the parameters of the similarity calculation network and the fusion network through online learning without retraining the feature extraction network. This makes the image recognition model have extremely high learning efficiency and strong online learning ability. Further, the similarity calculation network combines the output of the traditional neural network and the feature information obtained through fusion calculation, ensuring that the fusion features are closely related to the image recognition task. The label information of each input image corresponds one-to-one with its corresponding fusion feature. When images with the same label are input multiple times, the fused features will gradually store more feature differentiation values, which is similar to the mechanism in the human learning process where repeated learning makes certain neural connections thicker, and can more accurately generate the labels of the input images. Therefore, through the combination of the fixed feature extraction network and the flexible adjustment mechanism of the similarity calculation network in the present invention, it can quickly adapt to new image classification tasks, especially suitable for situations that require online learning and have a small amount of data.

[0063] Please refer to Figure 2 , the image label generation system 100 includes: an image acquisition module 110, a feature extraction module 120, a feature fusion module 130, and a label generation module 140. The above-mentioned image acquisition module 110 is used to acquire the image to be analyzed. The feature extraction module 120 is used to input the image into the feature extraction network of the image recognition model, perform multiple feature extractions on the image, and obtain multiple global features. The feature fusion module 130 is used to input the multiple global features into the fusion network of the image recognition model, determine the Gaussian weights of each global feature, and perform Gaussian weighted fusion on the corresponding global features based on the Gaussian weights to obtain fusion features. The label generation module 140 is used to input the fusion features into the similarity calculation network of the image recognition model, calculate the similarity between the fusion features and each feature in the feature library, and use the label of the feature whose similarity is greater than the preset feature threshold as the label of the image; wherein, each feature in the feature library corresponds to a preset label.

[0064] For the specific limitations of the image tag generation system, reference can be made to the limitations on the image tag generation method in the foregoing text, which will not be elaborated here. Each module in the above image tag generation system can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware format or independent of it, or stored in the memory of the computer device in software format, so as to facilitate the processor to call the operations corresponding to the above modules.

[0065] It should be noted that, in order to highlight the innovative part of the present invention, modules not closely related to solving the technical problems proposed by the present invention are not introduced in this embodiment, but this does not mean that there are no other modules in this embodiment.

[0066] Please refer to Figure 3 , the electronic device 1 may include a memory 12, a processor 13, and a bus, and may also include a computer program stored in the memory 12 and operable on the processor 13, such as an image tag generation program.

[0067] Among them, the memory 12 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as: SD or DX memory, etc.), magnetic memory, magnetic disk, optical disk, etc. The memory 12 can be an internal storage unit of the electronic device 1 in some embodiments, such as the mobile hard disk of the electronic device 1. The memory 12 can also be an external storage device of the electronic device 1 in other embodiments, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 1. Further, the memory 12 can also include both the internal storage unit and the external storage device of the electronic device 1. The memory 12 can be used not only to store the application software installed in the electronic device 1 and various types of data, such as the code for generating image tags, etc., but also to temporarily store the data that has been output or will be output.

[0068] In some embodiments, the processor 13 may be composed of an integrated circuit. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple packaged integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 13 is the control core (Control Unit) of the electronic device 1, connecting various components of the entire electronic device 1 through various interfaces and lines. By running or executing programs or modules stored in the memory 12 (such as the program for generating image tags, etc.), and by calling the data stored in the memory 12, it executes various functions of the electronic device 1 and processes data.

[0069] The processor 13 executes the operating system of the electronic device 1 and various installed application programs. The processor 13 executes the application programs to implement the steps in the above-mentioned method for generating image tags.

[0070] Exemplarily, the computer program may be divided into one or more modules. The one or more modules are stored in the memory 12 and executed by the processor 13 to complete this application. The one or more modules may be a series of computer program instruction segments capable of completing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device 1. For example, the computer program may be divided into an image acquisition module 110, a feature extraction module 120, a feature fusion module 130, and a tag generation module 140.

[0071] The above-mentioned integrated units implemented in the form of software function modules can be stored in a computer-readable storage medium. The computer-readable storage medium may be non-volatile or volatile. The above-mentioned software function modules are stored in a storage medium and include several instructions to enable a computer device (which may be a personal computer, a computer device, or a network device, etc.) or a processor to execute part of the functions of the method for generating image tags in various embodiments of this application.

[0072] In summary, a method, system, device, and medium for generating image tags disclosed by the present invention perform multiple feature extractions on an image. The generated global features contain the overall semantic information of the image, making the information expression of the image more abundant. Since the Gaussian weights are calculated based on the feature intensity or the spatial position of the feature in the image, the Gaussian weights can emphasize the more significant or critical regions in the image, thereby simulating the human attention to the significant regions. Through weighted fusion, not only the feature expression effect of the image is improved, but also the fused features have stronger adaptability in diverse tasks. Through similarity calculation, the precise matching between the fused features and the known features in the feature library is ensured, simulating the ability of humans to quickly generalize new information based on existing knowledge during the reasoning process, improving the classification accuracy, and reducing the probability of misclassification. This improves the problem that the prior art often relies on shallow feature extraction and complex feature fusion, but lacks the efficient feature summarization and reasoning capabilities like humans. In the feature extraction network of the present invention, the human eye texture feature-based human-like learning method is fixed in advance, greatly accelerating the learning speed of the network. Based on the human online learning ability, the proposed similarity calculation network structure has online learning ability and the function of human-like neuron learning, which will significantly improve the current neural network learning ability. Therefore, the present invention effectively overcomes various disadvantages in the prior art and has high industrial utilization value.

[0073] The above embodiments are only illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.

Claims

1. A method for generating an image label, characterized in that: The method comprises: Acquire an image to be analyzed; Inputting the image into a feature extraction network of an image recognition model, performing multiple feature extractions on the image to obtain multiple global features; Inputting multiple global features into the fusion network of the image recognition model, determining the Gaussian weight of each global feature, and performing Gaussian weighted fusion on the corresponding global features based on the Gaussian weight to obtain a fusion feature; The fused features are input into the similarity calculation network of the image recognition model, the similarity between the fused features and each feature in the feature library is calculated, and the labels of the features whose similarity is greater than a preset feature threshold are used as labels of the image; wherein each feature in the feature library corresponds to a preset label.

2. The method for generating an image tag according to claim 1, characterized in that: The image is input into a feature extraction network of an image recognition model, and multiple feature extractions are performed on the image to obtain multiple global features, including: Using the image as the initial area to be analyzed; Performing feature extraction on the area to be analyzed based on the filter of the first scale of the feature extraction network, obtaining and saving the global features of the area to be analyzed; Extracting features of the region to be analyzed based on a filter of a second scale of the feature extraction network, and obtaining and saving local features of the region to be analyzed; Detect whether there is a feature value greater than a preset feature threshold in the local features, and if so, use the area in the image corresponding to the feature value as a new area to be analyzed, and extract the global features and local features of the area to be analyzed again until there is no feature value greater than the feature threshold in the obtained local features.

3. The method for generating an image tag according to claim 1, characterized in that: The step of inputting a plurality of global features into the fusion network of the image recognition model, determining the Gaussian weights of the respective global features, and performing Gaussian weighted fusion on the corresponding global features based on the Gaussian weights to obtain fused features includes: For each global feature: Based on each feature carried in the global feature, dynamically identify the center point of the Gaussian function in the global feature; Determine the Gaussian weight corresponding to each global feature based on the center point of the Gaussian function and the position of the corresponding global feature in the image; Based on the Gaussian weights of each global feature, the corresponding global features are weighted fused to obtain the fused features.

4. The method for generating an image tag according to claim 1, characterized in that: The step of inputting a plurality of global features into the fusion network of the image recognition model, determining the Gaussian weights of the respective global features, and performing Gaussian weighted fusion on the corresponding global features based on the Gaussian weights to obtain fused features includes: For each global feature: based on the position of the global feature in the image, determine the center point of the Gaussian function corresponding to the position; wherein the center point of the Gaussian function is preset; Determine the Gaussian weight corresponding to each global feature based on the center point of the Gaussian function and the position of the corresponding global feature in the image; Based on the Gaussian weights of each global feature, the corresponding global features are weighted fused to obtain the fused features.

5. The method for generating an image tag according to claim 1, characterized in that: The step of inputting the fused features into the similarity calculation network of the image recognition model, calculating the similarity between the fused features and each feature in the feature library, and taking the labels of the features whose similarity is greater than a preset feature threshold as the labels of the image includes: Calculating the similarity between the fused feature and each feature in the feature library; For each similarity: Determine whether the similarity is greater than the preset similarity threshold: If so, the label of the corresponding feature is used as the label of the image; If not, continue to compare the next similarity until all similarities are compared and update information is generated.

6. The method for generating an image tag according to claim 5, characterized in that: After the update information is generated, the method further includes: saving the real label of the image and the fusion feature of the image into the feature library, and updating the feature library.

7. The method for generating an image tag according to claim 1, characterized in that: The image recognition model is obtained through training. When the image recognition model is trained, the parameters of the feature extraction network remain unchanged.

8. A system for generating image labels, characterized in that: The system comprises: An image acquisition module, used for acquiring an image to be analyzed; A feature extraction module, used for inputting the image into a feature extraction network of an image recognition model, performing multiple feature extractions on the image, and obtaining multiple global features; A feature fusion module, used to input multiple global features into the fusion network of the image recognition model, determine the Gaussian weight of each global feature, and perform Gaussian weighted fusion on the corresponding global features based on the Gaussian weight to obtain a fusion feature; A label generation module is used to input the fused features into the similarity calculation network of the image recognition model, calculate the similarity between the fused features and each feature in the feature library, and use the label of the feature whose similarity is greater than a preset feature threshold as the label of the image; wherein each feature in the feature library corresponds to a preset label.

9. An electronic device, characterized in that: The electronic device comprises: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the electronic device to implement the method for generating an image tag as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor of a computer, the computer is caused to execute the method for generating an image tag according to any one of claims 1 to 7.

Citation Information

Cited By

  • Data label automatic generation method and system

    CN120508830A