Image tag identification method and related equipment

Through multiple feature extraction and fusion, the problem of difficulty in identifying composite tags in the prior art is solved, and higher recognition accuracy and flexibility are achieved.

CN119992204APending Publication Date: 2025-05-13GUANGZHOU SHANGYUN NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510101358.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing image tag recognition technology is difficult to accurately identify composite tags.

Method used

By acquiring the visual features of the image, using the label recognition model to perform multiple different feature extractions, multiple single feature vectors and single tags are generated, and the composite tag of the image is determined by fusing these feature vectors.

Benefits of technology

It significantly improves the accuracy and flexibility of image tag recognition, and can more accurately identify composite tags.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992204A_ABST
    Figure CN119992204A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image tag identification method and related equipment. The method comprises the following steps: acquiring visual features of an image; performing multiple times of different feature extraction on the visual features by using a tag identification model to obtain a plurality of single feature vectors corresponding to the visual features, and determining a single tag corresponding to each single feature vector; fusing the plurality of single feature vectors by using the tag identification model to obtain a fused feature vector; based on the fusion feature vector, determining a composite label corresponding to the image; and obtaining a label identification result of the image according to the single label and the composite label. According to the invention, the accuracy of tag identification on the image can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of artificial intelligence technology, relates to image recognition technology, and in particular to an image label recognition method and related equipment. Background Art

[0002] Image label recognition technology is a technology that can identify relevant labels from images and is widely used in various fields. However, the relevant image label recognition technology is mainly based on convolutional neural networks. Although the network structure used by this technology is relatively simple and easy to implement, it is difficult to apply to the recognition of complex, non-explicit composite labels in practical applications. Therefore, the accuracy of the relevant image label recognition technology still needs to be improved. Summary of the invention

[0003] The embodiments of the present application provide an image label recognition method and related equipment, which can solve the problem that the existing image label recognition technology is difficult to accurately recognize the composite labels of images.

[0004] The first aspect of an embodiment of the present application provides an image label recognition method, comprising: acquiring visual features of an image; performing multiple different feature extractions on the visual features using a label recognition model to obtain multiple single feature vectors corresponding to the visual features, and determining a single label corresponding to each single feature vector; using the label recognition model to fuse the multiple single feature vectors to obtain a fused feature vector; based on the fused feature vector, determining a composite label corresponding to the image; and obtaining a label recognition result of the image based on the single label and the composite label.

[0005] According to an embodiment of the present application, obtaining the visual features of the image includes: dividing the image into multiple block images, performing linear projection on each block image, and obtaining projection features corresponding to each block image; performing global mapping on the projection features corresponding to all block images to obtain mapping features; performing context feature extraction on the mapping features to obtain context features corresponding to the image; and performing layer normalization processing, nonlinear transformation processing, and downsampling processing on the context features in sequence to obtain the visual features.

[0006] According to an embodiment of the present application, the label recognition model includes multiple first modules, and each feature extraction in multiple different feature extractions of the visual features includes: using each first module to perform a single feature extraction on the visual feature to obtain a single feature vector corresponding to the visual feature; and performing pooling processing, feature abstraction processing, straightening processing and linear activation operations on each single feature vector in turn to obtain the single label.

[0007] According to an embodiment of the present application, using each first module to perform a single feature extraction on the visual feature to obtain a single feature vector corresponding to the visual feature includes: using a convolution kernel of a first size to perform feature abstraction processing on the visual feature to obtain a first feature, wherein the feature abstraction processing includes convolution processing, normalization processing and nonlinear transformation processing; using a convolution kernel of a second size to perform feature abstraction processing on the first feature to obtain a second feature, wherein the second size is larger than the first size; using the convolution kernel of the first size to perform feature abstraction processing on the second feature to obtain a third feature, and using the third feature as the single feature vector.

[0008] According to an embodiment of the present application, the label recognition model includes a second module, and the use of the feature label recognition model to fuse the multiple single feature vectors to obtain a fused feature vector includes: using the second module to fuse the multiple single feature vectors to obtain a first fused tensor; performing dimension reorganization processing on the first fused tensor to obtain a second fused tensor, and the dimension of the second fused tensor is smaller than the dimension of the first fused tensor; performing feature fusion on the second fused tensor based on the converter structure of the second module to obtain a first fused feature; performing dimension reorganization processing on the first fused feature to obtain a second fused feature; performing average pooling processing on the second fused feature to obtain a third fused feature; flattening processing and multiple linear transformation processing on the third fused feature to obtain the fused feature vector.

[0009] According to an embodiment of the present application, the flattening processing and multiple linear transformation processing of the third fusion feature to obtain the fusion feature vector includes: flattening the third fusion feature to obtain a fourth fusion feature; linearly transforming the fourth fusion feature to obtain a fifth fusion feature; linearly transforming the fifth fusion feature to obtain a sixth fusion feature, and the dimension of the sixth fusion feature is smaller than the dimension of the fifth fusion feature; linearly transforming the sixth fusion feature to obtain a seventh fusion feature, and using the seventh fusion feature as the fusion feature vector.

[0010] According to an embodiment of the present application, the composite label includes multiple labels, and the dimension of the fused feature vector is equal to the number of labels included in the composite label.

[0011] According to an embodiment of the present application, determining the composite label corresponding to the image based on the fused feature vector includes: processing the fused feature vector using a preset activation function to obtain the composite label.

[0012] According to an embodiment of the present application, the label recognition model includes a first module and a second module, and the method also includes training the first module and the second module using a supervised training method, including: determining a first loss value of the first module based on a first loss function, and determining a second loss value of the second module based on a second loss function; training the first module according to the first loss value, and training the second module according to the second loss value.

[0013] According to a second aspect of an embodiment of the present application, there is provided an image label recognition device, the device comprising: an acquisition module for acquiring visual features of an image; a determination module for performing multiple different feature extractions on the visual features using a label recognition model to obtain multiple single feature vectors corresponding to the visual features, and determining a single label corresponding to each single feature vector; a fusion module for fusing the multiple single feature vectors using the label recognition model to obtain a fused feature vector; the determination module is further used to determine a composite label corresponding to the image based on the fused feature vector; and an output module for obtaining a label recognition result of the image based on the single label and the composite label.

[0014] A third aspect of an embodiment of the present application provides an electronic device, including: a memory, and a processor, wherein the processor executes computer-readable instructions stored in the memory to implement the image tag recognition method.

[0015] The image label recognition method provided by the embodiment of the present application can efficiently extract visual features based on the image input by the user, and perform feature extraction in multiple different ways through the label recognition model to generate multiple single feature vectors and corresponding single labels. These single feature vectors are further fused to form a fused feature vector, and the composite label corresponding to the image is determined accordingly. Combining single labels with composite labels, the label recognition results of the image can be obtained comprehensively and accurately, significantly improving the accuracy and flexibility of label recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0017] Figure 1 A schematic diagram of an application environment of an image tag recognition method provided in an embodiment of the present application.

[0018] Figure 2 A flowchart of an image tag recognition method provided in an embodiment of the present application.

[0019] Figure 3 A flowchart of a method for acquiring visual features of an image provided in an embodiment of the present application.

[0020] Figure 4 A flowchart of a method for determining a single label corresponding to each single feature vector provided in an embodiment of the present application.

[0021] Figure 5 A flowchart of a method for extracting a single feature from a visual feature provided in an embodiment of the present application.

[0022] Figure 6 This is an example diagram of the internal structure flow of the first module provided in an embodiment of the present application.

[0023] Figure 7 A schematic diagram of a flow chart of a method for fusing multiple single feature vectors provided in an embodiment of the present application.

[0024] Figure 8 This is an example diagram of the internal structure flow of the second module provided in an embodiment of the present application.

[0025] Fig. 9 This is an example diagram of a detailed process of step S706 provided in an embodiment of the present application.

[0026] Fig.10 This is an example diagram of the image tag recognition process provided in an embodiment of the present application.

[0027] Fig.11 This is an example diagram of the training process of the graph label recognition model provided in an embodiment of the present application.

[0028] Fig.12 A functional block diagram of an image tag recognition device provided in an embodiment of the present application.

[0029] Fig.13 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0030] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application is described in detail below with reference to the accompanying drawings and specific embodiments.

[0031] It should be noted that in this application, "at least one" means one or more, and "more than one" means two or more than two. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and drawings of this application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0032] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete way. The following embodiments and features in the embodiments may be combined with each other without conflict.

[0033] See also Figure 1 , is a schematic diagram of an application environment of an image tag recognition method provided by an embodiment of the present application. Figure 1 As shown, the user terminal 10 communicates with the server 20 through a network. The network can be a wired network communication or a wireless network communication. The wired network can be any one of a local area network, a metropolitan area network and a wide area network, a wireless fidelity (Wireless Fidelity, Wi-Fi), a self-organizing network wireless communication (ZigBee Wireless Networks, ZigBee) technology, an ultra-wideband (Ultra Wideband, UWB) technology, a wireless universal serial bus (Universal Serial Bus, USB) and the like.

[0034] The user terminal 10 may be an electronic device such as a mobile phone, a tablet computer, a multimedia player, a personal computer (PC), a wearable device, etc. The user terminal 10 may be a client with various applications installed, such as education applications, consulting applications, information broadcast applications, live broadcast applications, e-commerce applications, etc.

[0035] The server 20 is used to provide background services for the application in the user terminal 10. For example, the server 20 can be the background server of the above-mentioned e-commerce application. The server 20 can be an electronic device. In one embodiment of the present application, it can be a server, a server cluster composed of multiple servers, or a cloud computing service center.

[0036] Exemplarily, an e-commerce application is used as an example for explanation. The user performs an image input operation in a client (e.g., user terminal 10) on which an e-commerce application is installed, and the user terminal 10 sends the image to the server 20. The server 20 obtains the visual features of the image, and uses a label recognition model to perform multiple different feature extractions on the visual features to obtain multiple single feature vectors corresponding to the visual features, and determines a single label corresponding to each single feature vector. The server 20 uses a label recognition model to fuse multiple single feature vectors to obtain a fused feature vector. The server 20 determines the composite label corresponding to the image based on the fused feature vector. The server 20 obtains the label recognition result of the image based on the single label and the composite label. The server 20 sends the label recognition result to the user terminal 10, and the user terminal 10 presents the label recognition result to the user using a graphical user interface.

[0037] In this way, the server 20 can efficiently extract visual features based on the image input by the user, and perform multiple feature extractions in different ways through the label recognition model to generate multiple single feature vectors and corresponding single labels. These single feature vectors are further fused to form a fused feature vector, and the composite label corresponding to the image is determined accordingly. Combining single labels with composite labels can comprehensively and accurately obtain the label recognition results of the image, significantly improving the accuracy and flexibility of label recognition.

[0038] The following will be a computer program product running on a server (such as Figure 1 The server 20 shown in FIG. Figure 2 FIG. 1 is a flow chart of an image tag recognition method provided in an embodiment of the present application. In an embodiment of the present application, the method includes the following steps:

[0039] Step S201, obtaining visual features of an image.

[0040] In some embodiments of the present application, the image includes any image that the user determines to have a tag identified. In one example, the image may be an image of a product. If the product is clothing or apparel, the image may also be an image of a model wearing the product.

[0041] In some embodiments of the present application, an image may include a single object or multiple objects, and label recognition of the image may identify a single label (or single-type label) of a single object in the image. In one example, a single label usually has direct, clear and unique features, such as the brand name, specific color (such as "red"), material type (such as "cotton"), design style, etc. of the clothes shown in the image, and each label independently describes a certain attribute of the object.

[0042] In some embodiments of the present application, if an image contains multiple objects or a single object contained in the image combines multiple attributes, label recognition of the image can generate a composite label (or composite label). Compared with the intuitiveness and uniqueness of a single label, the feature combination of a composite label is more complex and non-obvious. Composite labels are usually based on single labels and use the correlation between labels to comprehensively or specifically describe the object. For example, for an image of clothing, "red-cotton-shirt" constitutes a composite label that combines the information of three single labels: color ("red"), material ("cotton"), and possible category ("shirt") to jointly describe the characteristics of the clothing.

[0043] In some embodiments of the present application, in order to realize label recognition of an image, the visual features (Visual Feature, VF) of the image can be extracted based on a preset computer vision model, and then the visual features can be used to predict and recognize the label. Among them, the computer vision model can include but is not limited to one or more of the following models: ConvNeXt (Convolutional Next), ResNet (Residual Network), VGG (Visual Geometry Group), Inception, ViT (Vision Transformer), Swin Transformer, etc. In one example, the method for obtaining the visual features of an image can refer to Figure 3 Flowchart shown.

[0044] In some embodiments of the present application, the image may be preprocessed before the visual features of the image are acquired. For example, the preprocessing includes one or more of the following processes: resizing and normalization.

[0045] In one example, resizing means unifying the size of the image to a preset dimension, and resizing can be achieved by scaling, cropping, or padding edges, etc. In this way, the resized images can have the same resolution and meet the size of the input image required by the computer vision model.

[0046] In one example, normalization processing means adjusting the pixel values ​​of an image to a specific range (e.g., 0 to 1, or -1 to 1). For example, if the image is an 8-bit image, normalization processing can be achieved by dividing the pixel values ​​of the image by 255 (the maximum pixel value of an 8-bit image). Alternatively, normalization processing can be achieved by a more complex normalization processing method, including but not limited to minimum-maximum normalization processing or Z-score normalization processing. In this way, it can help the model converge faster and improve the generalization ability of the model.

[0047] Step S202 , using the label recognition model to perform multiple different feature extractions on the visual features, obtain multiple single feature vectors corresponding to the visual features, and determine a single label corresponding to each single feature vector.

[0048] In some embodiments of the present application, the label recognition model includes multiple first modules, each of which includes but is not limited to multiple convolutional layers, normalization processing, nonlinear transformation processing, average pooling layers and straightening layers, and activation functions. These network layers and algorithms work together to input visual features, further abstract and extract the visual features, extract a single feature vector and determine the corresponding single label.

[0049] In some embodiments of the present application, the visual feature can be copied to obtain multiple identical visual features, and each first module can be used to extract features from each visual feature, thereby achieving multiple different feature extractions of the visual feature to obtain multiple single feature vectors. In one example, the tag recognition model includes N first modules, where N represents a preset positive integer greater than 1, and the first module can be a feature expert module. The visual feature can be copied to obtain N visual features, and each visual feature is input into a first module, and the corresponding feature vector is obtained through the output of the first module, and a total of N feature vectors are obtained.

[0050] In one example, a method for determining a single label corresponding to each single feature vector can refer to Figure 4 Flowchart shown.

[0051] In some embodiments of the present application, all single feature vectors obtained by all first modules constitute a single feature vector set FEs. In addition, each single feature vector obtained by the first module will be copied and input into the second module of the label recognition model, so that the recognition of composite labels based on single labels can be realized according to the single feature vector. For the introduction of the second module, please refer to the description in the subsequent embodiments.

[0052] Based on the above embodiment, a label recognition model including multiple first modules can be used to perform multiple different feature extractions on visual features to obtain multiple corresponding single feature vectors and their single labels. This process is achieved by copying the visual features and inputting them into each first module respectively, wherein the first module uses algorithms such as convolutional layers, normalization processing, and nonlinear transformation to deeply abstract the visual features. All single feature vectors constitute a feature vector set, which lays the foundation for subsequent composite label recognition based on single labels.

[0053] Step S203: using the label recognition model to fuse multiple single feature vectors to obtain a fused feature vector.

[0054] In some embodiments of the present application, the tag recognition model includes a second module, which includes but is not limited to a feature fusion layer, a dimensionality reorganization layer, a converter structure layer, a second dimensionality reorganization layer, an average pooling layer, a flattening layer, and multiple linear transformation layers. These layers or operations work together to extract and fuse useful information from multiple single feature vectors, establish associations between multiple single features, and generate a fused feature vector. In an example, the method of fusing multiple single feature vectors can refer to Figure 7 Flowchart shown.

[0055] Based on the above embodiment, useful information can be extracted and integrated from the third fused feature through flattening and multiple linear transformations, and finally a fused feature vector is obtained. Specifically, the multi-dimensional third fused feature is first converted into a one-dimensional fourth fused feature through a flattening operation; then, the fourth fused feature is input into the first linear transformation layer for feature extraction and integration to obtain the fifth fused feature; then, in order to reduce the computational complexity and avoid overfitting, the fifth fused feature is input into the second linear transformation layer for dimensionality reduction to obtain the sixth fused feature; finally, the sixth fused feature is input into the third linear transformation layer to obtain a fused feature vector that integrates multiple single feature vector information. This process not only effectively reduces the dimension of the feature, but also improves the representation ability and robustness of the feature, providing strong support for subsequent tasks such as label recognition.

[0056] Step S204: determining the composite label corresponding to the image based on the fused feature vector.

[0057] In some embodiments of the present application, the fused feature vector is processed using a preset activation function to obtain a composite label. For example, the activation function may include but is not limited to a Sigmoid function. The composite label includes multiple labels, and the number of labels included in the composite label is equal to the dimension of the fused feature vector.

[0058] In one example, the fused feature vector can be input into a neural network layer containing an activation function, and the probability value corresponding to the label feature contained in the fused feature vector can be obtained by transforming the activation function. According to the size of the label probability value, it can be determined whether each label exists in the composite label corresponding to the fused feature vector. For example, a probability threshold (such as 0.5) is set. When the probability value of a label exceeds the threshold, the label is considered to exist; otherwise, the label is considered not to exist. Among them, the choice of threshold can be adjusted according to the specific task and data characteristics to obtain the best classification effect. All determined labels are combined together to obtain a composite label, where the order of labels in the composite label can be sorted as needed or remain unchanged.

[0059] Based on the above embodiment, the fused feature vector can be processed using a preset activation function to obtain a composite label, which has significant advantages in multi-label classification tasks. By selecting a suitable activation function and label determination strategy, the category to which the instance belongs can be accurately determined, providing strong support for practical applications.

[0060] Step S205, obtaining a label recognition result of the image according to the single label and the compound label.

[0061] In some embodiments of the present application, a set of single labels and compound labels may be used as a label recognition result of an image. In this way, the label recognition result includes all possible label recognition results of the image, thereby achieving accurate label recognition of the image.

[0062] The tag identification method provided in the embodiment of the present application can be widely used in scenarios using secondary composite and event behavior tags, such as content review and compliance checking, search recommendations and intelligent classification of e-commerce platforms, supply chain management, virtual fitting, smart home and Internet of Things scenarios, and can significantly improve the accuracy, efficiency and intelligence level of product identification, thereby creating greater value for users and merchants.

[0063] In some embodiments of the present application, the label recognition model includes a first module and a second module, and the first module and the second module can be trained using a supervised training method, including: determining a first loss value of the first module based on a first loss function, determining a second loss value of the second module based on a second loss function, training the first module according to the first loss value, and training the second module according to the second loss value.

[0064] In one example, the first training data can be used to train the first module, and the first training data includes a sample image and a true single label corresponding to the sample image. The true single label can be obtained using an open source single label recognition model, or can be obtained by receiving user input. The sample visual features corresponding to the sample image can be input into the first module, and the first module is used to obtain a predicted single label based on the sample visual features. The difference between the true single label and the predicted single label is determined based on the first loss function, and the difference is used as the first loss value of the first module. The first loss value reflects the similarity or difference between the single label predicted by the first module and the true single label. The larger the first loss value, the worse the performance of the first module. If the first loss value is greater than the preset first loss threshold, the back propagation algorithm and optimizer (such as Adam, etc.) can be used to update the weight parameters of the first module to minimize the first loss value. Through multiple iterative training, the first module will gradually learn how to identify and obtain accurate single labels.

[0065] In one example, the second module can be trained using the second training data, the second training data including the sample image and the true compound label corresponding to the sample image. The true compound label can be obtained by receiving user input. The method of training the second module using the second training data is similar to the method of training the first module using the first training data, and the purpose is to minimize the second loss value so that the second module will gradually learn how to identify and obtain accurate compound labels.

[0066] Based on the above embodiment, the label recognition model can be trained by a supervised training method, so that the model can accurately recognize single labels and compound labels of images.

[0067] In some embodiments of the present application, reference Figure 3 As shown, the method for obtaining visual features of an image includes the following process.

[0068] Step S301 : dividing the image into a plurality of block images, performing linear projection on each block image, and obtaining projection features corresponding to each block image.

[0069] In some embodiments of the present application, taking the use of the ConvNext model to obtain the visual features of an image as an example, the ConvNext model can divide the image into multiple overlapping or non-overlapping block images (or image blocks or patches). The size of the block image can be a preset fixed value or dynamically determined according to actual needs. A linear projection operation can be applied to each block image through a predefined weight matrix (for example, the weight of a convolution kernel or a fully connected layer) to extract low-level features of each block image and obtain projection features corresponding to each block image.

[0070] Step S302 , globally mapping the projection features corresponding to all block images to obtain mapping features.

[0071] In some embodiments of the present application, the ConvNext model globally maps the projection features corresponding to all block images, indicating that all local features are integrated into a global feature vector. This can be achieved through splicing, average pooling, maximum pooling or other aggregation strategies. Fusion of local features into global features can facilitate subsequent steps to capture the overall structural information of the image.

[0072] Step S303: extract context features from the mapped features to obtain context features corresponding to the image.

[0073] In some embodiments of the present application, a convolutional neural network, a Transformer or other structure can be used to extract context features from the globally mapped features, so as to capture the spatial and temporal relationships between features, as well as higher-level abstract information. For example, the context features in an image can reflect the relationship between objects in the image, the scene layout, etc., which helps to understand the image content.

[0074] Step S304, performing layer normalization processing, nonlinear transformation processing and down-sampling processing on the context features in sequence to obtain visual features.

[0075] In some embodiments of the present application, layer normalization processing includes layer normalization (LayerNormalization) of context features, which can reduce internal covariate shift. Nonlinear transformation processing includes applying nonlinear activation functions (such as ReLU, GELU, etc.) to transform the normalized features to increase nonlinear expression capabilities. Downsampling processing includes downsampling the features after nonlinear transformation (such as average pooling, maximum pooling, or through compensation parameters in convolution operations), which can reduce the spatial dimension of the features while retaining key information.

[0076] In some embodiments of the present application, the visual features obtained in the above manner may be referred to as visual feature vectors, which may be used for subsequent tasks such as image label recognition.

[0077] Based on the above embodiments, the visual features of the image are effectively extracted through steps such as image block division, linear projection, global mapping, context feature extraction, and feature processing and output. The visual features of the image not only contain the local detail information of the image, but also integrate the global context information, which can provide strong support for subsequent image analysis.

[0078] In some embodiments of the present application, reference Figure 4 As shown, the method for determining a single label corresponding to each single feature vector includes the following process.

[0079] Step S401 : Utilize each first module to perform single feature extraction on the visual feature to obtain a single feature vector corresponding to the visual feature.

[0080] In some embodiments of the present application, the first module extracts a single feature of the visual feature through multiple convolution layers (including convolution kernels), normalization processing, nonlinear transformation processing, average pooling layer, and straightening layer to obtain a single feature vector of the visual feature. In one example, the method for extracting a single feature of the visual feature can refer to Figure 5 Flowchart shown.

[0081] Step S402, performing pooling processing, feature abstraction processing, straightening processing and linear activation operations on each single feature vector in sequence to obtain a single label.

[0082] In some embodiments of the present application, a single feature vector is subjected to average pooling (AvgPooling) and dimensionality reduction processing to obtain a fourth feature. The fourth feature is subjected to feature abstraction processing (hereinafter referred to as CBA processing) and dimensionality reduction processing using a convolution kernel of a first size to obtain a fifth feature. The fifth feature is flattened to obtain a sixth feature. The sixth feature is subjected to linear transformation processing using a preset linear activation function to obtain a linear transformation result.

[0083] In one example, reference Figure 6 As shown, it is an example diagram of the internal structure flow of the first module provided by an embodiment of the present application. The dimension of a single feature vector VF3 can be NF / / 4xHxW. Among them, NFxHxW represents the dimension or number of channels of the visual feature VF, for example, NF is 1024, H and W are the height and width of the feature map, for example, H=W=8, NF / / 4 means that NF is divisible by 4. After performing average pooling on the single feature vector, the dimension of the fourth feature VF4 obtained can be NF / / 4x1x1, and the spatial dimension of the feature can be reduced by the pooling operation.

[0084] In some embodiments of the present application, CBA processing refers to processing using a combination of convolution kernel (Conv) + batch normalization (BatchNormalization) + activation function. CBA processing includes convolution processing, normalization processing and nonlinear transformation processing, wherein the activation function is used to perform nonlinear transformation processing. The first size refers to the size of the convolution kernel (kernel size), for example, 1x1. Figure 6 As shown, VF4 is processed using CBA with a kernel size of 1x1 to obtain the fifth feature VF5 with a dimension of NF / / 8x1x1, thereby further extracting features and reducing the dimension of features.

[0085] In some embodiments of the present application, reference Figure 6 As shown, a Flatten operation is performed on VF5 to obtain the sixth feature VF6 with a dimension of NF / / 8, thereby converting the feature from a multi-dimensional space to a one-dimensional space for easy subsequent processing.

[0086] In some embodiments of the present application, VF6 is input into a linear activation operation, for example, the linear activation operation may be a sigmoid activation function or a softmax activation function, to obtain a single type label. The dimension of the single label result may be NF / / 8, and the dimension of all single label results output by all first modules may be NF / / 8->nc, where nc represents the number of categories of the single label corresponding to the first module.

[0087] Based on the above embodiment, the first module performs deep feature extraction on the visual features to obtain a single feature vector, and then performs average pooling processing on these vectors in turn to reduce the spatial dimension, uses convolution kernels to perform feature abstraction and dimensionality reduction processing, and straightens the features to convert them into one-dimensional space. Finally, a linear transformation is performed through a preset linear activation function to obtain a single label corresponding to each first module. This process can effectively realize the mapping from complex visual features to concise labels, and improve the accuracy and efficiency of label determination.

[0088] In some embodiments of the present application, reference Figure 5 As shown, the method for extracting a single feature from a visual feature includes the following process.

[0089] Step S501, using a convolution kernel of a first size to perform feature abstraction processing on the visual feature to obtain a first feature.

[0090] In some embodiments of the present application, reference Figure 6 As shown, the input data of the first module is the visual feature VF, for example, the dimension of VF is NFxHxW, and the dimension of VF1 is NF / / 2xHxW. Among them, NF is the input feature size, and different models have different NFs. For example, if the model used is ConvNext, NF is 1024. H and W are the height and width of the feature map, for example, H=W=8.

[0091] The first module uses a convolution kernel of a first size to perform feature abstraction processing (hereinafter referred to as CBA processing) on ​​the VF. For example, the first size represents the size of the convolution kernel (kernel size), such as 1x1. Among them, CBA processing means using convolution kernel (Conv) + batch normalization (Batch Normalization) + activation function to perform processing. CBA processing includes convolution processing, normalization processing and nonlinear transformation processing, in which the activation function is used for nonlinear transformation processing. Reference Figure 6 As shown, the output data of CBA processing is the first feature VF1, and the dimension is NF / / 2xHxW, where NF / / 2 means that NF is divisible by 2, thus achieving dimensionality reduction and preliminary abstraction of features.

[0092] Step S502: Use a convolution kernel of a second size to perform feature abstraction processing on the first feature to obtain a second feature.

[0093] In some embodiments of the present application, the first feature VF1 is further abstracted and extracted using a convolution kernel of a second size to obtain a second feature VF2. The second size is larger than the first size. For example, the first size is 1x1 and the second size is 3x3. The feature abstraction process represents the above-mentioned CBA process, and further extracts features through a larger convolution kernel. For example, referring to Figure 6 As shown, the dimension of the second feature VF2 is 512xHxW.

[0094] Step S503: Use a convolution kernel of the first size to perform feature abstraction processing on the second feature to obtain a third feature, and use the third feature as a single feature vector.

[0095] In some embodiments of the present application, the second feature VF2 is further abstracted and extracted using a convolution kernel of the first size to obtain a third feature VF3. For example, the second size is 3x3. For example, the dimension of VF2 is 512xHxW. Figure 6 As shown, VF2 is processed using CBA with a first size of 1x1 to obtain VF3 as a single feature vector with a dimension of NF / / 4xHxW, thereby achieving further integration and dimensionality reduction of features.

[0096] Based on the above embodiment, by using convolution kernels of different sizes in sequence to perform feature abstraction processing, first, a first-size convolution kernel is used to perform preliminary dimensionality reduction and abstraction to obtain a first feature, and then a second-size (larger than the first size) convolution kernel is used to further extract the feature to obtain a second feature. Finally, the first-size convolution kernel is used to abstract and integrate the second feature again to obtain a third feature as a single feature vector, thereby effectively realizing deep abstraction, dimensionality reduction and integration of features.

[0097] In some embodiments of the present application, reference Figure 7 As shown, the method for fusing multiple single feature vectors includes the following process.

[0098] Step S701, using the second module to fuse multiple single feature vectors to obtain a first fused tensor.

[0099] In some embodiments of the present application, the feature fusion layer of the second module fuses multiple single feature vectors. For example, the feature fusion layer can use a feature fusion algorithm to integrate these vectors to obtain a first fusion tensor. Among them, the feature fusion algorithm includes but is not limited to weighted average, maximum pooling, splicing, etc.

[0100] Following the above example, refer to Figure 8As shown, it is an example diagram of the internal structure flow of the second module provided in an embodiment of the present application. Collect each single feature vector FEs from the N first modules, and the dimension of each single feature vector is NF / / 4xHxW. For example, N is set to 16, indicating that there are 16 first modules and N single feature vectors. The collected N single feature vectors are merged into a four-dimensional tensor as the first fusion tensor FEs1, with a dimension of NxNF / / 4xHxW.

[0101] Step S702, re-dimensionalize the first fused tensor to obtain a second fused tensor.

[0102] In some embodiments of the present application, the dimension resizing layer of the second module performs dimension resizing processing (resize) on the first fused tensor to obtain a second fused tensor with a dimension smaller than that of the first fused tensor. For example, the dimension resizing processing may include but is not limited to adjusting the height, width, and number of channels (for image data) of the tensor or adjusting other dimensions, thereby changing the shape or dimension of the tensor to adapt to the input requirements of subsequent layers or specific algorithms.

[0103] Following the above example, refer to Figure 8 As shown, the first fused tensor Fes1 with a dimension of NxNF / / 4xHxW is reshaped to obtain a three-dimensional tensor as the second fused tensor FEs2 with a dimension of NxHWxNF / / 4, where HW is equal to the product of H and W.

[0104] Step S703: Perform feature fusion on the second fused tensor based on the converter structure of the second module to obtain a first fused feature.

[0105] In some embodiments of the present application, the transformer structure of the second module performs feature fusion on the second fused tensor. The transformer structure may include a self-attention mechanism and a feedforward neural network, which can capture the long-distance dependency in the second fused tensor and perform effective feature extraction to obtain the first fused feature.

[0106] Following the above example, refer to Figure 8 As shown, the second fused tensor FEs2 is input into the Transformer structure, which adopts a variant of the Llama model, including front RMS layer normalization, RoPE and other related technologies. The number of self-attention heads of the Transformer structure is equal to the number N of the first module. The Transformer module contains a total of M layers, for example, M is set to 8. The Transformer structure outputs the mixed and fused feature vector FE1 as the first fused feature, with a dimension of HWxNF / / 8.

[0107] Step S704: perform dimension reorganization on the first fused feature to obtain a second fused feature.

[0108] In some embodiments of the present application, the second dimension reorganization layer of the second module is similar to the above-mentioned dimension reorganization layer. This layer performs dimension reorganization on the first fusion feature to obtain the second fusion feature to meet the input requirements of the subsequent layer.

[0109] Following the above example, refer to Figure 8 As shown, the first fused feature FE1 is dimensionally reshaped to obtain the second fused feature FE2 with a dimension of NF / / 8xHxW.

[0110] Step S705, performing average pooling processing on the second fused features to obtain a third fused feature.

[0111] In some embodiments of the present application, the average pooling layer of the second module performs average pooling on the second fused features to obtain third fused features, thereby reducing the spatial dimension of the features, extracting global information, and increasing the robustness of the features.

[0112] Following the above example, refer to Figure 8 As shown, the second fusion feature FE2 is averaged and pooled to obtain the third fusion feature FE3, with a dimension of NF / / 8x1x1.

[0113] Step S706, flattening and multiple linear transformations are performed on the third fused feature to obtain a fused feature vector.

[0114] In some embodiments of the present application, the flattening layer of the second module converts the third fused feature from a multidimensional space into a one-dimensional space, that is, flattens it into a one-dimensional vector, so that the feature is input into a subsequent fully connected layer or a linear transformation layer. The linear transformation layer of the second module performs multiple linear transformations on the flattened features, such as multiplying the flattened features by a weight matrix and adding a bias vector, and then applying a nonlinear activation function (such as ReLU, sigmoid, etc.), so that the features can be further extracted and integrated to obtain a fused feature vector.

[0115] Based on the above embodiment, the second module realizes the fusion of multiple single feature vectors through a series of operations: first, the feature fusion layer integrates multiple single feature vectors to obtain the first fused tensor; then, the dimension reorganization layer adjusts the tensor dimension to adapt to subsequent processing; then, the transformer layer based on the Transformer structure captures long-distance dependencies and performs feature fusion to obtain the first fused feature; then, the dimension reorganization layer adjusts the dimension of the first fused feature again; then, the average pooling layer reduces the spatial dimension of the feature and extracts global information; finally, the flattening layer converts the feature into a one-dimensional space, and further extracts and integrates the feature through multiple linear transformations, and finally obtains the fused feature vector. This process can effectively fuse the information of multiple single feature vectors and improve the richness and accuracy of feature representation.

[0116] In one example, the detailed process of step S706 can refer to Fig. 9 Flowchart shown.

[0117] Step S901, flattening the third fused feature to obtain a fourth fused feature.

[0118] In some embodiments of the present application, the third fused feature is a multidimensional tensor whose dimensions include the height, width, and number of channels of the feature map. In order to input this multidimensional feature into the subsequent linear transformation layer, it can be flattened. The flattening operation converts the multidimensional tensor into a one-dimensional vector, namely the fourth fused feature FE4. Figure 8 As shown, FE3 is flattened to obtain FE4, with a dimension of NF / / 8.

[0119] Step S902, performing a linear transformation on the fourth fusion feature to obtain a fifth fusion feature.

[0120] In some embodiments of the present application, the fourth fused feature can be input into the first linear transformation layer for linear transformation. The linear transformation layer is usually composed of a weight matrix W and a bias vector b. The linear transformation layer can multiply the fourth fused feature FE4 with the weight matrix W and add the bias vector b to obtain the fifth fused feature FE5. The purpose of linear transformation is to further extract and integrate the features for subsequent processing. Figure 8 As shown, FE4 is input into the first linear layer (Linear1, including activation function), the input dimension is NF / / 8, and the output dimension is NF / / 4, and FE5 is obtained.

[0121] Step S903, performing a linear transformation on the fifth fused feature to obtain a sixth fused feature, wherein the dimension of the sixth fused feature is smaller than the dimension of the fifth fused feature.

[0122] In some embodiments of the present application, after obtaining the fifth fusion feature, the fifth fusion feature can be linearly transformed to further reduce its dimension so as to reduce computational complexity and avoid overfitting. For example, the fifth fusion feature FE5 is input into the second linear transformation layer, and its weight matrix W' has a smaller dimension than the weight matrix of the first linear transformation layer. The fifth fusion feature FE5 is multiplied by the weight matrix W', and a bias vector b' may be added to obtain the sixth fusion feature FE6. The dimension of the sixth fusion feature is smaller than that of the fifth fusion feature, thereby achieving dimensionality reduction.

[0123] Following the above example, refer to Figure 8 As shown, FE5 is input into the second linear layer (Linear2, including activation function), the input dimension is NF / / 4, and the output dimension is NF / / 8, and FE6 is obtained.

[0124] Step S904, performing a linear transformation on the sixth fusion feature to obtain a seventh fusion feature, and using the seventh fusion feature as a fusion feature vector.

[0125] In some embodiments of the present application, the sixth fused feature is linearly transformed to obtain a final fused feature vector. The purpose of this linear transformation may be to map the feature to a specific output space, or to perform operations such as splicing and comparison with other features. The sixth fused feature FE6 is input into the third linear transformation layer, FE6 is multiplied by the weight matrix W", and a bias vector b" may be added to obtain the seventh fused feature FE7 as the fused feature vector FE. The fused feature vector FE integrates the information of multiple single feature vectors and is extracted and integrated through multiple linear transformations.

[0126] Following the above example, refer to Figure 8 As shown, FE6 is input into the third linear layer (Linear3, including activation function), the input dimension is NF / / 8, the output dimension is nc (the number of composite labels), and the fused feature vector FE is obtained.

[0127] refer to Fig.10As shown, it is an example diagram of the image label recognition process provided by an embodiment of the present application. The image is input into the Backbone base model, and the Backbone base model represents a computer vision model, such as a ConvNeXt model; the visual features (Visual Feature, VF) of the image are obtained using the Backbone base model. The visual features are copied to obtain multiple identical visual features, and each visual feature is input into a first module, such as a feature expert (Feature Expert) module, and the corresponding feature vector (such as a single feature vector) is obtained through the output of the first module, and a total of N feature vectors are obtained. The N single feature vectors are input into a second module, such as a mixed fusion (Mix & fusion) module, and the second module is used to obtain a fused feature vector. Based on a preset activation function, such as a Sigmoid function, the fused feature vector is processed to obtain a composite label result, and the composite label result is output.

[0128] Based on the above embodiment, the image is first input into a computer vision model such as ConvNeXt (i.e., Backbone base model) to extract the visual features (VF) of the image. Subsequently, multiple copies of these visual features are copied, and each feature is input into the Feature Expert module to obtain N single feature vectors. Next, these N feature vectors are input into the Mix & Fusion module, through which a fused feature vector is generated. Finally, the fused feature vector is processed using a preset Sigmoid activation function to obtain a composite label result and output it. This process effectively improves the accuracy and efficiency of image label recognition.

[0129] refer to Fig.11 FIG. 1 is a diagram showing an example of a training process of a tag recognition model provided in an embodiment of the present application. Fig.10 In the method flow shown, during the training process of the label recognition model, each first module is trained separately, including obtaining the single feature label result of the image as the true single feature label result (ground truth, gt) according to the existing single feature label model, calculating the difference between the single feature label predicted by the first module and the result and gt based on the preset loss function, and training the first module by backpropagation according to the difference. Similarly, the second module is trained accordingly. In this way, the model can accurately identify the single label and compound label of the image.

[0130] See also Fig.12, is a principle block diagram of an image label recognition device provided in an embodiment of the present application. An image label recognition device is provided to meet one of the purposes of the present application, and is a functional embodiment of the image label recognition method of the present application. The image label recognition device includes: an acquisition module 91, which is used to acquire the visual features of the image; a determination module 92, which is used to perform multiple different feature extractions on the visual features using a label recognition model to obtain multiple single feature vectors corresponding to the visual features, and determine the single label corresponding to each single feature vector; a fusion module 93, which is used to fuse multiple single feature vectors using a label recognition model to obtain a fused feature vector; the determination module 92 is also used to determine the composite label corresponding to the image based on the fused feature vector; an output module 94, which is used to obtain the label recognition result of the image according to the single label and the composite label.

[0131] Another embodiment of the present application also provides an electronic device. Figure 1 The application environment is only an example. In other exemplary embodiments, the computer program product implementing the image tag recognition method of the embodiment of the present application can also be run on any electronic device with sufficient computing power (such as Fig.13 In the electronic device shown in the figure), the various steps of the image tag recognition method are executed to provide the image tag recognition function.

[0132] See also Fig.13 , is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Fig.13 As shown, in one embodiment of the present application, the electronic device 800 can be a mobile phone, a tablet computer, a smart wearable device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, a netbook, etc. The embodiment of the present application does not impose any restrictions on the specific type of the electronic device 800.

[0133] like Fig.13 As shown, the electronic device 800 may include, but is not limited to, a communication module 81, a memory 82, a processor 83, an input / output (I / O) interface 84, and a bus 85. The processor 83 is coupled to the communication module 81, the memory 82, and the I / O interface 84 through the bus 85.

[0134] Those skilled in the art will appreciate that the schematic diagram is merely an example of the electronic device 800 and does not constitute a limitation of the electronic device 800 , and may include more or fewer components than shown in the diagram, or a combination of certain components, or different components. For example, the electronic device 800 may also include a network access device, etc.

[0135] The communication module 81 may include a wired communication module and / or a wireless communication module. The wired communication module may provide one or more wired communication solutions such as Universal Serial Bus (USB), Controller Area Network (CAN), etc. The wireless communication module may provide one or more wireless communication solutions such as Wireless Fidelity (Wi-Fi), Bluetooth (BT), mobile communication network, Frequency Modulation (FM), near field communication technology (NFC), infrared technology (IR), etc.

[0136] The memory 82 can be used to store computer-readable instructions and / or modules. The processor 83 implements various functions of the electronic device 800 by running or executing the computer-readable instructions and / or modules stored in the memory 82 and calling the data stored in the memory 82. The memory 82 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device 800, etc. The memory 82 may include non-volatile and volatile memories, such as: a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other storage devices. The memory 82 may be an external memory and / or an internal memory of the electronic device 800. Further, the memory 82 may be a memory in a physical form, such as a memory stick, a TF card (Trans-flash Card), etc.

[0137] The processor 83 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor 83 is the computing core and control center of the electronic device 800, and uses various interfaces and lines to connect various parts of the entire electronic device 800, and execute the operating system of the electronic device 800 and various installed applications, program codes, etc.

[0138] Exemplarily, the computer-readable instructions may be divided into one or more modules / sub-modules / units, one or more modules / sub-modules / units are stored in the memory 82 and executed by the processor 83 to complete the present application. One or more modules / sub-modules / units may be a series of computer-readable instruction segments capable of completing a specific function, and the computer-readable instruction segments are used to describe the execution process of the computer-readable instructions in the electronic device 800. For example, the computer-readable instructions may be divided into the above-mentioned modules.

[0139] If the module / unit integrated in the electronic device 800 is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through computer-readable instructions, and the computer-readable instructions can be stored in a computer-readable storage medium. When the computer-readable instructions are executed by the processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer-readable instructions include computer-readable instruction codes, and the computer-readable instruction codes can be in source code form, object code form, executable files or some intermediate forms. Computer-readable media may include: any entity or device capable of carrying computer-readable instruction codes, recording media, USB flash drives, mobile hard disks, magnetic disks, optical disks, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory).

[0140] Combination Figures 2 to 10The memory 82 in the electronic device 800 stores computer-readable instructions, and the processor 83 can execute the computer-readable instructions stored in the memory 82 to implement the following Figures 2 to 10 The image label recognition method shown.

[0141] Specifically, the specific implementation method of the processor 83 for the above-mentioned computer readable instructions can refer to Figures 2 to 10 The description of the relevant steps in the corresponding embodiments will not be repeated here.

[0142] The I / O interface 84 is used to provide a channel for user input or output. For example, the I / O interface 84 can be used to connect various input and output devices, such as a mouse, keyboard, touch device, display screen, etc., so that the user can enter information or visualize information.

[0143] The bus 85 is at least used to provide a channel for mutual communication among the communication module 81 , the memory 82 , the processor 83 , and the I / O interface 84 in the electronic device 800 .

[0144] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of modules is only a logical function division, and there may be other division methods in actual implementation.

[0145] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0146] In addition, each functional module in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of hardware plus software functional modules.

[0147] Therefore, no matter from which point of view, the embodiments should be regarded as illustrative and non-restrictive, and the scope of the present application is limited by the appended claims rather than the above description, so it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present application. Any attached figure mark in the claims should not be regarded as limiting the claims involved.

[0148] In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices can also be implemented by one unit or device through software or hardware. The words first, second, etc. are used to indicate names, and do not indicate any specific order.

[0149] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present application and are not intended to limit it. Although the present application has been described in detail with reference to the preferred embodiments, a person of ordinary skill in the art should understand that the technical solution of the present application may be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present application.

Claims

1. A method for image label recognition, characterized in that: The method comprises: Obtain visual features of the image; Performing multiple different feature extractions on the visual feature using a label recognition model to obtain multiple single feature vectors corresponding to the visual feature, and determining a single label corresponding to each single feature vector; Using the label recognition model to fuse the multiple single feature vectors to obtain a fused feature vector; Based on the fused feature vector, determining a composite label corresponding to the image; A label recognition result of the image is obtained according to the single label and the composite label.

2. The image tag recognition method according to claim 1, characterized in that: The obtaining of visual features of the image comprises: Dividing the image into a plurality of block images, performing linear projection on each block image, and obtaining projection features corresponding to each block image; Perform global mapping on the projection features corresponding to all block images to obtain mapping features; Performing context feature extraction on the mapping feature to obtain context features corresponding to the image; The context features are sequentially subjected to layer normalization processing, nonlinear transformation processing and downsampling processing to obtain the visual features.

3. The image tag recognition method according to claim 1, characterized in that: The tag recognition model includes a plurality of first modules, and each feature extraction in performing a plurality of different feature extractions on the visual feature includes: Using each first module to extract a single feature from the visual feature to obtain a single feature vector corresponding to the visual feature; Each single feature vector is sequentially subjected to pooling processing, feature abstraction processing, straightening processing and linear activation operations to obtain the single label.

4. The image tag recognition method according to claim 3, characterized in that: The extracting a single feature from the visual feature using each first module to obtain a single feature vector corresponding to the visual feature comprises: Performing feature abstraction processing on the visual feature using a convolution kernel of a first size to obtain a first feature, wherein the feature abstraction processing includes convolution processing, normalization processing, and nonlinear transformation processing; Performing feature abstraction processing on the first feature using a convolution kernel of a second size to obtain a second feature, wherein the second size is larger than the first size; The second feature is subjected to feature abstraction processing by using the convolution kernel of the first size to obtain a third feature, and the third feature is used as the single feature vector.

5. The image tag recognition method according to claim 1, characterized in that: The label recognition model includes a second module, wherein the method of fusing the plurality of single feature vectors using the feature label recognition model to obtain a fused feature vector includes: fusing the plurality of single feature vectors using the second module to obtain a first fused tensor; Performing dimension reshaping on the first fused tensor to obtain a second fused tensor, wherein the dimension of the second fused tensor is smaller than the dimension of the first fused tensor; Performing feature fusion on the second fused tensor based on the converter structure of the second module to obtain a first fused feature; Performing dimensionality reorganization on the first fused features to obtain second fused features; Performing average pooling processing on the second fused features to obtain a third fused feature; The third fused feature is flattened and subjected to multiple linear transformations to obtain the fused feature vector.

6. The image tag recognition method according to claim 5, characterized in that: The step of flattening and performing multiple linear transformations on the third fused feature to obtain the fused feature vector includes: Flattening the third fused feature to obtain a fourth fused feature; Performing a linear transformation on the fourth fusion feature to obtain a fifth fusion feature; Performing a linear transformation on the fifth fusion feature to obtain a sixth fusion feature, wherein the dimension of the sixth fusion feature is smaller than the dimension of the fifth fusion feature; Perform a linear transformation on the sixth fusion feature to obtain a seventh fusion feature, and use the seventh fusion feature as the fusion feature vector.

7. The image tag recognition method according to claim 1, characterized in that: The compound label includes multiple labels, and the dimension of the fused feature vector is equal to the number of labels included in the compound label.

8. The image tag recognition method according to claim 1, characterized in that: The determining, based on the fused feature vector, a composite label corresponding to the image includes: The fused feature vector is processed using a preset activation function to obtain the composite label.

9. The image tag recognition method according to claim 1, characterized in that: The label recognition model includes a first module and a second module, and the method further includes training the first module and the second module using a supervised training method, including: Determine a first loss value of the first module based on a first loss function, and determine a second loss value of the second module based on a second loss function; The first module is trained according to the first loss value, and the second module is trained according to the second loss value.

10. An electronic device, characterized in that: include: Memory, and A processor, wherein the processor executes the computer-readable instructions stored in the memory to implement the image tag recognition method according to any one of claims 1 to 9.