Image classification method and device, computer equipment, electronic device and storage medium

By fusing the text description and image features of the images to be classified in image retrieval, and using a preset reference dictionary for feature enhancement, the problem of high image retrieval training is solved, and image classification in zero samples is achieved and retrieval accuracy is improved.

CN120032377APending Publication Date: 2025-05-23ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510115342.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In the prior art, image retrieval training is expensive and it is difficult to optimize the search results according to user intentions, especially when the user cannot accurately describe the intentions.

Method used

By obtaining the text description of the image to be classified and the original image features, the feature fusion is performed to obtain the original knowledge; then the preset reference dictionary is retrieved based on the original image features to obtain the search enhanced features; the search enhanced features are fused with the original image features to obtain the reference features; then the reference features are fused with the original text features to obtain the reference knowledge; finally the original knowledge and reference knowledge are fused, and the target fusion knowledge is determined to perform image classification.

Benefits of technology

It reduces the cost of image retrieval training, can realize image classification under zero sample training, and improves the accuracy and efficiency of image retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032377A_ABST
    Figure CN120032377A_ABST
Patent Text Reader

Abstract

The invention relates to an image classification method and device, computer equipment, an electronic device and a storage medium, and the method comprises the steps: carrying out the feature fusion of original image features and original text features of text description, and obtaining original knowledge; performing feature enhancement on the original image features by retrieving image features in a preset reference dictionary to obtain a plurality of retrieval enhancement features; performing feature fusion on the plurality of retrieval enhancement features and the original image features to obtain a plurality of reference features; respectively carrying out feature fusion on the plurality of reference features and the original text features to obtain a plurality of pieces of reference knowledge; and fusing the original knowledge and the plurality of reference knowledge to obtain a plurality of fused knowledge, determining target fused knowledge from the plurality of fused knowledge, and determining the image category corresponding to the target fused knowledge as the image category to which the to-be-classified image belongs, thereby solving the problem of high image retrieval training cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to image classification methods, devices, computer equipment, electronic devices and storage media. Background Art

[0002] Image retrieval has always been a core part of the computer vision field. One of the most challenging aspects of building an image retrieval system is the ability to accurately understand user intent. However, most image search engines are based on either image-to-image matching or image-to-text matching. The inherent disadvantage of these methods is that they cannot optimize the retrieved items based on the user's intent, especially when the user cannot accurately describe their intent through a single image or all keywords. Single-modality retrieval is increasingly difficult to meet people's needs.

[0003] Modality refers to the specific way people receive information. Since multimedia data is often a medium for transmitting multiple types of information (for example, a video often transmits text, visual, and auditory information at the same time), multimodal learning (Multimodal Deep Learning) has gradually developed into the main means of multimedia content analysis and understanding, and researchers at home and abroad have gradually achieved remarkable research results in the field of multimodal learning.

[0004] In recent years, deep learning has made breakthrough progress in image retrieval. Under certain experimental conditions, it has even surpassed human discrimination ability. However, most deep learning methods are supervised learning methods, which require a large number of labeled training samples when training deep models. In practical applications, most of the time, we cannot obtain a large number of labeled samples or require a lot of manpower and material resources to obtain such samples.

[0005] There is currently no effective solution to the problem of high image retrieval training cost in related technologies. Summary of the invention

[0006] In this embodiment, an image classification method, apparatus, computer equipment, electronic device and storage medium are provided to solve the problem of high image retrieval training cost in related technologies.

[0007] In a first aspect, an image classification method is provided in this embodiment, including:

[0008] Obtaining a text description for the image to be classified, and extracting original text features of the text description;

[0009] Extracting original image features of the image to be classified;

[0010] Performing feature fusion on the original image features and the original text features of the text description to obtain original knowledge;

[0011] According to the original image features, image features in a preset reference dictionary are retrieved to obtain a plurality of retrieval enhancement features; wherein the preset reference dictionary includes all categories of image feature-text feature pairs;

[0012] Fusing the plurality of retrieval enhancement features with the original image features to obtain a plurality of reference features;

[0013] Performing feature fusion on the reference features and the original text features respectively to obtain a plurality of reference knowledge;

[0014] Fusing the original knowledge and the plurality of reference knowledge respectively to obtain a plurality of fused knowledge;

[0015] According to a preset rule, target fusion knowledge is determined from the plurality of fusion knowledge, the image to be classified is classified according to the target fusion knowledge, and the image category to which the image to be classified belongs is determined.

[0016] In some of the embodiments, the image feature-text feature pairs in the preset reference dictionary are extracted by offline analysis in advance based on preset target category images.

[0017] In some of the embodiments, the image features in a preset reference dictionary are retrieved according to the original image features to obtain a plurality of retrieval enhancement features, including:

[0018] Calculating the similarity between the original image features and all image features in a preset reference dictionary,

[0019] Using image features in the reference dictionary that meet preset similarity conditions as retrieval-enhanced image features to obtain a plurality of retrieval-enhanced image features;

[0020] Acquire text features corresponding to the plurality of retrieval-enhanced image features to obtain a plurality of retrieval-enhanced text features,

[0021] The plurality of retrieval enhancement image features and the plurality of retrieval enhancement text features are combined into the plurality of retrieval enhancement features.

[0022] In some of the embodiments, the plurality of retrieval enhancement features are respectively fused with the original image features to obtain a plurality of reference features, including:

[0023] The plurality of retrieval enhancement features are used as input, the original image features are used as query conditions, and are sequentially input into a multi-head cross attention and fully connected feedforward neural network for feature fusion to obtain a plurality of reference features.

[0024] In some embodiments, the original knowledge and the plurality of reference knowledge are respectively fused to obtain a plurality of fused knowledge, including:

[0025] The original knowledge and the plurality of reference knowledge are fused by using a linear combination technique to obtain a plurality of fused knowledge.

[0026] In some of the embodiments, determining the target fusion knowledge from the plurality of fusion knowledge according to a preset rule includes:

[0027] The plurality of fusion knowledge are converted into corresponding probability distributions to obtain a plurality of probability distribution values, and the fusion knowledge corresponding to the maximum probability distribution value is determined as the target fusion knowledge.

[0028] In a second aspect, an image classification device is provided in this embodiment, including: an extraction module, an original knowledge acquisition module, a feature enhancement module, a reference feature acquisition module, a reference knowledge acquisition module, a fusion module and a classification module, wherein:

[0029] The extraction module is used to obtain a text description of the image to be classified and extract original text features of the text description; extract original image features of the image to be classified;

[0030] The original knowledge acquisition module is used to perform feature fusion on the original image features and the original text features of the text description to obtain original knowledge;

[0031] The feature enhancement module is used to retrieve image features in a preset reference dictionary based on the original image features to obtain a number of retrieval enhancement features; wherein the preset reference dictionary includes all categories of image feature-text feature pairs;

[0032] The reference feature acquisition module is used to perform feature fusion on the plurality of retrieval enhancement features and the original image features to obtain a plurality of reference features;

[0033] The reference knowledge acquisition module is used to perform feature fusion on the reference features and the original text features to obtain a plurality of reference knowledge;

[0034] The fusion module is used to fuse the original knowledge and the plurality of reference knowledge respectively to obtain a plurality of fused knowledge;

[0035] The classification module is used to determine target fusion knowledge from the plurality of fusion knowledge according to a preset rule, classify the image to be classified according to the target fusion knowledge, and determine the image category to which the image to be classified belongs.

[0036] In a third aspect, a computer device is provided in this embodiment, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the method described in the first aspect are implemented.

[0037] In a fourth aspect, an electronic device is provided in this embodiment, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the image classification method described in the first aspect when executing the computer program.

[0038] In a fourth aspect, a storage medium is provided in this embodiment, on which a computer program is stored, and when the program is executed by a processor, the image classification method described in the first aspect is implemented.

[0039] Compared with the related art, the image classification method provided in this embodiment obtains a text description for an image to be classified, and extracts the original text features of the text description; extracts the original image features of the image to be classified; performs feature fusion on the original image features and the original text features of the text description to obtain original knowledge; retrieves image features in a preset reference dictionary based on the original image features to obtain a number of retrieval enhancement features; wherein the preset reference dictionary includes all categories of image feature-text feature pairs; performs feature fusion on the several retrieval enhancement features and the original image features respectively to obtain a number of reference features; performs feature fusion on the several reference features and the original text features respectively to obtain a number of reference knowledge; fuses the original knowledge and the several reference knowledge respectively to obtain a number of fused knowledge; determines target fused knowledge from the several fused knowledge according to preset rules, classifies the image to be classified according to the target fused knowledge, and determines the image category to which the image to be classified belongs, thereby solving the problem of high image retrieval training cost.

[0040] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0042] Figure 1 is a hardware structure block diagram of a terminal of the image classification method of this embodiment;

[0043] Figure 2 is a flow chart of the image classification method of this embodiment;

[0044] Figure 3 is a schematic diagram of the fusion of the retrieval enhancement features and the original image features of the image classification of this embodiment;

[0045] Figure 4 is a flow chart of another image classification method of this embodiment;

[0046] Figure 5 It is a structural block diagram of the image classification device of this embodiment. DETAILED DESCRIPTION

[0047] In order to more clearly understand the purpose, technical solutions and advantages of the present application, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments.

[0048] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the general meaning understood by people with ordinary skills in the technical field to which this application belongs. The words "one", "a", "a", "the", "these" and the like in this application do not represent quantitative restrictions, and they can be singular or plural. The terms "include", "comprise", "have" and any variants thereof involved in this application are intended to cover non-exclusive inclusions; for example, a process, method and system, product or device comprising a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether directly or indirectly. The "multiple" involved in this application refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: A exists alone, A and B exist at the same time, and B exists alone. Usually, the character " / " indicates that the objects associated with each other are in an "or" relationship. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific ordering of the objects.

[0049] The method embodiment provided in this embodiment can be executed in a terminal, a computer or a similar computing device. For example, running on a terminal, Figure 1FIG. 1 is a hardware structure diagram of a terminal of the image classification method of this embodiment. Figure 1 As shown, the terminal may include one or more ( Figure 1 Only one is shown in the figure) processor 102 and memory 104 for storing data, wherein processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA. The above terminal may also include a transmission device 106 and an input and output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is for illustration only and does not limit the structure of the above terminal. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations shown.

[0050] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the image classification method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, to implement the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0051] The transmission device 106 is used to receive or send data via a network. The above network includes a wireless network provided by the communication provider of the terminal. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (Radio Frequency, referred to as RF) module, which is used to communicate with the Internet wirelessly.

[0052] In this embodiment, an image classification method is provided. Figure 2 is a flow chart of the image classification method of this embodiment, such as Figure 2 As shown, the process includes the following steps:

[0053] Step S201, obtaining a text description for the image to be classified, and extracting original text features of the text description.

[0054] The CLIP (Contrastive Language-Image Pre-Training) model is a multimodal pre-trained neural network released by OpenAI in 2021. It is an effective and scalable method for learning from natural language supervision. The core idea of ​​the model is to use a large amount of paired data of images and texts for pre-training to learn the alignment relationship between images and texts. The CLIP model has two modalities, one is the text modality and the other is the visual modality. The model consists of two main parts: a text encoder and an image encoder. Among them, the text encoder is used to extract text features, and the image encoder is used to extract image features.

[0055] The CLIP model builds a bridge between text and image, but few people actually use it for text and image retrieval, because the text encoder of CLIP cannot effectively perform semantic modeling on long texts, and text-text and image-image modality retrieval have not been explicitly trained. JinaCLIP uses a three-stage training method to optimize the CLIP model, which can support retrieval between text-text, image-image, and text-image modalities. In this embodiment, JinaCLIP is used to extract text features and image features.

[0056] First, for the image to be classified, a text description is constructed. It can be multiple or single. For example, to classify cats and dogs, two text descriptions can be constructed: "a picture of a cat" and "a picture of a dog". In this embodiment, "a picture of a dog" is taken as an example. The Jina CLIP text encoder is used to extract text features from the constructed text description, and the text description is converted into a text feature vector to obtain the original text feature F t .

[0057] Step S202: extracting original image features of the image to be classified.

[0058] Through the Jina CLIP image encoder, the image features of the image to be classified are extracted, and the image to be classified is converted into an image feature vector to obtain the original image feature f test .

[0059] Step S203, feature fusion is performed on the original image features and the original text features of the text description to obtain original knowledge.

[0060] The original image features are combined with the original text features of the text description, and the similarity between the two is calculated to obtain the similarity result, which is recorded as the original knowledge S org , and its calculation formula is as follows:

[0061]

[0062] Where T is the matrix transpose.

[0063] Step S204, based on the original image features, image features in a preset reference dictionary are retrieved to obtain a number of retrieval enhancement features; wherein the preset reference dictionary includes all categories of image feature-text feature pairs.

[0064] Specifically, prepare a set of image features and text description features of all categories as a reference dictionary in advance. For example, use 1000 categories of ImageNet, with at least 32 images in each category. The image features and text description features in the reference dictionary correspond to each other one by one, forming image feature-text feature pairs. test , search for similar image features in the reference dictionary D. Since the image features and text description features in the reference dictionary correspond one to one, the corresponding text description features can be obtained according to the obtained image features, and several similar image features and text description features with the highest similarity are selected as retrieval enhancement features. The specific number of selected features can be set according to the actual situation. In this embodiment, the top 20 image features and text description features with the highest similarity are selected as retrieval enhancement features.

[0065] Step S205 , fusing the plurality of retrieval enhancement features with the original image features to obtain a plurality of reference features.

[0066] The obtained retrieval enhancement features Respectively with the original image features f test Feature fusion is performed, wherein the feature fusion can use the attention mechanism to dynamically adjust the image features and text features to highlight important information; or it can use a simple splicing method to simply splice the image features and text features together to form a new feature vector; or it can use a weighted average method to perform a weighted average of the image features and text features according to specific weights to obtain a comprehensive representation. In this embodiment, the attention mechanism is used to dynamically fuse 20 retrieval enhancement features and original image features to obtain 20 reference features {f ref1 , f ref2 ...f ref20}.

[0067] Step S206 , performing feature fusion on a plurality of reference features and original text features respectively to obtain a plurality of reference knowledge.

[0068] Specifically, the 20 reference features {f ref1 , f ref2 ...fref20} respectively with the original text features F t The similarity is calculated by vector and matrix multiplication method, and the similarity values ​​of 20 reference features and original text features are obtained, which are recorded as reference knowledge {S ref1 , S ref2 ...S ref20}, and its calculation formula is as follows:

[0069]

[0070] Among them, T is the matrix transpose, and the matrix transpose is used to align the dimensions during multiplication.

[0071] Step S207, respectively fuse the original knowledge and several reference knowledge to obtain several fused knowledge; determine the target fused knowledge from the several fused knowledge according to the preset rules, classify the image to be classified according to the target fused knowledge, and determine the image category to which the image to be classified belongs.

[0072] The original knowledge S org and 20 references ref1 , S ref2 ...S ref20} are fused separately. Similarly, the fusion method can be simple addition or weighted average. After fusion, 20 fusion knowledge {S 1 , S 2 ...S 20 By comparing the sizes of the 20 fusion knowledge, the largest fusion knowledge is determined as the target fusion knowledge, and the image to be classified is classified according to the target fusion knowledge, so that the image to be classified can be determined as "a picture of a dog", thereby obtaining the image category to which the image to be classified belongs.

[0073] Through the above steps S201 to S207, a text description for the image to be classified is obtained, and the original text features of the text description are extracted; the original image features of the image to be classified are extracted; the original image features and the original text features of the text description are fused to obtain original knowledge; according to the original image features, the image features in the preset reference dictionary are retrieved to obtain a number of retrieval enhancement features; wherein the preset reference dictionary includes all categories of image feature-text feature pairs; the several retrieval enhancement features are fused with the original image features to obtain a number of reference features; the several reference features are fused with the original text features to obtain a number of reference knowledge; the original knowledge and the several reference knowledge are fused to obtain a number of fused knowledge; according to the preset rules, the target fused knowledge is determined from the several fused knowledge, the image to be classified is classified according to the target fused knowledge, and the image category to which the image to be classified belongs is determined. Compared with the prior art of performing image retrieval through training a large number of samples using a deep learning model, this embodiment, after extracting the original text features and the original image features, uses an offline constructed reference dictionary to perform feature enhancement on the extracted original image features of the image to be classified to obtain retrieval enhancement features, and then combines the original image features to obtain reference features through feature fusion; calculates the original text features and the reference features of the image to be classified to obtain reference knowledge; calculates the original image features of the image to be classified and the original knowledge of the original text features; fuses the reference knowledge and the original knowledge again to obtain the fused knowledge, determines the target fused knowledge from the fused knowledge, and classifies the image to be classified with the target fused knowledge, thereby achieving image classification. In this embodiment, there is no need to obtain corresponding training samples to train the model, and the classification of the target image can be achieved in the case of zero-sample training, thereby reducing the image retrieval training cost.

[0074] In some of the embodiments, the image feature-text feature pairs in the preset reference dictionary are extracted by offline analysis in advance based on preset target category images.

[0075] Specifically, in this embodiment, zero-sample image detection and classification is achieved by constructing a reference dictionary offline in advance. In order to improve the accuracy of image recognition, it is necessary to prepare in advance a set of images and text descriptions covering all categories, and extract image features and text features offline in advance through Jina CLIP image encoder and text encoder, wherein the image features and text features correspond one to one. For example, 1000 categories of ImageNet are used, with at least 32 pictures in each category, to form the images and texts of the original reference dictionary, from which image features and text features are extracted and stored in the reference dictionary. This process does not participate in the training process of the model recognition detection image, reducing the training cost of image retrieval.

[0076] In another embodiment, according to the original image features, image features in a preset reference dictionary are retrieved to obtain a number of retrieval enhancement features, including:

[0077] Calculate the similarity between the original image features and all the image features in the preset reference dictionary,

[0078] Using image features in the reference dictionary that meet preset similarity conditions as retrieval-enhanced image features to obtain a plurality of retrieval-enhanced image features;

[0079] The text features corresponding to the plurality of retrieval-enhanced image features are obtained to obtain a plurality of retrieval-enhanced text features, and the plurality of retrieval-enhanced image features and the plurality of retrieval-enhanced text features are combined into a plurality of retrieval-enhanced features.

[0080] Specifically, the original image features are enhanced by referring to the preset image features in the dictionary. First, the similarity between the original image features and all the image features in the reference dictionary is calculated, and the calculation formula is as follows:

[0081]

[0082] Among them, f i is the image feature in the reference dictionary, f test is the original image feature of the image to be classified, T is the matrix transpose, and the matrix transpose is used to align the dimensions during multiplication.

[0083] Through the TOPK rule, the image features in the reference dictionary corresponding to the first k similarities are selected as retrieval enhanced image features. In this embodiment, k=20. That is, 20 retrieval enhanced image features are obtained. According to the one-to-one correspondence between image features and text features in the reference dictionary, the corresponding text features are matched for the 20 retrieval enhanced image features to obtain 20 retrieval enhanced text features. The 20 retrieval enhanced image features and the 20 retrieval enhanced text features constitute the retrieval enhanced features after the image features of the image to be classified are enhanced. The reference dictionary stores accurate image and text features. By comparing the similarity of the original image features with the features in the reference dictionary, the characteristics of the original image features are enhanced, thereby improving the accuracy of subsequent image classification.

[0084] In some embodiments, several retrieval enhancement features are fused with original image features to obtain several reference features, including:

[0085] Taking several retrieval enhancement features as input and original image features as query conditions, they are sequentially input into multi-head cross attention and fully connected feedforward neural networks for feature fusion to obtain several reference features.

[0086] Specifically, in the above step S205, in the process of obtaining reference features by feature fusion, this embodiment adopts multi-head cross attention and fully connected feedforward neural network to perform, Figure 3 FIG. 1 is a schematic diagram of the fusion of the retrieval enhancement features and the original image features of the image classification of this embodiment. Figure 3 As shown, to retrieve enhanced image features As K, retrieve the enhanced text features As V, the original image feature ftest is used as the query condition Q, and is input into a 4-layer transformer block consisting of a multi-head cross attention MHCA and a fully connected feedforward neural network FFN to obtain the reference feature f ref , the specific formula is as follows:

[0087]

[0088] in, It is the fusion feature output of the i-1th layer. There are 4 layers of network in total, each layer consists of a multi-head cross attention MHCA and a fully connected feedforward neural network FFN.

[0089] In another embodiment, the original knowledge and several reference knowledge are fused respectively to obtain several fused knowledge, including:

[0090] The linear combination technology is used to fuse the original knowledge and several reference knowledge to obtain several fused knowledge.

[0091] Specifically, in step S207, when the original knowledge and the reference knowledge are integrated, the linear combination technology is used for integration in this embodiment, and the specific calculation formula is as follows:

[0092] S=λS org +(1-λ)S ref , where λ=0.5.

[0093] In some embodiments, target fusion knowledge is determined from a plurality of fusion knowledge according to a preset rule, including:

[0094] A number of fusion knowledge are converted into corresponding probability distributions to obtain a number of probability distribution values, and the fusion knowledge corresponding to the maximum probability distribution value is determined as the target fusion knowledge.

[0095] Specifically, in step S207, the target fusion knowledge is determined by the Softmax function, and several fusion knowledge are converted into probability distribution through the Softmax function to obtain several probability distribution values. The probability distribution values ​​are compared, and the fusion knowledge corresponding to the largest probability distribution value is used as the target fusion knowledge. The corresponding image category in the target fusion knowledge is used as the image category to which the image to be classified belongs, thereby realizing image classification.

[0096] This embodiment also provides an image classification method. Figure 4 is a flow chart of another image classification method of this embodiment. Figure 4 As shown, the process includes the following steps:

[0097] Step S401, obtaining a text description for the image to be classified, and extracting original text features of the text description;

[0098] Step S402, extracting original image features of the image to be classified;

[0099] Step S403, performing feature fusion on the original image features and the original text features of the text description to obtain original knowledge;

[0100] Step S404, constructing a reference dictionary, the reference dictionary includes image feature-text feature pairs extracted by offline analysis in advance according to a preset target category image;

[0101] Step S405, calculating the similarity between the original image features and all the image features in the reference dictionary, taking the image features in the reference dictionary that meet the preset similarity conditions as retrieval enhanced image features, and obtaining a plurality of retrieval enhanced image features; obtaining text features corresponding to the plurality of retrieval enhanced image features, and combining the plurality of retrieval enhanced image features and the plurality of text features into a plurality of retrieval enhanced features;

[0102] Step S406, using a multi-head cross attention and a fully connected feedforward neural network, respectively fuse a number of retrieval enhancement features with the original image features to obtain a number of reference features;

[0103] Step S407, performing feature fusion on a plurality of reference features and original text features respectively to obtain a plurality of reference knowledge;

[0104] Step S408, using linear combination technology to fuse the original knowledge and a plurality of reference knowledge to obtain a plurality of fused knowledge;

[0105] Step S409, converting a number of fusion knowledge into corresponding probability distributions to obtain a number of probability distribution values, determining the fusion knowledge corresponding to the largest probability distribution value as the target fusion knowledge, and taking the image category corresponding to the target fusion knowledge as the image category to which the image to be classified belongs.

[0106] Through the above steps S401 to S409, compared with the prior art of performing image retrieval through training with a large number of samples using a deep learning model, this embodiment, after extracting the original text features and the original image features, constructs a reference dictionary offline in advance, calculates the similarity between the extracted original image features of the image to be classified and the image characteristics of the reference dictionary, obtains retrieval enhancement features, improves the accuracy of the original image features, and then obtains reference features through feature fusion in combination with the original image features, obtains reference knowledge through calculation of the original text features and the reference features of the image to be classified; calculates the original knowledge of the original image features and the original text features of the image to be classified; fuses the reference knowledge and the original knowledge through a linear combination technique to obtain fused knowledge, and finally converts the fused knowledge into a probability distribution value, uses the fused knowledge corresponding to the maximum probability distribution value as the target fused knowledge, and uses the image category corresponding to the target fused knowledge as the target classification category of the image to be classified, thereby achieving image classification. In this embodiment, there is no need to obtain corresponding training samples to train the model, and the classification of the target image can be achieved in the case of zero-sample training, thereby reducing the image retrieval training cost.

[0107] In this embodiment, an image classification device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. The terms "module", "unit", "subunit", etc. used below can implement a combination of software and / or hardware of predetermined functions. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0108] Figure 5 is a structural block diagram of the image classification device of this embodiment, such as Figure 5 As shown, the device 50 includes: an extraction module 51, an original knowledge acquisition module 52, a feature enhancement module 53, a reference feature acquisition module 54, a reference knowledge acquisition module 55, a fusion module 56 and a classification module 57, wherein:

[0109] The extraction module 51 is used to obtain the text description of the image to be classified and extract the original text features of the text description; extract the original image features of the image to be classified;

[0110] The original knowledge acquisition module 52 is used to perform feature fusion on the original image features and the original text features of the text description to obtain the original knowledge;

[0111] The feature enhancement module 53 is used to retrieve the image features in the preset reference dictionary according to the original image features to obtain a plurality of retrieval enhancement features; wherein the preset reference dictionary includes all categories of image feature-text feature pairs;

[0112] A reference feature acquisition module 54 is used to fuse several retrieval enhancement features with original image features to obtain several reference features;

[0113] A reference knowledge acquisition module 55 is used to perform feature fusion on a plurality of reference features and original text features to obtain a plurality of reference knowledge;

[0114] A fusion module 56 is used to fuse the original knowledge and the plurality of reference knowledge respectively to obtain a plurality of fused knowledge;

[0115] The classification module 57 is used to determine target fusion knowledge from a plurality of fusion knowledge according to a preset rule, classify the image to be classified according to the target fusion knowledge, and determine the image category to which the image to be classified belongs.

[0116] It should be noted that the above modules can be functional modules or program modules, and can be implemented by software or hardware. For modules implemented by hardware, the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.

[0117] In this embodiment, a computer device is also provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method described in any of the above method embodiments are implemented.

[0118] In this embodiment, an electronic device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0119] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0120] Optionally, in this embodiment, the processor may be configured to perform the following steps through a computer program:

[0121] S1, obtaining a text description of the image to be classified and extracting original text features of the text description;

[0122] S2, extracting original image features of the image to be classified;

[0123] S3, feature fusion of the original image features and the original text features of the text description to obtain the original knowledge;

[0124] S4, according to the original image features, searching for image features in a preset reference dictionary to obtain a number of retrieval enhancement features; wherein the preset reference dictionary includes all categories of image feature-text feature pairs;

[0125] S5, fusing several retrieval enhancement features with original image features to obtain several reference features;

[0126] S6, performing feature fusion on a plurality of reference features and original text features respectively to obtain a plurality of reference knowledge;

[0127] S7, fusing the original knowledge and several reference knowledge respectively to obtain several fused knowledge;

[0128] S8, determining target fusion knowledge from a plurality of fusion knowledge according to a preset rule, classifying the image to be classified according to the target fusion knowledge, and determining the image category to which the image to be classified belongs.

[0129] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation modes, and will not be repeated in this embodiment.

[0130] In addition, in combination with the image classification method provided in the above embodiments, a storage medium may be provided in this embodiment to implement the image classification method. The storage medium stores a computer program, and when the computer program is executed by a processor, any image classification method in the above embodiments is implemented.

[0131] It should be understood that the specific embodiments described herein are only used to explain the application, rather than to limit it. Based on the embodiments provided in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the protection scope of this application.

[0132] Obviously, the drawings are only some examples or embodiments of the present application. For ordinary technicians in the field, the present application can also be applied to other similar situations based on these drawings without creative work. In addition, it is understandable that although the work done in this development process may be complicated and lengthy, for ordinary technicians in the field, certain changes in design, manufacturing or production based on the technical content disclosed in this application are only conventional technical means and should not be regarded as insufficient content disclosed in this application.

[0133] The term "embodiment" in this application refers to a specific feature, structure or characteristic described in conjunction with the embodiment that can be included in at least one embodiment of the present application. The appearance of this phrase in various locations in the specification does not necessarily mean the same embodiment, nor does it mean that it is mutually exclusive with other embodiments and is independent or optional. It is clearly or implicitly understood by those of ordinary skill in the art that the embodiments described in this application can be combined with other embodiments without conflict.

[0134] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.

[0135] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of patent protection. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the scope of protection of the present application. Therefore, the scope of protection of the present application shall be subject to the attached claims.

Claims

1. An image classification method, characterized in that: include: Obtaining a text description for the image to be classified, and extracting original text features of the text description; Extracting original image features of the image to be classified; Performing feature fusion on the original image features and the original text features of the text description to obtain original knowledge; According to the original image features, image features in a preset reference dictionary are retrieved to obtain a plurality of retrieval enhancement features; wherein the preset reference dictionary includes all categories of image feature-text feature pairs; Fusing the plurality of retrieval enhancement features with the original image features to obtain a plurality of reference features; Performing feature fusion on the reference features and the original text features respectively to obtain a plurality of reference knowledge; Fusing the original knowledge and the plurality of reference knowledge respectively to obtain a plurality of fused knowledge; According to a preset rule, target fusion knowledge is determined from the plurality of fusion knowledge, the image to be classified is classified according to the target fusion knowledge, and the image category to which the image to be classified belongs is determined.

2. The image classification method according to claim 1, characterized in that: The image feature-text feature pairs in the preset reference dictionary are extracted by offline analysis in advance based on preset target category images.

3. The image classification method according to claim 1, characterized in that: The method of retrieving image features in a preset reference dictionary based on the original image features to obtain a plurality of retrieval enhancement features includes: Calculating the similarity between the original image features and all image features in a preset reference dictionary, Using image features in the reference dictionary that meet preset similarity conditions as retrieval-enhanced image features to obtain a plurality of retrieval-enhanced image features; Acquire text features corresponding to the plurality of retrieval-enhanced image features to obtain a plurality of retrieval-enhanced text features, The plurality of retrieval enhancement image features and the plurality of retrieval enhancement text features are combined into the plurality of retrieval enhancement features.

4. The image classification method according to claim 1, characterized in that: The step of fusing the plurality of retrieval enhancement features with the original image features to obtain a plurality of reference features includes: The plurality of retrieval enhancement features are used as input, the original image features are used as query conditions, and are sequentially input into a multi-head cross attention and fully connected feedforward neural network for feature fusion to obtain a plurality of reference features.

5. The image classification method according to claim 1, characterized in that: The original knowledge and the plurality of reference knowledge are respectively fused to obtain a plurality of fused knowledge, including: The original knowledge and the plurality of reference knowledge are fused by using a linear combination technique to obtain a plurality of fused knowledge.

6. The image classification method according to claim 1, characterized in that: The step of determining target fusion knowledge from the plurality of fusion knowledge according to a preset rule includes: The plurality of fusion knowledge are converted into corresponding probability distributions to obtain a plurality of probability distribution values, and the fusion knowledge corresponding to the maximum probability distribution value is determined as the target fusion knowledge.

7. An image classification device, characterized in that: include: extraction module, original knowledge acquisition module, feature enhancement module, reference feature acquisition module, reference knowledge acquisition module, fusion module and classification module, wherein, The extraction module is used to obtain a text description of the image to be classified and extract original text features of the text description; extract original image features of the image to be classified; The original knowledge acquisition module is used to perform feature fusion on the original image features and the original text features of the text description to obtain original knowledge; The feature enhancement module is used to retrieve image features in a preset reference dictionary based on the original image features to obtain a number of retrieval enhancement features; wherein the preset reference dictionary includes all categories of image feature-text feature pairs; The reference feature acquisition module is used to perform feature fusion on the plurality of retrieval enhancement features and the original image features to obtain a plurality of reference features; The reference knowledge acquisition module is used to perform feature fusion on the reference features and the original text features to obtain a plurality of reference knowledge; The fusion module is used to fuse the original knowledge and the plurality of reference knowledge respectively to obtain a plurality of fused knowledge; The classification module is used to determine target fusion knowledge from the plurality of fusion knowledge according to a preset rule, classify the image to be classified according to the target fusion knowledge, and determine the image category to which the image to be classified belongs.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the image classification method according to any one of claims 1 to 6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the image classification method according to any one of claims 1 to 6 are implemented.

Citation Information

Cited By

  • Image classification method and device, equipment and medium

    CN121074917A