Data processing method and related device
During the computer vision model training process, the contrast loss is calculated using the distance and mean of the feature vector to achieve cross-modal semantic alignment, which solves the negative impact of semantic noise on model training, and improves image learning ability and model effect.
Patent Information
- Application Number
- CN202410064402.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-16
- Publication Date
- 2025-07-18
AI Technical Summary
Computer vision models are affected by semantic noise during training, resulting in a decline in image learning ability. It is difficult for the prior art to effectively eliminate the negative impact of this noise.
By retrieving keywords in the database, obtaining images and description text, generating eigenvectors, calculating distances and mean values of eigenvectors, and using contrast loss training to learn models, achieving cross-modal semantic alignment and eliminating the influence of semantic noise.
It effectively eliminates the negative impact of semantic noise on model training, improves image learning ability, and improves the training effect and generalization ability of the model.
Smart Images

Figure CN120336991A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method for data processing and related devices. Background Art
[0002] With the continuous development of computer vision technology (CV), the requirements for the image feature learning ability of computer vision models are getting higher and higher. However, the manifestation of the image feature learning ability and effect of computer vision models depends on the matching degree between the images and labels in the training set. In some specific application scenarios, technicians use Internet search engines to collect publicly available images to build a data set and use the data set for learning. The name of the concept category of interest is used as a search keyword in the Internet search engine to obtain a large number of Internet images, and the search keyword used during retrieval is used as the weak label of the large number of Internet images. Since the web images returned by the Internet search engine may not be related to the corresponding search keyword, there is a large amount of noise in the label data.
[0003] In order to suppress the influence of noise on the training result during the training of a computer vision model using a noisy data set, technicians have adopted techniques such as visual structures (such as neighbor information), clean sample sets, regularization means, confidence-based label bootstrap, and utilization of image additional meta-information. However, these methods are usually used to solve common label flip noise and out-of-domain (OOD) noise.
[0004] However, in addition to label flip noise and OOD noise, there is also a type of semantic noise that is often overlooked. For example, when using a search engine to obtain images, when there are polysemous or ambiguous situations in the search keywords, the images returned by the search will also contain a large amount of semantic meanings that are not expected, which means that most of the images retrieved at this time are noise data. Therefore, how to eliminate the influence of this type of noise on the image learning ability of computer vision models during the training of computer vision models has become an urgent problem for technicians to solve. Summary of the Invention
[0005] Embodiments of this application provide a method for data processing and related devices, which are used to eliminate the negative impact of semantic noise on the image learning ability of a computer vision model during the training of the computer vision model.
[0006] The first aspect of this application provides a method for data processing, including:
[0007] Generating N first feature vectors based on N labels, where N is a positive integer;
[0008] Search for N tags in the database to obtain N training images and descriptive texts of the N training images;
[0009] Based on the N training images and the descriptive texts of the N training images, generate N second feature vectors and N third feature vectors. There is a corresponding relationship between the N second feature vectors and the N training images, and there is a corresponding relationship between the N third feature vectors and the descriptive texts of the N training images;
[0010] Calculate the distances between the N third feature vectors and the N first feature vectors to obtain the K third feature vectors that are closest in distance to any one of the N first feature vectors. K is a positive integer and K is less than N;
[0011] Calculate the mean of the K second feature vectors to obtain a fourth feature vector. The K second feature vectors are the feature vectors corresponding to the K third feature vectors;
[0012] Use the contrast loss between the second feature vectors and the fourth feature vector to train the learning model.
[0013] The second aspect of this application provides a data processing device, including:
[0014] An acquisition unit for acquiring N tags, where N is a positive integer;
[0015] A generation unit for generating N first feature vectors based on the N tags;
[0016] A search unit for searching for N tags in the database to obtain N training images and descriptive texts of the N training images;
[0017] The generation unit is further configured to generate N second feature vectors and N third feature vectors based on the N training images and the descriptive texts of the N training images. There is a corresponding relationship between the N second feature vectors and the N training images, and there is a corresponding relationship between the N third feature vectors and the descriptive texts of the N training images;
[0018] A calculation unit for calculating the distances between the N third feature vectors and the N first feature vectors to obtain the K third feature vectors that are closest in distance to any one of the N first feature vectors. K is a positive integer and K is less than N;
[0019] The calculation unit is further configured to calculate the mean of the K second feature vectors to obtain a fourth feature vector. The K second feature vectors are the feature vectors corresponding to the K third feature vectors;
[0020] A training unit for using the contrast loss between the second feature vectors and the fourth feature vector to train the learning model.
[0021] In a possible implementation of the second aspect, the generating unit is specifically configured to:
[0022] Perform text augmentation on N labels to obtain N first texts;
[0023] Extract features from the N first texts to obtain N first feature vectors.
[0024] In a possible implementation of the second aspect, the generating unit is specifically configured to:
[0025] Extract features from N training images to obtain N second feature vectors;
[0026] Extract features from the description texts of N training images to obtain N fifth feature vectors;
[0027] Calculate the distances between the N second feature vectors and any one of the N second feature vectors, and obtain J second feature vectors that are the closest to any one of the N second feature vectors, where J is a positive integer less than N;
[0028] Calculate the mean of the J fifth feature vectors to obtain a third feature vector, where the J fifth feature vectors are the fifth feature vectors corresponding to the J second feature vectors, and the third feature vector is included in the N third feature vectors.
[0029] In a possible implementation of the second aspect, the apparatus further includes a deleting unit, configured to remove the training image corresponding to the second feature vector from the N training images to obtain N - 1 training images if the similarity between the second feature vector and the fourth feature vector does not meet the preset requirement;
[0030] The training unit is specifically configured to:
[0031] Generate updated second feature vectors and updated fourth feature vectors by using the N - 1 training images, the description texts of the N - 1 training images, and the N - 1 labels, where the N - 1 labels are the labels corresponding to the N - 1 training images;
[0032] Train the learning model by using the contrast loss between the updated second feature vectors and the updated fourth feature vectors.
[0033] In a possible implementation of the second aspect, the apparatus further includes an updating unit, configured to update the first feature vector corresponding to the second feature vector to obtain an updated first feature vector if the similarity between the second feature vector and the fourth feature vector meets the preset requirement;
[0034] The training unit is specifically configured to:
[0035] Calculate the updated fourth eigenvector by using the updated first, second, and third eigenvectors;
[0036] Train the learning model by using the contrastive loss between the second eigenvector and the updated fourth eigenvector.
[0037] In a possible implementation of the second aspect, the training unit is specifically configured to:
[0038] Train the learning model by using the contrastive loss between the second eigenvector and the fourth eigenvector, and the contrastive loss between the second eigenvector and the sixth eigenvector, where the sixth eigenvector is any one of the K second eigenvectors closest to the second eigenvector.
[0039] In a possible implementation of the second aspect, the calculation unit is specifically configured to:
[0040] Calculate the mean of the K second eigenvectors to obtain the seventh eigenvector;
[0041] Perform weighted summation on the seventh eigenvector and any one of the K second eigenvectors to obtain the fourth eigenvector.
[0042] In a possible implementation of the second aspect, the label type of any one of the N labels is included in C target label types, where C is a positive integer;
[0043] The generation unit is further configured to generate a first probability vector according to the similarity between the N second eigenvectors and the N fourth eigenvectors. The first probability vector is a C-dimensional vector, and the C-dimensional space of the first probability vector corresponds to C target label types;
[0044] The training unit is specifically configured to train the learning model by using the contrastive loss between the second eigenvector and the fourth eigenvector and the first probability vector.
[0045] In a possible implementation of the second aspect, the acquisition unit is further configured to acquire Q reference images, where Q is a positive integer;
[0046] The generation unit is further configured to perform feature extraction on the Q reference images to obtain Q first reference eigenvectors;
[0047] The generation unit is further configured to generate a second probability vector according to the similarity between the second eigenvector and the Q first reference eigenvectors, and the similarity between the Q first reference eigenvectors and the N fourth eigenvectors. The second probability vector is a C-dimensional vector, and the C-dimensional space of the second probability vector corresponds to C target label types;
[0048] A training unit, specifically used to train a learning model by using the contrast loss between the second feature vector and the fourth feature vector, and the sum of the losses of the first probability vector and the second probability vector.
[0049] In a possible implementation manner of the second aspect, a generation unit is further configured to generate a third probability vector based on the mapping between the first feature vector and C target label types. The third probability vector is a C-dimensional vector, and the C-dimensional space of the third probability vector corresponds to C target label types.
[0050] The training unit is further configured to train the learning model by using the contrast loss between the second feature vector and the fourth feature vector, the losses of the first probability vector and the second probability vector, and the sum of the logarithms of the third probability vector.
[0051] In a possible implementation manner of the second aspect, the training unit is specifically configured to train the learning model by using the contrast loss between the second feature vector and the fourth feature vector, the losses of the first probability vector and the second probability vector, and the weighted sum of the logarithms of the third probability vector.
[0052] The third aspect of the present application provides a computer device, including:
[0053] A memory, a transceiver, a processor, and a bus system;
[0054] Wherein, the memory is used to store programs;
[0055] The processor is configured to execute the programs in the memory, including executing the methods of the above aspects;
[0056] The bus system is used to connect the memory and the processor, so that the memory and the processor can communicate.
[0057] The fourth aspect of the present application provides a computer-readable storage medium, in which instructions are stored. When the instructions are run on a computer, the computer is made to execute the methods of the above aspects.
[0058] Another aspect of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above aspects.
[0059] From the above technical solutions, it can be seen that the embodiments of the present application have the following advantages:
[0060] The present application provides a data processing method and related devices. After determining the retrieval keyword, retrieve the image corresponding to the retrieval keyword and the description text of the image in the database based on the retrieval keyword, and use the distance between the feature vector of the description text and the feature vector of the retrieval keyword to obtain K description texts closely related to the retrieval keyword, calculate the mean value of the feature vectors of the K images to obtain the feature vector of the image prototype, where there is a corresponding relationship between the K images and the K description texts, and use the contrast loss between the feature vector of the image and the feature vector of the image prototype to train the learning model. Based on the understanding that images corresponding to similar retrieval keywords should also have similar features, find description texts similar to the retrieval keyword, use the images corresponding to these description texts to generate the feature vector of the image prototype, and consider that the feature vector of the image prototype carries the key features of the retrieval keyword, realizing cross-modal semantic alignment and largely eliminating the negative impact of semantic noise on model training. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 FIG. is a schematic architecture diagram of a data processing system provided by an embodiment of the present application;
[0062] Figure 2 FIG. is a schematic flow diagram of a data processing method provided by an embodiment of the present application;
[0063] Figure 3a FIG. is a schematic structural diagram of a data processing model provided by an embodiment of the present application;
[0064] Figure 3b FIG. is another schematic structural diagram of a data processing model provided by an embodiment of the present application;
[0065] Figure 4 FIG. is a schematic diagram of experimental data provided by an embodiment of the present application;
[0066] Figure 5 FIG. is another schematic diagram of experimental data provided by an embodiment of the present application;
[0067] Figure 6 FIG. is another schematic diagram of experimental data provided by an embodiment of the present application;
[0068] Figure 7 FIG. is another schematic diagram of experimental data provided by an embodiment of the present application;
[0069] Figure 8 FIG. is a schematic structural diagram of a data processing device provided by an embodiment of the present application;
[0070] Figure 9 FIG. is another schematic structural diagram of a data processing device provided by an embodiment of the present application;
[0071] Figure 10 This is a schematic structural diagram of the server provided by the embodiment of the present application. Specific implementation manners
[0072] The embodiment of the present application provides a data processing method and related devices, which are used to eliminate the negative impact of semantic noise on the image learning ability of a computer vision model when training the computer vision model.
[0073] Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "correspond to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0074] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, a theory, method, technology and application system that perceives the environment, acquires knowledge and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also to study the design principles and implementation methods of various intelligent machines, so that the machines have the functions of perception, reasoning and decision-making.
[0075] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model is also called the large model or the basic model, and can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0076] Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to machine vision that uses cameras and computers to replace human eyes for target recognition and measurement, etc., and further performs graphic processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technologies usually include image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc. technologies, and also include common biometric recognition technologies such as face recognition and fingerprint recognition.
[0077] To facilitate the understanding of the technical solutions provided by the embodiments of the present application, some key terms used in the embodiments of the present application are explained here first:
[0078] ResNet: A convolutional neural network in which the internal residual blocks use skip connections to alleviate the problem of vanishing gradients caused by increasing the depth in deep neural networks.
[0079] Instance-instance contrast loss: Different image enhancement methods are applied to an original image to obtain two input images from the same original image. Since image augmentation does not change the semantics of the image itself, it is considered that the feature representations of these two input images from the same original image should be as similar as possible (usually the cosine similarity is used for distance measurement), while the input images from different original images should be as far apart as possible.
[0080] Instance-prototype contrast loss: Replace the object of the instance-instance contrast loss mentioned above with the contrast loss between an instance and a prototype. Specifically, the prototype here refers to a sample vector maintained using multiple image instance samples belonging to the same category, representing the most prominent feature expression of this category.
[0081] Representation learning: A collection of techniques for learning a feature: converting the original data into a form that can be effectively exploited by machine learning. It avoids the trouble of manually extracting features, allows the computer to learn to use features while also learning how to extract features: learning how to learn.
[0082] Embedding space: Use a neural network model to map a high-dimensional space (for example, the three-dimensional coordinate point representation of an image in the RGB color space) to a low-dimensional space (the dimension is usually very small). In this low-dimensional space, each feature is no longer a point on the coordinate axis, but is scattered throughout the low-dimensional space. In machine learning and deep learning, this mapping process is called embedding, and the space after the mapping projection is called the embedding space.
[0083] Semantic alignment: The semantic alignment referred to in this patent mainly refers to the alignment of image modality and text modality, that is, multimodal data alignment. Specifically, given a picture and its corresponding category label and text, the information reflected in the image should be consistent with the information described in the text and consistent with the category label.
[0084] Label flipping noise: Given a set of categories A of interest, due to the Internet noise problem, the image sample x_i that originally belonged to category c_1 was incorrectly labeled as c_2, and at this time the c_2 label also belongs to A, so a "label flip" occurs.
[0085] Out-of-domain noise: Given a set of categories A of interest, due to the Internet noise problem, the image sample x_i that originally belonged to category c_1 was incorrectly labeled as c_2, and at this time the c_2 label does not belong to A. The true label of the image sample x_i is not within the scope of our interest and consideration, so it is called "OOD".
[0086] Semantic noise: Given a category c_1 that belongs to the set of interest A, the name of the category c_1 itself is ambiguous, resulting in a large number of sample images in the Internet graphs obtained based on this category name that do not match the semantics and whose content is inconsistent with the target semantics. These images occupy a large part of this category, making traditional noisy learning algorithms powerless.
[0087] The training of computer vision models requires the use of massive amounts of labeled data as training data. In order to increase the speed of collecting training data, you can use search engines or photo sharing software to search for keywords to obtain massive amounts of Internet images.
[0088] After collecting Internet images, a large number of high-quality samples can be obtained by cleaning the Internet images. The main cleaning basis is as follows:
[0089] 1) Images with higher retrieval rankings are usually more relevant to the search keywords and their image quality is usually better.
[0090] 2) Characteristics of the image itself, including image size, image information entropy and image format.
[0091] Based on the above cleaning basis, the keywords or keyword categories used when retrieving images are used as weak labels for Internet images behind the galaxy.
[0092] When using the cleaned Internet images to train a computer vision model, the cleaned Internet images can be used, or the cleaned Internet images can be further screened and the screened Internet images can be used to train the computer vision model.
[0093] In conventional solutions, methods such as using visual structures (such as neighbor information), clean sample sets, regularization means, confidence-based label bootstrapping, and utilization of image additional meta-information are used to suppress the negative impacts brought by label flipping noise or OOD noise on the training results of computer vision models.
[0094] However, during the process of training a computer vision model using a training data set, it is found that in addition to label flipping noise and OOD noise, there is also a type of noise that cannot be underestimated. This is semantic noise, which is caused by the inconsistency between the image content and the meaning represented by its attached text. For example, when retrieving images by using a search engine or a photo sharing software with keywords, if the retrieval keywords are polysemous or ambiguous words, the retrieved images, in addition to the expected images, also include images corresponding to semantic concepts not within the expectation. This means that most of the images corresponding to this retrieval keyword are noise. How to eliminate the negative impacts brought by semantic noise on model training has become an urgent problem to be solved currently.
[0095] Based on the above problems, this application proposes that after determining the retrieval keyword, the images corresponding to the retrieval keyword and the description text of the images can be retrieved from the database based on the retrieval keyword. By using the distance between the feature vector of the description text and the feature vector of the retrieval keyword, K description texts closely related to the retrieval keyword are obtained, and the mean value of the feature vectors of the K images is calculated to obtain the feature vector of the image prototype. Among them, there is a corresponding relationship between the K images and the K description texts. By using the contrast loss between the feature vector of the image and the feature vector of the image prototype, the learning model is trained. Based on the understanding that images corresponding to similar retrieval keywords should also have similar features, description texts similar to the retrieval keyword are found, and the feature vector of the image prototype is generated by using the images corresponding to these description texts. It is considered that the feature vector of the image prototype carries the key features of the retrieval keyword, realizing cross-modal semantic alignment and largely eliminating the negative impacts brought by semantic noise on model training.
[0096] For ease of understanding, please refer to Figure 1 , Figure 1 which is the application environment diagram of the data processing method in the embodiment of this application, as Figure 1As shown in the figure, the data processing method in the embodiment of the present application is applied to a data processing system. The data processing system includes: a server and a terminal device; where the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the embodiments of the present application do not limit this here.
[0097] The server first obtains N tags, where N is a positive integer;
[0098] Then, based on the N tags, N first feature vectors are generated;
[0099] Then, search for the N tags in the database to obtain N training images and description texts of the N training images;
[0100] Then, based on the N training images and the description texts of the N training images, N second feature vectors and N third feature vectors are generated. There is a corresponding relationship between the N second feature vectors and the N training images, and there is a corresponding relationship between the N third feature vectors and the description texts of the N training images;
[0101] Then, calculate the distances between the N third feature vectors and the N first feature vectors to obtain K third feature vectors whose distances are closest to any one of the N first feature vectors, where K is a positive integer and K is less than N;
[0102] Then, calculate the mean value of the K second feature vectors to obtain a fourth feature vector, and the K second feature vectors are the feature vectors corresponding to the K third feature vectors;
[0103] Finally, use the contrastive loss between the second feature vector and the fourth feature vector to train the learning model.
[0104] Next, from the perspective of the server, the data processing method in the present application will be introduced. Please refer to Figure 2 , the data processing method provided by the embodiment of the present application includes: steps S101 to S112. Specifically:
[0105] S101. Obtain N tags;
[0106] Obtain N tags typed by the target user through the interaction interface, where the N tags are N keywords used by the target user to search for training images in the database, and N is a positive integer.
[0107] Exemplarily, based on the search box of the search engine displayed in the interaction interface, the target user types in search keywords, such as "*cat", or, "*cat* continent". The N keywords typed by the target user are the N tags. For example, the keyword is "*cat", or, the keywords are "*cat" and "*continent".
[0108] Specifically, each of the N tags corresponds to a target tag type. Since there may be at least two tags with the same target tag type among the N tags, there can be C tags with different target tag types among the N tags, where C is a positive integer, and N is greater than or equal to C.
[0109] It can be understood that the description of the correspondence between the N tags and the C target tag types here is only an example. In actual applications, the relationship between the N tags and the C target tag types should be set in combination with the specific application scenario, and no limitation is made here.
[0110] S102. Generate N first feature vectors based on the N tags;
[0111] After obtaining the N tags, generate N first feature vectors according to the N tags.
[0112] Specifically, first perform text augmentation on the N tags to obtain N first texts, and then perform feature extraction on the N first texts to obtain N first feature vectors.
[0113] Exemplarily, taking the tag "*cat" as an example, searching for "*cat" in the semantic hierarchy tree (wordnet) obtains multiple results, including "a cat with striped fur, which is a domestic cat, pet cat, feline", "a medium-sized wild cat living in the central part of ** continent, belonging to the large felines in the forests of most regions", and "an endangered feline living in * continent, with black stripes and tawny fur". In this case, if the tag "*cat" is not further restricted, it will cause the problem of semantic ambiguity. Therefore, this application proposes that the definition text of the tag in word net can be used, and the context can be extended through the sister words (synonyms), sub-words (hyponyms), and parent words (hypernyms) of the tag to achieve text augmentation of the N tags and obtain N first texts.
[0114] In the foregoing example, the first text corresponding to "*cat" may be "a cat with striped fur, which is a domestic cat, a pet cat, a feline", or, "a medium-sized wild cat living in the central part of ** continent, belonging to the large felines in the forests of most regions", or, "an endangered feline living in * continent, with black stripes and tawny fur".
[0115] Then, feature extraction is performed on the N first texts to obtain N first feature vectors. Among them, the N labels can be represented by y i ∈{1, 2, …, N}, the N first feature vectors can be represented by s c , and the N first texts can be represented by t c .
[0116] It can be understood that the description of generating N first feature vectors based on N labels here is only an example. In actual applications, it should be set according to the specific application scenario, and there is no limitation here.
[0117] In the embodiments of the present application, in the process of text expansion of N labels to obtain N first texts, in the case where the N labels may be ambiguous, specific N first texts are obtained through text expansion, and feature extraction is performed on the N first texts to obtain N first feature vectors. The first feature vectors can be used for subsequent operations, and the ambiguity that the labels may bring has been eliminated in the first feature vectors. Semantic analysis is performed using word net, which solves the problem of semantic ambiguity that may be generated by the labels in the technical solution, thereby reducing the negative impact of semantic noise on the model training results.
[0118] It should be noted that there is no clear order between step S102 and step S103. The description in the embodiments of the present application is only an example. In actual applications, it should be set according to the specific application scenario and usage requirements, and there is no limitation here.
[0119] S103. Search for N labels in the database to obtain N training images and the description texts of the N training images;
[0120] After obtaining the N labels, search for the N labels in the database to obtain N training images and the description texts of the N training images. In the present application, it is considered that label A, training image A, and the description text A of training image A are a set of data, and label A is the label of training image A.
[0121] Among them, feature extraction of label A can obtain the first feature vector A, feature extraction of training image A can obtain the second feature vector A, find J second feature vectors similar to the second feature vector A, and calculate the mean of the J second feature vectors to obtain the third feature vector A.
[0122] For the above reasons, the present application believes that there is a corresponding relationship between any two of the first feature vector A, the second feature vector A, and the third feature vector A.
[0123] Exemplarily, N training images are retrieved from a search engine, along with the detailed pages corresponding to the N training images. The html tags, file format extensions, punctuation marks, numbers, stop words, and other characters in the detailed pages corresponding to the N training images are removed to obtain the description texts of the N training images.
[0124] Among them, the N training images can be represented by x i ∈{1, 2, …, N}, and the description texts of the N training images can be represented by t i ∈{1, 2, …, N}.
[0125] It can be understood that the description of obtaining the N training images and the description texts of the N training images here is only an example. In actual applications, it should be set according to the specific application scenarios, and there is no limitation here.
[0126] It should be noted that there is no clear sequence between step S102 and step S103. The description in the embodiments of the present application is only an example. In actual applications, it should be set according to the specific application scenarios and usage requirements, and there is no limitation here.
[0127] S104. Generate N second feature vectors and N third feature vectors based on the N training images and the description texts of the N training images;
[0128] After obtaining the N training images and the description texts of the N training images, generate N second feature vectors and N third feature vectors based on the N training images and the description texts of the N training images. Among them, the N second feature vectors can be represented by v i and the N third feature vectors can be represented by .
[0129] Exemplarily, feature extraction can be performed on the N training images to obtain N second feature vectors; feature extraction can be performed on the description texts of the N training images to obtain N fifth feature vectors (S i ).
[0130] Calculate the distances between the N second feature vectors to obtain the J feature vectors that are closest to any one of the N second feature vectors, where J is a positive integer less than N.
[0131] Specifically, the distance between the N second feature vectors can be obtained by calculating the cosine distance between the N second feature vectors. Since the N second feature vectors are the feature vectors of N images, that is, the cosine distance between the N second feature vectors can be calculated as the basis for judging the similarity between the N images. The specific formula is as follows:
[0132]
[0133] where v j is any one of the N second feature vectors except v i .
[0134] Since the neighbors ranked by the feature similarity evaluated using the cosine distance may not be the feature vectors with the highest similarity to the second feature vectors, a mutual k-nearest neighbor (K-reciprocal-NN) graph structure: G = {V, E} is constructed, and a reordering technique is used to evaluate the neighbor relationship. Each node in V represents an image sample, and the edge connectivity of E can be represented as an adjacency matrix A.
[0135]
[0136] where N(x i , j) and R(x i , j) = {x j |x j ∈N(x i , j) ∧ x i ∈N(x j , j)} represent the ordinary k-NN and k-reciprocal-NN graph structures respectively, and the k-NN graph structure is used to represent the distance distribution among the N training images.
[0137] N(x i , j) is for x i drawing the k-NN graph of x i with the N - 1 images other than x i , and the j images closest to x i in the k-NN graph are taken as the neighbor images of x i , and N(x i ) is taken as the neighbor image of x
[0138] Since x j ∈N(x i , j) means that x j is the neighbor image of x i , and x i ∈N(x j , j) means that x i is the neighbor image of x j , so R(xi , j) = {x j | x j ∈ N(x i , j) ∧ x i ∈ N(x j , j)} represents the case where x i and x j are neighbor images of each other.
[0139] Therefore, when x i and x j are neighbor images of each other, A ij = 1 - d(v i , v j );
[0140] When x i and x j are not neighbor images of each other, A ij = 0.
[0141] On this basis, the Jaccard distance can also be used to further calculate the distance (similarity) between N second feature vectors. The distance formula corrected by the Jaccard distance is as follows:
[0142]
[0143]
[0144]
[0145] where d * (v i , v j ) represents the distance between v i and v j , and can also be understood as the similarity between v i and v j . d * (v i , v j ) is the mean of d(v i , v j ) and d J (v i , v j ); d J (v i , v j ) is the Jaccard distance between v i and v j , which is used to describe the distance between two vectors using the cosine distance of the two vectors.
[0146] When x jis x i When it is a neighbor image of is v i and v j is the natural exponent of the negative cosine distance between;
[0147] When x j is a non - x i neighbor image, is 0.
[0148] v k is any one of the N second eigenvectors.
[0149] On this basis, in the way of graph convolution, calculate the mean value of the J fifth eigenvectors to obtain the third eigenvector. The J fifth eigenvectors are the fifth eigenvectors corresponding to the J second eigenvectors. The third eigenvector is included in the N third eigenvectors. The third eigenvector can be represented by represented.
[0150] The specific formula is as follows:
[0151]
[0152] Among them, is the adjacency matrix with self - connection, is the diagonal matrix, I N is the N - dimensional identity matrix.
[0153] In the embodiments of the present application, based on the characteristic that images with similar features will have similar description texts, first, use the distances between the second eigenvectors corresponding to the training images to find similar training images, and then calculate the mean value of the fifth eigenvectors to align the similarity of the description texts. Use the eigenvectors of different modalities as references for each other for feature alignment, and further ensure the consistency between the information carried by the eigenvectors of different modalities, thereby realizing multi - modal feature alignment.
[0154] It can be understood that the description of calculating the distances between the N second eigenvectors here is only an example. In actual applications, it should be set according to specific application scenarios and is not limited here.
[0155] S105. Calculate the distances between the N third eigenvectors and the N first eigenvectors, and obtain the K third eigenvectors with the closest distance to any one of the N first eigenvectors;
[0156] After generating N second feature vectors and N third feature vectors, calculate the distances between the N third feature vectors and the N first feature vectors, and obtain the K third feature vectors with the closest distances to any one of the N first feature vectors. Take the N feature vectors as text prototypes, and the third feature vectors with correct semantics can be selected by matching the query encoding and the sample key encoding to obtain the similarity.
[0157] Exemplarily, in order to further improve the accuracy of the solution, the method of calculating the distances between the N second feature vectors in the aforementioned step S104 is still used here to calculate the distances between the N third feature vectors and the N first feature vectors, and the specific operation method will not be elaborated here.
[0158] After calculating the distances between the N third feature vectors and the N first feature vectors, reorder the distances, sort the N images in ascending order according to the calculated distances, and in the sorting corresponding to each target label type, select the first K third feature vectors with the closest distances as the clean sample set D corresponding to the target label type. K 。
[0159]
[0160]
[0161] Among them, refers to the K smallest distances of the c-th category.
[0162] Since the N labels can be divided into C different categories, therefore, represents the images, text descriptions, and labels corresponding to the K third feature vectors with the closest distances to the feature vectors corresponding to the c-th category of labels.
[0163] x i is the training image, t i is the training text, y i is the label, and c is the label of the c-th category.
[0164] It can be understood that the description of obtaining the calculation of the distances between the N second feature vectors here is only an example. In actual applications, it should be set according to the specific application scenario, and no limitation is made here.
[0165] S106. Calculate the mean of the K second feature vectors to obtain the fourth feature vector;
[0166] After calculating the distances between the N third feature vectors and the N first feature vectors and obtaining the K third feature vectors with the closest distances to any one of the N first feature vectors, determine the K second feature vectors corresponding to the K third feature vectors. And calculate the mean of the K second feature vectors to obtain the fourth feature vector.
[0167] Exemplarily, the fourth eigenvector can be obtained through the mean of the clean sample subsets corresponding to each target label type. To further improve the accuracy of image feature extraction, it is also possible to represent z in the embedding space of all samples in to obtain the representation of the c-th target label type: i
[0168]
[0169] Among them, indicates the distance of the top K closest to the eigenvector corresponding to the label of the c-th category among the distance rankings of the eigenvectors corresponding to the labels of x i . The samples in are subjected to embedding space extraction to obtain z i .
[0170] Furthermore, based on the initialized image prototype (seventh eigenvector) z c , the fourth eigenvector can also be obtained by weighted summing any one of the K second eigenvectors and the seventh eigenvector according to the update criterion of the momentum method. The specific formula is as follows:
[0171]
[0172] Among them, is the initialization formula, is the correction formula, m p is a hyperparameter, and z i is the eigenvector obtained by performing embedding space extraction on the samples in .
[0173] In the embodiment of the present application, the seventh eigenvector is obtained by calculating the mean of the K second eigenvectors; the fourth eigenvector is obtained by weighted summing any one of the seventh eigenvector and the K second eigenvectors. The momentum method is used to update the fourth eigenvector, which improves the convergence speed of the fourth eigenvector and accelerates the approaching speed of the fourth eigenvector to the prototype eigenvector.
[0174] In the contrast learning between the second eigenvector and the fourth eigenvector, the training images of the same category can be made to approach their image prototypes, while the training images of different categories are pushed away in the embedding space. To improve the distinguishability in the feature dimension, a contrast loss between the second eigenvector and the sixth eigenvector is introduced to improve the distinguishability in the feature dimension. Among them, the sixth eigenvector is any one of the K second eigenvectors closest to the second eigenvector. The specific formula is as follows:
[0175]
[0176] is the loss function calculated for different tags. is the loss function calculated for different tags according to historical data.
[0177] where τ is the temperature coefficient; z i is for the embedding vector obtained by extracting the embedding space of the samples in is the representation of the training image of label y i in the embedding space, z c is for the representation of the training image with the C-th class label in the embedding space in i corresponding to z i z' i is x i after data augmentation to obtain x', i then, after feature extraction of x' i to obtain v', i after that, the mapping of v' i in the embedding space is z' i ; z' j corresponding to z j z' j is x j after data augmentation to obtain x' j then, after feature extraction of x' j to obtain v', j after that, v' j in the embedding space is z' j .
[0178] In the embodiments of the present application, by using the contrast loss between the second feature vector and the fourth feature vector, and the contrast loss between the second feature vector and the sixth feature vector, the learning model is trained. The sixth feature vector is any one of the K second feature vectors closest to the second feature vector. By introducing the contrast loss between the second feature vector and the sixth feature vector, the distinguishability in the feature dimension is improved, and the training effect of the learning model is improved.
[0179] It can be understood that the description of the method for obtaining the fourth feature vector here is only an example. In actual applications, it should be set according to specific application scenarios and is not limited here.
[0180] S107. Generate a third probability vector based on the mapping between the first feature vector and the C target label types;
[0181] After calculating the mean of the K second eigenvectors to obtain the fourth eigenvector, a third probability vector is generated based on the mapping between the first eigenvector and the C target label types.
[0182] Exemplarily, a third probability vector p is generated based on the mapping between the first eigenvector and the C target label types. i , where the third probability vector is a C-dimensional vector, and the C-dimensional spaces of the third probability vector correspond to the C target label types respectively. The C-dimensional space of the third probability vector represents the probability that the first eigenvector is mapped to the C-th target label type.
[0183] Specifically, since there may be damaged images or AI-generated images in the training images, etc., the present application proposes that the second eigenvector can be further subjected to feature extraction to obtain the embedding vector v corresponding to the second eigenvector. i , and only the core information related to the first eigenvector is included in the embedding vector corresponding to the second eigenvector. Therefore, the loss functions of the projector and the reconstructor corresponding to the feature extraction of the second eigenvector are as follows:
[0184]
[0185] Among them, is the probability vector of v i mapped to y i , is the probability vector of z in the embedding space i mapped to y i , is the eigenvector recovered by z i via the reconstructor.
[0186] In the embodiments of the present application, a third probability vector is generated based on the mapping between the first eigenvector and the C target label types. The third probability vector is a C-dimensional vector, and the C-dimensional spaces of the third probability vector correspond to the C target label types; the contrast loss between the second eigenvector and the fourth eigenvector, the loss between the first probability vector and the second probability vector, and the sum of the logarithms of the third probability vector are used to train the learning model. The probability distribution of the mapping between the first eigenvector and the C target label types is learned using the third probability vector, softening the labels and avoiding the difficulty in probability learning in the description method of either 0 or 1, thereby improving the learning efficiency of the model.
[0187] It can be understood that the description of generating the third probability vector here is only an example. In actual applications, it should be set according to the specific application scenarios, and no limitation is made here.
[0188] S108. Process the second eigenvector and the first eigenvector according to the similarity between the second eigenvector and the fourth eigenvector;
[0189] After calculating the mean of the K second eigenvectors to obtain the fourth eigenvector, the second eigenvector and the first eigenvector are processed according to the similarity between the second eigenvector and the fourth eigenvector.
[0190] Exemplarily, for the example introduced in the foregoing step S103, when the label A, the training image A, and the description text A of the training image A are a set of data, if the label A has multiple meanings, such as the label A is "*cat", which can either refer to "a cat with striped fur, which is a domestic cat, a pet cat, and a feline", or refer to "a medium-sized wild cat living in the central part of ** continent, belonging to the large felines in the forests of most regions".
[0191] When the correct semantics of the label A here is "a cat with striped fur, which is a domestic cat, a pet cat, and a feline", and the training image A is "a medium-sized wild cat living in the central part of ** continent, belonging to the large felines in the forests of most regions", the training image A is noise data.
[0192] Therefore, there will be such noise data in the second eigenvectors obtained by feature extraction of the training images.
[0193] At the same time, the first eigenvector is a text description prototype obtained directly by semantic analysis of N labels, and the fourth eigenvector is based on calculated, where. It is composed of the K third eigenvectors that are closest to the first eigenvector in distance. Therefore, the fourth eigenvector can represent the correct semantics of the label.
[0194] The distance between the correct data in the second eigenvector and the fourth eigenvector should be less than the distance between the noise data in the second eigenvector and the fourth eigenvector.
[0195] Therefore, based on the fourth eigenvector, the noise data in the second eigenvector can be effectively processed by combining the similarity between the feature vector (or embedding vector) of the training image and the feature vector (or embedding vector) of the visual prototype (the fourth eigenvector).
[0196] First, calculate the similarity between the second eigenvector and the fourth eigenvector, which is calculated by the following formula:
[0197]
[0198] r i(k) For the k-th label, the embedding vector z of the training image i and the embedding vector z of the visual prototype c between the similarities.
[0199] where τ is the temperature coefficient; z i is the feature vector obtained by extracting the embedding space of the samples in , z k is the representation of the training image of the k-th label in the embedding space, z c is the representation of the training image of the C-th label of the label in the embedding space.
[0200] Combined with the third probability vector, a target feature vector o is calculated to determine whether the similarity between the second feature vector and the fourth feature vector meets the preset conditions i , and the calculation formula is as follows:
[0201] o i = αp i + (1 - α)r i ;
[0202] where α is a hyperparameter, r i is the similarity between the embedding vector z of the training image i and the embedding vectors z of C visual prototypes c , p i is the probability vector of the labels of the embedding vector z of the training image i mapped to C categories.
[0203] After calculating the target feature vector o i , update the label corresponding to the training image according to the target feature vector to obtain the updated label
[0204] The update strategy of the updated label is as follows:
[0205]
[0206] where 0 ≤ γ ≤ 1.
[0207] When the training image is selected as the K third feature vectors closest to the target label type after being processed in the aforementioned steps S104 and S105, keep the label of the training image as the updated label;
[0208] When the maximum value of the target feature vector is higher than γ, the updated label is the label corresponding to the maximum value of the target feature vector;
[0209] When the confidence level of the output probability of any training image is higher than the average level, keep the label of the training image as the updated label;
[0210] Further, the updated label and the training image can be used as the inputs of step S101 and step 103, and the operations after step S104 proposed in this application are performed again;
[0211] The learning model is trained using the contrast loss between the second feature vector and the updated fourth feature vector.
[0212] In the embodiment of this application, when the similarity between the second feature vector and the fourth feature vector meets the preset requirements, the first feature vector corresponding to the second feature vector is updated to obtain the updated first feature vector; and the updated fourth feature vector is calculated using the updated first feature vector, the second feature vector, and the third feature vector; the learning model is trained using the contrast loss between the second feature vector and the updated fourth feature vector. During the training of the learning model, the label corresponding to the training image is continuously updated, improving the learning efficiency of the model.
[0213] Otherwise, the training image corresponding to the second feature vector is removed from the N training images to obtain N - 1 training images, and the updated second feature vector and the updated fourth feature vector are generated using the N - 1 training images, the description texts of the N - 1 training images, and the N - 1 labels, where the N - 1 labels are the labels corresponding to the N - 1 training images;
[0214] And the learning model is trained using the contrast loss between the updated second feature vector and the updated fourth feature vector.
[0215] In the embodiment of this application, when the similarity between the second feature vector and the fourth feature vector does not meet the preset requirements, the training image corresponding to the second feature vector is removed from the N training images to obtain N - 1 training images; the updated second feature vector and the updated fourth feature vector are generated using the N - 1 training images, the description texts of the N - 1 training images, and the N - 1 labels, where the N - 1 labels are the labels corresponding to the N - 1 training images; the learning model is trained using the contrast loss between the updated second feature vector and the updated fourth feature vector. During the training of the learning model, the noise in the training images is dynamically removed, further eliminating the noise in the noisy data and improving the learning efficiency and learning effect of the solution.
[0216] It can be understood that the description of processing the second feature vector and the first feature vector here is only an example. In practical applications, it should be set according to the specific application scenario, and no limitation is made here.
[0217] It should be noted that there is no clear sequence between step S108 and step S109. The description in the embodiment of this application is only an example. In practical applications, it should be set according to the specific application scenario and usage requirements, and no limitation is made here.
[0218] S109. Generate a first probability vector according to the similarity between N second feature vectors and N fourth feature vectors.
[0219] After calculating the mean of K second feature vectors to obtain the fourth feature vector, generate a first probability vector according to the similarity between N second feature vectors and N fourth feature vectors.
[0220] Exemplarily, generate a first probability vector q according to the similarity between N second feature vectors and N fourth feature vectors. i , the first probability vector is a C-dimensional vector. The C-dimensional space of the first probability vector corresponds to C target label types. The C-dimensional space of the first probability vector represents the probability that the second feature vector is mapped to the Cth target label type.
[0221] In the embodiments of the present application, by generating a first probability vector according to the similarity between N second feature vectors and N fourth feature vectors, the first probability vector is a C-dimensional vector, the C-dimensional space of the first probability vector corresponds to C target label types, and a learning model is trained using the contrast loss between the second feature vector and the fourth feature vector and the first probability vector. By calculating the similarity between the training image and the image prototype, the training image is pulled towards the corresponding image prototype, improving the classification effect of the learning model.
[0222] It can be understood that the description of generating the first probability vector here is only an example. In actual applications, it should be set according to the specific application scenario, and no limitation is made here.
[0223] It should be noted that there is no clear order between step S108 and step S109. The description in the embodiments of the present application is only an example. In actual applications, it should be set according to the specific application scenario and usage requirements, and no limitation is made here.
[0224] S110. Obtain Q reference images.
[0225] Obtain Q reference images from the visual dictionary. The visual dictionary is a set of the Q most recently accessed images maintained by the collective bootstrapping method, where Q is a positive integer.
[0226] Exemplarily, the visual dictionary stores the Q most recently accessed reference images. The embedding vectors (z′ j ) of the Q reference images are used as the keys of the visual dictionary, and the pseudo-labels obtained by the learning model predicting the Q reference images are the values of the visual dictionary. The embedding space representation (z i ) of the given training image is used as the query.
[0227] It is understandable that the description of obtaining Q reference images here is only an example. In actual applications, it should be set according to specific application scenarios and is not limited here.
[0228] S111. Generate a second probability vector according to the similarity between the second feature vector and Q first reference feature vectors, and the similarity between the Q first reference feature vectors and N fourth feature vectors;
[0229] After obtaining Q reference images, perform feature extraction on the Q reference images to obtain Q first reference feature vectors, and generate a second probability vector b according to the similarity between the second feature vector and the Q first reference feature vectors, and the similarity between the Q first reference feature vectors and N fourth feature vectors i , the second probability vector is a C-dimensional vector, the C-dimensional space of the second probability vector corresponds to C target label types, and the C-dimensional space of the second probability vector represents the probability that the second feature vector is mapped to the C-th target label type.
[0230] Since the learning model is prone to overfitting noise as the number of training rounds increases, therefore, this application proposes that a large number of labeled reference samples can be introduced to help evaluate the labels of current training samples to alleviate the occurrence of overfitting noise.
[0231] Exemplarily, based on the obtained Q reference images, calculate z by the bootstrap method i When using the Q reference data for bootstrap sampling, the probability vector b of matching the C class labels i The calculation formula is as follows:
[0232]
[0233]
[0234]
[0235] where α is a hyperparameter, z i is the embedding vector obtained by extracting the embedding space of the samples in ; z k is the representation of the training image of the k-th label in the embedding space, z c is the representation of the training image with the C-th label in the embedding space in i ; z′ j is the embedding space vector of the reference image, q′ i is the probability vector obtained by mapping the embedding vector z′ j to C label categories, and use the dot product calculation example between z′ i and z to calculate the contrast loss wij 。
[0236] Moreover, the optimization of the learning model is achieved by minimizing the KL loss function between the output of the learning model and the aforementioned bootstrap representation distribution. The formula of the loss function is as follows:
[0237]
[0238] In the embodiments of the present application, by introducing Q reference images in the visual dictionary, it helps to calculate z i During the bootstrap process, the distribution weights of the sampled labels are obtained to get z i which is the probability of C types of labels, improving the accuracy of the learning model in learning the labels corresponding to the image samples, avoiding the possible misguidance caused by the probability of either 0 or 1, and estimating the label of a current training image by referring to the labels of a large number of reference images, which can also improve the generalization of the model training results.
[0239] S112. Train the learning model by using the contrast loss between the second feature vector and the fourth feature vector, the loss between the first probability vector and the second probability vector, and the weighted sum of the logarithms of the third probability vector.
[0240] After generating the second probability vector according to the similarity between the second feature vector and Q first reference feature vectors, and the similarity between the Q first reference feature vectors and N fourth feature vectors, train the learning model by using the contrast loss between the second feature vector and the fourth feature vector, the loss between the first probability vector and the second probability vector, and the weighted sum of the logarithms of the third probability vector.
[0241] Exemplarily, in combination with the descriptions in the foregoing steps S101 to S111, the loss function of the learning model can also be the weighted sum result of multiple loss functions. The specific formula is as follows:
[0242]
[0243] In the embodiments of the present application, train the learning model by using the contrast loss between the second feature vector and the fourth feature vector, the loss between the first probability vector and the second probability vector, and the weighted sum of the logarithms of the third probability vector. Obtain the loss function of the learning model by weighting and summing multiple loss functions, and use the loss function of the learning model to train the learning model.
[0244] In the embodiments of the present application, after determining the retrieval keywords, the images corresponding to the retrieval keywords and the description texts of the images can be retrieved from the database based on the retrieval keywords. By using the distance between the feature vectors of the description texts and the feature vectors of the retrieval keywords, K description texts closely related to the retrieval keywords are obtained, and the mean value of the feature vectors of the K images is calculated to obtain the feature vector of the image prototype. Among them, there is a corresponding relationship between the K images and the K description texts. The learning model is trained using the contrastive loss between the feature vector of the image and the feature vector of the image prototype. Based on the understanding that images corresponding to similar retrieval keywords should also have similar features, description texts similar to the retrieval keywords are found, and the feature vector of the image prototype is generated using the images corresponding to these description texts. It is considered that the feature vector of the image prototype carries the key features of the retrieval keywords, achieving cross-modal semantic alignment and largely eliminating the negative impact of semantic noise on model training.
[0245] In addition, in the solution provided by the present application, the operation in step S104 ensures that similar images have similar text descriptions, which can not only eliminate semantic noise but also eliminate label flipping noise. And the operation in step S108 can eliminate out-of-domain noise. Therefore, the solution provided by the present application can also eliminate multiple types of noise from multiple perspectives.
[0246] In the foregoing Figure 2 introduced solution, the implementation process of the solution in the case of single-label classification was introduced. In the output of the single-label classification model, each image has one label or annotation.
[0247] In addition, there is also the possibility of multi-label classification. In a multi-label classification task, each image can contain (correspond to) multiple labels. Further, the image can also contain all labels.
[0248] Based on the above concept, it can be known from step S111 that the probability vector calculated for each training image is a vector of dimension C. Therefore, multi-label classification can be achieved by setting the number of output labels. For example, by setting the top 2 labels with the highest probabilities in the probability vector corresponding to the training image as the output, multi-label classification can be achieved.
[0249] For ease of understanding, a model for data processing in deep learning will be introduced below in combination with Figure 3a and Figure 3b The model is as Figure 3a shown and includes a data collector, a first text feature extractor, a second text feature extractor, a first image feature extractor, a text feature smoothing processor, a text aligner, a text prototype extractor, an image aligner, a second image feature extractor, a projector, a reconstructor, and a visual prototype extractor (image prototype extractor).
[0250] The data collector inputs the acquired image x i , the description text t of the image i and the extended text t of the label of the image c into the first image feature extractor, the second text feature extractor, and the first text feature extractor in sequence to obtain v i , S i and S c .
[0251] Input v i and S i into the text feature smoothing processor to obtain
[0252] Input and S c into the text aligner to obtain D K (x i , y i );
[0253] Input D K (x i , y i ) into the second image feature extractor to obtain vi;
[0254] Input v i into the projector, and after being processed by the projector and the reconstructor, obtain z i ;
[0255] According to z i , obtain the image prototype z c .
[0256] The process of the model for label assignment to the image is as Figure 3b shown, including: an image collector, a siamese feature encoder, a first projector, a second projector, an image prototype, and an image dictionary;
[0257] After the image collector acquires the image x i , it inputs the image into the first feature encoder and the second feature encoder in the siamese feature encoder. The first feature encoder is a feature encoder updated using the gradient method, and the second feature encoder is a feature encoder updated based on the momentum-based moving average method.
[0258] Input x i into the first feature encoder to obtain v i ;
[0259] Input the feature-enhanced x i ’ into the second feature encoder to obtain v i ’;
[0260] Input vi , obtain z i ;
[0261] Input v into the second projector i ’, obtain z i ’;
[0262] Utilize z i , z i ’ and the image prototype to calculate and obtain the contrast loss between z i and z c ;
[0263] Utilize z i , z i ’ and the image dictionary to calculate and obtain the contrast loss between z i and z j ’.
[0264] In addition, utilize the image prototype and the image dictionary to calculate and obtain the contrast loss between z j ’ and z c ;
[0265] On this basis, combine the image dictionary, the q j ’ output by the classifier and the contrast loss between z i and z j ’ to implement the collective bootstrap method to obtain the probability distribution b i ;
[0266] Combine b i and the probability distribution q i of z i calculated by the classifier to obtain the bootstrap loss D KL (q i || b i ).
[0267] In addition, utilize the probability distribution p i of v i calculated by the classifier, and utilize p i and the result of noise removal to obtain the classification loss
[0268] In addition, utilize v i and to obtain the reconstruction loss Obtained for reconstructing z i .
[0269] The following introduces the effects brought by the solution proposed in this application in combination with experimental data:
[0270] Please refer toFigure 4 For different models, the top1 / top5 accuracies of these models on the databases obtained by network sampling and the databases obtained by manual labeling are reported.
[0271] Due to the different choices of image encoders, the model of cross-modal noisy learning proposed in this application is compared with the model using the cross-entropy loss training method here. Figure 4 As can be seen, the model of cross-modal noisy learning proposed in this technical solution has achieved quite competitive performance on the database obtained by network sampling, with a 1.6% improvement (top1 accuracy) over the model using the cross-entropy loss training method. In addition, the model of cross-modal noisy learning has obtained a 1% gain (top1 accuracy) on the database obtained by manual labeling, proving that the model of cross-modal noisy learning is robust to the domain gap between the network and real-world datasets.
[0272] Please refer to Figure 5 , and the F1 (C-F1), overall F1 (O-F1), and mean average precision (mAP) of each model are compared.
[0273] Most previous multi-label methods were developed for true GT labels, while the solution provided in this application is web graph noisy learning. In this case, the model of cross-modal noisy learning is compared with the model using the cross-entropy loss training method that is trained using the database obtained by network sampling, and is evaluated on a clean test set. The model of cross-modal noisy learning has significantly increased by 1.5% (C-F1), 3.0% (O-F1), and 9.7% (mAP) compared to the base version used as a reference.
[0274] Please refer to Figure 6 , for the performance of the solution provided in this application in open-set recognition on the database obtained by single-label network sampling.
[0275] To verify whether the model of cross-modal noisy learning can identify anomalies of unknown classes, an experimental setup was carried out on the open-set detection task. Specifically, the model of cross-modal noisy learning and the model using the cross-entropy loss training method were trained on the database obtained by network sampling, and the database obtained by manual labeling was used as the test set for verification. The union of the image classes of the database obtained by network sampling and the image classes of the database obtained by manual labeling has 500 more image classes than the intersection of the image classes of the two, and these 500 image classes are considered to belong to the open classes. Figure 6 As can be seen, even in the presence of open classes, the model of cross-modal noisy learning still has a significant improvement in the (C-F1) index. Therefore, it can be considered that the model of cross-modal noisy learning has all-round superiority compared to the model using the cross-entropy loss training method.
[0276] Please refer to Figure 7 for the ablation experiment performances of text encoding, text augmentation, label bootstrapping, etc. involved in the cross-modal noisy learning model. Label bootstrapping can estimate the labels of training samples based on multiple labels that can be obtained from the training samples, obtain the distribution of the labels of the training samples, and perform statistical inference on the distribution of the labels of the training samples to obtain the weights (probabilities) of different labels mapped by the training images.
[0277] The ablation experiments for all the modules introduced in this technical solution show that the modules proposed in this technical solution are all effective. For example, semantic noise of lexical confusion can be identified and eliminated (such as food images in "drumstick" and team images in "*cat").
[0278] The following will describe the data processing device in this application in detail. Please refer to Figure 8 . Figure 8 It is a schematic diagram of an embodiment of the data processing device 10 in the embodiment of this application. The data processing device 10 includes:[[]]
[0279] An acquisition unit 110, configured to acquire N labels, where N is a positive integer;
[0280] A generation unit 120, configured to generate N first feature vectors based on the N labels;
[0281] A search unit 130, configured to search for N labels in the database to obtain N training images and descriptive texts of the N training images;
[0282] The generation unit 120 is further configured to generate N second feature vectors and N third feature vectors based on the N training images and the descriptive texts of the N training images. There is a corresponding relationship between the N second feature vectors and the N training images, and there is a corresponding relationship between the N third feature vectors and the descriptive texts of the N training images;
[0283] A calculation unit 140, configured to calculate the distances between the N third feature vectors and the N first feature vectors, and obtain K third feature vectors that are closest to any one of the N first feature vectors, where K is a positive integer and K is less than N;
[0284] The calculation unit 140 is further configured to calculate the mean value of the K second feature vectors to obtain a fourth feature vector, where the K second feature vectors are the feature vectors corresponding to the K third feature vectors;
[0285] A training unit 150, configured to train a learning model using the contrast loss between the second feature vectors and the fourth feature vectors.
[0286] Optionally, the generating unit 120 is specifically configured to:
[0287] Perform text augmentation on N labels to obtain N first texts;
[0288] Perform feature extraction on the N first texts to obtain N first feature vectors.
[0289] Optionally, the generating unit 120 is specifically configured to:
[0290] Perform feature extraction on N training images to obtain N second feature vectors;
[0291] Perform feature extraction on the description texts of the N training images to obtain N fifth feature vectors;
[0292] Calculate the distances between the N second feature vectors and any one of the N second feature vectors, and obtain J second feature vectors that are the closest to any one of the N second feature vectors, where J is a positive integer less than N;
[0293] Calculate the mean of the J fifth feature vectors to obtain a third feature vector, where the J fifth feature vectors are the fifth feature vectors corresponding to the J second feature vectors, and the third feature vector is included in the N third feature vectors.
[0294] In an alternative embodiment of the data processing apparatus provided in the corresponding embodiment of the present application, please refer to Figure 8 The data processing apparatus 10 further includes a deleting unit 160, configured to remove the training image corresponding to the second feature vector from the N training images to obtain N - 1 training images if the similarity between the second feature vector and the fourth feature vector does not meet the preset requirement; Figure 9 The training unit 150 is specifically configured to:
[0295] Generate an updated second feature vector and an updated fourth feature vector by using the N - 1 training images, the description texts of the N - 1 training images, and the N - 1 labels, where the N - 1 labels are the labels corresponding to the N - 1 training images;
[0296] Train the learning model by using the contrast loss between the updated second feature vector and the updated fourth feature vector.
[0297]
[0298] Optionally, the apparatus further includes an updating unit 170, configured to update the first feature vector corresponding to the second feature vector to obtain an updated first feature vector if the similarity between the second feature vector and the fourth feature vector meets the preset requirement;
[0299] The training unit 150 is specifically configured to:
[0300] Calculate the updated fourth eigenvector by using the updated first, second, and third eigenvectors.
[0301] Train the learning model by using the contrastive loss between the second eigenvector and the updated fourth eigenvector.
[0302] Optionally, the training unit 150 is specifically configured to:
[0303] Train the learning model by using the contrastive loss between the second eigenvector and the fourth eigenvector, and the contrastive loss between the second eigenvector and the sixth eigenvector, where the sixth eigenvector is any one of the K second eigenvectors closest to the second eigenvector.
[0304] Optionally, the calculation unit 140 is specifically configured to:
[0305] Calculate the mean of the K second eigenvectors to obtain the seventh eigenvector.
[0306] Perform a weighted sum of the seventh eigenvector and any one of the K second eigenvectors to obtain the fourth eigenvector.
[0307] Optionally, the label type of any one of the N labels is included in C target label types, where C is a positive integer.
[0308] The generation unit 120 is further configured to generate a first probability vector according to the similarity between the N second eigenvectors and the N fourth eigenvectors. The first probability vector is a C-dimensional vector, and the C-dimensional space of the first probability vector corresponds to the C target label types.
[0309] The training unit 150 is specifically configured to train the learning model by using the contrastive loss between the second eigenvector and the fourth eigenvector and the first probability vector.
[0310] Optionally, the acquisition unit 110 is further configured to acquire Q reference images, where Q is a positive integer.
[0311] The generation unit 120 is further configured to perform feature extraction on the Q reference images to obtain Q first reference eigenvectors.
[0312] The generation unit 120 is further configured to generate a second probability vector according to the similarity between the second eigenvector and the Q first reference eigenvectors, and the similarity between the Q first reference eigenvectors and the N fourth eigenvectors. The second probability vector is a C-dimensional vector, and the C-dimensional space of the second probability vector corresponds to the C target label types.
[0313] The training unit 150 is specifically configured to train the learning model by using the contrast loss between the second feature vector and the fourth feature vector, and the sum of the losses of the first probability vector and the second probability vector.
[0314] Optionally, the generating unit 120 is further configured to generate a third probability vector based on the mapping between the first feature vector and C target label types. The third probability vector is a C-dimensional vector, and the C-dimensional space of the third probability vector corresponds to C target label types.
[0315] The training unit 150 is specifically configured to train the learning model by using the contrast loss between the second feature vector and the fourth feature vector, the losses of the first probability vector and the second probability vector, and the sum of the logarithms of the third probability vector.
[0316] Optionally, the training unit 150 is specifically configured to train the learning model by using the contrast loss between the second feature vector and the fourth feature vector, the losses of the first probability vector and the second probability vector, and the weighted sum of the logarithms of the third probability vector.
[0317] Figure 10 FIG. is a schematic structural diagram of a server provided by an embodiment of the present application. The server 300 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 322 (for example, one or more processors) and a memory 332, and one or more storage media 330 (for example, one or more mass storage devices) for storing application programs 342 or data 344. Among them, the memory 332 and the storage media 330 may be transient storage or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 322 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the server 300.
[0318] The server 300 may further include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input / output interfaces 358, and / or one or more operating systems 341, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0319] The steps performed by the server in the above embodiments may be based on the Figure 10 server structure shown.
[0320] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0321] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the functions of that module or unit.
[0322] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.
[0323] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0324] In addition, the functional units in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0325] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0326] In the above, the above embodiments are only used to illustrate the technical solutions of this application, rather than to limit it; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of various embodiments of this application.
Claims
1. A method for data processing, characterized in that, Including: Generating N first feature vectors based on N tags, where N is a positive integer; Searching for the N tags in a database to obtain N training images and descriptive texts of the N training images; Generating N second feature vectors and N third feature vectors based on the N training images and the descriptive texts of the N training images, where there is a corresponding relationship between the N second feature vectors and the N training images, and there is a corresponding relationship between the N third feature vectors and the descriptive texts of the N training images; Calculating the distances between the N third feature vectors and the N first feature vectors to obtain K third feature vectors that are closest in distance to any one of the N first feature vectors, where K is a positive integer and K is less than N; Calculating the mean of the K second feature vectors to obtain a fourth feature vector, where the K second feature vectors are the feature vectors corresponding to the K third feature vectors; Training a learning model using the contrast loss between the second feature vectors and the fourth feature vector.
2. The method according to claim 1, wherein The generating N first feature vectors based on the N tags includes: Performing text augmentation on the N tags to obtain N first texts; Performing feature extraction on the N first texts to obtain the N first feature vectors.
3. The method according to claim 1 or 2, characterized in that, The generating N second feature vectors and third feature vectors based on the N training images and the descriptive texts of the N training images includes: Performing feature extraction on the N training images to obtain the N second feature vectors; Performing feature extraction on the descriptive texts of the N training images to obtain N fifth feature vectors; Calculating the distances between the N second feature vectors and any one of the N second feature vectors to obtain J second feature vectors that are closest in distance to any one of the N second feature vectors, where J is a positive integer less than N; Calculating the mean of the J fifth feature vectors to obtain a third feature vector, where the J fifth feature vectors are the fifth feature vectors corresponding to the J second feature vectors.
4. The method according to claim 1 or 2, characterized in that, The method further includes: If the similarity between the second feature vector and the fourth feature vector does not meet a preset requirement, removing the training image corresponding to the second feature vector from the N training images to obtain N - 1 training images; The training the learning model using the contrast loss between the second feature vectors and the fourth feature vector includes: Generating updated second feature vectors and an updated fourth feature vector using the N - 1 training images, the descriptive texts of the N - 1 training images, and N - 1 tags, where the N - 1 tags are the tags corresponding to the N - 1 training images; Training the learning model using the contrast loss between the updated second feature vectors and the updated fourth feature vector.
5. The method according to claim 4, wherein The method further includes: If the similarity between the second feature vector and the fourth feature vector meets the preset requirement, updating the first feature vector corresponding to the second feature vector to obtain an updated first feature vector; Training the learning model by using the contrastive loss between the second feature vector and the fourth feature vector includes: Calculating the updated fourth feature vector by using the updated first feature vector, the second feature vector, and the third feature vector; Training the learning model by using the contrastive loss between the second feature vector and the updated fourth feature vector.
6. The method according to claim 1 or 2, characterized in that The training of the learning model by using the contrastive loss between the second feature vector and the fourth feature vector includes: Training the learning model by using the contrastive loss between the second feature vector and the fourth feature vector and the contrastive loss between the second feature vector and the sixth feature vector, where the sixth feature vector is any one of the K second feature vectors closest to the second feature vector.
7. The method according to claim 1 or 2, characterized in that, The calculating of the mean of the K second feature vectors to obtain the fourth feature vector includes: Calculating the mean of the K second feature vectors to obtain the seventh feature vector; Weighted summing the seventh feature vector and any one of the K second feature vectors to obtain the fourth feature vector.
8. The method according to claim 1 or 2, characterized in that, The label type of any one of the N labels is included in C target label types, where C is a positive integer; The method further includes: Generating a first probability vector according to the similarity between the N second feature vectors and the N fourth feature vectors, where the first probability vector is a C-dimensional vector, and the C-dimensional space of the first probability vector corresponds to the C target label types; The training of the learning model by using the contrastive loss between the second feature vector and the fourth feature vector includes: Training the learning model by using the contrastive loss between the second feature vector and the fourth feature vector and the first probability vector.
9. The method according to claim 8, wherein The method further includes: Obtaining the Q reference images, where Q is a positive integer; Performing feature extraction on the Q reference images to obtain Q first reference feature vectors; Generating a second probability vector according to the similarity between the second feature vector and the Q first reference feature vectors and the similarity between the Q first reference feature vectors and the N fourth feature vectors, where the second probability vector is a C-dimensional vector, and the C-dimensional space of the second probability vector corresponds to the C target label types; The training of the learning model by using the contrastive loss between the second feature vector and the fourth feature vector and the first probability vector includes: Training the learning model by using the contrastive loss between the second feature vector and the fourth feature vector and the sum of the losses of the first probability vector and the second probability vector.
10. The method according to claim 9, characterized in that, The method further includes: Generating a third probability vector based on the mapping between the first feature vector and the C target label types, where the third probability vector is a C-dimensional vector, and the C-dimensional space of the third probability vector corresponds to the C target label types; The training of the learning model by using the contrastive loss between the second feature vector and the fourth feature vector and the sum of the losses of the first probability vector and the second probability vector includes: Train the learning model using the contrast loss between the second feature vector and the fourth feature vector, the loss between the first probability vector and the second probability vector, and the sum of the logarithms of the third probability vectors.
11. The method according to claim 10, characterized in that, The step of training the learning model using the contrast loss between the second feature vector and the fourth feature vector, the loss between the first probability vector and the second probability vector, and the sum of the logarithms of the third probability vectors includes: Train the learning model using the contrast loss between the second feature vector and the fourth feature vector, the loss between the first probability vector and the second probability vector, and the weighted sum of the logarithms of the third probability vectors.
12. A data processing device, characterized in that, Comprising: A generating unit, configured to generate N first feature vectors based on the N labels, where N is a positive integer; A searching unit, configured to search for the N labels in a database to obtain N training images and description texts of the N training images; The generating unit is further configured to generate N second feature vectors and N third feature vectors based on the N training images and the description texts of the N training images, where there is a corresponding relationship between the N second feature vectors and the N training images, and there is a corresponding relationship between the N third feature vectors and the description texts of the N training images; A calculating unit, configured to calculate the distances between the N third feature vectors and the N first feature vectors, and obtain K third feature vectors whose distances are closest to any one of the N first feature vectors, where K is a positive integer and K is less than N; The calculating unit is further configured to calculate the mean of the K second feature vectors to obtain a fourth feature vector, where the K second feature vectors are the feature vectors corresponding to the K third feature vectors; A training unit, configured to train the learning model using the contrast loss between the second feature vector and the fourth feature vector.
13. A computer device, characterized in that, Comprising: A memory, a transceiver, a processor, and a bus system; Wherein, the memory is used to store programs; The processor is configured to execute the programs in the memory, including executing the data processing method according to any one of claims 1 to 11; The bus system is used to connect the memory and the processor to enable communication between the memory and the processor.
14. A computer-readable storage medium, including instructions, which when running on a computer, cause the computer to execute the data processing method according to any one of claims 1 to 11.
15. A computer program product, comprising a computer program, characterized in that, The computer program is executed by the processor to perform the data processing method according to any one of claims 1 to 11.