An image recognition method, device, electronic device and storage medium
Through joint training, the image classification model of the embedding module and the deep semantic embedding module is learned, and the time-consuming and redundant problems caused by separate training of image features is solved, achieving more efficient and accurate image classification and retrieval.
Patent Information
- Application Number
- CN202111089047.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-16
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-09-16
AI Technical Summary
In the prior art, separate training of image feature extraction results in a long time-consuming training process and redundant model parameters, affecting the efficiency and accuracy of image classification.
The image classification joint model trained by the metric learning embedding module and the deep semantic embedding module is adopted to extract the shallow global features and deep semantic features of the image, and to fusion the feature to improve the accuracy and efficiency of image classification.
Through joint learning, the effect of feature extraction is enhanced, time-consuming, and the accuracy and efficiency of image classification are improved. It is suitable for image classification and retrieval tasks.
Smart Images

Figure CN114283316B_ABST
Abstract
Description
Background Art
[0002] With the rapid development of Internet technology and the rapid growth of multimedia resources, Internet content has gradually evolved from mainstream text information to various types of multimedia data; among them, with the continuous update of electronic devices and the continuous simplification and update of imaging equipment, image data resources have increased rapidly.
[0003] In fields such as image retrieval, it is often necessary to perform relevant search tasks based on the overall global features and main semantic features of an image. Therefore, how to efficiently and accurately extract the features of an image has become an urgent problem to be solved. Summary of the Invention
[0004] Embodiments of the present application provide an image recognition method, apparatus, electronic device, and storage medium to improve the accuracy and efficiency of image classification.
[0005] An image recognition method provided by an embodiment of the present application includes:
[0006] Based on the metric learning embedding module in the trained image classification joint model, extracting a first shallow global feature of the image to be detected, where the first shallow global feature characterizes the global basic information of the image to be detected; and
[0007] Based on the deep semantic embedding module in the image classification joint model, extracting a first deep semantic feature of the image to be detected, where the first deep semantic feature characterizes the image semantic information of the image to be detected;
[0008] Based on the semantic prediction module in the image classification joint model, performing feature fusion on the first shallow global feature and the first deep semantic feature, and performing multi-label classification on the image to be detected based on the obtained first fusion feature to obtain the classification result of the image to be detected;
[0009] Wherein, the image classification joint model is obtained by jointly training the metric learning embedding module and the deep semantic embedding module.
[0010] An image recognition apparatus provided by an embodiment of the present application includes:
[0011] A first feature extraction unit, configured to extract a first shallow global feature of the image to be detected based on the metric learning embedding module in the trained image classification joint model, where the first shallow global feature characterizes the global basic information of the image to be detected; and
[0012] A second feature extraction unit, configured to extract a first deep semantic feature of the image to be detected based on the deep semantic embedding module in the image classification joint model, where the first deep semantic feature characterizes the image semantic information of the image to be detected;
[0013] A classification prediction unit, configured to perform feature fusion on the first shallow global feature and the first deep semantic feature based on the semantic prediction module in the image classification joint model, and perform multi-label classification on the image to be detected based on the obtained first fusion feature to obtain the classification result of the image to be detected;
[0014] Wherein, the image classification joint model is obtained by jointly training the metric learning embedding module and the deep semantic embedding module.
[0015] Optionally, the apparatus further includes:
[0016] An image retrieval unit, configured to respectively extract the second shallow global features of each candidate image in the image library based on the metric learning embedding module; and respectively extract the second deep semantic features of the image to be detected based on the deep semantic embedding module;
[0017] Based on the semantic prediction module, perform feature fusion on each second shallow global feature and the corresponding second deep semantic feature to obtain the second fusion feature corresponding to each candidate image;
[0018] Based on the first fusion feature and the second fusion features corresponding to the respective candidate images, determine the similarity between each candidate image and the image to be detected;
[0019] Based on each similarity, determine the similar image corresponding to the image to be detected.
[0020] Optionally, the image classification joint model further includes: a basic convolution module; the apparatus further includes:
[0021] A basic processing unit, configured to input the image to be detected into the basic convolution module in the image classification joint model before the first feature extraction unit extracts the first shallow global feature of the image to be detected based on the metric learning embedding module in the trained image classification joint model, and perform global feature extraction on the image to be detected based on the basic convolution module to obtain the global feature map corresponding to the image to be detected;
[0022] The first feature extraction unit is specifically configured to:
[0023] Perform embedding processing on the global feature map based on the metric learning embedding module to obtain the first shallow global feature corresponding to the image to be detected.
[0024] Optionally, the deep semantic embedding module includes a convolutional layer and an embedding layer composed of a residual structure; the second feature extraction unit is specifically configured to:
[0025] Based on the convolutional layer in the deep semantic embedding module, feature extraction is performed on the semantic information in the global feature map, and based on the embedding layer in the deep semantic embedding module, embedding processing is performed on the extracted semantic information to obtain the first deep semantic feature corresponding to the image to be detected.
[0026] Optionally, the learning rate of the semantic prediction module is higher than that of other modules, and the other modules include the basic convolutional module, the metric learning embedding module, and the deep semantic embedding module.
[0027] Optionally, the basic convolutional module includes multiple convolutional layers; the network parameters corresponding to the convolutional layers in the basic convolutional module are obtained by initializing based on the parameters pre-trained on a specified sample library, and the network parameters corresponding to the convolutional layers in the deep semantic embedding module are obtained by random initialization.
[0028] Optionally, the apparatus further includes:
[0029] A training unit, configured to train the image classification joint model in the following manner:
[0030] Obtain a training sample data set, and select a training sample group from the training sample data set;
[0031] Input the selected training sample group into the trained image classification joint model, and obtain the third shallow global feature output by the metric learning embedding module in the image classification joint model, the third deep semantic feature output by the deep semantic embedding module, and the classification vector output by the semantic prediction module;
[0032] Construct an objective loss function based on the third shallow global feature, the third deep semantic feature, and the classification vector, and perform multiple adjustments on the network parameters of the image classification joint model based on the objective loss function until the image classification joint model converges, and output the trained image classification joint model.
[0033] Optionally, the training unit is specifically configured to:
[0034] Construct a first triplet loss function based on the third shallow global features corresponding to the respective training samples in the training sample group; and construct a second triplet loss function based on the third deep semantic features corresponding to the respective training samples;
[0035] Construct a multi-label loss function based on the respective classification vectors corresponding to the respective training samples and the corresponding multi-label vectors, where the classification vector represents the predicted probability of the training sample for each classification label, and the multi-label vector represents the true probability of the training sample for each classification label;
[0036] Obtain a classification entropy loss function based on the classification vectors corresponding to each of the training samples.
[0037] Perform a weighted sum of the first triple loss function, the second triple loss function, the multi-label loss function, and the classification entropy loss function to obtain an objective loss function.
[0038] Optionally, each training sample group includes three training samples: an anchor sample, a positive sample, and a negative sample; specifically, the training unit is configured to:
[0039] Determine a first distance between the third shallow global feature corresponding to the anchor sample in the training sample group and the third shallow global feature corresponding to the positive sample, and a second distance between the third shallow global feature corresponding to the anchor sample and the third shallow global feature corresponding to the negative sample;
[0040] Use the sum of the difference between the first distance and the second distance and a specified boundary value as a first target value, where the specified boundary value is used to represent the difference boundary of the similarity between the positive sample and the negative sample;
[0041] Use the maximum value among the first target value and the reference parameters as the first triple loss function.
[0042] Optionally, each training sample group includes three training samples: an anchor sample, a positive sample, and a negative sample; specifically, the training unit is configured to:
[0043] Based on a third distance between the third deep semantic feature corresponding to the anchor sample in the training sample group and the third deep semantic feature corresponding to the positive sample, and a fourth distance between the third deep semantic feature corresponding to the anchor sample and the third deep semantic feature corresponding to the negative sample;
[0044] Use the sum of the difference between the third distance and the fourth distance and a specified boundary value as a second target value, where the specified boundary value is used to represent the difference boundary of the similarity between the positive sample and the negative sample;
[0045] Use the maximum value among the second target value and the reference parameters as the second triple loss function.
[0046] Optionally, the training unit is further configured to:
[0047] Before adjusting the network parameters of the image classification joint model multiple times based on the objective loss function, perform a weighted sum of the first triple loss function, the multi-label loss function, and the classification entropy loss function to obtain an intermediate loss function;
[0048] Based on the intermediate loss function, the network parameters of the deep semantic module in the image classification joint model are adjusted a specified number of times.
[0049] Optionally, the device includes:
[0050] A sample construction unit, configured to obtain a training sample group in the training sample data set by the following method:
[0051] According to the similarity between the collected sample images, every two sample images with a similarity higher than a specified threshold are combined into a positive sample pair;
[0052] Based on each positive sample pair, sample recombination is performed to obtain multiple training sample groups for constituting the training sample data set. Each training sample group includes three training samples: an anchor sample, a positive sample, and a negative sample. Among them, the anchor sample and the positive sample are sample images with a similarity higher than the specified threshold, and the negative sample and the positive sample are sample images with a similarity not higher than the specified threshold.
[0053] Optionally, the sample construction unit is specifically configured to:
[0054] Select one sample image in a positive sample pair as a target sample, and respectively select one sample image from other sample pairs as candidate samples;
[0055] Respectively determine the distances between the target sample and each candidate sample;
[0056] Sort the candidate samples according to their respective corresponding distances, and select at least one candidate sample at a specified sorting position as the negative sample corresponding to the target sample;
[0057] Respectively combine each selected negative sample with the positive sample pair to form a training sample group, where the target sample in the positive sample pair is the positive sample in the training sample group, and the other sample in the positive sample pair is the anchor sample in the training sample group.
[0058] An electronic device provided by an embodiment of the present application includes a processor and a memory. Among them, the memory stores program code, and when the program code is executed by the processor, the processor is caused to execute the steps of any one of the above image recognition methods.
[0059] An embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the steps of any one of the above image recognition methods.
[0060] An embodiment of the present application provides a computer-readable storage medium, which includes program code. When the program product runs on an electronic device, the program code is used to cause the electronic device to execute the steps of any one of the above image recognition methods.
[0061] The beneficial effects of the present application are as follows:
[0062] An embodiment of the present application provides an image recognition method, device, electronic device, and storage medium. Since the present application proposes a model structure that is beneficial to the learning of two types of features by a metric learning embedding module for extracting shallow global features and a deep semantic embedding module for extracting deep semantic features, when classifying a to-be-detected image based on the image joint classification model in the present application, the image can be represented by jointly using both the underlying image representation and the deep semantic representation, enhancing the joint learning effect of the two, and reducing the time-consuming of feature extraction. Image classification can be performed based on the obtained image fusion features, and a more accurate classification result can be obtained, improving the accuracy and efficiency of image classification.
[0063] Other features and advantages of the present application will be described in the following specification, and, in part, will become apparent from the specification or will be understood by implementing the present application. The objectives and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the written specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0065] Figure 1 is a schematic diagram of an application scenario in an embodiment of the present application;
[0066] Figure 2 is a flowchart of the implementation of an image recognition method in an embodiment of the present application;
[0067] Figure 3 is a flowchart of the implementation of an image retrieval method in an embodiment of the present application;
[0068] Figure 4 is a schematic structural diagram of an image classification joint model in an embodiment of the present application;
[0069] Figure 5 is a schematic structural diagram of a residual module in an embodiment of the present application;
[0070] Figure 6It is a flowchart of another image recognition method in the embodiments of the present application;
[0071] Figure 7 It is a flowchart of a method for constructing a training sample data set in the embodiments of the present application;
[0072] Figure 8 It is a flowchart of a method for training an image classification joint model in the embodiments of the present application;
[0073] Figure 9 It is a flowchart of a method for calculating a loss function in the embodiments of the present application;
[0074] Figure 10 It is a schematic diagram of the composition structure of an image recognition device in the embodiments of the present application;
[0075] Figure 11 It is a schematic diagram of a hardware composition structure of an electronic device applying the embodiments of the present application;
[0076] Figure 12 It is a schematic diagram of a hardware composition structure of another electronic device applying the embodiments of the present application. Detailed implementation manners
[0077] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the technical solutions of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments recorded in this application document without creative efforts belong to the scope protected by the technical solutions of the present application.
[0078] Some concepts involved in the embodiments of the present application are introduced below.
[0079] Image recognition: It refers to category-level recognition, without considering specific instances of an object, only considering identifying the category of an object (such as a person, a dog, a cat, a bird, etc.) and giving the category to which the object belongs. A typical example is the recognition task in the large-scale general object recognition open-source data set Imagenet, identifying which one of 1000 categories a certain object is.
[0080] Image multi-label recognition: It is to identify whether an image has a combination of specified attribute labels through a computer. An image may have multiple attributes, and the multi-label recognition task is to determine which preset attribute labels a certain image has.
[0081] Shallow global features of an image: Represent the global basic information of the image, that is, the embedding representation of the basic feature information of the image, such as texture features, and a representation of Scale-invariant feature transform (SIFT) corner features. In deep learning, the output of a shallow neural network belongs to the underlying image representation. Among them, SIFT is a description used in the field of image processing. This description is scale-invariant, can detect key points in the image, and is a local feature descriptor.
[0082] Deep semantic features of an image: Represent the image semantic information of the image, that is, features with semantic information (semantic embedding). In this application, such features are output in the layer before the classification layer, that is, the semantic embedding can predict the semantic label of the image after passing through the classification layer.
[0083] Metric and metric learning: In mathematics, a metric (or distance function) is a function that defines the distance between elements in a set. A set with a metric is called a metric space. Metric learning, also known as similarity learning, is a common machine learning method for comparing and measuring the similarity between data. In computer vision, it has extensive applications and an extremely important position in major fields such as face recognition and image retrieval.
[0084] The embodiments of this application relate to artificial intelligence (AI) and machine learning technologies, and are designed based on computer vision technology and machine learning (ML) in artificial intelligence.
[0085] Artificial intelligence is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence.
[0086] Artificial intelligence is the study of the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology mainly includes several major directions such as computer vision technology, natural language processing technology, machine learning / deep learning, autonomous driving, and intelligent transportation. With the research and progress of artificial intelligence technology, artificial intelligence has been studied and applied in multiple fields, such as common smart homes, intelligent customer service, virtual assistants, smart speakers, intelligent marketing, driverless, autonomous driving, robots, intelligent healthcare, etc. It is believed that with the development of technology, artificial intelligence will be applied in more fields and play an increasingly important role.
[0087] Machine learning is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Compared with data mining, which looks for mutual characteristics among big data, machine learning pays more attention to the design of algorithms, enabling computers to automatically "learn" rules from data and use the rules to predict unknown data.
[0088] Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, and inductive learning. The image classification joint model in the embodiments of this application is trained using machine learning or deep learning technology. Based on the training method of the image classification joint model in the embodiments of this application, images can be classified, retrieved, etc.
[0089] The method for training the image classification joint model proposed in the embodiments of this application can be divided into two parts, including a training part and an application part; among them, the training part involves the technical field of machine learning. In the training part, the image classification joint model is trained using machine learning technology. Specifically, the training sample group in the training sample dataset given in the embodiments of this application is used to train the image classification joint model. After passing through the image classification joint model, the output result of the image classification joint model is obtained. Combining the output result, the model parameters are continuously adjusted, and the trained image classification joint model is output; the application part is used to classify, retrieve, etc. images using the image classification joint model trained in the training part.
[0090] The following briefly introduces the design concept of the embodiments of this application:
[0091] With the rapid development of Internet technology and the rapid growth of multimedia resources, Internet content has gradually evolved from the mainstream text information to various types of multimedia data; among them, with the continuous update of electronic devices and the continuous simplification and update of imaging equipment, image data resources have increased rapidly. Through a neural network model, embedding features can be extracted from image data, and some downstream tasks can be performed based on the extracted features.
[0092] When training image embedding features based on a deep neural network, it is often necessary to simultaneously represent the overall global features and the main semantic features of the image. Some downstream tasks have a high demand for semantic representation, such as the retrieval of similar category images by image search; some downstream tasks focus more on the overall features, such as simple image deduplication; and some tasks need to consider both features at the same time, such as semantic image deduplication.
[0093] However, in the deep learning stage of related technologies, generally different models are used to separately train and learn the extraction methods of the above two features. However, the method of separate training and learning of the two will cause certain troubles in the training process and application process. The separate learning of the two leads to time-consuming training and slow convergence of both. And since the underlying network structure is generally the same, there is a large amount of duplicate model parameter learning. Not only do two sets of models and parameters need to be maintained separately, but also the prediction of the two networks is very time-consuming when making predictions based on the models.
[0094] In view of this, the embodiments of the present application propose an image recognition method, device, electronic device and storage medium. Since the present application proposes a model structure that is beneficial to the learning of the two features by a metric learning embedding module for extracting shallow global features and a deep semantic embedding module for extracting deep semantic features, when classifying the image to be detected based on the image joint classification model in the present application, the image can be represented by jointly using the underlying image representation and the deep semantic representation, enhancing the joint learning effect of the two, and reducing the time consumption of feature extraction. Image classification can be performed based on the obtained image fusion features, and more accurate classification results can be obtained, improving the accuracy and efficiency of image classification.
[0095] The following describes the preferred embodiments of the present application with reference to the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application. And without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.
[0096] Such as Figure 1As shown, it is a schematic diagram of the application scenario of the embodiment of the present application. This application scenario diagram includes two terminal devices 110 and a server 120. In the terminal device 110 of the embodiment of the present application, a client can be installed, and the client is used for image classification, image retrieval, etc. The server 120 can include a server. The server is used to provide image materials for the client. For example, the image library in the embodiment of the present application can be located on the server side and stores multiple candidate images. Alternatively, the image library can also be located locally on the client. In addition, the server in the embodiment of the present application can also be used for image classification, image retrieval, etc., which is not specifically limited here.
[0097] In the embodiment of the present application, the image classification joint model can be deployed on the terminal device 110 for training, or can be deployed on the server 120 for training. A large number of training samples can be stored in the server 120, including multiple groups of training sample groups, which are used to train the image classification joint model. Optionally, after the image classification joint model is trained based on the training method in the embodiment of the present application, the trained image classification joint model can be directly deployed on the server 120 or the terminal device 110. Generally, the image classification joint model is directly deployed on the server 120. In the embodiment of the present application, the image classification joint model is commonly used for image classification, image retrieval, etc.
[0098] It should be noted that the image recognition method for training the image classification joint model provided in the embodiment of the present application can be applied to various application scenarios including image classification tasks and image retrieval tasks. The training samples used in different scenarios are different, and will not be listed one by one here.
[0099] Specifically, in the image classification scenario, the fusion features of the image can be obtained based on the image recognition method in the present application, etc. Based on this feature, it is judged which preset attribute labels the image has, and the judgment result is used as the classification result of the image.
[0100] In the image retrieval scenario, it is further divided into semantic retrieval, similarity retrieval, etc. These scenarios can all be applied. The shallow global features, deep semantic features, fusion features, etc. of the image can be obtained based on the image recognition method in the present application, and image retrieval is performed based on these features, such as image duplicate removal retrieval, similar subject retrieval in the image, etc.
[0101] In an alternative embodiment, the terminal device 110 and the server 120 can communicate through a communication network.
[0102] In an alternative embodiment, the communication network is a wired network or a wireless network.
[0103] In the embodiments of the present application, the terminal device 110 is a device used by a user, such as a personal computer, mobile phone, tablet computer, notebook, e - book reader, in - vehicle terminal, etc., which has a certain computing ability and runs instant messaging software and websites or social software and websites. Each terminal device 110 is connected to the server 120 through a wireless network. The server 120 is a single server or a server cluster or cloud computing center composed of several servers, or a virtualization platform.
[0104] It should be noted that Figure 1 The above are only examples. In fact, the number of terminal devices and servers is not limited and is not specifically defined in the embodiments of the present application.
[0105] Next, in combination with the above - described application scenario, the video detection method provided by the exemplary embodiments of the present application will be described with reference to the accompanying drawings. It should be noted that the above - described application scenario is only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited in this regard.
[0106] It should be noted that the image recognition method in the embodiments of the present application can be executed by the server or the terminal device alone, or jointly executed by the server and the terminal device, and is not specifically limited herein.
[0107] Referring to Figure 2 As shown, it is a flowchart of the implementation of an image recognition method provided by the embodiments of the present application. Taking the server as the execution entity as an example, the specific implementation process of the method is as follows:
[0108] S21: The server extracts the first - layer shallow global feature of the image to be detected based on the metric learning embedding module in the trained image classification joint model. The first - layer shallow global feature represents the global basic information of the image to be detected.
[0109] S22: The server extracts the first - layer deep semantic feature of the image to be detected based on the deep semantic embedding module in the image classification joint model. The first - layer deep semantic feature represents the image semantic information of the image to be detected.
[0110] It should be noted that the shallow global feature in the embodiments of the present application mainly refers to the overall information at the bottom layer of the image, and the deep semantic feature is extracted by further deep - layer feature abstraction of the embedding that describes the whole image. Therefore, the present application proposes a model that combines the bottom - layer feature and the deep semantic, that is, the image classification joint model in the present application. This model includes a metric learning embedding module and a deep semantic embedding module at the same time. After extracting the corresponding features based on these two modules respectively, image classification can be performed.
[0111] S23: The server performs feature fusion on the first shallow global feature and the first deep semantic feature based on the semantic prediction module in the image classification joint model, and performs multi-label classification on the image to be detected based on the obtained first fusion feature to obtain the classification result of the image to be detected.
[0112] Among them, the image classification joint model is obtained by jointly training the metric learning embedding module and the deep semantic embedding module.
[0113] In the above implementation, a model structure that is beneficial to the learning of the two features is obtained by using the metric learning embedding module for extracting shallow global features and the deep semantic embedding module for extracting deep semantic features. When classifying the image to be detected based on the image joint classification model in this application, the image can be represented by jointly using the underlying image representation and the deep semantic representation, enhancing the joint learning effect of the two, and reducing the time-consuming of feature extraction. Image classification based on the obtained image fusion feature can obtain a more accurate classification result, improving the accuracy and efficiency of image classification.
[0114] In an alternative implementation, this application can also perform image retrieval based on the output result of the image classification model. Refer to Figure 3 As shown, it is a flowchart of the implementation of an image retrieval method in an embodiment of this application, which specifically includes the following steps:
[0115] S301: The server respectively extracts the second shallow global features of each candidate image in the image library based on the metric learning embedding module; and respectively extracts the second deep semantic features of the image to be detected based on the deep semantic embedding module.
[0116] Among them, the content represented by the second shallow global feature is essentially the same as that represented by the first shallow global feature, that is, both represent the global basic information of the image; here, "first" and "second" are for the image. Among them, the first shallow global feature is for the image to be detected, and the second shallow global feature is for the candidate image. Similarly, the content represented by the second deep semantic feature and the first deep semantic feature is also essentially the same, both representing the image semantic information of the image. The same principle applies to the second fusion feature and the first fusion feature in the following text, and the repeated parts will not be elaborated.
[0117] S302: The server respectively performs feature fusion on each second shallow global feature and the corresponding second deep semantic feature based on the semantic prediction module to obtain the second fusion feature corresponding to each candidate image.
[0118] S303: The server respectively determines the similarity between each candidate image and the image to be detected based on the first fusion feature and the second fusion feature corresponding to each candidate image.
[0119] S304: The server determines the similar images corresponding to the image to be detected based on each similarity.
[0120] For example, in the scenario of image deduplication and retrieval, the similar images corresponding to the image to be detected in the image library can be determined based on the above method, including the images that are the same as or extremely similar to the image to be detected. Specifically, they can be candidate images with a similarity higher than a certain threshold, and then the images related to these can be removed. For another example, the above method can also be used for retrieving similar main bodies in the image. For example, when a user inputs a clothing image as the image to be detected, similar clothing images can be retrieved from the image library.
[0121] In practical applications, the global texture and semantic features of images are interrelated. For example, certain semantic categories only appear in specific global image environments. For example, polar bears are related to ice, while sheep and polar bears generally do not appear simultaneously. The image recognition method in this application makes good use of this correlation information to improve the learning effects of both, and can effectively improve the accuracy of image classification and image retrieval.
[0122] Optionally, the image classification joint model in the embodiments of the present application further includes a basic convolution module. That is, in the embodiments of the present application, the model structure can be divided into 4 modules: a basic convolution module, a metric learning embedding module, a deep semantic embedding module, a semantic prediction module, etc.
[0123] Refer to Figure 4 As shown, it is a schematic structural diagram of an image classification joint model in the embodiments of the present application. This model specifically includes: a basic convolution module 401, a metric learning embedding module 402, a deep semantic embedding module 403, and a semantic prediction module 404.
[0124] Optionally, the basic convolution module 401 includes multiple convolutional layers, such as Figure 4 the convolutional neural network (Convolutional Neural Network, CNN) shown in 401 in Figure 4 which can be a convolutional network composed of a residual structure. In addition, in addition to using the ResNet101 basic structure for global feature extraction, the image classification joint model in the present application also adds a deep semantic embedding module composed of a residual structure Conv6_x module, and a subsequent semantic multi-label classification module FC (i.e., the semantic prediction module). Among them, the deep semantic embedding module specifically includes a convolutional layer and an embedding layer composed of a residual structure. The convolutional layer composed of the residual structure belongs to the semantic basic part, such as Figure 4The Deep embedding shown in 403 is mainly exemplified in detail below by taking Deeper CNN including Embedding2 as an example.
[0125] This application considers that the residual structure can extract information better than ordinary convolutional layers. Therefore, the deep semantic embedding module still uses the residual structure. In practical applications, the number of stacked residual structures can be adjusted according to needs. In this article, there are 2 stacks.
[0126] Next, in combination with Table 1, Table 2, and Table 3, the network parameters of the image classification joint model in the embodiments of this application will be briefly described.
[0127] In the embodiments of this application, the network parameters corresponding to the convolutional layers in the basic convolutional module are initialized and determined based on the parameters pre-trained on a specified sample library. For example, the basic convolutional module is trained using ResNet101, and the parameters are shown in Table 1:
[0128] Table 1 Structure Table of ResNet101 Feature Module
[0129]
[0130]
[0131] Among them, the x3 blocks (residual modules) in Table 1 represent 3 such network structure stacks. Similarly, x2 blocks are similar. The basic convolutional module in this application includes five convolutional layers, as shown in Table 1, namely Conv1-Conv5_x. These five convolutional layers use the parameters of ResNet101 pre-trained on the Imagenet dataset as the initial network parameters.
[0132] It should be noted that the network results of the basic convolutional module listed in Table 1 are only for illustrative purposes. In addition to the network structure of ResNet101, different network structures and different pre-trained model weights can also be used as the basic convolutional module, such as ResNet50, Inceptionv4, etc. For large-scale data retrieval, small networks such as ResNet18 can be used, and the embedding dimension can be reduced, such as using 64 bits, etc., to reduce the feature storage space. In addition, in addition to using the Imagenet dataset to train the model, other large-scale datasets, such as the Open Image dataset, can also be considered.
[0133] Refer to Figure 5As shown, it is a schematic diagram of the structure of a residual module (block) in an embodiment of the present application. The residual module includes three convolutional layers, namely a 1x1 convolutional layer, a 3x3 convolutional layer, and a 1x1 convolutional layer. Modules of types such as Conv2_x and Conv3_x in Table 1 are all generated by stacking residual blocks under different parameters multiple times.
[0134] Next, a brief introduction to the network structure and parameters corresponding to the metric learning embedding module in the embodiment of the present application is given. Refer to Table 2 as follows:
[0135] Table 2 Structure Table of the Metric Learning Module Based on ResNet101
[0136] Layer name Output size Layer Pool_cr1 1x2048 Max pool Embedding 1x512 full connetction
[0137] Among them, 512 is the embedding dimension of the shallow global feature.
[0138] Optionally, in the embodiment of the present application, the network parameters corresponding to the convolutional layers in the deep semantic embedding module can be obtained based on random initialization. In this article, it is mainly illustrated by taking the Gaussian distribution with a variance of 0.01 and a mean of 0 for initialization. Specifically, refer to Table 3 for the network structure and parameters corresponding to the deep semantic embedding module and the semantic prediction module:
[0139] Table 3 Structure Table of the Classification Module Based on ResNet101
[0140]
[0141] Refer to Table 3. Among them, Conv6_x, Pool_cr1, and Embedding2 are deep semantic embedding modules, and FC is a semantic prediction module. Among them, 512 is the embedding dimension of the deep semantic feature, Embedding2 is a 512-dimensional semantic embedding, and Nclass is the number of multi-label categories, that is, the number of classification labels. When using Imagenet open-source image training, it is 1000 labels, and Nclass is 1000. Among them, Conv6_x is a newly added convolutional module for feature extraction after the global feature Conv5_x convolutional module. x2 blocks are similar to x3 blocks.
[0142] The convolutional layer in the deep semantic embedding module can be represented as Conv6_x, the embedding layer in the deep semantic embedding module can be represented as embedding2, and the semantic prediction module can be represented as a fully connected layer FC (full connection). Newly added layers such as Conv6_x, embedding, and FC layers are initialized with a Gaussian distribution with a variance of 0.01 and a mean of 0.
[0143] It should be noted that in the embodiments of the present application, other model structures may also be adopted for the branches of Tables 2 and 3. Only a feasible solution is given in the table, and no specific limitation is made here.
[0144] In the embodiments of the present application, a model structure beneficial to the learning of two types of features can be obtained by designing a deep residual feature extraction module on the conventional basic feature extraction network structure. Through multi-stage learning of pre-training and fine-tuning, the new module converges to the classification task, and then the iterative efficiency of different learning tasks is taken into account to achieve the convergence of the two features. Finally, the image is represented by jointly using the underlying image representation and the deep semantic representation for downstream retrieval or other related applications.
[0145] Optionally, in order to further improve the model learning effect, the present application also designs an asynchronous learning rate. The basic convolution module, the metric learning embedding module, and the deep semantic embedding module all adopt a learning rate of lr1 = 0.0005, while the multi-label classification FC layer (i.e., the semantic prediction module) adopts lr = 0.005 (= lr1 * 10), that is, the learning rate of the semantic prediction module is higher than that of other modules. Other modules include the basic convolution module, the metric learning embedding module, and the deep semantic embedding module.
[0146] Among them, the multi-label classification task is to make two samples with the same category output the same predicted category, that is, to fit the same learning objective. Under overfitting, the predictions of the labels for multiple dissimilar images are very similar. Since the FC layer is more likely to overfit the target, and this semantic overfitting will make the semantic embedding more likely to overfit, that is, the embeddings of two images of the same category are the same (that is, the embedding also overfits the classification target), resulting in no distinction between the embeddings of images. In the present application, the asynchronous learning rate is adopted, which can make the parameter update of the embedding layer slower than that of the FC layer (the update efficiency is 0.1 times that of the FC layer), effectively preventing the embedding from overfitting the multi-label target during the parameter update of each batch.
[0147] In the embodiments of the present application, considering that the semantic embedding is a deeper feature abstraction extracted from the embedding that describes the whole image, the present application proposes a joint model of underlying features and deep semantics. A model structure beneficial to the learning of two types of features is obtained by designing a deep residual feature extraction module on the conventional basic feature extraction network structure. Then, through multi-stage learning of pre-training and fine-tuning, the new module converges to the classification task. Finally, the iterative efficiency of different learning tasks is taken into account to achieve the convergence of the two features. Finally, the image is represented by jointly using the underlying image representation and the deep semantic representation for downstream retrieval or other related applications.
[0148] After introducing the structure of an image classification joint model in the embodiments of the present application, the following will introduce the process of image classification based on this image classification joint model in the embodiments of the present application in conjunction with Figure 4 .
[0149] Refer to Figure 6 shown in the figure, which is a schematic flowchart of an image classification method in the embodiments of the present application, specifically including the following steps:
[0150] S601: The server inputs the image to be detected into the basic convolution module in the image classification joint model, and performs global feature extraction on the image to be detected based on the basic convolution module to obtain the global feature map corresponding to the image to be detected.
[0151] S602: The server performs embedding processing on the global feature map based on the metric learning embedding module to obtain the first shallow global feature corresponding to the image to be detected.
[0152] S603: The server extracts the first deep semantic feature of the image to be detected based on the deep semantic embedding module in the image classification joint model.
[0153] S604: The server extracts the semantic information in the global feature map based on the convolutional layer in the deep semantic embedding module, and performs embedding processing on the extracted semantic information based on the embedding layer in the deep semantic embedding module to obtain the first deep semantic feature corresponding to the image to be detected.
[0154] S605: The server performs feature fusion on the first shallow global feature and the first deep semantic feature based on the semantic prediction module in the image classification joint model, and performs multi-label classification on the image to be detected based on the obtained first fusion feature to obtain the classification result of the image to be detected.
[0155] The following will introduce the model training process in the embodiments of the present application in detail:
[0156] The training samples in the embodiments of the present application are Figure 4 the triple samples shown in the figure. A training sample group includes three training samples: an anchor sample, a positive sample, and a negative sample, which can be represented as a triple (anchor, positive, negative), where anchor is the anchor sample, positive is the positive sample, and negative is the negative sample. anchor-positive is a similar or identical image, that is, the anchor sample and the positive sample are sample images with a similarity higher than the specified threshold, and anchor-negative is a dissimilar image, that is, the negative sample and the positive sample are sample images with a similarity not higher than the specified threshold.
[0157] It should be noted that this application supports the annotation of image sample pairs for triplet learning. Specifically, the training data for image metric learning can be annotated according to the following rules: The task of each annotation is to select 3 images that meet the rules from the full set of images to form a triplet (anchor, positive, negative).
[0158] However, considering that there are a large number of easy samples in the triplet data randomly generated, these samples are helpful for the model learning at the beginning, but soon (such as after the 3rd epoch), the model has established a good discrimination ability for easy samples. At this time, the loss of a large number of easy samples will be much greater than that of hard samples, so that hard samples are buried in easy samples. Therefore, obtaining enough hard samples in the middle and late stages has a greater impact on the model. Among them, easy samples refer to sample pairs that are easy to distinguish, and hard samples refer to samples that are not easy to distinguish, such as groups of similar samples with large feature differences, or groups of different samples with similar features.
[0159] Based on this, in the embodiments of this application, the collected sample images can be divided into training sample groups based on the following method. Refer to Figure 7 As shown, it is a flowchart of the implementation of a method for constructing a training sample data set in the embodiments of this application, which specifically includes the following steps:
[0160] S701: The server forms a positive sample pair for every two sample images with a similarity higher than a specified threshold according to the similarity between the collected sample images.
[0161] S702: The server performs sample recombination based on each positive sample pair to obtain multiple training sample groups for constructing the training sample data set.
[0162] That is, first, the collected sample images are divided into individual positive sample pairs. Since the two samples in each positive sample pair belong to the same or similar samples, one of them can be used as the anchor sample in the triplet sample, and the other can be used as the positive sample in the triplet sample. However, two samples from different positive sample pairs may belong to samples with low similarity. Based on this, sample recombination can be performed on each positive sample pair. Multiple training sample groups can be constructed based on one positive sample pair. Furthermore, for each positive sample pair, such recombination is performed, and enough training sample groups can be obtained to construct the training sample data set.
[0163] In an optional implementation manner, step S702 can be implemented based on the following process to realize the process of constructing multiple training sample groups based on one positive sample pair, which specifically includes the following steps:
[0164] S7021: The server selects one sample image from a positive sample pair as the target sample, and selects one sample image from each of the other sample pairs as a candidate sample.
[0165] S7022: The server respectively determines the distances between the target sample and each candidate sample.
[0166] S7023: The server sorts each candidate sample according to its corresponding distance, and selects at least one candidate sample at a specified sorting position as the negative sample corresponding to the target sample.
[0167] Among them, the specified sorting position can refer to the last N ones after sorting each candidate sample according to its corresponding distance from small to large, or the first N ones after sorting each candidate sample according to its corresponding distance from large to small, where N is a positive integer. That is, several candidate samples with a relatively large distance from the target sample are selected as negative samples.
[0168] S7024: The server respectively forms a training sample group by combining each selected negative sample with the positive sample pair. Among them, the target sample in the positive sample pair is the positive sample in the training sample group, and the other sample in the positive sample pair is the anchor sample in the training sample group.
[0169] In the above embodiment, each negative sample can form a training sample group with this sample pair; for a positive sample pair, if N negative samples are selected, N training sample groups can be formed. If there are n positive sample pairs, n×N training sample groups can be formed.
[0170] For example, for positive sample pair 1, this sample pair includes two sample images: sample image 1 and sample image 2. Assume that sample image 1 is selected as the target sample, and another 5 sample images are respectively selected from positive sample pairs 2-6 as candidate images, which are assumed to be: sample image 3, sample image 5, sample image 7, sample image 9, and sample image 11. Calculate the distances between sample image 1 and these five candidate images respectively. Assume that N is 3. Among them, the ones with the farthest distance from sample image 1 are: sample image 3, sample image 5, and sample image 7. Then, three training sample groups can be constructed based on these three candidate samples, which are {sample image 1, sample image 2, sample image 3}, {sample image 1, sample image 2, sample image 5}, and {sample image 1, sample image 2, sample image 7}.
[0171] When the embodiments of the present application are specifically implemented, the triple annotation is adjusted to: only annotate the positive sample pairs to obtain similar sample pairs. Specifically, the triples are obtained by the following mining in each batch of samples (batch-size, bs) pairs: for the sample x (i.e., the target sample) in a certain positive sample pair: from the remaining bs - 1 positive sample pairs, randomly select one image from each pair as a candidate sample, calculate the distance between each candidate sample and x respectively, sort them in ascending order of distance, and take the top 10 candidate samples as negative samples, and form triples with the positive samples in x respectively. Therefore, each sample generates 10 triples, and the entire batch obtains 10 * bs triples. Among them, bs needs to be set to a relatively large value, such as 256; the batch size is a hyperparameter used to define the number of samples to be processed before updating the internal model parameters.
[0172] Further, it is also necessary to perform semantic label annotation on the training samples. In the embodiments of the present application, only the above positive sample pair images can be multi-label annotated, and the labels include scene labels such as stage and restaurant; painting style labels such as game and anime; other labels such as selfie and other human actions, which are not specifically limited here. When an image contains a certain label, the true probability corresponding to the label is 1, and if it does not contain the label, the true probability corresponding to the label is 0; assume that the image classification joint model in the present application needs to learn 1000 multi-labels, and there is a situation where the image does not contain any labels, that is, the true probabilities corresponding to all labels are 0. When the true probabilities corresponding to these 1000 multi-labels are represented by vectors, they are the multi-label vectors in the present application, representing the true probabilities of the training samples corresponding to each classification label.
[0173] After introducing the preparation process of the training data in the embodiments of the present application, the training process of the image classification joint model in the embodiments of the present application will be introduced in detail based on the above training data:
[0174] In an alternative embodiment, the image classification joint model is trained in the following manner. Specifically, refer to Figure 8 As shown, it is a flowchart of a training method for an image classification joint model in the embodiments of the present application, including the following steps:
[0175] S801: The server obtains the training sample data set and selects a training sample group from the training sample data set.
[0176] S802: The server inputs the selected training sample group into the trained image classification joint model, and obtains the third shallow global feature output by the metric learning embedding module in the image classification joint model, the third deep semantic feature output by the deep semantic embedding module, and the classification vector output by the semantic prediction module.
[0177] S803: The server constructs an intermediate loss function based on the third shallow global feature and the classification vector, and based on the intermediate loss function, adjusts the network parameters of the deep semantic module in the image classification joint model for a specified number of times.
[0178] Specifically, based on the third shallow global feature and the classification vector, a first triplet loss function, a multi-label loss function, and a classification entropy loss function are constructed. The first triplet loss function, the multi-label loss function, and the classification entropy loss function are weighted and summed to obtain the intermediate loss function.
[0179] S804: The server constructs an objective loss function based on the third shallow global feature, the third deep semantic feature, and the classification vector, and based on the objective loss function, adjusts the network parameters of the image classification joint model multiple times until the image classification joint model converges, and outputs the trained image classification joint model.
[0180] Specifically, based on the third shallow global feature, the third deep semantic feature, and the classification vector, a first triplet loss function, a second triplet loss function, a multi-label loss function, and a classification entropy loss function are constructed. The first triplet loss function, the second triplet loss function, the multi-label loss function, and the classification entropy loss function are weighted and summed to obtain the objective loss function.
[0181] It should be noted that the above-listed training methods are only briefly introduced. In fact, based on the above training process, the full amount of data in the training sample dataset can be iterated for epoch rounds; each round of iteration processes the full amount of samples once; among them, the number of epochs is a hyperparameter that defines the number of times the learning algorithm works on the entire training dataset.
[0182] Specifically, the specific operations in each round of iteration are as follows: The full amount of samples are divided into Nb batches with each batch-size samples (triplet samples) as a batch. For each batch, the following operations are performed:
[0183] (1) Model forward: Set all the parameters of the model to the state that needs to be learned. During training, the neural network performs forward calculation on an input image to obtain the prediction result.
[0184] (2) Loss calculation: Calculate the triplet loss for em, including the first triplet loss function and the second triplet loss function; calculate the multi-label loss function (binary cross Entropy loss, bce loss) and the classification entropy loss function for FC. Sum them up to get the total loss. The specific calculation will be described later. Assume that the weights corresponding to the first triplet loss function, the second triplet loss function, the multi-label loss function, and the classification entropy loss function are: w1, w2, w3, w4.
[0185] (3) Model parameter update: Use the Stochastic Gradient Descent (SGD) method to perform backward gradient calculation on the loss of the previous step to obtain the updated values of all model parameters, and update the network.
[0186] It should be noted that step S803 in the embodiments of the present application is optional. During the training of the image classification joint model, after executing step S802, step S804 can be directly executed, that is, directly optimize the network parameters of the entire image classification joint model until the model converges. In this way, Conv6 has only one stage of learning; or, before step S804, execute step S803. In this way, Conv6 has two stages of learning.
[0187] In the case of executing step S803, the fully connected layer in the embodiments of the present application includes two stages of training: First, based on three loss functions (weighted sum), perform N iterations; then, based on four loss functions (weighted sum), perform unified iteration until convergence. Among them, the second stage is the adjustment of the overall model parameters, and the first stage is only the adjustment of the fully connected layer.
[0188] Take Figure 4 the shown image classification joint model as an example. In the embodiments of the present application, the parameters of Conv1-5 in the network are initialized with the parameters pre-trained on Imagenet, while Conv6 is randomly initialized. Different initializations result in different learning effects of the two modules. In order to enable Conv6 to effectively initialize the task, Conv6 can be divided into the following two stages of learning:
[0189] The first stage: Use three losses, namely the first triplet loss function, the multi-label loss function, and the classification entropy loss function, to train for a limited number of epochs (e.g., 10 epochs) to make the global embedding and classification converge preliminarily (weight settings 1, 1, 0.5). The purpose of learning here is to let Conv6 find a set of good parameters suitable for the current classification task. At the same time, since the global embedding needs to always maintain the metric learning features, it is always in the learning loss (because at this time the semantic embedding does not yet have good representation ability, and the semantic embedding and classification are not the same learning task, with a certain degree of conflict and being adjacent modules, which is likely to affect the parameter optimization of Conv6, so it is not added to the learning). The role of the entropy loss is to make other possibly related classes also have certain predicted values during semantic classification, but the predicted values do not need to be large enough to be recognized as those unlabeled classifications, so as to avoid overfitting of the network to the labeled classifications.
[0190] The second stage: Use all losses to train until convergence, so that the global embedding and the semantic embedding converge separately (weight settings of w1~w4 are 1:1:0.5:0.5). Subsequently, the weight of the semantic classification is lowered, and the two embeddings are preferentially learned.
[0191] It should be noted that the weights w1~w4 corresponding to each loss function in the embodiments of the present application, as well as some hyperparameters such as the margin during the calculation of Triplet-loss, can be adjusted as needed. Here is just an example and no specific limitations are made here.
[0192] The following will detail the calculation process of the loss function in the embodiments of the present application. The present application mainly involves three major types of loss functions, namely the triplet loss function, the classification entropy loss function, and the multi-label loss function. Among them, based on the shallow global features and the deep semantic features, the first triplet loss function and the second triplet loss function are respectively constructed. Therefore, the target loss function in the present application is mainly obtained by weighted summation of these four loss functions.
[0193] Refer to Figure 9 shown, which is a flowchart of a calculation method of a loss function in the embodiments of the present application, specifically including the following steps:
[0194] S901: The server constructs the first triplet loss function based on the third shallow global features corresponding to each training sample in the training sample group.
[0195] It should be noted that the third shallow global feature in the embodiments of this application has the same essence as the content represented by the first shallow global feature and the second shallow global feature in the above text, but only the corresponding objects are different. The third shallow global feature is for training samples, and the same principle applies to the third deep semantic feature and classification vector below.
[0196] Specifically, the first triplet loss function can be constructed based on the following process:
[0197] S9011: Determine the first distance between the third shallow global feature corresponding to the anchor sample in the training sample group and the third shallow global feature corresponding to the positive sample, and the second distance between the third shallow global feature corresponding to the anchor sample and the third shallow global feature corresponding to the negative sample.
[0198] S9012: Take the sum of the difference between the first distance and the second distance and the specified boundary value as the first target value. The specified boundary value is used to represent the difference boundary of the similarity between the positive sample and the negative sample.
[0199] S9013: Take the maximum value of the first target value and the reference parameter as the first triplet loss function.
[0200] For example, after finding the triplets (a, p, n) in the batch samples, calculate the first triplet loss function triplet loss1 based on the shallow global features of these triplet samples.
[0201] Among them, the calculation formula of the triplet loss function Triplet loss is as follows:
[0202] l tri = max(||x a - x p || - ||X a - x n || + α, 0)
[0203] Among them, α is the margin (specified boundary value), set to 4. Specifically, ||x a - x p || can represent the L2 distance between the third shallow global feature corresponding to the anchor sample and the third shallow global feature corresponding to the positive sample, that is, the first distance; ||x a - x n || can represent the L2 distance between the third shallow global feature corresponding to the anchor sample and the third shallow global feature corresponding to the negative sample, that is, the second distance.
[0204] In the above calculation formula of Triplet loss, the reference parameter is 0. If the first target value ||x a - x p||-||x a -x n ||If α is greater than 0, the first triplet loss function is the first target value. If the first target value||x a -x p ||-||x a -x n ||If α is not greater than 0, the first triplet loss function is 0.
[0205] In the embodiments of the present application, the purpose of triplet loss1 is to make the distance between the shallow global features of the anchor and the negative greater than the distance between the shallow global features of the anchor and the positive by 4 (i.e., greater than the margin). Based on this loss function, it can be ensured that the shallow global features of the anchor sample and the positive sample are similar, while the shallow global features of the anchor sample and the negative sample are vice versa, so as to ensure that there is a certain distance between the shallow global features of the positive sample and the negative sample, improving the classification accuracy.
[0206] S902: The server constructs a second triplet loss function based on the third deep semantic features corresponding to each training sample.
[0207] Correspondingly, the process of constructing the second triplet loss function is similar to the process of constructing the first triplet loss function, and specifically includes the following implementation steps:
[0208] S9021: Based on the third distance between the third deep semantic features corresponding to the anchor sample in the training sample group and the third deep semantic features corresponding to the positive sample, and the fourth distance between the third deep semantic features corresponding to the anchor sample and the third deep semantic features corresponding to the negative sample.
[0209] S9022: Take the sum of the difference between the third distance and the fourth distance and the specified boundary value as the second target value, where the specified boundary value is used to represent the difference boundary of the similarity between the positive sample and the negative sample.
[0210] S9023: Take the maximum value of the second target value and the reference parameter as the second triplet loss function.
[0211] In the embodiments of the present application, for the deep semantic features output by the deep semantic embedding module, the second triplet loss function triplet loss2 can also be calculated using the above calculation formula of Triplet loss, with the same margin.
[0212] Among them,||x a -x p ||can represent the L2 distance between the third deep semantic features corresponding to the anchor sample and the third deep semantic features corresponding to the positive sample, that is, the third distance;||x a-x n || can represent the L2 distance between the third deep semantic feature corresponding to the anchor sample and the third deep semantic feature corresponding to the negative sample, that is, the fourth distance.
[0213] In the embodiments of the present application, the purpose of the second triplet loss function triplet loss2 is to make the distance between the deep semantic features of the anchor and the negative greater than the distance between the deep semantic features of the anchor and the positive by 4. Based on this loss function, it can be ensured that the deep semantic features of the anchor sample and the positive sample are similar, while the deep semantic features of the anchor sample and the negative sample are vice versa, so as to ensure that there is a certain distance between the deep semantic features of the positive sample and the negative sample, improving the classification accuracy.
[0214] S903: The server constructs a multi-label loss function based on the classification vectors corresponding to each training sample and the corresponding multi-label vectors. The classification vector represents the predicted probability of the training sample for each classification label, and the multi-label vector represents the true probability of the training sample for each classification label.
[0215] In the embodiments of the present application, for the probability vector (i.e., the classification vector) output by the FC layer, the multi-label loss function bceloss between it and the multi-label vector obtained by multi-label annotation can be calculated, that is, Figure 4 the Class loss shown.
[0216] Specifically, for a certain sample i in the full amount of n samples, its true label vector t[i] (i.e., the multi-label classification vector) is a 0, 1 vector of 1*1000, and its predicted value o[i] (i.e., the classification vector) is the predicted probability for each of the 1*1000 labels. The multi-label loss is calculated according to the following formula:
[0217]
[0218] Among them, when the true value of a certain label bit is 1, the left side of the plus sign in the above formula takes effect, and when the true value is 0, the right side of the above formula takes effect. Therefore, the supervision information of a certain sample under all labels can be learned.
[0219] In the embodiments of the present application, the role of FC classification learning is to constrain the classification embedding to make it have classification information, and the classification information is very important for downstream tasks such as similar category image retrieval. On the other hand, since classification is prone to overfitting, resulting in little difference in the discrimination of the embedding, the learning rate of the FC layer is set to be 10 times larger than that of other layers in the above learning, so that the gradient of the FC layer only uses 0.1 of the original gradient when backpropagating and updating to the semantic embedding, thus avoiding overfitting of the embedding to the classification.
[0220] S904: The server obtains the classification entropy loss function based on the classification vectors corresponding to each training sample.
[0221] Among them, the classification entropy loss function is the Entropy loss in Figure 4 . In the embodiments of the present application, in order to avoid the overfitting problem of classification prediction, the Entropy loss is introduced, and the entropy is calculated for the prediction results of the FC layer, introducing the distribution loss. In the model learning process of the present application, it is hoped that the entropy is as large as possible, so that the model prediction results do not have a large prediction value for only one classification label, but allow multiple label prediction values to be large, but at the same time have a limited impact on the classification effect.
[0222] Specifically, for all 2n images in the batch (assuming there are n sample pairs in the batch), the following loss can be calculated for each image:
[0223]
[0224] Among them, p i is the prediction output of the FC layer of the i-th sample, and p ij is the prediction result of the i-th sample for the j-th category (i.e., the target category).
[0225] S905: The server performs weighted summation on the first triple loss function, the second triple loss function, the multi-label loss function, and the classification entropy loss function to obtain the target loss function.
[0226] Assume that the first triple loss function is expressed as L triplet1 , the second triple loss function is expressed as L triplet2 , the multi-label loss function is expressed as L bce , and the classification entropy loss function is expressed as L Entropy . The respective weights corresponding to these loss functions are: w1, w2, w3, w4. Accordingly, the target loss function is:
[0227] L total = w1L triplet1 + w2L triplet2 + w3L bce + w4L Entropy .
[0228] It should be noted that the specific calculation order of the above several loss functions may not be specifically limited, and the above is only an example.
[0229] In summary, the embodiments of the present application propose a network structure that supports joint learning of global image representation and semantic embedding. Semantic extraction is achieved by adding a deep residual structure to the underlying features of the image. At the same time, the underlying embedding representation and semantic embedding representation of the image are learned, reducing the extraction time-consuming in applications and improving the effects of both. In addition, learning strategies such as the anti-overfitting loss and the pre-training initialization method for the new parameter layer listed above are proposed to achieve effective learning of parameters. Through the above loss and learning method settings, more effective initialization of the overall network parameters (such as adding Conv6) is realized and the learning effect is ensured.
[0230] Based on the same inventive concept, the embodiments of the present application further provide an image recognition device. As Figure 10 shown, it is a schematic structural diagram of the image recognition device 1000, which may include:
[0231] A first feature extraction unit 1001, configured to extract a first shallow global feature of the image to be detected based on the metric learning embedding module in the trained image classification joint model, where the first shallow global feature represents the global basic information of the image to be detected; and
[0232] A second feature extraction unit 1002, configured to extract a first deep semantic feature of the image to be detected based on the deep semantic embedding module in the image classification joint model, where the first deep semantic feature represents the image semantic information of the image to be detected;
[0233] A classification prediction unit 1003, configured to perform feature fusion on the first shallow global feature and the first deep semantic feature based on the semantic prediction module in the image classification joint model, and perform multi-label classification on the image to be detected based on the obtained first fusion feature to obtain a classification result of the image to be detected;
[0234] Wherein, the image classification joint model is obtained by jointly training the metric learning embedding module and the deep semantic embedding module.
[0235] Optionally, the device further includes:
[0236] An image retrieval unit 1004, configured to extract a second shallow global feature of each candidate image in the image library based on the metric learning embedding module; and, extract a second deep semantic feature of the image to be detected based on the deep semantic embedding module;
[0237] Based on the semantic prediction module, perform feature fusion on each second shallow global feature and the corresponding second deep semantic feature to obtain a second fusion feature corresponding to each candidate image;
[0238] Based on the first fusion feature and the second fusion features corresponding to each candidate image, determine the similarity between each candidate image and the image to be detected;
[0239] Based on each similarity, determine the similar image corresponding to the image to be detected.
[0240] Optionally, the image classification joint model further includes: a basic convolution module; the apparatus further includes:
[0241] A basic processing unit 1005, configured to input the image to be detected into the basic convolution module in the image classification joint model before the first feature extraction unit 1001 extracts the first shallow global feature of the image to be detected based on the metric learning embedding module in the trained image classification joint model, and perform global feature extraction on the image to be detected based on the basic convolution module to obtain the global feature map corresponding to the image to be detected;
[0242] The first feature extraction unit 1001 is specifically configured to:
[0243] Based on the metric learning embedding module, perform embedding processing on the global feature map to obtain the first shallow global feature corresponding to the image to be detected.
[0244] Optionally, the deep semantic embedding module includes a convolutional layer and an embedding layer composed of residual structures; the second feature extraction unit 1002 is specifically configured to:
[0245] Based on the convolutional layer in the deep semantic embedding module, extract the semantic information in the global feature map, and based on the embedding layer in the deep semantic embedding module, perform embedding processing on the extracted semantic information to obtain the first deep semantic feature corresponding to the image to be detected.
[0246] Optionally, the learning rate of the semantic prediction module is higher than the learning rates of other modules, and the other modules include the basic convolution module, the metric learning embedding module, and the deep semantic embedding module.
[0247] Optionally, the basic convolution module includes multiple convolutional layers; the network parameters corresponding to the convolutional layers in the basic convolution module are initialized based on the parameters pre-trained on a specified sample library, and the network parameters corresponding to the convolutional layers in the deep semantic embedding module are randomly initialized.
[0248] Optionally, the apparatus further includes:
[0249] A training unit 1006, configured to train the image classification joint model in the following manner:
[0250] Obtain a training sample data set, and select a training sample group from the training sample data set;
[0251] Input the selected training sample group into the trained image classification joint model to obtain the third shallow global feature output by the metric learning embedding module in the image classification joint model, the third deep semantic feature output by the deep semantic embedding module, and the classification vector output by the semantic prediction module;
[0252] Construct an objective loss function based on the third shallow global feature, the third deep semantic feature, and the classification vector, and repeatedly adjust the network parameters of the image classification joint model based on the objective loss function until the image classification joint model converges, and output the trained image classification joint model.
[0253] Optionally, the training unit 1006 is specifically configured to:
[0254] Construct a first triplet loss function based on the third shallow global feature corresponding to each training sample in the training sample group; and construct a second triplet loss function based on the third deep semantic feature corresponding to each training sample;
[0255] Construct a multi-label loss function based on the classification vector corresponding to each training sample and the corresponding multi-label vector, where the classification vector represents the predicted probability of the training sample for each classification label, and the multi-label vector represents the true probability of the training sample for each classification label;
[0256] Obtain a classification entropy loss function based on the classification vector corresponding to each training sample;
[0257] Perform a weighted sum of the first triplet loss function, the second triplet loss function, the multi-label loss function, and the classification entropy loss function to obtain the objective loss function.
[0258] Optionally, each training sample group includes three training samples: an anchor sample, a positive sample, and a negative sample; the training unit 1006 is specifically configured to:
[0259] Determine the first distance between the third shallow global feature corresponding to the anchor sample in the training sample group and the third shallow global feature corresponding to the positive sample, and the second distance between the third shallow global feature corresponding to the anchor sample and the third shallow global feature corresponding to the negative sample;
[0260] Use the sum of the difference between the first distance and the second distance and a specified boundary value as the first target value, where the specified boundary value is used to represent the difference boundary of the similarity between the positive sample and the negative sample;
[0261] Use the maximum value of the first target value and the reference parameter as the first triplet loss function.
[0262] Optionally, each training sample group includes three training samples: an anchor sample, a positive sample, and a negative sample; the training unit 1006 is specifically configured to:
[0263] Based on the third distance between the third deep semantic feature corresponding to the anchor sample in the training sample group and the third deep semantic feature corresponding to the positive sample, and the fourth distance between the third deep semantic feature corresponding to the anchor sample and the third deep semantic feature corresponding to the negative sample;
[0264] Take the sum of the difference between the third distance and the fourth distance and a specified boundary value as the second target value, where the specified boundary value is used to represent the difference boundary of the similarity between the positive sample and the negative sample;
[0265] Take the second target value and the maximum value in the reference parameters as the second triple loss function.
[0266] Optionally, the training unit 1006 is further configured to:
[0267] Before adjusting the network parameters of the image classification joint model multiple times based on the target loss function, perform a weighted sum of the first triple loss function, the multi-label loss function, and the classification entropy loss function to obtain an intermediate loss function;
[0268] Based on the intermediate loss function, adjust the network parameters of the deep semantic module in the image classification joint model a specified number of times.
[0269] Optionally, the device includes:
[0270] A sample construction unit 1007, configured to obtain the training sample groups in the training sample dataset in the following manner:
[0271] According to the similarity between the collected sample images, form a positive sample pair for every two sample images with a similarity higher than a specified threshold;
[0272] Based on the positive sample pairs, perform sample recombination to obtain multiple training sample groups for constructing the training sample dataset. Each training sample group includes three training samples: an anchor sample, a positive sample, and a negative sample, where the anchor sample and the positive sample are sample images with a similarity higher than the specified threshold, and the negative sample and the positive sample are sample images with a similarity not higher than the specified threshold.
[0273] Optionally, the sample construction unit 1007 is specifically configured to:
[0274] Select one sample image from a positive sample pair as the target sample, and select one sample image from other sample pairs as the candidate samples respectively;
[0275] Determine the distances between the target sample and each candidate sample respectively;
[0276] Sort each candidate sample according to its corresponding distance, and select at least one candidate sample at a specified sorting position as the negative sample corresponding to the target sample;
[0277] For each selected negative sample and the positive sample pair, form a training sample group, where the target sample in the positive sample pair is the positive sample in the training sample group, and the other sample in the positive sample pair is the anchor sample in the training sample group.
[0278] Since the present application proposes a model structure that is beneficial to the learning of two types of features by a metric learning embedding module for extracting shallow global features and a deep semantic embedding module for extracting deep semantic features, when classifying the image to be detected based on the image joint classification model in the present application, the image can be represented by jointly using both the underlying image representation and the deep semantic representation, enhancing the joint learning effect of the two, and reducing the time-consuming for feature extraction. Image classification based on the obtained image fusion features can obtain more accurate classification results, improving the accuracy and efficiency of image classification.
[0279] For the convenience of description, the above parts are divided into various modules (or units) according to their functions and described separately. Of course, when implementing the present application, the functions of the various modules (or units) can be implemented in the same or multiple software or hardware.
[0280] After introducing the image recognition method and device of the exemplary embodiment of the present application, next, an image recognition device according to another exemplary embodiment of the present application is introduced.
[0281] Those skilled in the art can understand that various aspects of the present application can be implemented as a system, a method, or a program product. Therefore, various aspects of the present application can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.
[0282] In some possible implementation manners, the image recognition device according to the present application may at least include a processor and a memory. Among them, the memory stores program code, and when the program code is executed by the processor, the processor executes the steps in the image recognition method according to various exemplary embodiments of the present application described in this specification. For example, the processor may execute as Figure 2 the steps shown in.
[0283] Based on the same inventive concept as the above method embodiment, an electronic device is further provided in the embodiment of the present application. In one embodiment, the electronic device may be a server, such asFigure 1 Server 120 as shown. In this embodiment, the structure of the electronic device may be as Figure 11 shown, including a memory 1101, a communication module 1103, and one or more processors 1102.
[0284] The memory 1101 is used to store the computer programs executed by the processor 1102. The memory 1101 may mainly include a program storage area and a data storage area. Among them, the program storage area may store the operating system, programs required to run the instant messaging function, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc.
[0285] The memory 1101 may be a volatile memory, such as a random-access memory (RAM); the memory 1101 may also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or the memory 1101 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1101 may be a combination of the above memories.
[0286] The processor 1102 may include one or more central processing units (CPUs) or be a digital processing unit, etc. The processor 1102 is used to implement the above image recognition method when calling the computer programs stored in the memory 1101.
[0287] The communication module 1103 is used to communicate with the terminal device and other servers.
[0288] In the embodiments of the present application, the specific connection medium between the above memory 1101, communication module 1103, and processor 1102 is not limited. In the embodiments of the present application Figure 11 it is connected between the memory 1101 and the processor 1102 through a bus 1104. The bus 1104 is described in thick lines in Figure 11 The connection methods between other components are only for illustrative purposes and are not limited thereto. The bus 1104 may be divided into an address bus, a data bus, a control bus, etc. For the convenience of description, Figure 11 only one thick line is used to describe it in
[0289] The computer storage medium is stored in the memory 1101, and computer-executable instructions are stored in the computer storage medium. The computer-executable instructions are used to implement the image recognition method of the embodiments of the present application. The processor 1102 is used to execute the above-mentioned image recognition method, as Figure 2 shown.
[0290] In another embodiment, the electronic device may also be other electronic devices, such as Figure 1 shown in the terminal device 110. In this embodiment, the structure of the electronic device may be as Figure 12 shown, including: communication component 1210, memory 1220, display unit 1230, camera 1240, sensor 1250, audio circuit 1260, Bluetooth module 1270, processor 1280 and other components.
[0291] The communication component 1210 is used to communicate with the server. In some embodiments, it may include a Wireless Fidelity (WiFi) module. The WiFi module belongs to short-range wireless transmission technology, and the electronic device can help users send and receive information through the WiFi module.
[0292] The memory 1220 can be used to store software programs and data. The processor 1280 executes various functions and data processing of the terminal device 110 by running the software programs or data stored in the memory 1220. The memory 1220 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. The memory 1220 stores an operating system that enables the terminal device 110 to operate. In the present application, the memory 1220 can store the operating system and various application programs, and can also store the code for executing the image recognition method of the embodiments of the present application.
[0293] The display unit 1230 can also be used to display the information input by the user or the information provided to the user, as well as the graphical user interface (GUI) of various menus of the terminal device 110. Specifically, the display unit 1230 may include a display screen 1232 disposed on the front of the terminal device 110. Among them, the display screen 1232 can be configured in the form of a liquid crystal display, a light-emitting diode, etc. The display unit 1230 can be used to display the image classification results, image retrieval results, etc. in the embodiments of the present application.
[0294] The display unit 1230 can also be used to receive input digital or character information and generate signal inputs related to the user settings and function control of the terminal device 110. Specifically, the display unit 1230 may include a touch screen 1231 disposed on the front of the terminal device 110, which can collect touch operations of the user thereon or nearby, such as clicking a button, dragging a scroll bar, etc.
[0295] Among them, the touch screen 1231 can cover the display screen 1232, or the touch screen 1231 and the display screen 1232 can be integrated to implement the input and output functions of the terminal device 110. After integration, it can be simply referred to as a touch display screen. In this application, the display unit 1230 can display application programs and corresponding operation steps.
[0296] The camera 1240 can be used to capture static images, and the user can post comments through the images captured by the camera 1240 via an application. The camera 1240 can be one or multiple. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the processor 1280 to convert it into a digital image signal.
[0297] The terminal device may further include at least one sensor 1250, such as an acceleration sensor 1251, a distance sensor 1252, a fingerprint sensor 1253, and a temperature sensor 1254. The terminal device may also be configured with other sensors such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, a light sensor, and a motion sensor.
[0298] The audio circuit 1260, the speaker 1261, and the microphone 1262 can provide an audio interface between the user and the terminal device 110. The audio circuit 1260 can transmit the electrical signal converted from the received audio data to the speaker 1261, and the speaker 1261 converts it into a sound signal for output. The terminal device 110 may also be configured with volume buttons for adjusting the volume of the sound signal. On the other hand, the microphone 1262 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1260 and then converted into audio data, and then the audio data is output to the communication component 1210 to be sent to, for example, another terminal device 110, or the audio data is output to the memory 1220 for further processing.
[0299] The Bluetooth module 1270 is used to interact with other Bluetooth devices with Bluetooth modules through the Bluetooth protocol. For example, the terminal device can establish a Bluetooth connection with a wearable electronic device (such as a smart watch) that also has a Bluetooth module through the Bluetooth module 1270, so as to perform data interaction.
[0300] The processor 1280 is the control center of the terminal device, connecting various parts of the entire terminal through various interfaces and circuits. By running or executing software programs stored in the memory 1220 and calling data stored in the memory 1220, it executes various functions of the terminal device and processes data. In some embodiments, the processor 1280 may include one or more processing units; the processor 1280 may also integrate an application processor and a baseband processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the baseband processor mainly processes wireless communication. It can be understood that the above baseband processor may not be integrated into the processor 1280. In this application, the processor 1280 can run the operating system, application programs, user interface display and touch response, as well as the image recognition method of the embodiments of this application. In addition, the processor 1280 is coupled to the display unit 1230.
[0301] In some possible implementation manners, various aspects of the image recognition method provided in this application can also be implemented in the form of a program product, which includes program code. When the program product runs on an electronic device, the program code is used to cause the electronic device to execute the steps in the image recognition method according to various exemplary embodiments of this application described above in this specification. For example, the electronic device can execute the steps as shown in Figure 2 shown.
[0302] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0303] A program product of an embodiment of the present application may be a portable compact disc read-only memory (CD-ROM) and include program code, and may be run on a computing device. However, the program product of the present application is not limited thereto. In this document, a readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with a command execution system, apparatus, or device.
[0304] A readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A readable signal medium may also be any readable medium other than a readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with a command execution system, apparatus, or device.
[0305] The program code contained on a readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the foregoing.
[0306] The program code for performing the operations of the present application may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).
[0307] It should be noted that although several units or subunits of the apparatus are mentioned in the above detailed description, such a division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more of the above-described units may be embodied in one unit. Conversely, the features and functions of one unit described above may be further divided and embodied by multiple units.
[0308] In addition, although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the shown operations must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.
[0309] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0310] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present application.
[0311] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.
Claims
1. An image recognition method, characterized in that, The method includes: extracting first shallow global features of the image to be detected based on a metric learning embedding module in a trained image classification joint model, where the first shallow global features represent global basic information of the image to be detected; and extracting first deep semantic features of the image to be detected based on a deep semantic embedding module in the image classification joint model, where the first deep semantic features represent image semantic information of the image to be detected; performing feature fusion on the first shallow global features and the first deep semantic features based on a semantic prediction module in the image classification joint model, and performing multi-label classification on the image to be detected based on the obtained first fusion features to obtain a classification result of the image to be detected; wherein, the image classification joint model is obtained by jointly training the metric learning embedding module and the deep semantic embedding module; a training sample data set for training the image classification joint model includes a plurality of training sample groups, and the plurality of training sample groups are obtained by the following method: According to the similarity between each collected sample image, every two sample images with a similarity higher than a specified threshold are formed into a positive sample pair; based on each positive sample pair, sample recombination is performed to obtain a plurality of training sample groups for constituting the training sample data set, and each training sample group includes three training samples: an anchor sample, a positive sample, and a negative sample, where the anchor sample and the positive sample are sample images with a similarity higher than the specified threshold, and the negative sample and the positive sample are sample images with a similarity not higher than the specified threshold.
2. The method according to claim 1, wherein The performing multi-label classification on the image to be detected based on the obtained first fusion features further includes: extracting second shallow global features of each candidate image in the image library based on the metric learning embedding module; and, extracting second deep semantic features of the image to be detected based on the deep semantic embedding module; performing feature fusion on each second shallow global feature and the corresponding second deep semantic feature based on the semantic prediction module to obtain a second fusion feature corresponding to each candidate image; respectively determining the similarity between each candidate image and the image to be detected based on the first fusion feature and the second fusion feature corresponding to each candidate image; determining a similar image corresponding to the image to be detected based on each similarity.
3. The method according to claim 1, wherein The image classification joint model further includes: a basic convolution module; before extracting first shallow global features of the image to be detected based on the metric learning embedding module in the trained image classification joint model, it further includes: inputting the image to be detected into the basic convolution module in the image classification joint model, and performing global feature extraction on the image to be detected based on the basic convolution module to obtain a global feature map corresponding to the image to be detected; The extracting first shallow global features of the image to be detected based on the metric learning embedding module in the trained image classification joint model includes: Based on the metric learning embedding module, perform embedding processing on the global feature map to obtain the first shallow global feature corresponding to the image to be detected.
4. The method according to claim 3, wherein The deep semantic embedding module includes a convolutional layer and an embedding layer composed of residual structures; Based on the deep semantic embedding module in the image classification joint model, extract the first deep semantic feature of the image to be detected, Based on the convolutional layer in the deep semantic embedding module, extract the semantic information features in the global feature map, and based on the embedding layer in the deep semantic embedding module, perform embedding processing on the extracted semantic information to obtain the first deep semantic feature corresponding to the image to be detected.
5. The method according to claim 3, wherein The learning rate of the semantic prediction module is higher than that of other modules, and the other modules include the basic convolutional module, the metric learning embedding module, and the deep semantic embedding module.
6. The method according to claim 4, wherein The basic convolutional module includes multiple convolutional layers; the network parameters corresponding to the convolutional layers in the basic convolutional module are initialized based on the parameters pre-trained on a specified sample library, and the network parameters corresponding to the convolutional layers in the deep semantic embedding module are initialized based on random initialization.
7. The method according to claim 1, wherein The image classification joint model is trained in the following manner: Select a training sample group from the training sample dataset; Input the selected training sample group into the trained image classification joint model to obtain the third shallow global feature output by the metric learning embedding module in the image classification joint model, the third deep semantic feature output by the deep semantic embedding module, and the classification vector output by the semantic prediction module. The classification vector represents the predicted probability of the training sample for each classification label; Construct an objective loss function based on the third shallow global feature, the third deep semantic feature, and the classification vector, and based on the objective loss function, adjust the network parameters of the image classification joint model multiple times until the image classification joint model converges, and output the trained image classification joint model.
8. The method according to claim 7, wherein The construction of the loss function based on the third shallow global feature, the third deep semantic feature, and the classification vector includes: Construct a first triplet loss function based on the third shallow global features corresponding to each training sample in the training sample group; and construct a second triplet loss function based on the third deep semantic features corresponding to each training sample; Construct a multi-label loss function based on the classification vector corresponding to each training sample and the corresponding multi-label vector. The multi-label vector represents the true probability of the training sample for each classification label; Obtain a classification entropy loss function based on the classification vector corresponding to each training sample; Perform weighted summation on the first triplet loss function, the second triplet loss function, the multi-label loss function, and the classification entropy loss function to obtain the objective loss function.
9. The method according to claim 8, wherein The construction of the first triplet loss function based on the third shallow global features corresponding to each training sample in the training sample group includes: Determine the first distance between the third shallow global feature corresponding to the anchor sample in the training sample group and the third shallow global feature corresponding to the positive sample, and the second distance between the third shallow global feature corresponding to the anchor sample and the third shallow global feature corresponding to the negative sample; Take the sum of the difference between the first distance and the second distance and a specified boundary value as the first target value, where the specified boundary value is used to represent the difference boundary of the similarity between the positive sample and the negative sample; Take the maximum value of the first target value and the reference parameter as the first triplet loss function.
10. The method according to claim 8, characterized in that Constructing the second triplet loss function based on the third deep semantic features corresponding to the respective training samples includes: Based on the third distance between the third deep semantic feature corresponding to the anchor sample in the training sample group and the third deep semantic feature corresponding to the positive sample, and the fourth distance between the third deep semantic feature corresponding to the anchor sample and the third deep semantic feature corresponding to the negative sample; Take the sum of the difference between the third distance and the fourth distance and a specified boundary value as the second target value, where the specified boundary value is used to represent the difference boundary of the similarity between the positive sample and the negative sample; Take the maximum value of the second target value and the reference parameter as the second triplet loss function.
11. The method according to claim 8, wherein Before adjusting the network parameters of the image classification joint model multiple times based on the target loss function, it further includes: Perform weighted summation on the first triplet loss function, the multi-label loss function, and the classification entropy loss function to obtain an intermediate loss function; Based on the intermediate loss function, adjust the network parameters of the deep semantic module in the image classification joint model a specified number of times.
12. The method according to claim 1, wherein The sample recombination based on positive sample pairs to obtain multiple training sample groups for constituting the training sample data set includes: Select one sample image from a positive sample pair as the target sample, and select one sample image from other sample pairs as the candidate sample respectively; Determine the distances between the target sample and each candidate sample respectively; Sort the candidate samples according to their respective corresponding distances, and select at least one candidate sample at a specified sorting position as the negative sample corresponding to the target sample; Form a training sample group by combining each selected negative sample with the positive sample pair respectively, where the target sample in the positive sample pair is the positive sample in the training sample group, and the other sample in the positive sample pair is the anchor sample in the training sample group.
13. An image recognition device, characterized in that, It includes: A first feature extraction unit for extracting the first shallow global feature of the image to be detected based on the metric learning embedding module in the trained image classification joint model, where the first shallow global feature characterizes the global basic information of the image to be detected; And A second feature extraction unit for extracting the first deep semantic feature of the image to be detected based on the deep semantic embedding module in the image classification joint model, where the first deep semantic feature characterizes the image semantic information of the image to be detected; A classification prediction unit, configured to perform feature fusion on the first shallow global feature and the first deep semantic feature based on the semantic prediction module in the image classification joint model, and perform multi-label classification on the image to be detected based on the obtained first fusion feature, so as to obtain a classification result of the image to be detected; wherein, the image classification joint model is obtained by jointly training the metric learning embedding module and the deep semantic embedding module; the training sample data set for training the image classification joint model includes multiple training sample groups, and the multiple training sample groups are obtained by the XX unit in the following manner: According to the similarity between each pair of collected sample images, every two sample images with a similarity higher than a specified threshold are combined into a positive sample pair; based on each positive sample pair, sample recombination is performed to obtain multiple training sample groups for constituting the training sample data set, and each training sample group includes three training samples: an anchor sample, a positive sample, and a negative sample, wherein the anchor sample and the positive sample are sample images with a similarity higher than the specified threshold, and the negative sample and the positive sample are sample images with a similarity not higher than the specified threshold.
14. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores program codes, and when the program codes are executed by the processor, the processor is caused to execute the steps of any one of claims 1 to 12.
15. A computer-readable storage medium, characterized in that, It includes program codes, and when the storage medium runs on an electronic device, the program codes are used to cause the electronic device to execute the steps of any one of claims 1 to 12.
16. A computer program product, characterized in that, It includes a computer program, characterized in that when the computer program is executed by a processor, the steps of any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Target detection method and device based on scene semantics, equipment and storage medium
CN112966697A
Vehicle appearance part identification method and device, electronic equipment and medium
CN113191364A