Model training methods, retrieval methods, models, devices, and media
By performing interference processing and Riemann optimization algorithm on the basic image, training images are generated and a target detection model is formed, which solves the problem of low recognition accuracy in vehicle identification and achieves effective detection and recognition in different scenarios.
Patent Information
- Application Number
- CN202111590804.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-23
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2041-12-23
AI Technical Summary
Existing image processing technology in vehicle recognition pays too much attention to the individual samples of the target to be detected, resulting in low recognition accuracy for changes in angle, light, etc., and an imbalance between global and local information, which affects recognition accuracy.
By interfering with the basic image and simulating different shooting angles, lighting, occlusion and other scenes, training images are generated, and the features are classified using the Riemann optimization algorithm of the neural network until the loss function converges to form a target detection model.
The vehicle recognition accuracy in different scenarios is improved, and effective detection and recognition of the target to be detected is achieved.
Smart Images

Figure CN114462479B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a model training method, a retrieval method, a model, a device, and a medium. Background Art
[0002] Generally, in image processing, object recognition refers to determining the class of a given image, while re-identification refers to finding images of the same class as a given image. Examples include pedestrian re-identification, vehicle re-identification, and pedestrian-vehicle mutual recognition.
[0003] In real-world scenarios, vehicle images are inevitably affected by lighting, occlusion, and shooting angles, and the vehicle itself presents different visual relationships from the front, back, and sides. For example, different shooting angles result in different degrees of distortion of the target being detected.
[0004] However, the existing system algorithm pays too much attention to the individual samples of the target to be detected, resulting in the inability to identify the target to be detected as the same type of object due to changes in angle, light, etc. during the recognition process, resulting in low recognition accuracy. In addition, due to the excessive attention paid to the local part of the target to be detected, the global and local information of the target to be detected are unbalanced, thereby reducing the recognition accuracy of the target to be detected. Summary of the Invention
[0005] In order to solve the above technical problems, the technical solution adopted in the first aspect of this application is to provide a target detection model training method, which includes: performing interference processing on multiple basic images to obtain multiple training images; inputting the multiple training images into the neural network to be trained, and classifying the features of the multiple training images based on the correlation information between the multiple training images, until the trained neural network is obtained when the loss function of the neural network is determined to converge based on the classification result; and determining the trained neural network as the target detection model.
[0006] To solve the above technical problems, the technical solution adopted in the second aspect of this application is to provide a method for retrieving image targets, which includes: obtaining an image to be retrieved; inputting the image to be retrieved into the target detection model as described in the first aspect, wherein the target detection model stores multiple base library images; using the target monitoring model, determining the cosine similarity between the image to be retrieved and the multiple base library images; and determining the base library images whose cosine similarity meets preset requirements.
[0007] To solve the above technical problems, the technical solution adopted in the third aspect of this application is to provide a target retrieval model, which includes:
[0008] An interference module, used for performing interference processing on multiple basic images to obtain multiple training images;
[0009] An input module, used to input multiple training images into the neural network to be trained;
[0010] a processing module, configured to classify features of the plurality of training images based on association information between the plurality of training images, and obtain a trained neural network when it is determined based on the classification result that the loss function of the neural network converges;
[0011] The determination module is used to determine the trained neural network as a target detection model.
[0012] In order to solve the above technical problems, the technical solution adopted in the fourth aspect of this application is to provide an electronic device, which includes: a processor and a memory, the memory stores a computer program, and the processor is used to execute the computer program to implement the method of the first aspect or the second aspect of this application.
[0013] In order to solve the above technical problems, the technical solution adopted in the fifth aspect of this application is to provide a computer-readable storage medium, which stores a computer program, and the computer program can implement the method of the first aspect or the second aspect of this application when executed by a processor.
[0014] The beneficial effect of the present application is that the present application can simulate various scenes such as different shooting angles, different lighting, different distances, different occlusions, etc. by performing interference processing on the basic image, so that the target to be detected presents different visual presentations, and then inputs the multiple training images obtained after the interference processing into the neural network to be trained, and mines the features of the multiple training images to realize the differences of the same type of targets for classification and division, until the trained neural network is obtained when the loss function of the neural network is determined to converge based on the classification result, forming a target detection model, and using the target monitoring model to realize detection and identification of the target to be detected in different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0016] Figure 1 This is a flow chart of an embodiment of the model training method of the present application;
[0017] Figure 2 yes Figure 1 A schematic diagram of a specific implementation process of step S12;
[0018] Figure 3 yes Figure 2A schematic diagram of a specific implementation process of step S23 and step S24;
[0019] Figure 4 yes Figure 2 A schematic diagram of a specific implementation process of step S22;
[0020] Figure 5 yes Figure 4 A schematic diagram of a specific implementation process of step S43;
[0021] Figure 6 This is a diagram of the structure distribution using Riemann optimization training during training;
[0022] Figure 7 This is a flowchart of an embodiment of the image target retrieval method of the present application;
[0023] Figure 8 This is a schematic block diagram of the structure of an embodiment of the target detection model of the present application;
[0024] Figure 9 This is a schematic block diagram of the structure of an electronic device embodiment of the present application;
[0025] Figure 10 This is a schematic block diagram of a circuit of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION
[0026] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0027] It will be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0028] It should also be understood that the terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0029] It should be further understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0030] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0031] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0032] In order to illustrate the technical solution of this application, the following is an explanation through specific embodiments. This application provides a target detection model training method. Figure 1 , Figure 1 This is a flow chart of an embodiment of the model training method of the present application, which specifically includes the following steps:
[0033] S11: performing interference processing on multiple basic images to obtain multiple training images;
[0034] Typically, in real-world scenarios, the base image corresponding to the target to be detected is inevitably affected by lighting, occlusion, and shooting angle, and the vehicle itself presents different visual relationships from the front, back, and sides. For example, different shooting angles lead to different degrees of distortion in the target to be detected, different lighting conditions lead to different brightness and pixel counts in the target to be detected, different distances lead to different sizes of the target to be detected, and different occlusions lead to different completeness rates of the target to be detected.
[0035] Therefore, in order to better simulate actual scenes, it is necessary to perform interference processing on multiple basic images. For example, multiple basic images are randomly sampled according to probability and randomly selected parts are erased to simulate occlusion effects, and they are rotated at random angles to force the network to learn rotation-invariant features. In other words, a series of interference processing is performed on multiple basic images, and the interference processing method can be flexibly set based on business needs. Specifically, at least one interference processing of erasing local processing, blurring processing, or occlusion processing is performed on multiple basic images to obtain multiple training images. Among them, the basic image uses a quick data set, and the clean data set refers to a collection of pictures without occlusion, damage, etc.
[0036] The probability can be preset manually, for example, 20% is erased and 80% remains unchanged. Then, a random number generated in the range of 0-20% indicates that the base image needs data enhancement. Since the target is compact in the base image, random erasing means that a region can be randomly selected to remove pixels based on the preset upper and lower limits of the rectangle's width and height and the aspect ratio. Specifically, the original data set corresponding to the base image is divided into a training set, a validation set, and a test set. Multiple base images in the training set are interfered with to obtain multiple training images.
[0037] S12: Inputting a plurality of training images into the neural network to be trained, and classifying features of the plurality of training images based on correlation information between the plurality of training images, until a loss function of the neural network is determined to converge based on the classification result, thereby obtaining a trained neural network;
[0038] The same target to be detected may be divided into multiple categories. For example, a car can be divided into doors, wheels, front, and rear categories. Each category can be set to represent it to distinguish it from other categories. For example, a car viewed from the front and a car viewed from the back are obviously different, but the difference is not significant in the basic image. They can be classified into the same category but have certain differences between them.
[0039] Among them, if the target to be detected is a vehicle, in order to train the model for each part of the vehicle, the vehicle category can be divided into 700 categories, and the number of set identifiers can be the number of categories marked in the initial training set.
[0040] Specifically, multiple training images are input into the neural network to be trained, and the features of the multiple training images are classified by obtaining the correlation information between the multiple training images. The convergence of the loss function of the neural network is calculated according to the classification results, and the trained neural network is determined by the convergence of the loss function.
[0041] S13: Determine the trained neural network as a target detection model.
[0042] By classifying the features of multiple training images, a classification result can be obtained, and a trained neural network can be obtained. Specifically, because a neural network often includes multiple parameters, and the trained neural network indicates that multiple parameters in the neural network also meet the preset conditions, it can be used as a target monitoring model. In a specific embodiment, the target detection model can also be exported.
[0043] Therefore, this application can simulate various scenes such as different shooting angles, different lighting, different distances, different occlusions, etc. by performing interference processing on the basic image, so that the target to be detected presents different visual presentations, and then input the multiple training images obtained after interference processing into the neural network to be trained, and mine the features of the multiple training images to realize the differences of similar targets for classification and division, until the trained neural network is obtained when the loss function of the neural network is determined to converge based on the classification results, forming a target detection model, and using this target monitoring model to realize detection and identification of the target to be detected in different scenarios.
[0044] The correlation information includes at least the Riemann distance. Based on the correlation information between multiple training images, the features of multiple training images are classified. Figure 2 , Figure 2 yes Figure 1 A specific implementation flow diagram of step S12 includes:
[0045] S21: using the neural network to be trained to perform feature extraction on the plurality of training images respectively, to obtain features corresponding to the plurality of training images respectively;
[0046] In this application, since the neural network includes a backbone and multi-scale branches connected to the backbone, both the backbone and the branches contain convolutional layers and pooling layers. The backbone is a ResNet series network pre-trained on ImageNet. For the sake of computational efficiency, the parameters of the backbone layer are fixed, so Riemann optimization is mainly optimized on the multi-scale branches based on the backbone. Of course, when appropriate, such as when the loss function has better convergence, the convolutional layers and pooling layers of the backbone and their parameters can be adaptively adjusted.
[0047] Specifically, multiple training images are processed separately according to the convolutional layer in the neural network to be trained to obtain a feature map corresponding to each training image.
[0048] The feature maps corresponding to each training image are processed according to the pooling layer in the neural network to be trained. Specifically, features of different granularities can be obtained by extracting the feature maps through horizontal segmentation, and the features corresponding to the training images can be obtained based on the features of different granularities. Horizontal segmentation is also called horizontal segmentation. For example, if an 8*8 image is segmented 0 and 1 times horizontally, one 8*8 image and two 4*8 images are obtained, which are the feature sizes of different granularities. Local information comes from local pooling, which is used to reduce the network's field of view and prevent the neural network from focusing too much on global information.
[0049] S22: Classify each obtained feature based on the Riemann distance between multiple training images to obtain multiple feature categories and class centers corresponding to each feature;
[0050] Specifically, the short thread distance of the vertices of each training image on the N-dimensional sphere is regarded as the Riemann distance, which represents the size or length of the edge connecting two vertices, wherein N is a positive integer greater than or equal to 2, such as N=256. The features of multiple training images are classified according to Riemann optimization. The obtained features can be classified by using the Riemann gradient descent method on the 256-dimensional unit sphere to obtain multiple feature categories and the class center corresponding to each feature.
[0051] The class center corresponding to a feature is determined based on the feature category to which the feature belongs, and after classifying the features of multiple training images, the distance between each feature and the class center needs to be further judged.
[0052] S23: normalize each feature according to its corresponding class center;
[0053] Specifically, normalization is a method of simplifying calculations by transforming dimensionless expressions into scalars. This was primarily proposed to facilitate data processing. Mapping data to a standard threshold makes processing more efficient and effective, and should be considered part of standard processing.
[0054] For example, if the image pixel threshold is [0, 256], all other pixels collected in the image can be converted to this range for subsequent comparison and calculation. Here, since the class center corresponding to each feature is determined based on the feature category to which the feature belongs, a class center represents the threshold standard for that category, and the features of that category can be normalized. Similarly, for different class centers, each feature can be normalized.
[0055] S24: Based on the result of the normalization processing, determine the loss value of the loss function, and determine whether the loss function converges based on the loss value.
[0056] Specifically, the normalized result is obtained, such as input into the loss function to calculate the loss value, so as to shorten the distance between the class center and each feature, cluster each feature, and determine whether the loss function converges based on the loss value.
[0057] For further information, see Figure 3 , Figure 3 yes Figure 2 A schematic diagram of a specific implementation flow of step S23 and step S24, wherein the specific step of normalizing each feature according to the class center corresponding to each feature is referred to in step S31, and the specific step of determining the loss value of the loss function based on the result of the normalization process is referred to in steps S32 and S33, which are specifically as follows:
[0058] S31: Perform dot products between the features corresponding to the multiple training images and the class centers corresponding to the features to obtain the cosine distance;
[0059] Since the features corresponding to the training images are eigenvectors, and the class center is also called the center vector, the cosine distance between the eigenvector and the center vector can be obtained by doing a dot product between the eigenvector and the center vector corresponding to the feature.
[0060] Since there is one central vector in a category and multiple eigenvectors in a category, multiple cosine distances can be obtained by performing dot products between multiple eigenvectors and the central vector to determine the distance between the eigenvector and the central vector.
[0061] S32: Input the cosine distance into the normalization layer for normalization;
[0062] Specifically, the cosine distance is input into the normalization layer for normalization, where the normalization layer includes scale and offset parameters. The specific meaning of normalization is as follows Figure 2 Step S23 in is not described again here.
[0063] S33: Input the normalized result into the loss function layer for loss calculation to obtain the cross entropy, and obtain the loss value based on the cross entropy.
[0064] Specifically, the normalized result is input into the loss function layer for loss calculation, where the loss function layer can use the cross entropy loss function. The cross entropy loss function is often used in classification problems, especially when neural networks are used for classification problems. Cross entropy is also often used as a loss function. In addition, since cross entropy involves calculating the probability of each category, cross entropy almost always appears together with the sigmoid (or softmax) function.
[0065] Therefore, the probability output is obtained through the softmax function, and the normalized result is obtained. The normalized result is then processed through the cross entropy loss function to obtain the cross entropy, and the loss value is obtained based on the cross entropy. The smaller the loss value, the better the classification effect on the training set.
[0066] When using gradient descent to update parameters, the speed at which a neural network learns depends on two factors: the learning rate and the partial derivative. Generally speaking, using the softmax logistic function to calculate probabilities and combining it with cross-entropy as the loss function results in faster learning when the neural network performs poorly and slower learning when it performs well.
[0067] Furthermore, based on the Riemann distance between multiple training images, the obtained features are classified to obtain multiple feature categories and the class center corresponding to each feature. According to the Riemann distance between multiple training images, the weight of the edge is obtained; among them, the Riemann distance is a relatively general concept, specifically referring to the short-thread distance between two vertices on the high-dimensional sphere of the training image, and the "edge weight" refers to the size or length between two vertices of the training image.
[0068] See also Figure 4 , Figure 4 yes Figure 2 A specific implementation flow diagram of step S22 in FIG. 1 specifically includes the following steps:
[0069] S41: Perform K-nearest neighbor optimization calculation on multiple training images to obtain the Riemann distance;
[0070] Specifically, because the pooling layer includes multiple global pooling layers and multiple local pooling layers, based on Riemann optimization, the feature vectors are optimized and calculated separately to obtain multiple subclass centers, which are used to represent the centers of multiple classifications. In addition, the Riemann distance between feature vectors can be obtained. Specifically, this can be calculated using the formula function corresponding to the Riemann distance.
[0071] Among them, after inputting unlabeled data into K-Nearest Neighbor optimization, each feature in the new data is compared with the corresponding feature of the data in the sample set, and the classification label of the data with the most similar features (nearest neighbor) in the sample set is extracted.
[0072] S42: Use Riemann distance as edge weight;
[0073] Since the Riemann distance is actually very similar to the edge weight, in a sense, the edge weight is another way of saying the Riemann distance. Therefore, the Riemann distance will be used as the edge weight here.
[0074] Specifically, the Riemann distance is similar to the geodesic distance. For example, for short-range lines, two-dimensional space is a straight-line distance, three-dimensional space is an arc distance, higher dimensions are curved spaces, and so on.
[0075] S43: Compare the weight of the edge with a preset threshold, and classify each feature obtained using the comparison result to obtain multiple feature categories and the class center corresponding to each feature.
[0076] Specifically, the weight of the edge is compared with the preset threshold, and the comparison results are used to classify the features corresponding to multiple training images, and the multiple categories to which the features belong and the class centers corresponding to each feature are obtained. Figure 5 , Figure 5 yes Figure 4 A specific implementation flow diagram of step S43 in FIG. 4 includes the following steps:
[0077] S51: Whether the weight of the edge is less than the preset threshold;
[0078] To compare edge weights with a preset threshold, the neural network also uses a threshold, such as pi / 4, pi / 8, or pi / 16, to compare the difference between edge weights and determine whether two samples belong to the same category based on the size of the difference. Specifically, the comparison can be performed by determining whether the edge weight is less than a given threshold.
[0079] Because the sample feature points are located on the unit hypersphere, the Riemann distance is limited to the interval [0, pi]. After the neural network is given, the threshold selection is only related to the network input, so the optimal threshold can be found through multiple experiments. The unit hypersphere refers to the surface of an n-dimensional sphere with a radius of 1, where n is a positive integer greater than or equal to 2, such as 256 dimensions.
[0080] When the edge weight is less than the preset threshold, the process proceeds to step S52, which assigns the features corresponding to the training image to the same category. The assignment here refers to the logical merging of different subcategories. For example, category 1 is divided into 1-1 and 1-2, both of which are considered 1 when calculating the loss. When the edge weight is equal to or greater than the preset threshold, the process proceeds to step S53, which assigns the features corresponding to the training image to another category. This means that a new category can be constructed.
[0081] S54: performing classification optimization according to the category to which the feature corresponding to the training image belongs, and obtaining the class center corresponding to the feature;
[0082] Specifically, according to the divided categories, the class center on the high-dimensional sphere surface can be optimized, that is, classification optimization is performed according to the category to which the features corresponding to the training image belong to obtain the class center corresponding to the feature.
[0083] S55: Count all the class centers to serve as the classification heads of the neural network to be trained, and form a corresponding relationship between the classification heads and the target categories.
[0084] Specifically, all class centers are counted as the classification head of the neural network to be trained, and finally the classification head and the corresponding category record information are output, that is, the correspondence between the classification head and the target category is formed.
[0085] Among them, the number of training times can be selected as needed. This application only describes the embodiment based on one training, and the number of iterations can be 200 times. The empirical value debugged manually can also be Riemann optimization at 80 times, 120 times and 150 times, so as to obtain the Riemann distance, and then mine the features of multiple training images to achieve the difference between similar targets for class division.
[0086] In addition, when a part of the backbone is released, the parameters of the backbone can also participate in the training and optimization process of the target detection model, and the training time can also be extended, which is not limited here.
[0087] For further information, see Figure 6 , Figure 6 This is a diagram of the structure distribution using Riemann optimization training during the training process. Figure 6 In the figure, the dotted box indicates that the pre-training parameters are shared. It is the backbone part, including a 5-layer backbone network. The five layers 01234 of the backbone network correspond to Figure 6 Layers 1 to 5 in the [1] . If a total of 200 iterations (epochs) are performed, the multi-center subroutine can be executed at the 80th, 120th, and 150th epochs. This subroutine takes multiple base images as input, inputs the backbone, and generates feature maps. Pooling is performed to obtain features from multiple training images, and then multiple class centers are calculated using a KNN optimization algorithm. The resulting features are then dot-producted with the class centers to obtain cosine distances, which are then passed through a softmax layer with scale and offset parameters and fed into the loss function layer. During training, the backbone network is first fixed and subsequent branches are optimized. After 20 epochs, the backbone is also added to the iterations. Training lasts for 200 epochs, and the learning rate is annealed at 0.1 at the 80th, 120th, and 150th epochs.
[0088] Pretrained parameter sharing is a form of knowledge transfer, meaning that weights are preloaded during training. These weights are typically obtained through training on classification tasks. Pooling layer connections combine their outputs and send them to the next module. Class center connections involve KNN calculating the class center for each pooling layer output.
[0089] Alternatively, the Alternating Direction Method of Multipliers (ADMM) can be used as an alternative to the Riemann optimization algorithm. This computational framework is used to solve separable convex optimization problems. Due to its fast processing speed and good convergence, ADMM is suitable for solving distributed convex optimization problems, particularly statistical learning problems. It is primarily used when the solution space is large, requiring a block-by-block solution, and the absolute accuracy of the solution is not too high.
[0090] The specific difference is that Riemann optimization executes the gradient descent algorithm directly on the manifold, while the alternating direction multiplier method is executed in the Euclidean space.
[0091] In order to illustrate the technical solution of the present application, the following is an explanation through specific embodiments. The present application also provides an image target retrieval method, which is based on a model training method and is specifically applied to a target detection model. Figure 7 , Figure 7 FIG1 is a flow chart of an embodiment of a method for retrieving an image target of the present application, which specifically includes the following steps:
[0092] S61: Acquire the image to be retrieved;
[0093] Typically, the image to be retrieved can be obtained by actual shooting as a retrieval image; the image to be retrieved can also be obtained by obtaining a certain frame of an image video as a retrieval image. Of course, those skilled in the art can also obtain the image to be retrieved by other means, which are not limited here.
[0094] S62: Inputting the image to be retrieved into the target detection model, wherein the target detection model stores a plurality of base images;
[0095] Specifically, based on the target detection model obtained by the model training method, the image to be retrieved is input into the trained target detection model to perform cosine similarity calculation on multiple base library images to obtain multiple similarities.
[0096] Specifically, the images in the training set constitute the base library images, where the target detection model stores multiple base library images. The base library images can be constructed by the user himself or added before the product leaves the factory. The user can also continue to expand the number of images in the base library.
[0097] S63: Determine the cosine similarity between the image to be retrieved and the multiple base images using the target detection model;
[0098] Specifically, the image to be retrieved and multiple base images can be input into the cosine function to calculate the cosine similarity, so that the cosine similarity between the image to be retrieved and the multiple base images can be obtained, which is used to compare the matching degree between the image to be retrieved and the base images.
[0099] S64: Determine the base library images whose cosine similarity meets the preset requirements.
[0100] Specifically, a preset cosine similarity value can be manually preset in advance to filter the multiple calculated cosine similarities. For example, the base images corresponding to the cosine similarities before the preset ranking can be extracted as the search results. In other words, the base images with the highest similarity ranking with the search results can be extracted as the search results as needed.
[0101] Therefore, this application provides a multi-center vehicle re-identification method based on Riemann optimization, wherein the multi-center Riemann optimization algorithm realizes adaptive adjustment of the number of classification heads.
[0102] In this way, by adopting an end-to-end approach as a whole, local multi-center calculations are performed using subroutines, and the computational complexity and storage complexity are reduced compared to the global end-to-end complexity; multi-center calculations are not only applicable to vehicles, but also to other comparable situations.
[0103] In order to illustrate the technical solution of this application, this application also provides a retrieval model, please refer to Figure 8 , Figure 8 This is a schematic block diagram of the structure of an embodiment of the target detection model of the present application. The target detection model 7 includes:
[0104] An interference module 71 is used to perform interference processing on multiple basic images to obtain multiple training images;
[0105] An input module 72, configured to input a plurality of training images into the neural network to be trained;
[0106] A processing module 73 is configured to classify features of the plurality of training images based on correlation information between the plurality of training images, and obtain a trained neural network when it is determined based on the classification result that the loss function of the neural network has converged;
[0107] The determination module 74 is used to determine the trained neural network as a target detection model.
[0108] Therefore, this application can simulate various scenes such as different shooting angles, different lighting, different distances, different occlusions, etc. by performing interference processing on the basic image, so that the target to be detected presents different visual presentations, and then input the multiple training images obtained after interference processing into the neural network to be trained, and mine the features of the multiple training images to realize the differences of similar targets for classification and division, until the trained neural network is obtained when the loss function of the neural network is determined to converge based on the classification results, forming a target detection model, and using this target monitoring model to realize detection and identification of the target to be detected in different scenarios.
[0109] In order to illustrate the technical solution of the present application, the present application also provides an electronic device, which can be a computer or a mobile phone, etc., without specific limitation, please refer to Figure 9 , Figure 9 This is a schematic block diagram of the structure of an electronic device embodiment of the present application. The electronic device 8 includes: a processor 81 and a memory 82. The memory 82 stores a computer program 821. The processor 81 is used to execute the computer program 821 to implement the method of the first aspect or the second aspect of the embodiment of the present application, which will not be repeated here.
[0110] In addition, this application also provides a computer-readable storage medium, see Figure 10 , Figure 10 This is a circuit schematic block diagram of an embodiment of a computer-readable storage medium of the present application. The computer-readable storage medium 9 stores a computer program 91. When the computer program 91 is executed by a processor, it can implement the method of the first aspect or the second aspect of the embodiment of the present application, which will not be repeated here.
[0111] If it is implemented in the form of a software functional unit and sold or used as an independent product, it can also be stored in a device with a storage function. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage device, including a number of instructions (program data) to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor (processor) to execute all or part of the steps of the various embodiments of the present invention. The aforementioned storage device includes various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and electronic devices such as computers, mobile phones, laptops, tablet computers, cameras, etc. having the above-mentioned storage media.
[0112] The description of the execution process of program data in a device with a storage function can be referred to the description in the above-mentioned embodiment of the model training method of the present application, and will not be repeated here.
[0113] The above description is merely an embodiment of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A target detection model training method, characterized in that: The method comprises: Perform interference processing on multiple basic images to obtain multiple training images; Inputting the plurality of training images into a neural network to be trained, and classifying features of the plurality of training images based on association information between the plurality of training images, until a loss function of the neural network is determined to converge based on a result of the classification, thereby obtaining a trained neural network; Determine the trained neural network as the target detection model; The association information includes Riemann distance, and the classifying the features of the plurality of training images based on the association information between the plurality of training images includes: Using the neural network to be trained, feature extraction is performed on each of the plurality of training images to obtain features corresponding to each of the plurality of training images; Based on the Riemann distance between the plurality of training images, the obtained features are classified to obtain a plurality of feature categories and class centers corresponding to the features, wherein the class center corresponding to one of the features is determined based on the feature category to which the feature belongs.
2. The method according to claim 1, characterized in that After classifying the features of the plurality of training images, the method further includes: Normalizing each feature according to the class center corresponding to each feature; Based on a result of the normalization process, a loss value of the loss function is determined, and based on the loss value, it is determined whether the loss function converges.
3. The method according to claim 2, characterized in that The normalizing of the features according to the class centers corresponding to the features includes: Performing dot products between the features corresponding to the plurality of training images and the cluster centers corresponding to the features to obtain cosine distances; Inputting the cosine distance into a normalization layer for normalization processing, wherein the normalization layer includes scale and offset parameters; Determining the loss value of the loss function based on the result of the normalization processing includes: The normalized result is input into the loss function layer for loss calculation to obtain cross entropy, and the loss value is obtained based on the cross entropy.
4. The method according to claim 1, wherein The method of classifying the obtained features based on the Riemann distance between the plurality of training images to obtain a plurality of feature categories and class centers corresponding to the features includes: Obtaining edge weights according to the Riemann distances between the plurality of training images; The weight of the edge is compared with a preset threshold, and the obtained features are classified using the comparison result to obtain multiple feature categories and the class center corresponding to each feature.
5. The method according to claim 4, characterized in that Obtaining edge weights according to the Riemann distances between the plurality of training images includes: Performing K-nearest neighbor optimization calculation on the plurality of training images to obtain the Riemann distance; The Riemann distance is used as the weight of the edge.
6. The method according to claim 4, characterized in that The weight of the edge is compared with a preset threshold, and the features corresponding to the plurality of training images are classified using the comparison result to obtain a plurality of categories to which the features belong and the class center corresponding to each feature, including: Compare the weight of the edge with a preset threshold, where the preset threshold is pi / 8; When the weight of the edge is less than a preset threshold, the features corresponding to the training images are classified into the same category; When the weight of the edge is equal to or greater than a preset threshold, classifying the feature corresponding to the training image as another category; Performing classification optimization according to the category to which the features corresponding to the training images belong, and obtaining the class center corresponding to the features; All class centers are counted to serve as classification heads of the neural network to be trained, and a corresponding relationship between the classification heads and target classes is formed.
7. The method according to any one of claims 1 to 6, characterized in that The neural network includes a backbone and multi-scale branches connected to the backbone, wherein the backbone and the branches both include convolutional layers and pooling layers; The step of extracting features from the plurality of training images using the neural network to be trained to obtain features corresponding to the plurality of training images includes: Processing the plurality of training images separately according to the convolutional layer in the neural network to be trained to obtain a feature map corresponding to each of the training images; The feature map corresponding to each of the training images is processed according to the pooling layer in the neural network to be trained, the feature map is extracted by horizontal segmentation to obtain features of different granularities, and the features corresponding to the training image are obtained based on the features of different granularities.
8. The method according to any one of claims 1 to 6, characterized in that The interference processing is performed on the multiple basic images to obtain multiple training images, including: At least one interference processing of erasing local processing, blurring processing or occlusion processing is performed on the multiple basic images to obtain the multiple training images.
9. A method for retrieving an image target, characterized in that: Get the image to be retrieved; Inputting the image to be retrieved into the target detection model according to any one of claims 1 to 6, wherein the target detection model stores a plurality of base images; The target detection model is used to determine the cosine similarity between the image to be retrieved and the plurality of base images; and base images whose cosine similarity meets a preset requirement are determined.
10. A target detection model, characterized in that: include: An interference module, used for performing interference processing on multiple basic images to obtain multiple training images; An input module, configured to input a plurality of the training images into the neural network to be trained; a processing module configured to classify features of the plurality of training images based on association information between the plurality of training images, until a trained neural network is obtained when it is determined based on the classification results that the loss function of the neural network has converged; wherein the association information includes Riemann distance, and the processing module is specifically configured to use the neural network to be trained to extract features from the plurality of training images to obtain features corresponding to the plurality of training images; and classify each obtained feature based on the Riemann distance between the plurality of training images to obtain a plurality of feature categories and class centers corresponding to each feature, wherein the class center corresponding to one of the features is determined based on the feature category to which the feature belongs; The determination module is used to determine the trained neural network as a target detection model.
11. An electronic device, characterized in that: include: A processor and a memory, wherein the memory stores a computer program, and the processor is configured to execute the computer program to implement the method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 9 can be implemented.
Citation Information
Patent Citations
Vehicle re-identification method in multi-view environment based on multi-center measurement loss
CN111814584A
Target re-identification method, network training method thereof and related device
CN111814655A