Feature rearrangement method and device, storage medium and electronic device
By determining the adjacency matrix of the sample set and calculating the similarity of the feature elements in the graph search technology, the problems of complex and low validity of the search results in the prior art are solved, and more efficient image retrieval is achieved.
Patent Information
- Application Number
- CN202110833210.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-22
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-07-22
AI Technical Summary
In the prior art, the implementation method of image search results using image search results is complex, the search results are interfering with and low effectiveness.
By determining the adjacency matrix of the sample set, the similarity between feature elements is calculated, and the feature elements are rearranged according to the similarity, a graph convolution network is constructed to improve the effectiveness of the search results.
The processing efficiency of searching pictures with pictures is improved, the interference of search results is reduced, and the effectiveness of search results is improved.
Smart Images

Figure CN113505843B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision and video image processing technology, and in particular to a feature rearrangement method and device, a storage medium and an electronic device. Background Art
[0002] With the development of information technology and social progress, a large amount of video and image information is recorded every moment, but the organization and search of this huge amount of information is a difficult problem. Image search technology can alleviate this pain point to a great extent. Image search mainly extracts the features of images or videos, compares the extracted features with the feature base library, and obtains similar targets to achieve the purpose of image and video retrieval. It has broad application prospects in the fields of Internet search engines, public security, etc.
[0003] In the related technology, when re-identifying pedestrians, two networks, one large and one small, are designed and trained at the same time. During the training process, the small network is allowed to approach the effect of the large network to improve the performance of the small network. However, since two different networks need to be designed and the stronger representation ability of the large network is used to improve the effect of the small network, the small network needs to approach the characteristics of the large network while learning label information. Increasing the amount of calculation will also lead to unstable training, and greater requirements are placed on the design of the network.
[0004] In addition, for image search methods, related technologies use neural networks to simultaneously extract high-level and low-level features, and then fuse them to further reduce the image feature dimension. The reduced features are used as image features for retrieval, which improves the retrieval effect without significantly increasing the retrieval time. However, the fusion of high-level semantic features with low-level features is easily affected by factors such as environmental background and lighting. Fusion of low-level features with high-level features on a large-scale data set will introduce uncertainty, which will interfere with the retrieval efficiency of the main target.
[0005] Furthermore, related technologies have proposed using an adversarial network to generate a set of images of different poses, and then fusing the generated images with the original images at the bottom layer and sending them to the feature extraction network to improve the representation ability of features and the effect of pedestrian re-identification. However, since the adversarial generative network is used to generate pseudo data and use pseudo data information to enhance features, the pseudo data generated by this solution usually has pseudo textures, which will affect the effectiveness of feature extraction and ultimately interfere with the effectiveness of retrieval.
[0006] In view of the above problems, the existing technology has complex implementation methods for image search retrieval results, interference in retrieval results, low effectiveness, and other problems, and no effective solution has been proposed so far. Summary of the invention
[0007] The embodiments of the present invention provide a feature rearrangement method and device, a storage medium and an electronic device to at least solve the technical problems in the prior art that the implementation of image search retrieval results is complex, the retrieval results have interference, and the effectiveness is low.
[0008] According to one aspect of an embodiment of the present invention, a feature rearrangement method is provided, comprising: determining an adjacency matrix of a first sample set, wherein the first sample set includes a plurality of feature elements and samples corresponding to the feature elements; if it is determined according to the adjacency matrix that a first feature element and a second feature element having a connection relationship exist in the first sample set, then determining the similarity between the first feature element and the second feature element, wherein the connection relationship is used to indicate that the sample corresponding to the first feature element and the sample corresponding to the second feature element have the same target sample; and rearranging the feature elements in the first sample set according to the similarity.
[0009] Optionally, before determining the adjacency matrix of the first sample set, the method further includes: inputting the second sample set to be retrieved and queried into the first target neural network for feature extraction to obtain a first feature of the second sample set; comparing the first feature with a preset second feature in the database to obtain a first set with similar third features, wherein the third feature is used to indicate a first feature existing in the preset second feature whose similarity meets a preset threshold; performing association determination on each feature element in the first set in the second sample set to be retrieved and queried to determine a target number of samples corresponding to each feature element to obtain a second set; merging the first set with the second set to obtain the first sample set.
[0010] Optionally, after rearranging the feature elements in the first sample set according to the similarity, the method further includes: constructing input features of the graph convolutional network according to the feature elements in the first sample set, wherein the feature function corresponding to the input features is: ,in, represents the sample features corresponding to the first sample set extracted using the first target neural network, The standard features extracted by the feature extraction network same as the first target neural network are used for the feature elements in the first sample set. is a cosine distance function, which is used to indicate the similarity between the sample features of the sample set to be queried and the standard features; a target graph convolutional network is established according to the input features and the adjacency matrix; the target graph convolutional network is iteratively updated to determine a second target neural network for rearranging feature elements.
[0011] Optionally, after iteratively updating the target graph convolutional network to determine the second target neural network for rearranging feature elements, the method further includes: setting a training label for the second target neural network, wherein the training label is used to indicate the similarity classification result of the second target neural network for the first sample set; when the training label is a first label, determining that the sample feature corresponding to the training label and the standard feature corresponding to the feature element belong to the same category or the same object identifier, and the first label is used to characterize that there are similar feature elements in the similarity classification result of the first sample set; when the training label is a second label, determining that the sample feature corresponding to the training label and the standard feature corresponding to the feature element do not belong to the same category or the same object identifier; the second label is used to characterize that there are no similar feature elements in the similarity classification result of the first sample set.
[0012] Optionally, the second sample set to be retrieved is input into the first target neural network for feature extraction to obtain the first feature of the second sample set, including: determining the object identifiers contained in the second sample set, wherein the object identifiers are used to indicate the confirmed unique identifier of the attribute information of each object; classifying the second sample set according to the object identifiers, and preprocessing different samples under the same object identifier, wherein the preprocessing includes at least one of the following: random erasing of data in the samples, random rotation of data in the samples, random flipping of data in the samples, and random graying of data in the samples; and according to the classification results, the preprocessed second sample set is sequentially input into the first target neural network for feature extraction to obtain the first feature of the second sample set.
[0013] Optionally, the above method also includes: reorganizing the preprocessed second sample set to obtain a third sample set; inputting the third sample set into the first target neural network for feature extraction to obtain a fourth feature; and determining the similarity loss function corresponding to the first target neural network processing the second sample set based on the difference between the fourth feature and the first feature.
[0014] Optionally, before obtaining a second sample set of queries to be retrieved and inputting it into the first target neural network for feature extraction to obtain the first feature of the second sample set, the method further includes: training the training neural network using a preset fourth sample set to obtain the first target neural network, wherein the target loss function used in the process of training the training neural network is a loss function determined based on the first loss function and / or the similarity loss function, the first loss function is obtained by weighted summing a plurality of basic loss functions corresponding to the fourth sample set using a first preset weight coefficient, the plurality of basic loss functions include at least one of the following: a cross entropy loss function, a triplet loss function, and a circle loss function; the similarity loss function is obtained by weighted summing a self-similarity loss function and a mutual similarity loss function corresponding to the fourth sample set using a second preset weight coefficient, the self-similarity loss function is used to indicate the mean square error of the feature differences before and after the fourth sample set is reorganized, and the mutual similarity loss function is used to indicate the correlation coefficient matrix corresponding to the features before and after the fourth sample set is reorganized.
[0015] According to another aspect of an embodiment of the present invention, a feature rearrangement method device is further provided, comprising: a first determination module, used to determine an adjacency matrix of a first sample set, wherein the first sample set includes multiple feature elements and samples corresponding to the feature elements; a second determination module, used to determine the similarity between the first feature element and the second feature element if it is determined according to the adjacency matrix that there are a first feature element and a second feature element having a connection relationship in the first sample set, wherein the connection relationship is used to indicate that the sample corresponding to the first feature element and the sample corresponding to the second feature element have the same target sample; and a rearrangement module, used to rearrange the feature elements in the first sample set according to the similarity.
[0016] In an exemplary embodiment, the above-mentioned device also includes: a comparison module, which inputs the second sample set to be retrieved and queried into the first target neural network for feature extraction to obtain the first feature of the second sample set; compares the first feature with the preset second feature in the database to obtain a first set of similar third features, wherein the third feature is used to indicate the first feature existing in the preset second feature whose similarity meets the preset threshold; associates each feature element in the first set in the second sample set to be retrieved and queried, and determines the target number of samples corresponding to each feature element to obtain the second set; merges the first set with the second set to obtain the first sample set.
[0017] In an exemplary embodiment, the apparatus further includes: a construction module, configured to construct an input feature of the graph convolutional network according to the feature elements in the first sample set, wherein the feature function corresponding to the input feature is: ,in, represents the sample features corresponding to the first sample set extracted using the first target neural network, The standard features extracted by the feature extraction network same as the first target neural network are used for the feature elements in the first sample set. is a cosine distance function, which is used to indicate the similarity between the sample features of the sample set to be queried and the standard features; a target graph convolutional network is established according to the input features and the adjacency matrix; the target graph convolutional network is iteratively updated to determine a second target neural network for rearranging feature elements.
[0018] In an exemplary embodiment, the above-mentioned construction module also includes: a label unit, used to set a training label for the second target neural network, wherein the training label is used to indicate the similarity classification result of the second target neural network for the first sample set; when the training label is a first label, it is determined that the sample feature corresponding to the training label and the standard feature corresponding to the feature element belong to the same category or the same object identification, and the first label is used to characterize that there are similar feature elements in the similarity classification result of the first sample set; when the training label is a second label, it is determined that the sample feature corresponding to the training label and the standard feature corresponding to the feature element do not belong to the same category or the same object identification; the second label is used to characterize that there are no similar feature elements in the similarity classification result of the first sample set.
[0019] In an exemplary embodiment, the above-mentioned comparison module is also used to determine the object identifiers contained in the second sample set, wherein the object identifier is used to indicate the confirmed unique identifier of the attribute information of each object; the second sample set is classified according to the object identifier, and different samples under the same object identifier are preprocessed, wherein the preprocessing includes at least one of the following: random erasing of data in the sample, random rotation of data in the sample, random flipping of data in the sample, and random graying of data in the sample; according to the classification result, the preprocessed second sample set is sequentially input into the first target neural network for feature extraction to obtain the first feature of the second sample set.
[0020] In an exemplary embodiment, the above-mentioned comparison module also includes: a recombination unit, which is used to reorganize the preprocessed second sample set to obtain a third sample set; input the third sample set into the first target neural network for feature extraction to obtain a fourth feature; and determine the similarity loss function corresponding to the first target neural network processing the second sample set based on the difference between the fourth feature and the first feature.
[0021] In an exemplary embodiment, the above-mentioned comparison module also includes: a training unit, which is used to train the training neural network using a preset fourth sample set to obtain a first target neural network, wherein the target loss function used in the process of training the training neural network is a loss function determined based on the first loss function and / or the similarity loss function, and the first loss function is obtained by weighted summing multiple basic loss functions corresponding to the fourth sample set using a first preset weight coefficient, and the multiple basic loss functions include at least one of the following: a cross entropy loss function, a triplet loss function, and a circle loss function; the similarity loss function is obtained by weighted summing the self-similarity loss function and the mutual similarity loss function corresponding to the fourth sample set using a second preset weight coefficient, and the self-similarity loss function is used to indicate the mean square error of the feature differences before and after the fourth sample set is reorganized, and the mutual similarity loss function is used to indicate the correlation coefficient matrix corresponding to the features before and after the fourth sample set is reorganized.
[0022] According to another aspect of the embodiments of the present invention, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute a method in any one of the method embodiments when running.
[0023] According to another aspect of an embodiment of the present invention, there is provided an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute a method in any one of the above method embodiments through the computer program.
[0024] In an embodiment of the present invention, an adjacency matrix of a first sample set is determined, wherein the first sample set includes a plurality of feature elements and samples corresponding to the feature elements; if it is determined according to the adjacency matrix that a first feature element and a second feature element having a connection relationship exist in the first sample set, the similarity between the first feature element and the second feature element is determined, wherein the connection relationship is used to indicate that the sample corresponding to the first feature element and the sample corresponding to the second feature element have the same target sample; the feature elements in the first sample set are rearranged according to the similarity, thereby achieving the purpose of rearranging the retrieval results according to the element features corresponding to the retrieval results, thereby achieving the purpose of improving the effectiveness of the retrieval results, thereby achieving the technical effect of improving the processing efficiency of image search, thereby solving the technical problems in the prior art that the implementation method of image search retrieval results is complex, the retrieval results are interfered, and the effectiveness is low. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0026] Figure 1 is a hardware structure block diagram of a computer terminal according to a feature rearrangement method according to an embodiment of the present invention;
[0027] Figure 2 is a flow chart of a feature rearrangement method according to an embodiment of the present invention;
[0028] Figure 3 A schematic diagram of a process of searching for images by image according to an optional embodiment of the present invention;
[0029] Figure 4 A process schematic diagram of a neural network training method based on self-similarity constraints according to an optional embodiment of the present invention;
[0030] Figure 5 A schematic diagram of updating parameters of a feature extraction neural network according to an optional embodiment of the present invention;
[0031] Figure 6 is a schematic diagram of similarity association between samples according to an optional embodiment of the present invention;
[0032] Figure 7 It is a structural schematic diagram of a feature rearrangement device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0033] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0035] The method embodiments provided in the embodiments of the present application can be executed in a computer terminal, a mobile terminal or a similar computing device. Taking running on a computer terminal as an example, Figure 1 FIG. 1 is a hardware structure block diagram of a computer terminal of a feature rearrangement method according to an embodiment of the present invention. Figure 1 As shown, the computer terminal 10 may include one or more ( Figure 1 Only one is shown in the figure) a processor 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Optionally, the computer terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above-mentioned computer terminal. Figure 1 More or fewer components as shown, or with Figure 1 Equivalent functions or comparisons shown Figure 1 A different configuration with more features is shown.
[0036] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the feature rearrangement method in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, to implement the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0037] The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0038] According to one aspect of an embodiment of the present invention, a feature rearrangement method is provided, which is applied to the above-mentioned computer terminal, such as Figure 2 As shown, the feature rearrangement method includes:
[0039] Step S202, determining an adjacency matrix of a first sample set, wherein the first sample set includes a plurality of characteristic elements and samples corresponding to the characteristic elements;
[0040] Step S204: if it is determined according to the adjacency matrix that a first characteristic element and a second characteristic element having a connection relationship exist in the first sample set, then the similarity between the first characteristic element and the second characteristic element is determined, wherein the connection relationship is used to indicate that the sample corresponding to the first characteristic element and the sample corresponding to the second characteristic element have the same target sample;
[0041] Step S206: rearrange the characteristic elements in the first sample set according to the similarity.
[0042] Through the above steps, the adjacency matrix of the first sample set is determined, wherein the first sample set includes multiple feature elements and samples corresponding to the feature elements; if it is determined according to the adjacency matrix that there are a first feature element and a second feature element with a connection relationship in the first sample set, the similarity between the first feature element and the second feature element is determined, wherein the connection relationship is used to indicate that the sample corresponding to the first feature element and the sample corresponding to the second feature element have the same target sample; the feature elements in the first sample set are rearranged according to the similarity, thereby achieving the purpose of rearranging the retrieval results according to the element features corresponding to the retrieval results, thereby achieving the purpose of improving the effectiveness of the retrieval results, thereby achieving the technical effect of improving the processing efficiency of image search, thereby solving the technical problems in the prior art that the implementation method of image search retrieval results is complex, the retrieval results are interfered, and the effectiveness is low.
[0043] That is to say, there may be multiple feature elements in the first sample set at the same time, and different feature elements correspond to different numbers of samples. The condition for determining that there is an association between feature elements is that there are consistent samples between the feature elements. For example, the first sample set may be a set of multiple suspected images queried from an image database for a photo of a vehicle passing by a traffic intersection to be queried, and the photo of the vehicle passing by has feature elements such as vehicle, logo, location, time, etc. Each individual feature element corresponds to multiple suspected images that meet the requirements, and then by comparing the similarity between the feature elements, the number of occurrences of similar suspected images is reduced, thereby improving the accuracy of the search results of image search, and further rearrange the multiple suspected images that meet the requirements according to the similarity confirmation results, so that the interference in the retrieval results is greatly reduced.
[0044] In an exemplary embodiment, before determining the adjacency matrix of the first sample set, the method further includes: inputting the second sample set to be retrieved and queried into the first target neural network for feature extraction to obtain a first feature of the second sample set; comparing the first feature with a preset second feature in the database to obtain a first set of similar third features, wherein the third feature is used to indicate a first feature existing in the preset second feature whose similarity meets a preset threshold; performing association determination on each feature element in the first set in the second sample set to be retrieved and queried to determine a target number of samples corresponding to each feature element to obtain a second set; merging the first set with the second set to obtain the first sample set.
[0045] It should be noted that the second sample set to be retrieved can be a video or an image. After being processed by the first target neural network, the features obtained are features that can characterize the input video or image information, which are used for subsequent feature comparison.
[0046] For example, for any second sample set q to be searched, the first feature of the second sample set q is extracted through a feature extraction network (equivalent to the first target neural network implemented in the present invention), and the first feature extracted from the second sample set q is compared with the second features of all samples in the database g to find the most similar samples, forming a set ,gather indicates that q is found in database g The most similar samples, Sort by similarity. Each element in , find the closest one in the second sample set q samples, and get the set ,in Sort by similarity and finally In order and Merge: When merging, the sets are taken as the union, that is, ,in , and finally get a set of K samples .
[0047] In an exemplary embodiment, after rearranging the feature elements in the first sample set according to the similarity, the method further includes: constructing input features of the graph convolutional network according to the feature elements in the first sample set, wherein the feature function corresponding to the input feature is: ,in, represents the sample features corresponding to the first sample set extracted using the first target neural network, The standard features extracted by the feature extraction network same as the first target neural network are used for the feature elements in the first sample set. is a cosine distance function, which is used to indicate the similarity between the sample features of the sample set to be queried and the standard features; the target graph convolutional network is iteratively updated to determine a second target neural network for rearranging feature elements.
[0048] For example, for the set The element i in constructs its input feature, and its input feature is expressed as , where Fq represents the features of the query sample q extracted using the feature extraction network; Fi represents the features extracted by the element i in the set using the same feature extraction network; Dcos(Fq, Fi) is the cosine distance function, which is used to indicate the similarity between the features of sample q and the elements i in the set. To construct the adjacency matrix A of the graph convolution, where A is a symmetric matrix, the construction logic of A is as follows:
[0049] ;
[0050] in, Representing a collection Whether there is a connection between elements i and j in the set Find element i in The most similar samples ( ), if element j is i One of the samples, ,on the contrary = 0. It represents the similarity between element i and element j. Finally, a graph convolutional network is constructed based on the adjacency matrix A and the input features, and then a second target neural network is obtained, which can be used to rearrange samples by similarity.
[0051] It should be noted that the graph convolutional network (GCN) is a neural network that processes graph data in deep learning, and is mainly used to process graph structured data.
[0052] In an exemplary embodiment, after iteratively updating the target graph convolutional network to determine a second target neural network for rearranging feature elements, the method further includes: setting a training label for the second target neural network, wherein the training label is used to indicate a similarity classification result of the second target neural network for the first sample set; when the training label is a first label, determining that the sample feature corresponding to the training label and the standard feature corresponding to the feature element belong to the same category or the same object identifier, and the first label is used to characterize that similar feature elements exist in the similarity classification result of the first sample set; when the training label is a second label, determining that the sample feature corresponding to the training label and the standard feature corresponding to the feature element do not belong to the same category or the same object identifier; and the second label is used to characterize that similar feature elements do not exist in the similarity classification result of the first sample set.
[0053] In short, in order to ensure that the second target neural network can better perform feature rearrangement classification, the second target neural network is updated in an iterative manner to obtain the second target neural network used for similarity rearrangement network. Then, in the training process of the second target neural network, the similarity rearrangement label corresponding to the second target neural network is set according to the confirmed similarity. It is expressed as ,in:
[0054] ;
[0055] when When it is 1, it is equivalent to the first label in the embodiment of the present invention. When it is 0, it is equivalent to the second label in the embodiment of the present invention, and then the training samples used to train the second target neural network are labeled, so that the recognition of the trained second target neural network is more accurate, and then the first feature corresponding to the first target neural network and the preset second feature in the database are compared through the trained second target neural network. The similarity is rearranged to obtain the final image sample result, thereby improving the effectiveness of image retrieval.
[0056] In an exemplary embodiment, a second sample set to be retrieved is obtained and input into the first target neural network for feature extraction to obtain a first feature of the second sample set, including: determining object identifiers contained in the second sample set, wherein the object identifier is used to indicate a confirmed unique identifier for attribute information of each object; classifying the second sample set according to the object identifier, and preprocessing different samples under the same object identifier, wherein the preprocessing includes at least one of the following: random erasing of data in the sample, random rotation of data in the sample, random flipping of data in the sample, and random graying of data in the sample; and inputting the preprocessed second sample set into the first target neural network in sequence according to the classification result for feature extraction to obtain the first feature of the second sample set.
[0057] In an exemplary embodiment, the method further includes: reorganizing the preprocessed second sample set to obtain a third sample set; inputting the third sample set into the first target neural network for feature extraction to obtain a fourth feature; and determining the similarity loss function corresponding to the first target neural network processing the second sample set based on the difference between the fourth feature and the first feature.
[0058] It is understandable that since there are certain differences in the features obtained between the sample sets before and after the reorganization after being processed by the first target neural network, the similarity loss function of the first target neural network for feature extraction can be determined based on these differences, and then the loss function of the first target neural network can be optimized, so that the feature extraction through the first target neural network is more accurate, the feature representation ability is enhanced, and the retrieval effect is improved.
[0059] In an exemplary embodiment, a second sample set of queries to be retrieved is obtained and input into the first target neural network for feature extraction. Before obtaining the first feature of the second sample set, the method further includes: training the training neural network using a preset fourth sample set to obtain the first target neural network, wherein the target loss function used in the process of training the training neural network is a loss function determined according to the first loss function and / or the similarity loss function, the first loss function is obtained by weighted summing a plurality of basic loss functions corresponding to the fourth sample set using a first preset weight coefficient, the plurality of basic loss functions including at least one of the following: a cross entropy loss function, a triplet loss function, and a circle loss function; the similarity loss function is obtained by weighted summing a self-similarity loss function and a mutual similarity loss function corresponding to the fourth sample set using a second preset weight coefficient, the self-similarity loss function is used to indicate the mean square error of the feature difference before and after the fourth sample set is reorganized, and the mutual similarity loss function is used to indicate the correlation coefficient matrix corresponding to the features before and after the fourth sample set is reorganized.
[0060] For example, the first loss function of image search corresponding to the first target neural network is Contains cross entropy loss function , triplet loss function And Circle loss function ,Right now: ,in , as well as is the weight coefficient. Similar loss function Mainly includes self-similarity loss function And the mutual similarity loss function ,Right now: ,in and is the weight coefficient. Among them, the self-similarity loss function And the cross-correlation loss function The definition is as follows: Input data D passes through model A to obtain feature F, where , the input data D is reorganized to obtain , Through neural network After getting the features ,in ,but: ; is the mean square error corresponding to the features of the input data D (equivalent to the fourth sample set in the embodiment of the present invention) before and after reorganization;
[0061] ; is the correlation coefficient matrix corresponding to the features before and after the input data D is reorganized; The correlation matrix is calculated as follows:
[0062] ;
[0063] ;
[0064] Then we can determine the loss function of the entire first target neural network training process as: .
[0065] In order to better understand the technical solutions of the embodiments of the present invention and the optional embodiments, the process of the above-mentioned feature rearrangement method is explained below with reference to examples, but it is not intended to limit the technical solutions of the embodiments of the present invention.
[0066] An optional embodiment of the present invention proposes a method for searching images by image based on similarity constraints, such as Figure 3 FIG. 1 is a flow chart of searching for images by image in an optional embodiment of the present invention, which specifically includes the following steps:
[0067] Step 1: Extract features from the input image or video through a feature extraction network;
[0068] Step 2: Perform feature comparison on the extracted input image or video features in combination with the stored content in the database to obtain a similarity ranking result.
[0069] Step 3: Rearrange the features based on the results of feature comparison.
[0070] As an optional implementation, it can be understood that the feature extraction network is a neural network, and its main function is to extract features that can characterize the input video or image information. The feature is used for subsequent feature comparison. Since the feature extraction network is a neural network, the optional embodiment of the present invention proposes a neural network training method based on self-similarity constraints (equivalent to the first target neural network in the embodiment of the present invention), and the specific scheme is as follows:
[0071] Training sample collection: The samples for image search are images or videos with different identification IDs. Each ID collects multiple sets of image or video data taken by different cameras at different angles and under different lighting conditions.
[0072] Model building and training, according to the input sample type is video or image, the corresponding neural network is designed, for the convenience of description, it is represented by neural network A. After completing the network design, the network training can be carried out. First, during the training process, each network iteration randomly extracts N ids without replacement, and each id randomly selects M groups of videos or images as input data D, which can be expressed as ,in Represents the M samples of the i-th id.
[0073] Optionally, the network forward process is as follows Figure 4 As shown in the figure, it is divided into two branches. In the upper branch, the input data passes through the neural network A, and then the image search loss is calculated according to the label id information; in the lower branch, the input data is first reorganized, and the reorganized data passes through the neural network, and the similarity loss is calculated together with the output of the upper branch.
[0074] The data reorganization rules in the above process are as follows:
[0075] 1) Randomly shuffle the input data within the id: , It can be expressed as: ,in It indicates the arrangement of the M samples in the first id after random shuffling.
[0076] 2) Use random erasing, random rotation, random flipping, random grayscale and other operations to transform the samples.
[0077] It should be noted that neural networks The structural parameters of the neural network A are exactly the same as the image search loss mentioned in the above process. Contains cross entropy loss function , triplet loss function And Circle loss function ,Right now: ,in , as well as The optional embodiment of the present invention combines the three loss functions to improve the difference of features extracted by the image search feature extraction model and improve the generalization performance of the model.
[0078] Optionally, the optional embodiment of the present invention also proposes a similar loss , to further improve the representation ability of features, similarity loss mainly includes self-similarity loss and mutual similarity loss ,Right now: ,in and is the weight coefficient. Among them, the self-similarity loss With cross-correlation loss The definition is as follows:
[0079] Input data D passes through model A to obtain feature F, where , the input data D is reorganized to obtain , Through neural network After getting the features ,in ,but:
[0080] ;
[0081] ;
[0082] in, The correlation matrix is calculated as follows:
[0083] ;
[0084] ;
[0085] The loss function during the entire network training process for: After obtaining the loss function, stochastic gradient descent SGD is used for optimization, and the model parameters corresponding to neural network A are iteratively updated (that is, the loss function of neural network A is updated in combination with a similar loss function) until the loss of the model corresponding to neural network A converges.
[0086] Optional, Figure 5 Schematic diagram of parameter updating of a feature extraction neural network according to an optional embodiment of the present invention. In the parameter updating stage, the neural network It does not participate in parameter updating, that is, it only passes through the model corresponding to neural network A during back propagation, and only updates the parameters of the model corresponding to neural network A. Then, when training again, the neural network synchronizes the updated parameter information of the model corresponding to neural network A.
[0087] As an optional implementation, feature rearrangement is to extract video or image features through the feature extraction network after obtaining the feature extraction network, and compare them with the features previously extracted by the feature extraction network in the database to obtain the similarity ranking result. This result is usually displayed as the final retrieval result of image search, but in actual application, due to the differences in shooting angle, lighting, etc. between the samples in the database and the samples to be queried, the similarity between some samples in the database and the samples to be queried may be relatively low, such as Figure 6 As shown, it is a schematic diagram of the similarity association between samples. The sample to be checked has a high similarity with similar samples B and C in the database, but has a low similarity with similar sample A. Usually, sample A is not considered to be of the same type or with the same ID as the sample to be checked, but A may have a high similarity with samples B and C, and the results are rearranged based on this information. Based on the above principle, an optional embodiment of the present invention proposes a similarity rearrangement network to improve this problem.
[0088] Optionally, a similarity rearrangement network (equivalent to the second target neural network in the embodiment of the present invention) is constructed as follows:
[0089] Step 1: Training sample construction: For any query sample q to be retrieved, extract features in the database g through the above feature extraction network and compare the features extracted from q with the features of all samples in g to find the most similar samples, forming a set , indicating that q is found in g The most similar samples, Sort by similarity. Each element in , find the closest one in q samples, and get the set ,in Sort by similarity. In order and Merge: When merging, the sets are taken as the union, that is, ,in Finally, we get a set of K samples , if the union of the first n samples is greater than or equal to K, then the following ones. After obtaining the samples, a graph convolutional network is constructed. For each sample To construct the adjacency matrix A of the graph convolution, where A is The construction logic of the symmetric matrix A is as follows:
[0090] ;
[0091] in, Representing a collection Whether there is a connection between elements i and j in the set Find element i in The most similar samples ( ), if element j is i One of the samples, ,on the contrary = 0. Represents the similarity between element i and element j. In an optional embodiment of the present invention, cosine distance is used to measure the similarity between elements.
[0092] Step 2: After constructing the adjacency matrix A, it is necessary to construct the input features of the graph convolutional network. The element i in constructs its input feature, and its input feature is expressed as , where Fq represents the features of the query sample q extracted using the feature extraction network; Fi represents the features extracted by the element i in the set using the same feature extraction network; Dcos(Fq, Fi) is the cosine distance function, which is used to indicate the similarity between the features of the sample q and the element i in the set. After obtaining the adjacency matrix and the input features, the graph convolution network can be constructed.
[0093] Step 3: Update the network iteratively to obtain the final similarity rearrangement network. The labels in the training process are represented as ,in:
[0094] ;
[0095] After obtaining the similarity rearrangement network, the preliminary comparison results are rearranged by similarity through this network to obtain the final result, thereby improving the effectiveness of the retrieval.
[0096] Through the above embodiments, by proposing a similarity loss function and a feature rearrangement network for training a new feature extraction network, through the similarity loss function, on the basis of the image search loss function, the relationship between samples of the same type and the same ID is emphasized, the relationship is added to the loss constraint, the representation ability of the feature is enhanced, the retrieval effect is improved, the query image is processed for feature, and the self-similarity constraint is used. Without increasing the difficulty of network design, the correlation between samples of the same ID is used to improve the model feature extraction ability, and the correlation relationship between samples is used to rearrange the retrieval results, and by constructing a similarity connection graph structure, a neural network is used to mine potential low-similarity samples of the same type and the same ID, thereby improving the effectiveness of the retrieval results and obtaining better image search retrieval results. This makes the application scenarios more extensive.
[0097] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.
[0098] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0099] According to another aspect of the embodiments of the present invention, there is also provided a feature rearrangement device for implementing the feature rearrangement method described above. The device is used to implement the above embodiments and preferred implementation modes, and will not be described again. As used below, the term "module" may be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware, is also possible and conceivable. Figure 7 As shown, the device comprises:
[0100] A first determination module 72, configured to determine an adjacency matrix of a first sample set, wherein the first sample set includes a plurality of characteristic elements and samples corresponding to the characteristic elements;
[0101] A second determination module 74 is configured to determine the similarity between the first characteristic element and the second characteristic element if it is determined according to the adjacency matrix that there are a first characteristic element and a second characteristic element having a connection relationship in the first sample set, wherein the connection relationship is used to indicate that the sample corresponding to the first characteristic element and the sample corresponding to the second characteristic element have the same target sample;
[0102] The rearrangement module 76 is configured to rearrange the feature elements in the first sample set according to the similarity.
[0103] Through the above-mentioned device, an adjacency matrix of a first sample set is determined, wherein the first sample set includes a plurality of feature elements and samples corresponding to the feature elements; if it is determined according to the adjacency matrix that a first feature element and a second feature element having a connection relationship exist in the first sample set, the similarity between the first feature element and the second feature element is determined, wherein the connection relationship is used to indicate that the sample corresponding to the first feature element and the sample corresponding to the second feature element have the same target sample; the feature elements in the first sample set are rearranged according to the similarity, thereby achieving the purpose of rearranging the retrieval results according to the element features corresponding to the retrieval results, thereby achieving the purpose of improving the effectiveness of the retrieval results, thereby achieving the technical effect of improving the processing efficiency of image search, thereby solving the technical problems in the prior art that the implementation method of image search retrieval results is complex, the retrieval results are interfered, and the effectiveness is low.
[0104] That is to say, there may be multiple feature elements in the first sample set at the same time, and different feature elements correspond to different numbers of samples. The condition for determining that there is an association between feature elements is that there are consistent samples between the feature elements. For example, the first sample set may be a set of multiple suspected images queried from an image database for a photo of a vehicle passing by a traffic intersection to be queried, and the photo of the vehicle passing by has feature elements such as vehicle, logo, location, time, etc. Each individual feature element corresponds to multiple suspected images that meet the requirements, and then by comparing the similarity between the feature elements, the number of occurrences of similar suspected images is reduced, thereby improving the accuracy of the search results of image search, and further rearrange the multiple suspected images that meet the requirements according to the similarity confirmation results, so that the interference in the retrieval results is greatly reduced.
[0105] In an exemplary embodiment, the above-mentioned device also includes: a comparison module, which inputs the second sample set to be retrieved and queried into the first target neural network for feature extraction to obtain the first feature of the second sample set; compares the first feature with the preset second feature in the database to obtain a first set of similar third features, wherein the third feature is used to indicate the first feature existing in the preset second feature whose similarity meets the preset threshold; associates each feature element in the first set in the second sample set to be retrieved and queried, and determines the target number of samples corresponding to each feature element to obtain the second set; merges the first set with the second set to obtain the first sample set.
[0106] It should be noted that the second sample set to be retrieved can be a video or an image. After being processed by the first target neural network, the features obtained are features that can characterize the input video or image information, which are used for subsequent feature comparison.
[0107] For example, for any second sample set q to be searched, the first feature of the second sample set q is extracted through a feature extraction network (equivalent to the first target neural network implemented in the present invention), and the first feature extracted from the second sample set q is compared with the second features of all samples in the database g to find the most similar samples, forming a set ,gather indicates that q is found in database g The most similar samples, Sort by similarity. Each element in , find the closest one in the second sample set q samples, and get the set ,in Sort by similarity and finally In order and Merge: When merging, the sets are taken as the union, that is, ,in , and finally get a set of K samples .
[0108] In an exemplary embodiment, the apparatus further includes: a construction module, configured to construct an input feature of the graph convolutional network according to the feature elements in the first sample set, wherein the feature function corresponding to the input feature is: ,in, represents the sample features corresponding to the first sample set extracted using the first target neural network, The standard features extracted by the feature extraction network same as the first target neural network are used for the feature elements in the first sample set. is a cosine distance function, which is used to indicate the similarity between the sample features of the sample set to be queried and the standard features; a target graph convolutional network is established according to the input features and the adjacency matrix; the target graph convolutional network is iteratively updated to determine a second target neural network for rearranging feature elements.
[0109] For example, for the set The element i in constructs its input feature, and its input feature is expressed as , where Fq represents the features of the query sample q extracted using the feature extraction network; Fi represents the features extracted by the element i in the set using the same feature extraction network; Dcos(Fq, Fi) is the cosine distance function, which is used to indicate the similarity between the features of sample q and the elements i in the set. To construct the adjacency matrix A of the graph convolution, where A is a symmetric matrix, the construction logic of A is as follows:
[0110] ;
[0111] in, Representing a collection Whether there is a connection between elements i and j in the set Find element i in The most similar samples ( ), if element j is i One of the samples, ,on the contrary = 0. It represents the similarity between element i and element j. Finally, a graph convolutional network is constructed based on the adjacency matrix A and the input features, and then a second target neural network is obtained, which can be used to rearrange samples by similarity.
[0112] It should be noted that the graph convolutional network (GCN) is a neural network that processes graph data in deep learning, and is mainly used to process graph structured data.
[0113] In an exemplary embodiment, the above-mentioned construction module also includes: a label unit, used to set a training label for the second target neural network, wherein the training label is used to indicate the similarity classification result of the second target neural network for the first sample set; when the training label is a first label, it is determined that the sample feature corresponding to the training label and the standard feature corresponding to the feature element belong to the same category or the same object identification, and the first label is used to characterize that there are similar feature elements in the similarity classification result of the first sample set; when the training label is a second label, it is determined that the sample feature corresponding to the training label and the standard feature corresponding to the feature element do not belong to the same category or the same object identification; the second label is used to characterize that there are no similar feature elements in the similarity classification result of the first sample set.
[0114] In short, in order to ensure that the second target neural network can better perform feature rearrangement classification, the second target neural network is updated in an iterative manner to obtain the second target neural network used for similarity rearrangement network. Then, in the training process of the second target neural network, the similarity rearrangement label corresponding to the second target neural network is set according to the confirmed similarity. It is expressed as in:
[0115] ;
[0116] when When it is 1, it is equivalent to the first label in the embodiment of the present invention. When it is 0, it is equivalent to the second label in the embodiment of the present invention, and then the training samples used to train the second target neural network are labeled, so that the recognition of the trained second target neural network is more accurate, and then the first feature corresponding to the first target neural network and the preset second feature in the database are compared through the trained second target neural network. The similarity is rearranged to obtain the final image sample result, thereby improving the effectiveness of image retrieval.
[0117] In an exemplary embodiment, the above-mentioned comparison module is also used to determine the object identifiers contained in the second sample set, wherein the object identifier is used to indicate the confirmed unique identifier of the attribute information of each object; the second sample set is classified according to the object identifier, and different samples under the same object identifier are preprocessed, wherein the preprocessing includes at least one of the following: random erasing of data in the sample, random rotation of data in the sample, random flipping of data in the sample, and random graying of data in the sample; according to the classification result, the preprocessed second sample set is sequentially input into the first target neural network for feature extraction to obtain the first feature of the second sample set.
[0118] In an exemplary embodiment, the above-mentioned comparison module also includes: a recombination unit, which is used to reorganize the preprocessed second sample set to obtain a third sample set; input the third sample set into the first target neural network for feature extraction to obtain a fourth feature; and determine the similarity loss function corresponding to the first target neural network processing the second sample set based on the difference between the fourth feature and the first feature.
[0119] It is understandable that since there are certain differences in the features obtained between the sample sets before and after the reorganization after being processed by the first target neural network, the similarity loss function of the first target neural network for feature extraction can be determined based on these differences, and then the loss function of the first target neural network can be optimized, so that the feature extraction through the first target neural network is more accurate, the feature representation ability is enhanced, and the retrieval effect is improved.
[0120] In an exemplary embodiment, the above-mentioned comparison module also includes: a training unit, which is used to train the training neural network using a preset fourth sample set to obtain a first target neural network, wherein the target loss function used in the process of training the training neural network is a loss function determined based on the first loss function and / or the similarity loss function, and the first loss function is obtained by weighted summing multiple basic loss functions corresponding to the fourth sample set using a first preset weight coefficient, and the multiple basic loss functions include at least one of the following: a cross entropy loss function, a triplet loss function, and a circle loss function; the similarity loss function is obtained by weighted summing the self-similarity loss function and the mutual similarity loss function corresponding to the fourth sample set using a second preset weight coefficient, and the self-similarity loss function is used to indicate the mean square error of the feature differences before and after the fourth sample set is reorganized, and the mutual similarity loss function is used to indicate the correlation coefficient matrix corresponding to the features before and after the fourth sample set is reorganized.
[0121] For example, the first loss function of image search corresponding to the first target neural network is Contains cross entropy loss function , triplet loss function And Circle loss function ,Right now: ,in , as well as is the weight coefficient. Similar loss function Mainly includes self-similarity loss function And the mutual similarity loss function ,Right now: ,in and is the weight coefficient. Among them, the self-similarity loss function And the cross-correlation loss function The definition is as follows: Input data D passes through model A to obtain feature F, where , the input data D is reorganized to obtain , Through neural network After getting the features ,in ,but:
[0122] ; is the mean square error corresponding to the features of the input data D (equivalent to the fourth sample set in the embodiment of the present invention) before and after reorganization;
[0123] ; is the correlation coefficient matrix corresponding to the features before and after the input data D is reorganized; The correlation matrix is calculated as follows:
[0124] ;
[0125] ;
[0126] Then the loss function of the entire first target neural network training process can be determined for: .
[0127] It should be noted that the above modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or the above modules are located in different processors in any combination.
[0128] An embodiment of the present invention further provides a storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above method embodiments when running.
[0129] Optionally, in this embodiment, the storage medium may be configured to store a computer program for performing the following steps:
[0130] S1. Determine an adjacency matrix of a first sample set, wherein the first sample set includes a plurality of characteristic elements and samples corresponding to the characteristic elements;
[0131] S2. If it is determined according to the adjacency matrix that a first characteristic element and a second characteristic element having a connection relationship exist in the first sample set, then the similarity between the first characteristic element and the second characteristic element is determined, wherein the connection relationship is used to indicate that the sample corresponding to the first characteristic element and the sample corresponding to the second characteristic element have the same target sample;
[0132] S3. Rearrange the characteristic elements in the first sample set according to the similarity.
[0133] Optionally, in this embodiment, the above-mentioned storage medium may include but is not limited to: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and other media that can store computer programs.
[0134] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.
[0135] An embodiment of the present invention further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0136] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0137] Optionally, in this embodiment, the processor may be configured to perform the following steps through a computer program:
[0138] S1. Determine an adjacency matrix of a first sample set, wherein the first sample set includes a plurality of characteristic elements and samples corresponding to the characteristic elements;
[0139] S2. If it is determined according to the adjacency matrix that a first characteristic element and a second characteristic element having a connection relationship exist in the first sample set, then the similarity between the first characteristic element and the second characteristic element is determined, wherein the connection relationship is used to indicate that the sample corresponding to the first characteristic element and the sample corresponding to the second characteristic element have the same target sample;
[0140] S3. Rearrange the characteristic elements in the first sample set according to the similarity.
[0141] Optionally, in this embodiment, a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing hardware related to the terminal device through a program, and the program may be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0142] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0143] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling one or more computer devices (which can be personal computers, servers or network devices, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention.
[0144] In the above embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0145] In the several embodiments provided in the present application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0146] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0147] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0148] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A feature rearrangement method, characterized in that: Applications to computer vision including graph convolutional networks and / or video image processing, including: Determine an adjacency matrix of a first sample set, wherein the first sample set includes a plurality of characteristic elements and samples corresponding to the characteristic elements; If it is determined according to the adjacency matrix that a first characteristic element and a second characteristic element having a connection relationship exist in the first sample set, then determining the similarity between the first characteristic element and the second characteristic element, wherein the connection relationship is used to indicate that the sample corresponding to the first characteristic element and the sample corresponding to the second characteristic element have the same target sample; Rearranging the characteristic elements in the first sample set according to the similarity; The plurality of characteristic elements exist on a target image, the sample is a suspected image queried from an image database according to the target image, and the first sample set is a set consisting of a plurality of suspected images; The adjacency matrix is a K×K symmetric matrix composed of αij*βij, αij represents the connection relationship between the first characteristic element i and the second characteristic element j, βij represents the similarity between the first characteristic element i and the second characteristic element j; the value of αij is 0 indicating that there is no connection relationship or 1 indicating that there is a connection relationship; Wherein, before determining the adjacency matrix of the first sample set, the method also includes: inputting the second sample set to be retrieved and queried into the first target neural network for feature extraction to obtain the first feature of the second sample set; comparing the first feature with the preset second feature in the database to obtain a first set of similar third features, wherein the third feature is used to indicate the first feature existing in the preset second feature whose similarity meets a preset threshold; performing association determination on each feature element in the first set in the second sample set to be retrieved and queried, and determining the target number of samples corresponding to each feature element to obtain a second set; merging the first set with the second set to obtain the first sample set.
2. The method according to claim 1, characterized in that After rearranging the characteristic elements in the first sample set according to the similarity, the method further includes: The input features of the graph convolutional network are constructed according to the feature elements in the first sample set, wherein the feature function corresponding to the input features is: , wherein the represents the sample features corresponding to the first sample set extracted using the first target neural network, The standard features extracted by the feature extraction network same as the first target neural network are used for the feature elements in the first sample set. is a cosine distance function, which is used to indicate the similarity between the sample features of the sample set to be queried and the standard features; Establishing a target graph convolutional network according to the input features and the adjacency matrix; The target graph convolutional network is iteratively updated to determine a second target neural network for rearranging feature elements.
3. The method according to claim 2, characterized in that After iteratively updating the target graph convolutional network to determine a second target neural network for rearranging feature elements, the method further includes: Setting a training label of the second target neural network, wherein the training label is used to indicate a similarity classification result of the second target neural network for the first sample set; When the training label is a first label, it is determined that the sample feature corresponding to the training label and the standard feature corresponding to the feature element belong to the same category or the same object identifier, and the first label is used to characterize that similar feature elements exist in the similarity classification results of the first sample set; When the training label is a second label, it is determined that the sample feature corresponding to the training label and the standard feature corresponding to the feature element do not belong to the same category or the same object identifier; the second label is used to characterize that there are no similar feature elements in the similarity classification results of the first sample set.
4. The method according to claim 1, characterized in that: Inputting the second sample set to be searched into the first target neural network for feature extraction to obtain the first feature of the second sample set includes: Determine the object identifiers included in the second sample set, wherein the object identifier is used to indicate a unique identifier for confirming attribute information of each object; Classifying the second sample set according to the object identifier, and preprocessing different samples under the same object identifier, wherein the preprocessing includes at least one of the following: random erasing of data in the sample, random rotation of data in the sample, random flipping of data in the sample, and random graying of data in the sample; According to the classification result, the preprocessed second sample set is sequentially input into the first target neural network for feature extraction to obtain the first feature of the second sample set.
5. The method according to claim 4, characterized in that The method further comprises: Recombining the preprocessed second sample set to obtain a third sample set; Inputting the third sample set into the first target neural network for feature extraction to obtain a fourth feature; Determine a similarity loss function corresponding to processing the second sample set by the first target neural network based on the difference between the fourth feature and the first feature.
6. The method according to claim 1, characterized in that Before obtaining a second sample set of queries to be searched and inputting it into the first target neural network for feature extraction and obtaining a first feature of the second sample set, the method further includes: The training neural network is trained using a preset fourth sample set to obtain the first target neural network, wherein the target loss function used in the process of training the training neural network is a loss function determined based on the first loss function and / or the similarity loss function, the first loss function is obtained by weighted summing a plurality of basic loss functions corresponding to the fourth sample set using a first preset weight coefficient, the plurality of basic loss functions include at least one of the following: a cross entropy loss function, a triplet loss function, and a circle loss function; the similarity loss function is obtained by weighted summing a self-similarity loss function and a mutual similarity loss function corresponding to the fourth sample set using a second preset weight coefficient, the self-similarity loss function is used to indicate the mean square error of the feature difference before and after the fourth sample set is reorganized, and the mutual similarity loss function is used to indicate the correlation coefficient matrix corresponding to the features before and after the fourth sample set is reorganized.
7. A feature rearrangement device, characterized in that: Applications to computer vision and / or video image processing involving graph convolutional networks, including: A first determination module, configured to determine an adjacency matrix of a first sample set, wherein the first sample set includes a plurality of characteristic elements and samples corresponding to the characteristic elements; a second determination module, configured to determine the similarity between the first characteristic element and the second characteristic element if it is determined according to the adjacency matrix that there are a first characteristic element and a second characteristic element having a connection relationship in the first sample set, wherein the connection relationship is used to indicate that the sample corresponding to the first characteristic element and the sample corresponding to the second characteristic element have the same target sample; a rearrangement module, configured to rearrange the feature elements in the first sample set according to the similarity; The plurality of characteristic elements exist on a target image, the sample is a suspected image queried from an image database according to the target image, and the first sample set is a set consisting of a plurality of suspected images; The adjacency matrix is a K×K symmetric matrix composed of αij*βij, αij represents the connection relationship between the first characteristic element i and the second characteristic element j, βij represents the similarity between the first characteristic element i and the second characteristic element j; the value of αij is 0 indicating that there is no connection relationship or 1 indicating that there is a connection relationship; The device also includes: a comparison module, which is used to input the second sample set to be retrieved and queried into the first target neural network for feature extraction to obtain the first feature of the second sample set; compare the first feature with the preset second feature in the database to obtain a first set of similar third features, wherein the third feature is used to indicate the first feature existing in the preset second feature whose similarity meets a preset threshold; associate each feature element in the first set with the second sample set to be retrieved and queried, and determine the target number of samples corresponding to each feature element to obtain a second set; merge the first set with the second set to obtain a first sample set.
8. A computer-readable storage medium, the computer-readable storage medium comprising a stored program, wherein: When the program is executed, the method described in any one of claims 1 to 6 is executed.
9. An electronic device, comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 6 through the computer program.
Citation Information
Patent Citations
Target recognition model training method, device and equipment and storage medium
CN111626119A