Method, device, equipment and medium for searching and sorting images collected across devices
By utilizing similarity retrieval and graph propagation processing in cross-device image retrieval and sorting to adjust image similarity values, the problem of low accuracy in cross-device image sorting is solved, and higher sorting accuracy is achieved.
Patent Information
- Application Number
- CN202311817357.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-26
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2043-12-26
AI Technical Summary
In the retrieval and ranking of images collected across devices, the ranking results are less accurate, mainly because factors such as the installation height, angle, and lighting of different devices lead to large differences in images of the same person, resulting in errors in calculating sample similarity.
By performing similarity retrieval from a preset human image set, K retrieval images are obtained, an image sequence is constructed based on similarity sorting, the similarity values between the images are calculated, a first matrix is constructed and the similarity values are adjusted according to the acquisition device information, graph propagation processing is performed, and re-sorting is performed to improve accuracy.
The accuracy of similarity between images is improved, the accuracy of sorting results is enhanced, the human body image features of the same person are ensured to have more similarities, and the error impact between images collected by different devices is reduced.
Smart Images

Figure CN118035478B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a retrieval and sorting method, apparatus, device, and medium for cross-device image acquisition. Background Art
[0002] With the advancement and development of artificial intelligence technology and the growing demands of enterprise management and public safety, human re-identification technology has been widely used in various aspects of social life due to its ability to track, match, and identify target individuals across time and space. It has also been a research hotspot in the field of computer vision in recent years. Re-identification essentially involves retrieving images of a pedestrian from a surveillance camera, sorting the images captured across devices based on their similarity to the pedestrian image. Based on the sorted results, images of the same person are found. However, due to factors such as the height, angle, and lighting of different acquisition devices, the images captured of the same person may vary significantly. Images may be captured from the front, back, side, or even half the body. This can lead to errors in calculating sample similarity, resulting in low ranking accuracy. Therefore, improving the accuracy of cross-device image retrieval and ranking is a pressing issue. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a method, apparatus, device, and medium for searching and ranking images collected across devices, so as to solve the problem of low accuracy of ranking results in searching and ranking images collected across devices.
[0004] In a first aspect, an embodiment of the present invention provides a method for retrieving and ranking images collected across devices, the method comprising:
[0005] Performing similarity retrieval on the acquired target human image from a preset human image set to obtain K retrieval images, and sorting the target human image and the K retrieval images based on similarity to obtain an image sequence, where K is an integer greater than zero;
[0006] calculating, based on original image features of each image in the image sequence, a similarity between any two images in the image sequence to obtain a first similarity value, and constructing a first matrix based on the first similarity value, wherein row numbers and column numbers of elements corresponding to the first similarity value in the first matrix correspond to the sorting numbers of the any two images in the image sequence;
[0007] Determining, based on acquisition device information of the target human image and the K search images, a first image and a second image acquired by the same acquisition device, and setting elements corresponding to the first similarity values between the first image and the second image in the first matrix to preset values to obtain a second matrix;
[0008] Performing graph propagation processing on the original image features of each image according to the first matrix and the second matrix to obtain target image features of each image;
[0009] According to the similarity between the target image feature of the target human image and the target image features of the K search images, the K search images are re-ranked to obtain a ranking result.
[0010] In a second aspect, an embodiment of the present invention provides a retrieval and ranking device for collecting images across devices, the retrieval and ranking device comprising:
[0011] a retrieval module configured to perform similarity retrieval on the acquired target human image from a preset human image set to obtain K retrieval images, and to sort the target human image and the K retrieval images based on similarity to obtain an image sequence, where K is an integer greater than zero;
[0012] a construction module, configured to calculate, based on original image features of each image in the image sequence, a similarity between any two images in the image sequence to obtain a first similarity value, and construct a first matrix based on the first similarity value, wherein the row numbers and column numbers of the elements corresponding to the first similarity value in the first matrix correspond to the sorting numbers of the any two images in the image sequence;
[0013] an obtaining module, configured to determine, based on acquisition device information of the target human image and the K search images, a first image and a second image acquired by the same acquisition device, and set elements corresponding to the first similarity values between the first image and the second image in the first matrix to preset values to obtain a second matrix;
[0014] a graph propagation module, configured to perform graph propagation processing on the original image features of each image according to the first matrix and the second matrix to obtain target image features of each image;
[0015] The sorting module is used to re-sort the K search images according to the similarity between the target image features of the target human image and the target image features of the K search images to obtain a sorting result.
[0016] In a third aspect, an embodiment of the present invention provides a computer device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the retrieval and ranking method as described in the first aspect when executing the computer program.
[0017] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the retrieval and ranking method as described in the first aspect is implemented.
[0018] Compared with the prior art, the present invention has the following beneficial effects:
[0019] From a preset human image set, a similarity search is performed on the acquired target human image to obtain K retrieval images. The target human image and the K retrieval images are sorted based on similarity to obtain an image sequence, where K is an integer greater than zero. Based on the original image features of each image in the image sequence, the similarity between any two images in the image sequence is calculated to obtain a first similarity value. Based on the first similarity value, a first matrix is constructed, where the row number and column number of the element corresponding to the first similarity value in the first matrix correspond to the sorting number of any two images in the image sequence. Based on the acquisition device information of the target human image and the K retrieval images, the first image and the second image acquired by the same acquisition device are determined. The elements corresponding to the first similarity value between the first image and the second image in the first matrix are set to preset values to obtain a second matrix. Graph propagation processing is performed on the original image features of each image to obtain the target image features of each image. Based on the similarity between the target image features of the target human image and the target image features of the K retrieval images, the K retrieval images are re-sorted to obtain a sorting result. In the present application, a similarity search is performed on the acquired target human body image, and the similarity between any two images in the retrieval image and the target human body image is calculated, and the similarity between any two images is transferred to the image features of the target human body image and the retrieval image, so that the human body image features of the same person have more similarities, and the target features of the target human body image and the retrieval image are determined. Based on the target image features of the target human body image and the target image features of the retrieval image, the similarity between each retrieval image and the target human body image is recalculated to obtain the target similarity, thereby improving the accuracy of the similarity between images, and reordering the retrieval images based on the target similarity, thereby improving the accuracy of the order. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0021] Figure 1 This is a schematic diagram of an application environment of a method for retrieving and ranking images collected across devices provided by an embodiment of the present invention;
[0022] Figure 2 This is a flowchart of a method for retrieving and ranking images collected across devices provided by one embodiment of the present invention;
[0023] Figure 3 This is a schematic structural diagram of a device for retrieving and ranking images collected across devices, provided by one embodiment of the present invention;
[0024] Figure 4 It is a structural diagram of a computer device provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0026] In the following description, specific details such as particular system structures and techniques are provided for purposes of illustration, not limitation, to facilitate a thorough understanding of the embodiments of the present invention. However, it will be apparent to those skilled in the art that the present invention may be practiced in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted so as not to obscure the description of the present invention with unnecessary detail.
[0027] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0028] It will also be understood that the term "and / or" used in the present description and appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0029] As used in the present specification and the appended claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" may be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0030] In addition, in the description of the present specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0031] References to "one embodiment" or "some embodiments" in the present specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present invention. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0032] Embodiments of the present invention can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results.
[0033] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0034] It should be understood that the order of execution of the steps in the following embodiments does not necessarily mean the order in which they are executed. The order in which each process is executed should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0035] In order to illustrate the technical solution of the present invention, specific embodiments are provided below.
[0036] An embodiment of the present invention provides a retrieval and ranking method for images collected across devices, which can be applied in the following situations: Figure 1 In an application environment, a client communicates with a server. The client includes but is not limited to PDAs, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cloud computing devices, personal digital assistants (PDAs), and other computer devices. The server can be a standalone server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0037] See also Figure 2 , is a flow chart of a method for searching and ranking images collected across devices provided by an embodiment of the present invention. The above method for searching and ranking images collected across devices can be applied to Figure 1 The server in Figure 2 As shown, the retrieval and ranking method for images collected across devices may include the following steps.
[0038] S201: Perform similarity retrieval on the acquired target human image from a preset human image set to obtain K retrieval images, and sort the target human image and the K retrieval images based on similarity to obtain an image sequence, where K is an integer greater than zero.
[0039] In step S201, the preset human body image set is a pre-collected human body image set, including human body images of different human bodies collected by multiple acquisition devices. Each human body image includes a corresponding human body identifier, so that the identity information can be determined when a matching human body image is retrieved for the target human body image in the subsequent process. A similarity search is performed on the acquired target human body image to obtain K retrieval images, where the K retrieval images include the K human body images in the preset human body image set that have the greatest similarity to the target human body image. The target human body image and the K retrieval images are constructed into an image sequence, where each image in the image sequence has a corresponding sort number, and the image sequence includes K+1 images.
[0040] In this embodiment, a similarity search is performed on the acquired target human image from a preset human image set to obtain K search images. During the similarity detection, the similarity between the target human image and each image in the preset human image set is calculated. According to the similarity, K search images are determined. The target human image and the K search images are sorted based on the similarity to obtain an image sequence. During the sorting, the K search images can be sorted from large to small based on the similarity, and the target human image can be placed in front of or behind the K search images.
[0041] It should be noted that when calculating the similarity between the target human image and each image in the preset human image set, the similarity between the target human image and each image in the preset human image set can be obtained by calculating the cosine value between the target human image and the original image features in each image in the preset human image set. It should be noted that other methods can also be used to calculate the similarity between the target human image and each image in the preset human image set, and this embodiment is not limited thereto.
[0042] It should be noted that before calculating the cosine value between the target human image and the original image features of each image in the preset human image set, it is necessary to extract the original image features of the target human image and each image in the preset human image set. The original image features can be extracted using a feature extraction network. For example, the feature extraction network can be a convolutional network. The convolutional network can include a pooling layer, a 1×1 convolution layer, a batch normalization layer, a ReLU (Rectified Linear Unit) activation function layer, a 1×1 convolution layer, and a batch normalization layer. The step size of the pooling layer here is 1, and the image features are not downsampled. The purpose of introduction is to perform local feature fusion interaction on the image features. The 1x1 convolution layer increases the dimension of the image feature channels to twice the original number, and is then connected to a batch normalization layer to normalize the image features. The ReLU activation function layer introduces a nonlinear transformation into the network as a nonlinear function.
[0043] In this embodiment, a similarity search is performed on the acquired target human image from a preset human image set to obtain K search images. The similarity search can retrieve human images similar to the target human image, so as to determine the human image of the same person as the target human image from the similar human images. The target human image and the K search images are sorted based on the similarity to obtain an image sequence. The image sequence is constructed to determine the sorting number of each image so as to construct the first matrix and the second matrix in the subsequent process according to the corresponding sorting number.
[0044] Optionally, similarity retrieval is performed on the acquired target human image from a preset human image set to obtain K retrieval images, including:
[0045] Obtaining preset human body image features of a preset human body image set and image features of a target human body image;
[0046] Calculating similarity values between the images in the preset human image set and the target human image based on the preset human image features and the image features;
[0047] According to the similarity value, K retrieval images similar to the target human image are retrieved from the preset human image set.
[0048] In this embodiment, preset human body image features of a preset human body image set are obtained, and feature extraction is performed on a target human body image to obtain image features of the target human body image. During feature extraction, the target human body image can be processed by dividing the target human body image into multiple human body sub-images of the same size in equal proportions, and arranged according to a specified order to obtain a human body image sequence corresponding to the target human body image. The specified order can be understood as expanding in the same direction. Convolution is performed on each human body sub-image separately to extract the sub-image feature vector with a specified dimension corresponding to each human body sub-image, and then the sub-image feature vectors are arranged according to the specified order of the human body image sequence to obtain the image features of the target human body image.
[0049] In this embodiment, when convolving each human body sub-image and extracting the feature vector of each sub-image, a convolution kernel with the same size as the human body sub-image is used for convolution, so that when extracting the sub-image feature vector, the features of each human body sub-image can be fully extracted, and it helps to avoid the convolution kernel size being too large or too small to affect the extraction quality of the sub-image feature vector, making the extraction of the sub-image feature vector more reasonable.
[0050] Based on the preset human body image features and the image features, the similarity between the images in the preset human body image set and the target human body image is calculated. When calculating the similarity, the similarity can be obtained by calculating the cosine value between the preset human body image features and the image features. If the cosine value between the preset human body image features and the image features is larger, the similarity between the images in the preset human body image set and the target human body image is greater. If the cosine value between the preset human body image features and the image features is smaller, the similarity between the images in the preset human body image set and the target human body image is smaller.
[0051] Based on the similarity, K retrieval images similar to the target human image are retrieved from the preset human image set. During the retrieval, the K images with the greatest similarity to the target human image in the preset human data set are used as the K retrieval images.
[0052] In this embodiment, by using preset human image features and image features, the similarity between the images in the preset human image set and the target human image is calculated, and K images with greater similarity are selected as K retrieval images, so as to determine the human body image of the same person as the target human body image from the K retrieval images.
[0053] In another embodiment, the similarity between the images in the preset human image set and the target human image is calculated based on the preset human image features and the image features. When calculating the similarity, the feature distance between the preset human image features and the image features can also be used as the similarity size. If the distance between the preset human image features and the image features is larger, it is considered that the similarity between the images in the preset human image set and the target human image is smaller. If the distance between the preset human image features and the image features is smaller, it is considered that the similarity between the images in the preset human image set and the target human image is greater.
[0054] Optionally, the target human image and the K search images are sorted based on similarity to obtain an image sequence, including:
[0055] The target human image is ranked first, and the K search images are sorted from large to small according to their similarity. The sorting numbers of the target human image and the K search images are determined to obtain an image sequence.
[0056] In this embodiment, when constructing an image sequence, the target human body image is placed in the first place, and the K search images are placed behind the target human body image. When sorting the K search images, they are sorted in sequence according to the similarity between the K search images and the target human body image, and the sorting numbers of the target human body image and the K search images are determined. For example, the sorting number of the target human body image is 1, and the detection image with the greatest similarity to the target human body image in the K search images is placed in the second place, and the resulting sorting number is 2. The detection image with the least similarity to the target human body image in the K search images is placed in the last place, and the resulting sorting number is K+1. According to the sorting rules, the corresponding image sequence is obtained.
[0057] S202: Calculate the similarity between any two images in the image sequence based on the original image features of each image in the image sequence to obtain a first similarity value, and construct a first matrix based on the first similarity value, where the row numbers and column numbers of the elements corresponding to the first similarity values in the first matrix correspond to the sorting numbers of the any two images in the image sequence.
[0058] In step S202, the first similarity value is calculated based on the original image features of the image, and a first matrix is constructed according to the first similarity value, wherein the original image features of the image are the original image features extracted using the feature extraction network, and the elements in the rows and columns of the first matrix are composed of the first similarity values, and the row numbers and column numbers of the elements corresponding to the first similarity values in the first matrix correspond to the sorting numbers of any two images in the image sequence.
[0059] In this embodiment, the similarity between any two images in the image sequence is calculated based on the original image features of each image in the image sequence, where the any two images can be any two images in the image sequence. For example, a first image is selected from the image sequence, where the first image can be any image in the image sequence, and then a second image is selected from the image sequence, where the second image can be any image in the image sequence other than the first image. If the first image is the image numbered 2 in the image sequence, the second image can be the image numbered 7 in the image sequence, and so on. By calculating the similarity between any two images in the image sequence, K*(K+1) first similarity values can be obtained.
[0060] It should be noted that when constructing the first matrix, the first similarity value on the diagonal of the first matrix is set to the similarity value between the image and the image itself, that is, it is set to 1, so the first matrix is a matrix of (K+1)×(K+1) size, including (K+1)*(K+1) first similarity values.
[0061] It should be noted that the elements in the rows and columns of the first matrix are composed of first similarity values, and the row and column numbers of the elements corresponding to the first similarity values in the first matrix correspond to the ranking numbers of any two images in the image sequence. For example, the element in the i-th row and j-th column of the first matrix is the first similarity value between the image with ranking number i and the image with ranking number j in the image sequence, and the element in the i+1-th row and j+1-th column of the first matrix is the first similarity value between the image with ranking number i+1 and the image with ranking number j+1 in the image sequence.
[0062] It should be noted that the elements in the j-th row and the i-th column of the first matrix are the first similarity values between the image with the sorting number j and the image with the sorting number i in the image sequence. Since the first similarity value between the image with the sorting number i and the image with the sorting number j in the image sequence is equal to the first similarity value between the image with the sorting number j and the image with the sorting number i in the image sequence, the first matrix is a symmetric matrix.
[0063] It should be noted that, since the image with the first sorting number in the image sequence is the target human body image, the first row in the first matrix is the first similarity value between the K detection images and the target human body image, and the first row and first column are the first similarity values between the target human body image and the image itself, and the first similarity value is 1.
[0064] In this embodiment, a first matrix is constructed based on the first similarity values, and the row numbers and column numbers of the elements corresponding to the first similarity values in the first matrix correspond to the sorting numbers of any two images in the image sequence, so as to facilitate determining the first similarity values of any two images based on the first matrix and obtaining an image with a greater similarity to each image.
[0065] In another embodiment, when calculating the similarity between any two images in the image sequence based on the original image features of each image in the image sequence, the distance between the original image features of the any two images may be calculated, and the distance between the original image features of the any two images may be used as the first similarity value. The distance calculation formula between the original image features of the any two images is as follows:
[0066]
[0067] Among them, W1 i,j is the distance between the original image features of the image with sorting number i and the image with sorting number j in the image sequence, that is, the first similarity value between the image with sorting number i and the image with sorting number j in the image sequence, f i is the original image feature of the image numbered i in the image sequence, f j is the original image feature of the image numbered j in the image sequence, is the L2 norm and γ is a parameter.
[0068] S203: Determine the first image and the second image captured by the same acquisition device based on the acquisition device information of the target human image and the K search images, set the elements corresponding to the first similarity values between the first image and the second image in the first matrix to preset values, and obtain a second matrix.
[0069] In step S203, based on the acquisition device information of the target human image and the K search images, the first image and the second image acquired by the same acquisition device are determined, and the elements corresponding to the first similarity values between the first image and the second image in the first matrix are set to preset values to obtain the second matrix, and the images acquired by the same acquisition device are determined, so as to facilitate processing of the first similarity values between the images acquired by the same acquisition device and eliminate the problem of excessive similarity between images corresponding to different human bodies acquired by the same acquisition device.
[0070] In this embodiment, based on the acquisition device information of the target human image and the K search images, the first image and the second image captured by the same acquisition device are determined. The first image and the second image can be any two images in the image sequence. For example, the first image can be the image numbered 3 in the image sequence, and the second image can be the image numbered 5 in the image sequence. The first image and the second image captured by the same acquisition device have the same acquisition height and acquisition angle, resulting in a high first similarity value between the first image and the second image. However, if the first image and the second image captured by the same acquisition device are images of different human bodies, the similarity value between the first image and the second image may be greater than the similarity value between human body images of the same person captured by different acquisition devices. To avoid connectivity between images captured by the same acquisition device, the elements corresponding to the first similarity values between the first image and the second image in the first matrix are set to preset values, thereby obtaining a second matrix. When setting the elements corresponding to the first similarity values between the first image and the second image in the first matrix to preset values, the corresponding elements can be determined based on the ranking numbers of the first image and the second image in the image sequence. Other methods can also be used to determine the corresponding elements, which are not limited in this embodiment.
[0071] It should be noted that if the images captured by the same acquisition device are two images, the first image and the second image are respectively any one of the two images. For example, if the two images captured by the same acquisition device are sorted 3 and 6 in the image sequence, the first image is 3 or 6. If the first image is 3, the second image is 6. If the first image is 6, the second image is 3. The element corresponding to the first similarity value between the image sorted 3 and the image sorted 6 in the image sequence is set to a preset value, that is, the element in the 3rd row and 6th column and the element in the 6th row and 3rd column in the first matrix are set to the preset value.
[0072] It should be noted that if the images captured by the same acquisition device are three or more images, any two of the three or more images are used as the first image and the second image, and all elements corresponding to the first similarity values of the first image and the second image are set to preset values. For example, if the three images captured by the same acquisition device are sorted 3, 6, and 8 in the image sequence, any two of the three images sorted 3, 6, and 8 are used as the first image and the second image. If the image sorted 3 is the first image and the image sorted 6 is the second image, then the elements corresponding to the first similarity values of the image sorted 3 and the image sorted 6 in the image sequence are set to preset values, that is, the elements in the 3rd row and 6th column and the elements in the 6th row and 3rd column in the first matrix are set to preset values. Since the three images with sorting numbers of 3, 6 and 8 can form any two images in three cases, namely the first image and the second image composed of the images with sorting numbers of 3 and 6, the first image and the second image composed of the images with sorting numbers of 3 and 8, and the first image and the second image composed of the images with sorting numbers of 6 and 8, it is necessary to set the elements in the 3rd row and 6th column, the elements in the 6th row and 3rd column, the elements in the 3rd row and 8th column, the elements in the 8th row and 3rd column, the elements in the 6th row and 8th column, and the elements in the 8th row and 6th column in the first matrix to preset values.
[0073] It should be noted that the preset value in this embodiment is set to 0, and the preset value can also be set to other values, which is not limited in this embodiment.
[0074] It should be noted that the preset value in this embodiment is set to 0, which can completely avoid the influence of the same acquisition device on the similarity between images.
[0075] In this embodiment, based on the image acquisition device information, it is determined whether any two images in the image sequence were acquired by the same acquisition device. If the first and second images in any two images were acquired by the same acquisition device, the element corresponding to the first similarity value between the first and second images in the first matrix is set to a preset value, resulting in a second matrix. In this embodiment, the preset value is set to 0. The second matrix only contains the first similarity values between any two images in the image sequence acquired by different acquisition devices. This facilitates the propagation and fusion of only the image features of the same person acquired by different acquisition devices during the subsequent image propagation process, thereby improving the similarity between images of the same person acquired by different acquisition devices and enhancing the accuracy of cross-device image retrieval.
[0076] In another embodiment, an element corresponding to the first similarity value between the first image and the second image in the first matrix is set to a preset value, where the preset value may also be a value obtained by scaling down the first similarity value between the first image and the second image. For example, if the first similarity value between the first image and the second image in the first matrix is a, the first similarity value a between the first image and the second image in the first matrix is scaled down and used as a new first similarity value to obtain the second matrix.
[0077] In this embodiment, the element corresponding to the first similarity value between the first image and the second image in the first matrix is set to a value after the first similarity value is reduced. After the first similarity value is reduced, the new first similarity value obtained is smaller. The new first similarity value can also weaken the connection between images captured by the same acquisition device.
[0078] Optionally, setting elements corresponding to the first similarity values between the first image and the second image in the first matrix to preset values to obtain a second matrix includes:
[0079] Determining, according to the sorting numbers of the first image and the second image in the image sequence, a target row number and a target column number corresponding to a first similarity value calculated between the first image and the second image in the first matrix;
[0080] The elements corresponding to the target row number and the target column number are set to preset values to obtain a second matrix.
[0081] In this embodiment, the target row number and target column number corresponding to the first similarity value calculated from the first image and the second image in the first matrix are determined based on the sorting numbers of the first image and the second image in the image sequence. For example, if the sorting number of the first image is 3 and the sorting number of the second image is 4, then the target row number corresponding to the first similarity value calculated from the first image and the second image is either 3 or 4. If the target row number is 3, then the corresponding target row number is 4, and if the target row number is 4, then the corresponding target column number is 3. The elements corresponding to the target row number and target column number are set to preset values to obtain a second matrix. For example, if the target row number is 3 and the target column number is 4, then the elements in the third row and fourth column of the first matrix are set to preset values.
[0082] It should be noted that the first similarity value between any two images in the image sequence corresponds to two elements in the first matrix. Therefore, when determining the target row number and the target column number, two target row numbers and two target column numbers need to be determined, and the two elements corresponding to the two target row numbers and the two target column numbers are symmetrical in the first matrix.
[0083] In this embodiment, based on the sorting numbers of the first image and the second image in the image sequence, the target row number and the target column number corresponding to the first similarity value calculated from the first image and the second image in the first matrix are determined, and the elements corresponding to the target row number and the target column number are set to preset values to obtain the second matrix. The target row number and the target column number are directly determined based on the sorting numbers in the image sequence, thereby improving the efficiency of determining the target row number and the target column number.
[0084] S204: Performing graph propagation processing on the original image features of each image according to the first matrix and the second matrix to obtain target image features of each image.
[0085] In step S204, graph propagation processing is performed on the original image features of each image according to the first matrix and the second matrix to obtain the target image features of each image, wherein the graph propagation processing refers to transferring the first similarity value in the first matrix and the new first similarity value after reset in the second matrix to the original image features in each image, thereby increasing the similarity of the human body image features of the same person and increasing the difference of the human body image features of different people.
[0086] In this embodiment, graph propagation processing is performed on the original image features of each image based on the first matrix and the second matrix, that is, the first similarity value in the first matrix and the new first similarity value after reset in the second matrix are transferred to the original image features in each image. For example, the graph propagation processing can transfer the first similarity value in the first matrix and the new first similarity value after reset in the second matrix to the original image features of each image, for example, calculating the mean of the first similarity values of each row or column in the first matrix, and calculating the mean of the new first similarity values after reset for each row or column in the second matrix, obtaining the first mean corresponding to each row or column in the first matrix and the second mean corresponding to each row or column in the second matrix, adding the first mean and the second mean in the same row or column, and using the result of the addition as the weight value of the original image feature of the sorted numbered image corresponding to the row or column, weighting the weight value and the corresponding original image feature, and calculating the target image feature of each image.
[0087] For example, the first similarity values of the second row in the first matrix are averaged to obtain the first mean of the first matrix, the new first similarity values after reset in the second matrix are averaged to obtain the second mean of the second matrix, the first mean and the second mean are added to obtain the result of the addition of the second row, the result of the addition of the second row is used as the weight value of the original image feature of the image with sorting number 2, the weight value is multiplied by the original image feature of the image with sorting number 2 to obtain the target image feature.
[0088] In this embodiment, graph propagation processing is performed on the original image features of each image based on the first matrix and the second matrix, and the first similarity value in the first matrix and the reset new first similarity value in the second matrix are transferred to the original image features to obtain target image features, wherein the second matrix only contains the similarities between images captured by different acquisition devices, and the obtained target image features weaken the connection between images captured by the same acquisition device, thereby making the obtained target image features more robust.
[0089] Optionally, graph propagation processing is performed on the original image features of each image according to the first matrix and the second matrix to obtain target image features of each image, including:
[0090] Performing graph propagation processing on the original image features of each image through the first matrix and the second matrix respectively to obtain the first image features of each image and the second image features of each image;
[0091] The first image feature and the second image feature are weightedly summed to obtain the target image feature of each image.
[0092] In this embodiment, graph propagation processing is performed on the original image features of each image through the first matrix and the second matrix respectively to obtain the first image features of each image and the second image features of each image. When performing graph propagation, the first similarity value in the first matrix and the new first similarity value after reset in the second matrix are transferred to the original image features of each image.
[0093] In this embodiment, when the original image features of each image are subjected to graph propagation processing through the first matrix, the first similarity values of each row in the first matrix are averaged to obtain the first mean of each row, and the first mean of each row is used as the weight value of the image with a sorting number equal to the row number of the row, and the weight value is weighted with the original image features of the image with the sorting number to obtain the first image feature of the image. For example, the first similarity values of the second row in the first matrix are averaged to obtain the first mean of the second row in the first matrix, and the first mean of the second row is used as the weight value of the original image feature of the image with a sorting number of 2, and the weight value is weighted with the original image feature of the image with the sorting number of 2 to obtain the first image feature of the image with the sorting number of 2.
[0094] When the original image features of each image are subjected to graph propagation processing through the second matrix, the new first similarity values after reset of each row in the second matrix are averaged to obtain the second mean of each row, the second mean of each row is used as the weight value of the image with the same sorting number as the row number of the row, the weight value is weighted with the original image features of the image with the sorting number to obtain the second image features of the image, for example, the new first similarity values after reset of the second row in the second matrix are averaged to obtain the second mean of the second row in the second matrix, the second mean of the second row is used as the weight value of the original image features of the image with the sorting number 2, the weight value is weighted with the original image features of the image with the sorting number 2 to obtain the second image features of the image with the sorting number 2.
[0095] The first image feature and the second image feature are weighted and summed to obtain the target image feature of each image. The calculation formula of the target image feature is as follows:
[0096] F * =αF1+(1-α)F2
[0097] Among them, F * is the target image feature of each image, F1 is the first image feature of each image, F2 is the second image feature of each image, and α is a parameter less than 1 and greater than 0.
[0098] In this embodiment, graph propagation processing is performed on the original image features of each image through the first matrix and the second matrix respectively to obtain the first image features of each image and the second image features of each image. The first image features and the second image features of each image are the fusion of similarity features between other images and the image, thereby strengthening the association between images corresponding to the same person, weakening the association between images corresponding to different people, and improving the accuracy of the target image features.
[0099] In another embodiment, graph propagation processing is performed on the original image features of each image using the first matrix and the second matrix, respectively. Before obtaining the first image features and the second image features of each image, the first similarity value in the first matrix and the new first similarity value reset in the second matrix can also be processed. The elements in each row of the first matrix and the second matrix are the similarity values between the image with the same sorting number as the row and other images. The greater the similarity value, the greater the possibility that the image is of the same person, and the smaller the similarity, the smaller the possibility that the image is of the same person. In order to strengthen the association between images corresponding to the same person and weaken the association between images corresponding to different people, a similarity threshold is set, and similarity values less than the similarity threshold are set to 0, that is, the remaining images with less similarity to the image are considered not to be images of the same person.
[0100] In this embodiment, similarity values less than the similarity threshold in the first matrix and the second matrix are set to 0 to obtain a new first matrix and a new second matrix. The new first matrix and the new second matrix strengthen the association between images corresponding to the same person and weaken the association between images corresponding to different people. The image of the same person can be distinguished from images of other different people, thereby increasing the distinctiveness of images between different people.
[0101] After obtaining the new first matrix and the new second matrix, graph propagation processing is performed on the original image features of each image using the new first matrix and the new second matrix respectively.
[0102] Optionally, performing graph propagation processing on the original image features of each image using the first matrix and the second matrix to obtain the first image features of each image and the second image features of each image includes:
[0103] Normalizing the rows of the first matrix and the second matrix respectively to obtain a first normalized matrix of the first matrix and a second normalized matrix of the second matrix;
[0104] weighting the original image features of each image according to the first normalization matrix to obtain a first image feature of each image;
[0105] The original image features of each image are weighted according to the second normalization matrix to obtain the second image features of each image.
[0106] In this embodiment, the first matrix and the second matrix are normalized to obtain a first normalized matrix of the first matrix and a second normalized matrix of the second matrix. During the normalization process, the first matrix and the second matrix are converted into diagonal matrices. The calculation formula for normalizing the first matrix and the second matrix is as follows:
[0107]
[0108]
[0109] in, is the first normalized matrix, is the second normalized matrix, D1 and D2 are diagonal matrices, W1 is the first matrix, and W2 is the second matrix. D1 and D2 are matrices of size (K+1)×(K+1). Among them, the first normalized matrix and the second normalized matrix are diagonal matrices. According to the first normalized matrix, the original image features of each image are weighted to obtain the first image features of each image. According to the second normalized matrix, the original image features of each image are weighted to obtain the second image features of each image. That is, the diagonal elements corresponding to each row in the first normalized matrix and the second normalized matrix are multiplied with the original image features of the images with the same sorting number as the row to obtain the first image features of each image and the second image features of each image. Among them, the calculation formulas for the first image features of each image and the second image features of each image are as follows:
[0110]
[0111]
[0112] Among them, F1 is the feature matrix of the first image features of all images in the image sequence. The size of F1 is (K+1)×N, where (K+1) is the number of images in the image sequence, N is the feature dimension corresponding to the original image features of each image, and each row in the F1 feature matrix is the first image feature of the image with the same sorting number as the corresponding row number. is the first normalized matrix, is the second normalized matrix, F is the feature matrix of the original image features of all images in the image sequence, and the size of F is (K+1)×N, where (K+1) is the number of images in the image sequence, N is the feature dimension corresponding to the original image features of each image, and each row in the feature matrix F is the original image features of the images with the same sorting number as the corresponding row number.
[0113] In this embodiment, the first matrix and the second matrix are normalized. The normalization process can transfer the similarity values in each row of the first matrix and the second matrix to the diagonal values of that row in the diagonal matrix. The original image features of each image are weighted according to the first normalized matrix and the second normalized matrix to obtain the first image features and the second image features of each image. That is, the similarity values in each row of the first matrix and the second matrix are transferred to the original image features of the image with the same sorting number as that row, so that the original features can aggregate the similarity features of different images, thereby improving the accuracy of the first image features and the second image features of each image.
[0114] S205: Reorder the K search images according to the similarity between the target image features of the target human image and the target image features of the K search images to obtain a ranking result.
[0115] In step S205, the K search images are reordered based on the similarity between the target image features of the target human image and the target image features of the K search images to obtain a ranking result. The reordering is based on descending similarity, so the target image features contain more feature information, thereby improving the accuracy of the ranking.
[0116] In this embodiment, the K search images are reordered based on the similarity between the target image feature of the target human image and the target image feature of each search image. When calculating the similarity, the feature distance between the target image feature of the target human image and the target image feature of each search image can be used as the similarity value. If the distance between the target image feature of the target human image and the target image feature of each search image is larger, it is considered that the similarity between the target human image and the search image is smaller. If the distance between the target image feature of the target human image and the target image feature of each search image is smaller, it is considered that the similarity between the target human image and the search image is greater. The K search images are reordered based on the similarity values to obtain a sorting result. When sorting, the images are sorted from large to small based on the similarity values, that is, the image most similar to the target human image is placed at the front.
[0117] In another embodiment, the target image features of each image can be spliced with the original image features of the corresponding image to obtain the spliced features of each image. According to the similarity between the spliced features of the target human body image and the spliced features of each search image, the K search images are re-sorted to obtain the sorting result. The length of the spliced feature is the sum of the lengths of the original image features and the target image features. Increasing the length of the spliced feature allows the spliced feature to contain more feature information.
[0118] In this embodiment, the target image features of each image are spliced with the original image features of the corresponding image to obtain the spliced features of each image, wherein the length of the spliced features is 2m, and m is the feature dimension of the original image features and the target image features in each image. Based on the similarity between the spliced features of the target human body image and the spliced features of each search image, when calculating the similarity between the spliced features of the target human body image and the spliced features of each search image, the cosine function formula can be used to calculate the cosine value between the spliced features of the target human body image and the spliced features of each search image, and the cosine value is used as the similarity value between the spliced features of the target human body image and the spliced features of each search image, wherein the larger the cosine value, the greater the similarity value, and the smaller the cosine value, the smaller the similarity value. The K search images are re-sorted according to the similarity value to obtain a sorting result.
[0119] In this embodiment, the original image features of each image are spliced with the target image features to obtain a spliced feature. The length of the spliced feature is the sum of the lengths of the original image features and the target image features. The increased length of the spliced feature allows the spliced feature to contain more feature information. The similarity between the target human image and each retrieval image is calculated based on the spliced features of the target human image and the spliced features of each retrieval image, thereby improving the accuracy of the similarity calculation. The K retrieval images are re-sorted according to the similarity to obtain a sorting result, thereby improving the accuracy of the sorting. The retrieval images of the same person as the target human image can be arranged in front, and the retrieval images of different people can be arranged in the back. The similarity between corresponding images of the same person is strengthened, and the similarity between corresponding images of different people is weakened, which is more conducive to truncation of retrieval images of the same person.
[0120] Optionally, the K search images are reordered according to similarities between target image features of the target human image and target image features of the K search images to obtain a ranking result, including:
[0121] Calculate the cosine similarity between the target image features of the target human image and the target image features of the K retrieval images;
[0122] According to the cosine similarity value, the K search images are re-ranked to obtain the ranking result.
[0123] In this embodiment, when calculating the similarity between the target image feature of the target human image and the target image feature of each retrieval image, a cosine function formula can be used to calculate the cosine value between the target image feature of the target human image and the target image feature of each retrieval image, and the cosine value is used as the similarity value between the target image feature of the target human image and the target image feature of each retrieval image, wherein the larger the cosine value, the greater the similarity value, and the smaller the cosine value, the smaller the similarity value. The K retrieval images are reordered according to the similarity value to obtain a sorting result, wherein, when sorting, the images are sorted from large to small according to the similarity value, that is, the image most similar to the target human image is placed at the front.
[0124] From a preset human image set, a similarity search is performed on the acquired target human image to obtain K retrieval images. The target human image and the K retrieval images are sorted based on similarity to obtain an image sequence, where K is an integer greater than zero. Based on the original image features of each image in the image sequence, the similarity between any two images in the image sequence is calculated to obtain a first similarity value. Based on the first similarity value, a first matrix is constructed, where the row number and column number of the element corresponding to the first similarity value in the first matrix correspond to the sorting number of any two images in the image sequence. Based on the acquisition device information of the target human image and the K retrieval images, the first image and the second image acquired by the same acquisition device are determined. The elements corresponding to the first similarity value between the first image and the second image in the first matrix are set to preset values to obtain a second matrix. Graph propagation processing is performed on the original image features of each image to obtain the target image features of each image. Based on the similarity between the target image features of the target human image and the target image features of the K retrieval images, the K retrieval images are re-sorted to obtain a sorting result. In the present application, a similarity search is performed on the acquired target human body image, and the similarity between any two images in the retrieval image and the target human body image is calculated, and the similarity between any two images is transferred to the image features of the target human body image and the retrieval image, so that the human body image features of the same person have more similarities, and the target features of the target human body image and the retrieval image are determined. Based on the target image features of the target human body image and the target image features of the retrieval image, the similarity between each retrieval image and the target human body image is recalculated to obtain the target similarity, thereby improving the accuracy of the similarity between images, and reordering the retrieval images based on the target similarity, thereby improving the accuracy of the order.
[0125] See also Figure 3 , Figure 3 This is a schematic diagram of the structure of a cross-device image acquisition retrieval and sorting device provided by an embodiment of the present invention. In this embodiment, the terminal includes various units for executing Figure 2 Each step in the corresponding embodiment. Please refer to Figure 2 For the convenience of explanation, only the parts related to this embodiment are shown. Figure 3 The retrieval and ranking device 30 includes: a retrieval module 31, a construction module 32, an acquisition module 33, a graph propagation module 34, and a ranking module 35.
[0126] A retrieval module 31 is configured to perform a similarity search on the acquired target human image from a preset human image set to obtain K retrieval images, and to sort the target human image and the K retrieval images based on similarity to obtain an image sequence, where K is an integer greater than zero;
[0127] a construction module 32 configured to calculate the similarity between any two images in the image sequence based on the original image features of each image in the image sequence to obtain a first similarity value, and to construct a first matrix based on the first similarity value, wherein the row numbers and column numbers of the elements corresponding to the first similarity value in the first matrix correspond to the ordering numbers of the any two images in the image sequence;
[0128] an obtaining module 33 for determining, based on acquisition device information of the target human image and the K search images, the first image and the second image acquired by the same acquisition device, and setting elements corresponding to the first similarity values between the first image and the second image in the first matrix to preset values to obtain a second matrix;
[0129] A graph propagation module 34 is configured to perform graph propagation processing on the original image features of each image according to the first matrix and the second matrix to obtain target image features of each image;
[0130] The sorting module 35 is configured to re-sort the K search images according to the similarity between the target image features of the target human image and the target image features of the K search images to obtain a sorting result.
[0131] Optionally, the retrieval module 31 includes:
[0132] The acquiring unit is used to acquire the preset human body image features of the preset human body image set and the image features of the target human body image.
[0133] The first calculation unit is used to calculate the similarity value between the image in the preset human image set and the target human image according to the preset human image feature and the image feature.
[0134] The retrieval unit is used to retrieve K retrieval images similar to the target human image from a preset human image set according to the similarity value.
[0135] Optionally, the retrieval module 31 includes:
[0136] The first determining unit is used to rank the target human image first, sort the K search images from large to small according to similarity, determine the sorting numbers of the target human image and the K search images, and obtain an image sequence.
[0137] Optionally, the obtaining module 33 includes:
[0138] The second determining unit is configured to determine, according to the sorting numbers of the first image and the second image in the image sequence, a target row number and a target column number corresponding to the first similarity value calculated from the first image and the second image in the first matrix.
[0139] The obtaining unit is used to set the elements corresponding to the target row number and the target column number to preset values to obtain the second matrix.
[0140] Optionally, the graph propagation module 34 includes:
[0141] The graph propagation unit is used to perform graph propagation processing on the original image features of each image through the first matrix and the second matrix respectively, to obtain the first image features of each image and the second image features of each image.
[0142] The weighting unit is used to perform weighted sum calculation on the first image feature and the second image feature to obtain the target image feature of each image.
[0143] Optionally, the graph propagation unit includes:
[0144] The normalization subunit is used to perform normalization processing on the first matrix and the second matrix respectively to obtain a first normalized matrix of the first matrix and a second normalized matrix of the second matrix.
[0145] The first weighting subunit is used to weight the original image features of each image according to the first normalization matrix to obtain the first image features of each image.
[0146] The second weighting subunit is configured to weight the original image features of each image according to the second normalization matrix to obtain the second image features of each image.
[0147] Optionally, the sorting module 35 includes:
[0148] The second calculation unit is used to calculate the cosine similarity value between the target image feature of the target human image and the target image features of the K search images.
[0149] The sorting unit is used to re-sort the K search images according to the cosine similarity values to obtain a sorting result.
[0150] It should be noted that the information interaction, execution process, etc. between the above-mentioned modules, units, and sub-units are based on the same concept as the embodiment of the method of the present invention. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0151] Figure 4 A schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Figure 4 As shown, the computer device of this embodiment includes: at least one processor ( Figure 4Only one is shown), a memory, and a computer program stored in the memory and executable on at least one processor, wherein when the processor executes the computer program, the steps in any of the above-mentioned embodiments of the method for retrieving and ranking images collected across devices are implemented.
[0152] The computer device may include, but is not limited to, a processor and a memory. It will be understood by those skilled in the art that Figure 4 This is merely an example of a computer device and does not constitute a limitation on the computer device. The computer device may include more or fewer components than shown in the figure, or a combination of certain components, or different components.
[0153] The processor may be a CPU, or other general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. A general-purpose processor may be a microprocessor, or any conventional processor.
[0154] The memory includes a readable storage medium, an internal memory, etc., wherein the internal memory can be the memory of a computer device, and the internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The readable storage medium can be the hard disk of the computer device, and in other embodiments, it can also be an external storage device of the computer device, for example, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the computer device. Furthermore, the memory can also include both the internal storage unit of the computer device and the external storage device. The memory is used to store the operating system, application programs, boot loaders (BootLoader), data, and other programs, such as the program code of the computer program. The memory can also be used to temporarily store data that has been output or is about to be output.
[0155] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other and are not used to limit the scope of protection of the present invention. The specific working process of the units and modules in the above-mentioned device can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of the above-mentioned method embodiment. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include at least: any entity or device capable of carrying computer program code, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.
[0156] The present invention may implement all or part of the processes in the above-mentioned method embodiments, and may also be completed through a computer program product. When the computer program product runs on a computer device, the computer device can implement the steps in the above-mentioned method embodiments when executing the computer program product.
[0157] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0158] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0159] In the embodiments provided by the present invention, it should be understood that the disclosed apparatus / computer equipment and methods can be implemented in other ways. For example, the apparatus / computer equipment embodiments described above are merely illustrative. For example, the division of modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0160] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0161] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A retrieval and ranking method for images collected across devices, characterized in that: The retrieval and ranking method comprises: Performing similarity retrieval on the acquired target human image from a preset human image set to obtain K retrieval images, and sorting the target human image and the K retrieval images based on similarity to obtain an image sequence, where K is an integer greater than zero; calculating, based on original image features of each image in the image sequence, a similarity between any two images in the image sequence to obtain a first similarity value, and constructing a first matrix based on the first similarity value, wherein row numbers and column numbers of elements corresponding to the first similarity value in the first matrix correspond to the sorting numbers of the any two images in the image sequence; Determining, based on acquisition device information of the target human image and the K search images, a first image and a second image acquired by the same acquisition device, and setting elements corresponding to the first similarity values between the first image and the second image in the first matrix to preset values to obtain a second matrix; Performing graph propagation processing on the original image features of each image according to the first matrix and the second matrix to obtain target image features of each image; According to the similarity between the target image feature of the target human image and the target image features of the K search images, the K search images are re-ranked to obtain a ranking result.
2. The search and ranking method according to claim 1, wherein: The step of setting the elements corresponding to the first similarity values between the first image and the second image in the first matrix to preset values to obtain a second matrix includes: determining, according to the sorting numbers of the first image and the second image in the image sequence, a target row number and a target column number corresponding to a first similarity value calculated from the first image and the second image in the first matrix; The elements corresponding to the target row number and the target column number are set to preset values to obtain a second matrix.
3. The search and ranking method according to claim 1, wherein: The performing graph propagation processing on the original image features of each image according to the first matrix and the second matrix to obtain the target image features of each image includes: Performing graph propagation processing on the original image features of each image using the first matrix and the second matrix to obtain a first image feature of each image and a second image feature of each image; A weighted sum is performed on the first image feature and the second image feature to obtain a target image feature for each image.
4. The search and ranking method according to claim 3, wherein: The method of performing graph propagation processing on the original image features of each image using the first matrix and the second matrix to obtain the first image features of each image and the second image features of each image includes: performing normalization processing on the first matrix and the second matrix respectively to obtain a first normalized matrix of the first matrix and a second normalized matrix of the second matrix; weighting the original image features of each image according to the first normalization matrix to obtain a first image feature of each image; The original image features of each image are weighted according to the second normalization matrix to obtain the second image features of each image.
5. The search and ranking method according to claim 1, wherein: The reordering of the K search images according to the similarity between the target image feature of the target human image and the target image features of the K search images to obtain a sorting result includes: Calculating cosine similarity values between target image features of the target human image and target image features of K retrieval images; The K search images are reordered according to the cosine similarity values to obtain a sorting result.
6. The search and ranking method according to claim 1, wherein: The method of performing similarity retrieval on the acquired target human body image from the preset human body image set to obtain K retrieval images includes: Acquire preset human body image features of the preset human body image set and image features of the target human body image; Calculating a similarity value between the image in the preset human image set and the target human image based on the preset human image feature and the image feature; According to the similarity value, K retrieval images similar to the target human image are retrieved from the preset human image set.
7. The search and ranking method according to claim 1, wherein: The step of sorting the target human body image and the K search images based on similarity to obtain an image sequence includes: The target human image is ranked first, the K search images are sorted from large to small according to the similarity, and the sorting numbers of the target human image and the K search images are determined to obtain an image sequence.
8. A retrieval and ranking device for collecting images across devices, characterized in that: The retrieval and sorting device comprises: a retrieval module configured to perform similarity retrieval on the acquired target human image from a preset human image set to obtain K retrieval images, and to sort the target human image and the K retrieval images based on similarity to obtain an image sequence, where K is an integer greater than zero; a construction module, configured to calculate, based on original image features of each image in the image sequence, a similarity between any two images in the image sequence to obtain a first similarity value, and construct a first matrix based on the first similarity value, wherein the row numbers and column numbers of the elements corresponding to the first similarity value in the first matrix correspond to the sorting numbers of the any two images in the image sequence; an obtaining module, configured to determine, based on acquisition device information of the target human image and the K search images, a first image and a second image acquired by the same acquisition device, and set elements corresponding to the first similarity values between the first image and the second image in the first matrix to preset values to obtain a second matrix; a graph propagation module, configured to perform graph propagation processing on the original image features of each image according to the first matrix and the second matrix to obtain target image features of each image; The sorting module is used to re-sort the K search images according to the similarity between the target image features of the target human image and the target image features of the K search images to obtain a sorting result.
9. A computer device, characterized in that: The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the retrieval and ranking method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the retrieval and ranking method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Image retrieval result reordering method and device, equipment and medium
CN113407751A
Method of clustering using encoder-decoder model based on attention mechanism and storage medium for image recognition
US20230186600A1