Target retrieval method and device, terminal and computer storage medium
By generating extended-dimensional retrieval vectors through a visual encoder and performing similarity calculations directly in the vector database, the problems of slow image retrieval speed and inaccurate positioning in existing technologies are solved, achieving fast and accurate target positioning and retrieval.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA MOBILE GRP GUANGDONG CO LTD
- Filing Date
- 2025-12-25
- Publication Date
- 2026-04-17
AI Technical Summary
In existing technologies, image retrieval methods have high computational complexity and slow retrieval speed, which cannot meet the real-time retrieval needs of massive image data. Furthermore, they lack the resolvability of retrieval results and cannot locate the specific position of the target in the image.
The first visual encoder generates the first visual feature vector of the candidate target, and combines it with the second visual encoder to generate an extended-dimensional retrieval vector. Similarity is calculated directly in the vector database, and candidate target and coordinate information are stored together. Text description and reference image retrieval are supported.
It improves retrieval speed and efficiency, enhances the parsability of retrieval results, enables quick location of target coordinates in images, and improves user experience.
Smart Images

Figure CN121880591A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of data processing technology, and in particular relates to a target retrieval method, apparatus, terminal and computer storage medium. Background Technology
[0002] With the development of the Internet and the popularization of multimedia technology, image data has exploded. Retrieving search results that correspond to user needs from massive amounts of image data is a very important step.
[0003] In existing technologies, there are retrieval methods that analyze search text to perform image retrieval, and there are retrieval methods that extract the image features of the entire image and then calculate similarity from the features stored in the database. However, the text-based retrieval method is prone to inaccurate or incomplete label coverage, while the retrieval method that extracts the image features of the entire image and calculates similarity has high computational complexity and slow retrieval speed, making it difficult to meet the real-time retrieval needs of large amounts of image data, and lacking the resolvability of retrieval results. Summary of the Invention
[0004] This application provides a target retrieval method, apparatus, terminal, and computer storage medium, which can improve retrieval speed and efficiency by extending the visual feature vector, allowing the extended visual feature vector to be directly used for similarity calculation, obtaining the coordinate information of the retrieval target in the image, and enhancing the resolvability of the retrieval results.
[0005] In a first aspect, embodiments of this application provide a target retrieval method, the method comprising: The first visual encoder performs candidate target recognition processing on the image to be entered into the database, and obtains the first visual feature vectors of multiple candidate targets in the image to be entered into the database and the coordinate information of the candidate targets in the image to be entered into the database. The first visual encoder calculates the first visual feature vector based on the scalar of the candidate target and the second visual encoder outputting the second visual feature vector of the candidate target. The scalar is the vector adjustment coefficient obtained by the multilayer perceptron in the vector matching training of the candidate target. The candidate targets, their coordinate information, and the first visual feature vector are associated and stored in the vector database; Receive a retrieval instruction and generate a retrieval vector with extended dimensions from the text description and / or reference image contained in the retrieval instruction through a second visual encoder; The inner product of the retrieval vector and the first visual feature vectors of each candidate target in the vector database is calculated to retrieve at least one first visual feature vector whose similarity to the retrieval vector exceeds a preset similarity threshold. The retrieved first visual feature vector is associated with the stored candidate target and coordinate information to determine the retrieval result.
[0006] Secondly, embodiments of this application provide a target retrieval device, the device comprising: An extension module is used to perform candidate target recognition processing on the image to be entered into the database through a first visual encoder, to obtain the first visual feature vectors of multiple candidate targets in the image to be entered into the database and the coordinate information of the candidate targets in the image to be entered into the database. The first visual encoder calculates the first visual feature vector based on the scalar of the candidate target and the second visual feature vector output by the second visual encoder of the candidate target. The scalar is the vector adjustment coefficient obtained by the multilayer perceptron in the vector matching training of the candidate target. The storage module is used to associate and store candidate targets, their coordinate information, and the first visual feature vector into a vector database. The receiving module is used to receive retrieval instructions and generate a retrieval vector containing extended dimensions from the text description and / or reference image contained in the retrieval instructions through a second visual encoder. The retrieval module is used to perform an inner product calculation between the retrieval vector and the first visual feature vectors of each candidate target in the vector database, so as to retrieve at least one first visual feature vector whose similarity to the retrieval vector exceeds a preset similarity threshold. The results module is used to associate the retrieved first visual feature vector with the stored candidate target and coordinate information to determine the retrieval results.
[0007] Thirdly, embodiments of this application provide a terminal device, the device including: a processor and a memory storing computer program instructions; the processor reads and executes the computer program instructions to implement the target retrieval method as described in the first aspect.
[0008] Fourthly, embodiments of this application provide a computer storage medium storing computer program instructions, which, when executed by a processor, implement the target retrieval method as described in the first aspect.
[0009] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed, implements the target retrieval method as described in the first aspect.
[0010] The target retrieval method, apparatus, and computer storage medium of this application embodiment perform candidate target recognition processing on the image to be entered into the database using a first visual encoder. This obtains first visual feature vectors of multiple candidate targets in the image to be entered into the database and the coordinate information of the candidate targets in the image. Specifically, the first visual encoder calculates the first visual feature vector based on the scalar of the candidate target and the second visual feature vector output by the second visual encoder. The scalar is the vector adjustment coefficient obtained by the multilayer perceptron during vector matching training of the candidate targets. The candidate targets, their coordinate information, and the first visual feature vectors are associated and stored in a vector database. The first visual feature vector is obtained by transforming the second visual feature vector, ensuring consistency between the first and second visual feature vectors in the matching process. This facilitates the retrieval of the target target in the vector database using the first visual feature vector. The system directly calculates similarity between visual feature vectors to obtain search results. Upon receiving a search instruction, a second visual encoder generates a search vector with extended dimensions from the text description and / or reference image included in the search instruction. The search vector and the first visual feature vectors of each candidate target in the vector database are then multiplied together to retrieve at least one first visual feature vector whose similarity to the search vector exceeds a preset similarity threshold. The retrieved first visual feature vector is associated with the stored candidate target and coordinate information to determine the search result. This enables image-based or text-based search of the target. The coordinate information included in the search results increases the parseability of the results, improving the user experience. Direct retrieval within a vector database containing first visual feature vectors via the search vector accelerates search efficiency, reduces search latency, and enhances the user experience. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a schematic flowchart of a target retrieval method provided in an embodiment of this application; Figure 2 This is a schematic diagram of candidate targets included in an image to be added to the database, provided in an embodiment of this application. Figure 3 This is a schematic diagram of a reference frame image provided in an embodiment of this application; Figure 4 This is another flowchart of target retrieval provided in the embodiments of this application; Figure 5 This is a schematic diagram of a device structure provided in an embodiment of this application; Figure 6This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation
[0013] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.
[0014] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.
[0015] In existing technologies, image retrieval can be performed through text. This involves analyzing text information related to the image to be retrieved, such as tags, titles, or descriptions, to locate and extract the corresponding image data. Image retrieval through text relies on the relevance between the text content and the image to be retrieved, which can achieve effective retrieval and classification of the image. However, during the retrieval process, the workload of labeling the image to be retrieved is large, the keyword extraction is inaccurate, and the quality of the retrieval results is highly dependent on the accuracy and completeness of the description of the image to be retrieved, resulting in poor retrieval performance.
[0016] Image retrieval can also be performed based on content. This involves extracting image features from the image to be retrieved, such as color, texture, and shape, and then calculating the similarity between these image features and features stored in the database. The advantage of this method is that it directly utilizes information from the image itself, avoiding problems caused by incorrect or missing labels.
[0017] Image retrieval can also be performed using deep learning, which involves extracting image features through convolutional neural networks and then performing retrieval using similarity metrics. This improves the retrieval results to some extent. However, the overall computational complexity is high and the retrieval speed is slow, making it difficult to meet the real-time retrieval needs of massive image data. Furthermore, the retrieval process can only be performed by comparing image features and cannot be performed by using text descriptions.
[0018] Image retrieval can also be performed based on visual-language contrastive learning. This involves encoding images using a visual encoder to obtain visual vector data, which is then stored in a vector database. During the retrieval process, a text encoder obtains text vectors, which are then compared with the visual vectors in the vector database to retrieve the top few search results. However, since each image to be retrieved generates a global visual feature vector, if the target to be retrieved occupies a small proportion of the overall image to be retrieved, this small target image is difficult to be fully represented in the feature vector, resulting in poor retrieval performance for small targets.
[0019] Furthermore, existing image retrieval methods lack target localization capabilities when retrieving images, meaning they cannot further indicate the specific location of the target within the image being retrieved, resulting in a lack of resolvability of the retrieval results. The target retrieval method provided in this application can achieve both text-based and image-based image retrieval. By pre-vectorizing and storing massive amounts of data in a vector database, candidate targets can be quickly retrieved from the vector database using text descriptions and reference images. Simultaneously, the specific coordinate information of the candidate targets within the image is obtained, thereby improving the retrieval precision, retrieval capability, and retrieval efficiency.
[0020] To address the problems in the prior art, embodiments of this application provide a target retrieval method, apparatus, terminal, and computer storage medium.
[0021] The target retrieval method provided in the embodiments of this application will be introduced first below.
[0022] Figure 1 A flowchart illustrating a target retrieval method according to an embodiment of this application is shown. Figure 1 As shown, the method may include the following steps: S101: The first visual encoder performs candidate target recognition processing on the image to be entered into the database, and obtains the first visual feature vectors of multiple candidate targets in the image to be entered into the database and the coordinate information of the candidate targets in the image to be entered into the database.
[0023] The first visual encoder calculates the first visual feature vector based on the scalar of the candidate target and the second visual encoder outputting the second visual feature vector of the candidate target. The scalar is the vector adjustment coefficient obtained by the multilayer perceptron in the vector matching training of the candidate target.
[0024] Specifically, the image to be added to the database is input to the first visual encoder for candidate target recognition processing. The first visual encoder can call the trained multilayer perceptron (MLP). The MLP uses the second visual feature vector extracted by the second visual encoder in the previous stage for the same candidate target as input to obtain the vector adjustment coefficient, i.e., the scalar, corresponding to the candidate target.
[0025] Among them, the scalar is learned in large-scale vector matching training and is used to adjust the expressive power of the feature vector and enhance the discriminative power of the feature vector, so that the second visual feature vector is transformed into the first visual feature vector that can be directly used for similarity calculation in the vector database after adaptation.
[0026] This application embodiment adjusts the second visual feature vector by using a scalar, embedding confidence information into the first visual feature vector generated by the first visual encoder. In the subsequent similarity calculation process, there is no need to calculate the confidence value separately, reducing the amount of computation. At the same time, the scalar enhances the expressive power of the first visual feature vector of small targets, improving the accuracy of retrieval.
[0027] S102: Associate and store the candidate target, its coordinate information, and the first visual feature vector in the vector database.
[0028] Specifically, the image to be added to the database is processed by the candidate target recognition of the first visual encoder into multiple data units. Each data unit includes the candidate target and the corresponding first visual feature vector and coordinate information. The data units are then associated and stored in the vector database.
[0029] The associated storage data may also include other information corresponding to the candidate target, such as the image path. Associated storage means binding and storing the candidate target's first visual feature vector, coordinate information, and image path to ensure the consistency and integrity of the stored data. The first visual feature vector can be used as an index for storage.
[0030] Vector databases are used to store and retrieve feature vectors, which can be used for similarity calculation. This allows for direct calculation of similarity within the vector database, simplifying the calculation process and reducing computational complexity.
[0031] This application embodiment, through associated storage, can directly determine the complete information of candidate targets during the retrieval process. At the same time, by storing candidate targets and their corresponding coordinate information and first visual feature vectors in a vector database, the retrieval can be completed directly in the vector database through the efficient indexing mechanism supported by the vector database, shortening the retrieval time and improving the retrieval efficiency.
[0032] S103: Receive a retrieval instruction and generate a retrieval vector containing extended dimensions from the text description and / or reference image contained in the retrieval instruction using a second visual encoder.
[0033] Specifically, upon receiving a retrieval instruction, a second visual encoder is used to generate a retrieval vector including extended dimensions from the text description and / or reference image in the retrieval instruction. For the text description, a text feature vector can be obtained through a text encoder corresponding to the second visual encoder. For the reference image, a reference visual feature vector is obtained through the second visual encoder. The extended dimensions can be obtained by appending a preset constant value to the end of the text feature vector or the reference visual feature vector, thus ensuring that the dimension of the retrieval vector is consistent with the dimension of the first visual feature vector.
[0034] This application embodiment generates a retrieval vector that includes extended dimensions, making the dimensions of the retrieval vector the same as the dimensions of the first visual feature vector. This allows for direct similarity calculation with the first visual feature vector in the vector database during the retrieval process, improving retrieval efficiency and ensuring the accuracy of retrieval results.
[0035] S104: Perform an inner product calculation between the retrieval vector and the first visual feature vectors of each candidate target in the vector database to retrieve at least one first visual feature vector whose similarity to the retrieval vector exceeds a preset similarity threshold.
[0036] Specifically, the retrieval vector is submitted to the vector database, and the inner product of the retrieval vector and the first visual feature vector of each candidate target in the vector database is calculated. The value of the inner product represents the similarity between the retrieval vector and the first visual feature vector of each candidate target.
[0037] Based on the value calculated by the inner product, at least one first visual feature vector with a similarity exceeding a preset similarity threshold is retrieved. This can be done by sorting the values calculated by the inner product and selecting at least one first visual feature vector that exceeds the preset similarity threshold or is ranked first.
[0038] This application embodiment obtains a first visual feature vector by extracting and modulating the second visual feature vector of the candidate target, and stores the first visual feature vector in a vector database before retrieval. Thus, after the retrieval vector is determined, similarity calculation can be performed directly in the vector database based on the retrieval vector and the first visual feature vector, reducing computational complexity and improving retrieval efficiency.
[0039] S105: The retrieved first visual feature vector is associated with the stored candidate target and coordinate information to determine the retrieval result.
[0040] Specifically, the vector database searches for corresponding candidate targets and coordinate information based on the retrieved first visual feature vector, and then determines the candidate targets and coordinate information as the search results.
[0041] Once the search results are determined, candidate targets can be marked in the corresponding image to be added to the database according to the coordinate information, and displayed as "In the current image, the search results are located within the coordinate range of (x1, y1, x2, y2)".
[0042] This application embodiment obtains search results by returning complete information of candidate targets, which enhances the flexibility and practicality of target retrieval and improves user experience.
[0043] This application embodiment obtains a first visual encoder by adjusting the second visual feature vector with a scalar, and obtains a first visual feature vector based on the first visual encoder. This allows the first visual feature vector to have prior content, transferring the computation in the retrieval process to the process of obtaining the first visual feature vector. This simplifies the computation in the retrieval process, enables real-time response under massive data scale, improves computational efficiency, and supports text description retrieval and reference image retrieval methods. The obtained retrieval results include coordinate information, improving the interpretability of the retrieval results.
[0044] In some embodiments, the scalar includes a first scalar and a second scalar. The first visual feature vector is an extended dimension vector formed by multiplying the second visual feature vector corresponding to the second visual encoder by the first scalar and then concatenating the product of the first scalar and the second scalar. The first scalar is a scaling factor used to adjust the scaling ratio of each candidate target in the inner product calculation, and the second scalar is a bias factor used to adjust the offset of each candidate target in the inner product calculation.
[0045] Specifically, the scalar includes a first scalar and a second scalar. The first scalar is a scaling factor obtained through the multilayer perceptron, used to adjust the scaling ratio of each candidate target in the inner product calculation. The second scalar is a bias factor obtained through the multilayer perceptron, used to adjust the offset of each candidate target in the inner product calculation. Each candidate target has a corresponding first scalar and a second scalar, so that the similarity obtained through the second visual feature vector is consistent with the similarity obtained through the first visual feature vector during the similarity calculation process.
[0046] The first visual feature vector is obtained through a specific mathematical transformation. The second visual feature vector is multiplied by the first scalar to scale the second visual feature vector. Then, the product of the first and second scalars is determined. The scaled second visual feature vector is then concatenated with the above product to obtain an extended dimension vector with increased dimensions, which is the first visual feature vector.
[0047] For example, the open set target detection algorithm OWL-ViT is used to process the images to be added to the database. That is, the second visual encoder is used to process the images to be added to the database to obtain the second visual feature vector. In this case, an image to be added to the database contains 576 candidate targets, and the coordinate information of each candidate target in the image to be added to the database is determined. The second visual feature vector is a 512-dimensional vector.
[0048] In the retrieval process using the second visual feature vector, i.e., the similarity with the text description or reference image, the second visual feature vector of each candidate target is multiplied by the text feature vector or reference visual feature vector. Then, a bias coefficient is added to the result of the inner product, and then multiplied by the scaling factor to obtain the similarity between the candidate target and the text description or reference image. This process can be expressed mathematically as follows: Where y represents similarity, μ represents the first scalar (scaling ratio), q represents the text feature vector or reference image vector, v represents the second visual feature vector, and β represents the second scalar (bias coefficient). The inner product operation is represented by β=f1(v) and μ=f2(v), where f1 and f2 represent the multilayer perceptron functions.
[0049] As can be seen from the above formula, in the process of calculating similarity through the second visual feature vector, the inner product between the second visual feature vector and the reference visual feature vector / text feature vector is not directly calculated. Instead, it requires mathematical operations of addition and multiplication. That is, it is not possible to achieve fast target retrieval by directly performing inner product operations between the reference visual feature vector / text feature vector submitted to the vector database and the second visual feature vector already existing in the vector database during the calculation process.
[0050] Therefore, by transforming the above formula, we can obtain: ; in, This means that by appending a scalar 1 to the end of the 512-dimensional vector q represented by the reference visual feature vector / text feature vector, a 513-dimensional vector is obtained. The first visual feature vector is calculated by multiplying the 512-dimensional vector represented by the second visual feature vector by a first scalar, and then concatenating the product of the first and second scalars to obtain a 513-dimensional vector containing extended dimensions, which is the first visual feature vector.
[0051] As can be seen from the above transformation formula, under the premise of keeping its mathematical logic unchanged, the similarity between the candidate target and the input text or reference image can be obtained by calculating the similarity between two 513-dimensional vectors. The similarity calculation result y is completely consistent with that before the transformation, which makes it convenient to use a vector database to implement similarity calculation.
[0052] This application embodiment learns the scaling ratio and bias coefficient in scalar form and embeds them in advance into the first visual feature vector, which simplifies the original conditional similarity calculation process into a standard dot product between two extended vectors, thereby improving the computational efficiency of the retrieval stage.
[0053] In some embodiments, before storing the candidate target and its coordinate information, along with the first visual feature vector, in a vector database, the method further includes: Classify each candidate target and determine a target category corresponding to each candidate target from multiple preset target categories; For multiple candidate targets of the same target category, redundant candidate targets are removed from the multiple candidate targets of the same target category according to a preset redundancy removal method.
[0054] Specifically, before the candidate target, its corresponding first visual feature vector, and coordinate information are associated and stored in the vector database, the candidate target can be optimized, thereby reducing the amount of data stored in the vector database and improving the quality of the data stored in the vector database.
[0055] Each candidate target is classified to determine its semantic affiliation. From multiple preset target categories, a target category corresponding to each candidate target is determined. The target category can be a common general target category, such as animal, person, vehicle, plant, tree, fruit, furniture, beverage, machinery, office supplies, vessel, tool, electrical equipment, equipment, dress, jewelry, traffic sign, building, etc. The above target categories can be narrowed or expanded according to different application scenarios.
[0056] in addition, Figure 2 This is a schematic diagram of candidate targets included in an image to be added to the database, as provided in an embodiment of this application. Figure 2 As shown, since the object detection model generates a large number of overlapping or highly similar candidate targets, if all candidate targets are associated and stored in the vector database, it will lead to serious data redundancy in the vector database, wasting storage space and introducing a large number of repeated or invalid similarity calculations during the retrieval process, causing a large amount of retrieval delay. The candidate targets in each target category can be simplified by removing redundancy.
[0057] For multiple candidate targets of the same target category, redundant candidate targets are removed from the multiple candidate targets of the same target category according to a preset redundancy removal method. The preset redundancy removal method is a screening mechanism based on overlap.
[0058] This application embodiment improves the semantic level of detection results by determining the target category of each candidate target, and at the same time provides a basis for preset redundancy removal methods. By using preset redundancy removal methods, data redundancy is reduced, storage costs and retrieval latency are reduced, the problem of the same target being detected multiple times is avoided, and the quality of retrieval results and user experience are improved.
[0059] In some embodiments, classifying candidate targets and determining a target category corresponding to each candidate target from a plurality of preset target categories includes: The text of multiple pre-defined target categories is input into the text encoder to obtain the text feature vector corresponding to each target category; The matching degree between the first visual feature vector corresponding to the candidate target and each text feature vector is determined, so as to determine a target category corresponding to the candidate target based on the maximum matching degree.
[0060] Specifically, multiple preset target categories and corresponding category texts are obtained. The category texts are then input into the text encoder to obtain the text feature vectors corresponding to the target categories. The text encoder and the second visual encoder share the same multimodal semantic space, which can convert abstract category texts into text feature vectors with explicit mathematical representations.
[0061] For each candidate target, the matching degree between the first visual feature vector and the text feature vector is determined. The matching degree can be calculated by vector inner product or cosine similarity.
[0062] Iterate through the matching degree between the first visual feature vector and each text feature vector corresponding to the current candidate target, and determine the target category corresponding to the maximum matching degree as the target category corresponding to the current candidate target.
[0063] For example, after determining the application scenario, a preset target category is obtained, and each candidate target is classified. For each candidate target in the current image to be added to the database, the first visual feature vector of the candidate target and the matching degree corresponding to each target category are determined. If it is less than the preset category matching degree threshold, the candidate target can be discarded. The preset category matching degree threshold can be set to 0.1. Setting it to a small value can avoid filtering out candidate targets belonging to the target category.
[0064] This application embodiment improves classification accuracy by inputting the category text of the preset target category into a text encoder to generate a text feature vector, and determines the target category corresponding to each candidate target based on the matching degree. At the same time, it provides accurate input for the subsequent preset redundancy removal method, and also provides support for semantics in the retrieval process, thereby improving overall performance and user experience.
[0065] In some embodiments, for multiple candidate targets of the same target category, redundant candidate targets are removed from the multiple candidate targets of the same target category according to a preset redundancy removal method, including: For the same target category, the candidate target with the highest matching degree in the target category is determined as the first benchmark candidate target; Determine the candidate intersection-union ratio (CIU) of the first baseline candidate target and other candidate targets in the current target category; If the candidate crossover ratio exceeds a preset candidate crossover ratio threshold, other candidate targets used for comparison are removed. If the candidate crossover ratio does not exceed the preset candidate crossover ratio threshold, other candidate targets used for comparison are marked as targets to be screened, and other candidate targets with the highest matching degree among the targets to be screened are determined as the second benchmark candidate targets; Repeat the steps of determining the candidate intersection-union ratio of the first baseline candidate target and other candidate targets in the current target category until redundancy removal is completed for all candidate targets in the target category.
[0066] Specifically, for the same target category, all candidate targets included in that target category are traversed, and the candidate target with the highest matching degree is determined as the first baseline candidate target, where the matching degree represents the similarity between the candidate target and the text features of its category.
[0067] Determine the candidate intersection-union ratio (CIU) between the first baseline candidate target and other candidate targets in the current target category, i.e., determine the degree of overlap between the first baseline candidate target and other candidate targets. If the CIU exceeds a preset CIU threshold, remove other candidate targets that are being compared with the current first baseline candidate target. This is because if the CIU exceeds the preset CIU threshold, it means that the other candidate targets and the first baseline candidate target almost represent the same target. Therefore, retain the one with the higher matching degree, i.e., retain the first baseline candidate target, and treat the other candidate targets as redundant and remove them.
[0068] If the candidate crossover ratio (CWR) does not exceed the preset candidate CWR threshold, it indicates that the other candidate targets used for comparison are spatially independent and may correspond to different targets. These are recorded as targets to be screened. After completing the screening based on the first benchmark candidate target, the target with the highest matching degree among the targets to be screened is determined as the second benchmark candidate target. The second benchmark candidate target is used as the new benchmark, and the above process is repeated. That is, the candidate CWR of the second benchmark candidate target and other candidate targets is determined, and the other candidate target is judged and removed or retained according to the preset candidate CWR threshold.
[0069] Repeat the above process. In each iteration, a portion of redundant candidate annual targets that overlap with the baseline candidate targets of the current iteration will be removed, and the baseline candidate targets and targets to be screened for the next iteration will be selected. This process continues until all candidate targets in the current target category have been processed, that is, all candidate targets have been removed or become the baseline candidate targets for a certain iteration. These remaining candidate targets are those that do not overlap highly in space and have a high degree of matching with the current target category. The redundancy removal process can eliminate spatial redundancy to the greatest extent.
[0070] This application embodiment removes redundancy from multiple candidate targets of the same target category, reducing data redundancy while improving the accuracy of search results. After redundancy removal, the amount of data stored in the vector database is reduced, lowering storage costs and search latency, and improving overall performance and user experience.
[0071] In some embodiments, receiving a retrieval instruction and generating a retrieval vector containing extended dimensions from the text description and / or reference image contained in the retrieval instruction via a second visual encoder includes: When the retrieval instruction includes a reference image, the reference image is input into the second visual encoder to obtain multiple reference candidate targets and the reference visual feature vectors corresponding to the reference candidate targets; Receive the reference bounding box image in the reference image, and among all the reference candidate targets, determine the reference candidate target corresponding to the maximum reference intersection-union ratio with the reference bounding box image as the standard candidate target, and determine the reference visual feature vector corresponding to the standard candidate target as the initial retrieval vector; The initial retrieval vector is expanded by extending the dimensions to generate a new retrieval vector.
[0072] Specifically, when the retrieval instruction includes a reference image, the reference image is input into the second visual encoder to obtain multiple reference candidate targets in the reference image and the reference visual feature vectors corresponding to the reference candidate targets.
[0073] Additionally, a reference bounding box image is received from the reference image. The reference bounding box image can be generated by interactively selecting a bounding box in the reference image to clearly mark the target area to be searched. If no interaction is received, the entire reference image is regarded as the reference bounding box image.
[0074] Among all the reference candidate targets, the reference intersection-union ratio with the reference bounding box image is determined, and the reference candidate target with the largest reference intersection-union ratio is determined as the standard candidate target. The reference visual feature vector associated with the standard candidate target is extracted as the initial retrieval vector, that is, the initial retrieval vector is obtained by the second visual encoder.
[0075] Furthermore, in order to enable the vector database to directly perform inner product operations based on the retrieval vector and the first visual feature vector, the initial retrieval vector is expanded by extending its dimensions. This makes the retrieval vector have the same dimension as the first visual feature vector in the vector database, facilitating the input of the retrieval vector into the vector database and the inner product operation with the first visual feature vector to obtain the retrieval result.
[0076] The reference frame image refers to a region on the reference image that the user specifies in order to clearly search for the target. It can be a rectangle, a blurred area, an outline, or a click pointing to a point in the image.
[0077] Figure 3 This is a schematic diagram of a reference frame image provided in an embodiment of this application, such as... Figure 3 As shown, the reference image includes a reference bounding box image. The reference image is processed by a second visual encoder to obtain 576 candidate targets. Each candidate target corresponds to a 512-dimensional feature vector and a reference bounding box image in the reference image. The reference intersection-union ratio (CIU) of the above 576 candidate targets and the reference bounding box image is calculated. The reference visual feature vector of the candidate target corresponding to the largest CIU is used as the initial retrieval vector of the reference bounding box image. Then, a scalar 1 is concatenated to the initial retrieval vector to form a 513-dimensional feature vector, i.e., the retrieval vector. The retrieval vector is input into the vector database for retrieval to obtain the retrieval results.
[0078] In this embodiment, an initial retrieval vector is obtained through a reference image and a reference bounding box image. The initial retrieval vector is then expanded into a retrieval vector by extending the dimension. This allows the retrieval vector to be directly multiplied by the first visual feature vector after being input into the vector database, thereby obtaining the retrieval target and improving retrieval accuracy and efficiency.
[0079] In some embodiments, a retrieval vector containing extended dimensions is generated from the text description and / or reference image contained in the retrieval instruction by a second visual encoder, including: When the retrieval instruction includes a text description, the text description is input into the text encoder to obtain the reference text feature vector; The retrieval vector is generated by expanding the feature vector of the reference text through dimensional expansion.
[0080] Specifically, the text description is input into the text encoder to obtain the reference text feature vector. The dimension of the reference text feature vector is different from the dimension of the first visual feature in the vector database. Therefore, the reference text feature vector is expanded by extending the dimension to generate the retrieval vector, so that the feature dimension of the retrieval vector is the same as the feature dimension of the first visual feature vector.
[0081] The expansion dimension can be 1.
[0082] For example, if the reference text feature dimension is 512, a 1 is appended to it to form a 513-dimensional retrieval vector.
[0083] This application embodiment obtains an initial retrieval vector through text description, and expands it into a retrieval vector by extending the dimension reference text feature vector. This allows the retrieval vector to be directly input into the vector database and then directly perform inner product calculation with the first visual feature vector to obtain the retrieval target, thereby improving retrieval accuracy and efficiency.
[0084] Figure 4This is another target retrieval flowchart provided in the embodiments of this application, such as... Figure 4 As shown, the image to be entered into the database is input into the decoupled visual encoder, namely the first visual encoder. Decoupling means decoupling the feature vector of the candidate target and the feature vector to be retrieved, thereby obtaining the feature vector and position information corresponding to multiple (N) candidate targets. The candidate targets are then filtered by a preset redundancy removal method to obtain the feature vector and position information corresponding to more than N (M) candidate targets, and these are associated and stored in the vector database.
[0085] During the retrieval phase, target retrieval can be performed using the retrieved image or the retrieved text. If the submitted image is received, it is input into the visual encoder, i.e., the second visual encoder, to obtain multiple reference candidate targets and corresponding reference image feature vectors. The optimal target matching process is then performed. The optimal target matching process includes: if there is a reference bounding box image in the retrieved image, determining the reference intersection-union ratio (CIU) between the reference candidate target and the reference bounding box image, and determining the reference candidate target with the largest CIU as the optimal target.
[0086] The feature vector of the reference image corresponding to the optimal target is expanded by extending the dimension, and the expanded reference image feature vector is determined as the visual vector. The visual vector is input into the vector database, and the inner product is calculated between the feature vector and the visual vector corresponding to the candidate target to retrieve the retrieval results. The retrieval results are the image paths and the position information of the targets in the images corresponding to the top k (top-k) targets. The images are the images to be entered into the database.
[0087] When a text is submitted for retrieval, it is input into a text editor to obtain an initial text vector. The initial text vector is then expanded by extending the dimension to obtain a new text vector. This text vector is then input into a vector database. The inner product of the feature vectors corresponding to the candidate targets and the text vector is calculated to retrieve the retrieval results. The retrieval results include the image paths corresponding to the top-k targets and the location information of the targets in the images. The images are the images to be entered into the database.
[0088] Figure 5 This is a schematic diagram of a device structure provided in an embodiment of this application. Figure 5 As shown, the device may include an expansion module 510, a storage module 520, a receiving module 530, a retrieval module 540, and a result module 550.
[0089] The extension module 510 is used to perform candidate target recognition processing on the image to be entered into the database through the first visual encoder, and to obtain the first visual feature vectors of multiple candidate targets in the image to be entered into the database and the coordinate information of the candidate targets in the image to be entered into the database. The first visual encoder calculates the first visual feature vector based on the scalar of the candidate target and the second visual feature vector output by the second visual encoder of the candidate target. The scalar is the vector adjustment coefficient obtained by the multilayer perceptron in the vector matching training of the candidate target. Storage module 520 is used to associate and store candidate targets and their coordinate information, as well as the first visual feature vector, into a vector database; The receiving module 530 is used to receive a retrieval instruction and generate a retrieval vector containing extended dimensions from the text description and / or reference image contained in the retrieval instruction through a second visual encoder. The retrieval module 540 is used to perform an inner product calculation between the retrieval vector and the first visual feature vectors of each candidate target in the vector database, so as to retrieve at least one first visual feature vector whose similarity to the retrieval vector exceeds a preset similarity threshold. The result module 550 is used to associate the retrieved first visual feature vector with the stored candidate target and coordinate information to determine the retrieval result.
[0090] In some embodiments, the scalar includes a first scalar and a second scalar. The first visual feature vector is an extended dimension vector formed by multiplying the second visual feature vector corresponding to the second visual encoder by the first scalar and then concatenating the product of the first scalar and the second scalar. The first scalar is a scaling factor used to adjust the scaling ratio of each candidate target in the inner product calculation, and the second scalar is a bias factor used to adjust the offset of each candidate target in the inner product calculation.
[0091] In some embodiments, before storing the candidate target and its coordinate information along with the first visual feature vector in the vector database, the storage module 520 is further configured to: Classify each candidate target and determine a target category corresponding to each candidate target from multiple preset target categories; For multiple candidate targets of the same target category, redundant candidate targets are removed from the multiple candidate targets of the same target category according to a preset redundancy removal method.
[0092] In some embodiments, the storage module 520 classifies each candidate target and determines a target category corresponding to each candidate target from a plurality of preset target categories, for the purpose of: The text of multiple pre-defined target categories is input into the text encoder to obtain the text feature vector corresponding to each target category; The matching degree between the first visual feature vector corresponding to the candidate target and each text feature vector is determined, so as to determine a target category corresponding to the candidate target based on the maximum matching degree.
[0093] In some embodiments, the storage module 520 removes redundant candidate targets from multiple candidate targets of the same target category according to a preset redundancy removal method, for the following purposes: For the same target category, the candidate target with the highest matching degree in the target category is determined as the first benchmark candidate target; Determine the candidate intersection-union ratio (CIU) of the first baseline candidate target and other candidate targets in the current target category; If the candidate crossover ratio exceeds a preset candidate crossover ratio threshold, other candidate targets used for comparison are removed. If the candidate crossover ratio does not exceed the preset candidate crossover ratio threshold, other candidate targets used for comparison are marked as targets to be screened, and other candidate targets with the highest matching degree among the targets to be screened are determined as the second benchmark candidate targets; Repeat the steps of determining the candidate intersection-union ratio of the first baseline candidate target and other candidate targets in the current target category until redundancy removal is completed for all candidate targets in the target category.
[0094] In some embodiments, the receiving module 530 receives a retrieval instruction and generates a retrieval vector containing extended dimensions from the text description and / or reference image contained in the retrieval instruction using a second visual encoder, for: When the retrieval instruction includes a reference image, the reference image is input into the second visual encoder to obtain multiple reference candidate targets and the reference visual feature vectors corresponding to the reference candidate targets; Receive the reference bounding box image in the reference image, and among all the reference candidate targets, determine the reference candidate target corresponding to the maximum reference intersection-union ratio with the reference bounding box image as the standard candidate target, and determine the reference visual feature vector corresponding to the standard candidate target as the initial retrieval vector; The initial retrieval vector is expanded by extending the dimensions to generate a new retrieval vector.
[0095] In some embodiments, the receiving module 530 receives a retrieval instruction and generates a retrieval vector containing extended dimensions from the text description and / or reference image contained in the retrieval instruction using a second visual encoder, for: When the retrieval instruction includes a text description, the text description is input into the text encoder to obtain the reference text feature vector; The retrieval vector is generated by expanding the feature vector of the reference text through dimensional expansion.
[0096] Figure 6A schematic diagram of the hardware structure of the terminal device provided in an embodiment of this application is shown.
[0097] The terminal device may include a processor 601 and a memory 602 storing computer program instructions.
[0098] Specifically, the processor 601 may include a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0099] Memory 602 may include mass storage for data or instructions. For example, and not limitingly, memory 602 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. In one instance, memory 602 may include removable or non-removable (or fixed) media, or memory 602 may be non-volatile solid-state memory. Memory 602 may be internal or external to the integrated gateway disaster recovery device.
[0100] In one instance, memory 602 may be read-only memory (ROM). In one instance, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or flash memory, or a combination of two or more of these.
[0101] Memory 602 may include read-only memory (ROM), random access memory (RAM), disk storage media device, optical storage media device, flash memory device, electrical, optical, or other physical / tangible memory storage device. Therefore, generally, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to the first aspect of this disclosure.
[0102] The processor 601 reads and executes computer program instructions stored in the memory 602 to achieve... Figure 1 The target retrieval method in the illustrated embodiment.
[0103] In one example, the terminal device may further include a communication interface 603 and a bus 604. Wherein, for example... Figure 6 As shown, the processor 601, memory 602, and communication interface 603 are connected through bus 604 and complete communication with each other.
[0104] The communication interface 603 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0105] Bus 604 includes hardware, software, or both, that couples components of an online data traffic metering device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 604 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, this application contemplates any suitable bus or interconnect.
[0106] Furthermore, in conjunction with the target retrieval methods in the above embodiments, this application embodiment can provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the target retrieval methods in the above embodiments.
[0107] This application also provides a computer program product, including a computer program that, when executed by a processor, implements any of the target retrieval methods described in the above embodiments.
[0108] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0109] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, read-only memory (ROM), flash memory, erasable read-only memory (EROM), floppy disks, compact disc read-only memory (CD-ROM), optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0110] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0111] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.
[0112] The above description is merely a specific implementation of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.
Claims
1. A target retrieval method, characterized in that, include: The first visual encoder performs candidate target recognition processing on the image to be entered into the database, and obtains the first visual feature vector of multiple candidate targets in the image to be entered into the database and the coordinate information of the candidate targets in the image to be entered into the database. The first visual encoder calculates the first visual feature vector based on the scalar of the candidate target and the second visual feature vector output by the second visual encoder for the candidate target. The scalar is the vector adjustment coefficient obtained by the multilayer perceptron in the vector matching training of the candidate target. The candidate target, its coordinate information, and the first visual feature vector are associated and stored in a vector database; Upon receiving a retrieval instruction, the second visual encoder generates a retrieval vector containing extended dimensions from the text description and / or reference image contained in the retrieval instruction. The inner product of the retrieval vector and the first visual feature vectors of each candidate target in the vector database is calculated to retrieve at least one first visual feature vector whose similarity to the retrieval vector exceeds a preset similarity threshold. The retrieved first visual feature vector is associated with the stored candidate target and the coordinate information to determine the retrieval result.
2. The method according to claim 1, characterized in that, The scalar includes a first scalar and a second scalar. The first visual feature vector is an extended dimension vector formed by multiplying the second visual feature vector corresponding to the second visual encoder by the first scalar and then concatenating the product of the first scalar and the second scalar. The first scalar is a scaling factor used to adjust the scaling ratio of each candidate target in the inner product calculation, and the second scalar is a bias factor used to adjust the offset of each candidate target in the inner product calculation.
3. The method according to claim 1, characterized in that, Before associating and storing the candidate target, its coordinate information, and the first visual feature vector in the vector database, the method further includes: The candidate targets are classified, and a target category corresponding to each candidate target is determined from a plurality of preset target categories; For multiple candidate targets of the same target category, redundant candidate targets are removed from the multiple candidate targets of the same target category according to a preset redundancy removal method.
4. The method according to claim 3, characterized in that, The step of classifying the candidate targets, and determining a target category corresponding to each candidate target from a preset plurality of target categories, includes: The text of multiple preset target categories is input into the text encoder to obtain the text feature vector corresponding to each target category; The matching degree between the first visual feature vector corresponding to the candidate target and each of the text feature vectors is determined, so as to determine a target category corresponding to the candidate target based on the maximum matching degree.
5. The method according to claim 4, characterized in that, The process of removing redundant candidate targets from multiple candidate targets of the same target category according to a preset redundancy removal method includes: For the same target category, the candidate target with the highest matching degree in the target category is determined as the first benchmark candidate target; Determine the candidate intersection-union ratio (CIU) of the first benchmark candidate target and other candidate targets in the current target category; If the candidate crossover ratio exceeds a preset candidate crossover ratio threshold, remove the other candidate targets used for comparison. If the candidate crossover ratio does not exceed the preset candidate crossover ratio threshold, the other candidate targets used for comparison are marked as targets to be screened, and the other candidate targets with the highest matching degree among the targets to be screened are determined as second benchmark candidate targets; Repeat the step of determining the candidate intersection-union ratio of the first benchmark candidate target and other candidate targets in the current target category until redundancy removal of all candidate targets in the target category is completed.
6. The method according to claim 1, characterized in that, The step of receiving a retrieval instruction, and generating a retrieval vector with extended dimensions from the text description and / or reference image contained in the retrieval instruction using the second visual encoder, includes: When the retrieval instruction includes the reference image, the reference image is input to the second visual encoder to obtain multiple reference candidate targets and reference visual feature vectors corresponding to the reference candidate targets; Receive the reference bounding box image in the reference image, and among all the reference candidate targets, determine the reference candidate target corresponding to the maximum reference intersection-union ratio of the reference bounding box image as the standard candidate target, and determine the reference visual feature vector corresponding to the standard candidate target as the initial retrieval vector; The initial retrieval vector is expanded by extending its dimensions to generate a retrieval vector.
7. The method according to claim 1, characterized in that, Receiving a retrieval instruction, the step of generating a retrieval vector with extended dimensions from the text description and / or reference image contained in the retrieval instruction through the second visual encoder includes: If the search instruction includes the text description, the text description is input into the text encoder to obtain the reference text feature vector; The reference text feature vector is expanded by extending its dimensions to generate a retrieval vector.
8. A target retrieval device, characterized in that, The device includes: An extension module is used to perform candidate target recognition processing on the image to be entered into the database through a first visual encoder, to obtain the first visual feature vectors of multiple candidate targets in the image to be entered into the database and the coordinate information of the candidate targets in the image to be entered into the database. The first visual encoder calculates the first visual feature vector based on the scalar of the candidate target and the second visual feature vector output by the second visual encoder for the candidate target. The scalar is the vector adjustment coefficient obtained by the multilayer perceptron in the vector matching training of the candidate target. A storage module is used to associate and store the candidate target, the coordinate information of the candidate target, and the first visual feature vector into a vector database; The receiving module is used to receive a retrieval instruction and generate a retrieval vector containing extended dimensions from the text description and / or reference image contained in the retrieval instruction through the second visual encoder. The retrieval module is used to perform an inner product calculation between the retrieval vector and the first visual feature vectors of each candidate target in the vector database, so as to retrieve at least one first visual feature vector whose similarity to the retrieval vector exceeds a preset similarity threshold. The results module is used to associate the retrieved first visual feature vector with the stored candidate target and the coordinate information to determine the retrieval result.
9. A terminal device, characterized in that, The device includes: a processor and a memory storing computer program instructions; the processor reads and executes the computer program instructions to implement the target retrieval method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer storage medium stores computer program instructions, which, when executed by a processor, implement the target retrieval method as described in any one of claims 1 to 7.