Store image retrieval method, apparatus and device, and storage medium

By performing object recognition, image cropping and feature extraction on store images, combined with database retrieval and similarity calculation, the misidentification problem caused by chain stores is solved, and the accuracy and security of store image retrieval is improved.

CN119992558APending Publication Date: 2025-05-13SHENZHEN XIAOYUDIAN DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411811984.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the search for store images, the current technology has a high misidentification rate due to the similarity of chain stores, which affects the search accuracy and increases the risk of fraud.

Method used

By identifying the target image, extracting signboard text and non-signature text, cropping the image, extracting vectors and local features respectively, combining database retrieval and similarity calculation, screening candidate images.

Benefits of technology

It improves the retrieval accuracy of store images, reduces mismatch caused by high image similarity, and enhances the identity recognition ability of stores.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992558A_ABST
    Figure CN119992558A_ABST
Patent Text Reader

Abstract

The invention relates to a store image retrieval method and device, equipment and a storage medium, and the method comprises the steps: carrying out the target recognition, image cutting and local feature extraction of a target image, obtaining a signboard character, a non-signboard character, a target sub-image and the local feature of the target image, carrying out the vector extraction of the target sub-image and the target image, and obtaining the local feature of the target image; obtaining a target sub-image vector and a target image vector, and retrieving a database based on the signboard characters, the non-signboard characters and the target image vector to obtain a candidate image set; calculating a similarity score of the target image and each candidate image based on the target image vector, the target sub-image vector and the signboard characters, and performing local feature matching on the target image and each candidate image based on local features of the target image to obtain a matching feature number; obtaining similar images of the target image in the candidate image set based on the similarity score and the matching feature number; compared with the prior art, the technical scheme of the invention can improve the retrieval precision of the store image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image processing, and in particular to a store image retrieval method, device, equipment and storage medium. Background Art

[0002] During the credit review process for microfinance store business premises, applicants may impersonate other customer store information that has passed the credit review, and then submit business-related information that does not meet the entry requirements for entry, resulting in a certain risk of fraud.

[0003] In order to avoid the situation where an applicant impersonates a customer's store information that has passed the credit review, the store images submitted by the applicant are generally searched using traditional image retrieval methods in the prior art.

[0004] However, due to the existence of chain stores, many stores often have highly similar names and exterior decorations, which leads to a high misrecognition rate when traditional image retrieval methods extract store text for retrieval based on image vector conversion or OCR technology, thereby affecting retrieval accuracy and increasing potential fraud risks. Summary of the invention

[0005] The present application provides a store image retrieval method, device, equipment and storage medium, which can improve the retrieval accuracy of store images.

[0006] In a first aspect, the present application provides a store image retrieval method, which acquires a target image, performs target recognition on the target image, and obtains signboard text and non-signboard text on the target image; performs image cropping on the target image to obtain a target sub-image, performs vector extraction on the target sub-image and the target image respectively to obtain a target sub-image vector and a target image vector, and performs local feature extraction on the target image to obtain local features of the target image; based on the signboard text, the non-signboard text and the target image vector, a pre-constructed database is searched to obtain a candidate image set; based on the target image vector, the target sub-image vector and the signboard text, a similarity score between the target image and each candidate image in the candidate image set is calculated respectively, and based on the local features of the target image, local feature matching processing is performed on the target image and each candidate image in the candidate image set to obtain the number of matching features; based on the similarity score and the number of matching features, the candidate image set is screened to obtain similar images of the target image.

[0007] In a possible implementation, target recognition is performed on the target image to obtain the signboard text and non-signboard text on the target image, specifically including: based on a pre-trained signboard recognition model, performing signboard recognition on the target image to obtain the signboard area coordinates; based on a pre-trained text detection model, performing text detection on the target image to obtain the text area, and the text area coordinates corresponding to the text area; based on the pre-trained text recognition model, performing text recognition on the text area to obtain text data, and the text coordinates corresponding to the text data; based on the signboard area coordinates and the text coordinates, classifying the text data to obtain the signboard text and non-signboard text on the target image.

[0008] In a possible implementation, vector extraction is performed on the target sub-image and the target image respectively to obtain a target sub-image vector and a target image vector, which specifically includes: performing image increment processing on the target sub-image to obtain an incremental target sub-image set, and performing vector extraction on each incremental target sub-image in the target sub-image set to obtain an incremental target sub-image vector corresponding to each incremental target sub-image; performing averaging processing on all incremental target sub-image vectors to obtain an incremental target sub-image mean vector, and using the incremental target sub-image mean vector as the target sub-image vector corresponding to the target sub-image; performing image increment processing on the target image to obtain an incremental target image set, and performing vector extraction on each incremental target image in the target image set to obtain an incremental target image vector corresponding to each incremental target image; performing averaging processing on all incremental target image vectors to obtain an incremental target image mean vector, and using the incremental target image mean vector as the target image vector corresponding to the target image.

[0009] In a possible implementation, local features of the target image are extracted to obtain local features of the target image, specifically including: obtaining a pre-trained DELG model, wherein the DELG model includes a local network and a Resnet model, and the output end of the third-layer feature map of the Resnet model is connected to the input end of the local network; performing image increment processing on the target image to obtain multiple incremental target images, inputting each incremental target image into the DELG model respectively, so that the Resnet model in the DELG model outputs a target feature map, and inputting the target feature map into the local network to obtain a feature score corresponding to each target feature point in the target feature map; normalizing the target feature map to obtain a normalized original feature value, and multiplying the normalized original feature value and the feature score to obtain a feature score corresponding to each target feature point in the target feature map. The method comprises the following steps: determining a feature value corresponding to a target feature point; setting a feature score threshold, performing feature point screening processing on the target feature map based on the feature score threshold, obtaining a target feature point index that meets a preset number of feature points, and segmenting the target feature map based on a preset step size to obtain a plurality of target feature boxes; performing splicing processing on the feature score, the feature value and the target image box corresponding to the target feature point index in each incremental target image to obtain a spliced ​​total feature corresponding to each incremental target image; performing non-maximum suppression processing on the spliced ​​total feature of each incremental target image based on the feature score and the target image box corresponding to each incremental target image to obtain a plurality of retained feature point indexes; obtaining and determining the target image local feature of the target image based on the first feature score, the first feature value and the first target image box corresponding to each of the plurality of retained feature point indexes in each incremental target image.

[0010] In a possible implementation, a pre-constructed database is searched based on the signboard text, the non-signboard text and the target image vector to obtain a candidate image set, which specifically includes: obtaining sample signboard text and sample non-signboard text corresponding to each sample image in the pre-constructed database, and based on the signboard text and the non-signboard text, performing a search process on the sample signboard text and the sample non-signboard text corresponding to each sample image in the database, and obtaining a first candidate image set based on the search result; obtaining a sample image vector corresponding to each sample image in the database, respectively calculating the image similarity between the target image vector and the sample image vector corresponding to each sample image in the database, and determining a second candidate image set based on the image similarity; and integrating the first candidate image set and the second candidate image set to obtain a candidate image set.

[0011] In a possible implementation, the similarity scores between the target image and each candidate image in the candidate image set are calculated based on the target image vector, the target sub-image vector and the signboard text, specifically including: obtaining the candidate image vector, the candidate sub-image vector and the candidate signboard text corresponding to each candidate image in the candidate image set; calculating the first similarity between the image vector of the target image and the selected image vector corresponding to each candidate image in the candidate image set, and calculating the second similarity between the signboard text of the target image and the candidate signboard text corresponding to each candidate image in the candidate image set, and calculating the third similarity between the target sub-image vector of the target image and the candidate sub-image vector corresponding to each candidate image in the candidate image set; weighting the first similarity, the second similarity and the third similarity to obtain the similarity score between the target image and each candidate image in the candidate image set.

[0012] In a possible implementation, based on the local features of the target image, performing local feature matching processing on the target image and each candidate image in the candidate picture set to obtain the number of matching features specifically includes: obtaining the candidate image local features corresponding to each candidate image in the candidate image set; calculating the feature distance between the target image and the candidate image local features corresponding to each candidate image in the candidate picture set, and performing feature point screening processing on the target image and each candidate image in the candidate picture set based on the feature distance to obtain a first local feature point set corresponding to the target image and a second local feature point set corresponding to each candidate image in the candidate picture set; and performing feature point screening processing on the first local feature point set and the second local feature point set. The second local feature point set is filtered to obtain a first filtered local feature point set and a second filtered local feature point set; the first filtered local feature point set and the second filtered local feature point set are classified to obtain the number of target signboard area feature points and the number of target non-signboard area feature points corresponding to the first filtered local feature point set, as well as the number of candidate signboard area feature points and the number of candidate non-signboard area feature points corresponding to the second filtered local feature point set; the number of target signboard area feature points and the number of target non-signboard area feature points corresponding to the target image, as well as the number of candidate signboard area feature points and the number of candidate non-signboard area feature points corresponding to each candidate image in the candidate picture set are taken as the matching feature number.

[0013] In a second aspect, the present application provides a store image retrieval device, including: a target image acquisition module, a target image feature acquisition module, a candidate image retrieval module, an image comparison module and a similar image determination module; wherein the target image acquisition module is used to acquire a target image, perform target recognition on the target image, and obtain the signboard text and non-signboard text on the target image; the target image feature acquisition module is used to perform image cropping on the target image to obtain a target sub-image, perform vector extraction on the target sub-image and the target image respectively to obtain a target sub-image vector and a target image vector, and perform local feature extraction on the target image to obtain a target image local feature; the candidate image retrieval module The module is used to search a pre-constructed database based on the signboard text, the non-signboard text and the target image vector to obtain a candidate image set; the image comparison module is used to calculate the similarity scores between the target image and each candidate image in the candidate image set based on the target image vector, the target sub-image vector and the signboard text, and based on the local features of the target image, perform local feature matching processing on the target image and each candidate image in the candidate image set to obtain the number of matching features; the similar image determination module is used to screen the candidate image set based on the similarity scores and the number of matching features to obtain similar images of the target image.

[0014] In a third aspect, an embodiment of the present application further provides a computer device, which includes a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the above method when executing the computer program.

[0015] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, wherein the storage medium stores a computer program, and the computer program can implement the above method when executed by a processor.

[0016] The present invention provides a method which has the following advantages over the prior art:

[0017] By performing target recognition on the acquired target image, the signboard text and non-signboard text on the target image are obtained; the store name can be fully utilized for identification at the character level. This helps to distinguish stores with similar appearance and decoration but different names, and enhances the ability to distinguish based on text information; the target image is cropped to obtain a target sub-image; by extracting vector information from the target image and the target sub-image respectively, and performing local feature extraction, the specific content and regional features of the image can be effectively identified, which is especially suitable for situations where the appearance of stores is slightly different and difficult to distinguish only through the overall image; based on the signboard text, the non-signboard text and the target image vector, a pre-constructed database is searched to obtain a set of candidate images; and based on the target image vector, the target sub-image vector and the signboard text, the candidate images are divided. The similarity scores between the target image and each candidate image in the candidate image set are calculated separately, and based on the local features of the target image, local feature matching processing is performed on the target image and each candidate image in the candidate image set to obtain the number of matching features; in the scenario of chain stores or stores with similar appearance, it can help reduce false matches caused by high image similarity; finally, based on the similarity scores and the number of matching features, the candidate image set is further screened to ensure that the similar images finally screened out are more in line with the actual situation of the target image; this not only reduces the interference of irrelevant images, but also improves the accuracy of the retrieval system. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0019] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0020] One or more embodiments are exemplarily described by pictures in the corresponding drawings, and these exemplified descriptions do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, and unless otherwise stated, the figures in the drawings do not constitute proportional limitations.

[0021] Figure 1 It is a flowchart of an embodiment of a store image retrieval method provided by the present application;

[0022] Figure 2 It is a structural schematic diagram of an embodiment of a store image retrieval device provided by the present application;

[0023] Figure 3 It is a structural schematic diagram of an electronic device provided by this application. DETAILED DESCRIPTION

[0024] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0025] The disclosure below provides many different embodiments or examples to realize the different structures of the present application. In order to simplify the disclosure of the present application, the parts and settings of specific examples are described below. Of course, they are only examples, and the purpose is not to limit the present application. In addition, the present application can repeat reference numbers and / or letters in different examples. This repetition is for the purpose of simplification and clarity, and does not itself indicate the relationship between the various embodiments and / or settings discussed.

[0026] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprises" indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.

[0027] It should also be understood that the terms used in this application specification are only for the purpose of describing specific embodiments and are not intended to limit the application. As used in this application specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.

[0028] It should be further understood that the term “and / or” used in the specification and appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0029] As used in this specification and the appended claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if [described condition or event] is detected" may be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0030] Example 1, see Figure 1 , Figure 1 is a flow chart of an embodiment of a store image retrieval method provided by the present application, such as Figure 1 As shown, the method includes steps 101 to 105, which are specifically as follows:

[0031] Step 101: Acquire a target image, perform target recognition on the target image, and obtain signboard text and non-signboard text on the target image.

[0032] In one embodiment, based on a pre-trained signboard recognition model, signboard recognition is performed on the target image to obtain the coordinates of the signboard area.

[0033] Specifically, since the YOLOv8 model can recognize objects in an image and give the bounding box coordinates of the object, the sign recognition model is trained based on the YOLOv8 model.

[0034] Specifically, when the signboard recognition model is trained, a batch of sample images containing signs can be obtained, and a labeling tool can be used to label the signboards in the sample images with bounding boxes and category labels to generate labeling information, wherein the labeling information includes the coordinates and category information of the bounding box. The sample image is used as the model output of the signboard recognition model, and the labeling information corresponding to the sample image is used as the model output. The signboard recognition model is trained to obtain a trained signboard recognition model.

[0035] Specifically, after the target image is input into the sign recognition model, the sign recognition model will output the bounding box coordinates of the detected sign, and use the bounding box coordinates as the sign area coordinates, wherein the sign area coordinates include the coordinates of the upper left corner of the sign area and the coordinates of the lower right corner of the sign area.

[0036] In one embodiment, text detection is performed on the target image based on a pre-trained text detection model to obtain a text area and text area coordinates corresponding to the text area.

[0037] Specifically, the text detection model is the open source text detection model PP-OCRv4 in PaddleOCR.

[0038] Specifically, when using a text detection model to detect text on a target image, the target image is preprocessed, including grayscale, binarization, denoising and other operations to improve the accuracy of text detection; using the target image after morphological transformation, a contour detection method is used to locate areas that may contain text, and then the coordinates of the text area are determined by methods such as screening area, contour approximation and minimum circumscribed rectangle. Based on the detected text area coordinates, the text area is cropped out from the preprocessed target image to obtain the text area and the text area coordinates corresponding to the text area.

[0039] In one embodiment, based on a pre-trained text recognition model, text recognition is performed on the text area to obtain text data and text coordinates corresponding to the text data.

[0040] Specifically, the text recognition model is the open source text recognition model SVTRv2 in PaddleOCR.

[0041] Specifically, the text recognition model is used to receive the text area output by the text detection model.

[0042] Specifically, when the text recognition model performs text recognition on the text area, the text area image is divided into small blocks, preliminary features are extracted from each small block, and the features of each small block are mixed, merged and combined; this step is performed in a hierarchical manner, and each level further integrates the features of the previous level. For example, the first level may only process local features of the characters, such as strokes, while the second level begins to integrate the entire character, and the third level begins to integrate the relationship between multiple characters. Subsequently, the features of the entire image are integrated to understand the relative position and relationship between the characters, and refine the details within each character to ensure that the text recognition model can recognize the precise features of a single character; through multi-level processing, the model can not only recognize each character, but also understand the relationship between characters, thereby accurately recognizing the entire text content.

[0043] Specifically, the text recognition model outputs the recognition results of each text area, including the recognized text string and the coordinates of each text character, that is, the text data, and the text coordinates corresponding to the text data. These coordinates can help us understand the position of each character in the original image.

[0044] In one embodiment, based on the sign area coordinates and the text coordinates, the text data is classified and processed to obtain the sign text and non-sign text on the target image.

[0045] Specifically, it is determined whether the text coordinates corresponding to the target text data are located within the sign area coordinates. If so, it is determined that the target text data is the sign text. Otherwise, it is considered that the target text data is non-sign text.

[0046] Step 102: perform image cropping on the target image to obtain a target sub-image, perform vector extraction on the target sub-image and the target image respectively to obtain a target sub-image vector and a target image vector, and perform local feature extraction on the target image to obtain local features of the target image.

[0047] In one embodiment, the coordinates of a signboard area in the target image are obtained, wherein the coordinates of the signboard area include the coordinates of the upper left corner of the signboard area and the coordinates of the lower right corner of the signboard area; based on the coordinates of the upper left corner of the signboard area, the left horizontal coordinate and the left vertical coordinate of the signboard area are determined, and based on the lower right corner of the signboard area, the right horizontal coordinate and the right vertical coordinate of the signboard area are determined; based on the absolute difference between the right vertical coordinate of the signboard area and the left vertical coordinate of the signboard area, the height of the target sub-image is determined.

[0048] In one embodiment, the height of the target sub-image is set as the right vertical coordinate of the target sub-image, and the left vertical coordinate of the target sub-image is set to zero; the left horizontal coordinate of the signboard area and the left vertical coordinate of the target sub-image are set as the lower left coordinate of the target sub-image, and the right horizontal coordinate of the signboard area and the right vertical coordinate of the target sub-image are set as the upper right coordinate of the target sub-image; based on the lower left coordinate of the target sub-image and the upper right coordinate of the target sub-image, the target image is cropped to obtain the target sub-image.

[0049] In one embodiment, image increment processing is performed on the target sub-image to obtain an incremental target sub-image set.

[0050] Specifically, the DOLG model is used to perform image compression and image expansion processing on the target sub-image in a ratio of [0.5, 0.7071, 1.0, 1.4142] to obtain incremental target sub-images of different ratios, and all incremental target sub-images of different ratios are integrated to obtain an incremental target sub-image set.

[0051] The example illustrates that when image compression processing is performed on the target sub-image based on a ratio of 0.5, the image length and image width corresponding to the target sub-image are compressed to 0.5 times of the original, and an incremental target sub-image corresponding to the ratio of 0.5 is obtained; when image expansion processing is performed on the target sub-image based on a ratio of 1.4142, the image length and image width corresponding to the target sub-image are expanded to 1.4142 times of the original, and an incremental target sub-image corresponding to the ratio of 1.4142 is obtained.

[0052] In one embodiment, vector extraction is performed on each incremental target sub-image in the target sub-image set to obtain an incremental target sub-image vector corresponding to each incremental target sub-image.

[0053] Preferably, when performing vector extraction on each incremental target sub-image in the target sub-image set, models trained on large-scale data sets, such as VGG, ResNet, Inception, etc., can also be used. These models have learned rich image features, and the intermediate layer outputs of these models can be used as feature vectors.

[0054] In one embodiment, all incremental target sub-image vectors are averaged to obtain an incremental target sub-image mean vector, and the incremental target sub-image mean vector is used as the target sub-image vector corresponding to the target sub-image.

[0055] In one embodiment, image increment processing is performed on the target image to obtain an incremental target image set, and vector extraction is performed on each incremental target image in the target image set to obtain an incremental target image vector corresponding to each incremental target image, all incremental target image vectors are averaged to obtain an incremental target image mean vector, and the incremental target image mean vector is used as the target image vector corresponding to the target image.

[0056] Specifically, the manner of performing image increment processing and row vector extraction on the target image is the same as the manner of performing image increment processing and vector extraction on the target sub-image described above, and will not be described in detail herein.

[0057] In one embodiment, local features of the target image are extracted based on a pre-trained DELG model to obtain local features of the target image.

[0058] Specifically, a pre-trained DELG model is obtained, wherein the DELG model includes a local network and a Resnet model, and the output end of the third layer feature map of the Resnet model is connected to the input end of the local network; the DELG model is a unified model for local and global image features, which improves the accuracy of image recognition by integrating global features and local features.

[0059] In one embodiment, the target image is subjected to image increment processing to obtain a plurality of incremental target images, and each incremental target image is respectively input into the DELG model so that the Resnet model in the DELG model outputs a target feature map, and the target feature map is input into the local network to obtain a feature score corresponding to each target feature point in the target feature map.

[0060] Specifically, when performing image increment processing on the target image, the same method is used.

[0061] [0.5, 0.7071, 1.0, 1.4142] ratios are used to perform image compression and image expansion processing to obtain incremental target images of different ratios, and then all incremental target images of different ratios are integrated to obtain an incremental target image set.

[0062] Specifically, in the DELG model, the target feature map output by the third layer of ResNet is used as the input of the local network. The local network receives the target feature map of the third layer of ResNet as input, and then further extracts features through the convolution layer. In this process, convolution kernels of different sizes may be used to capture features of different scales. Then, the target feature map is normalized, which is usually to reduce internal covariate shift and improve the generalization ability of the model. The normalization layer calculates the mean and variance of each channel and uses these statistics to adjust the distribution of the feature map. After the normalization step, the feature map passes through the convolution layer again to further abstract and extract features. Finally, the feature map passes through the Softplus activation function. Softplus is a smooth ReLU activation function that can provide nonlinear capabilities to the model while maintaining numerical stability. After the above processing, the local network outputs new feature scores, which represent the importance or significance of each target feature point in the target feature map.

[0063] In one embodiment, the target feature map is normalized to obtain a normalized original feature value, and the normalized original feature value and the feature score are multiplied to obtain a feature value corresponding to each target feature point in the target feature map.

[0064] Specifically, each target feature point in the target feature map is normalized to obtain a normalized original feature value corresponding to each target feature point, and the normalized original feature value corresponding to each target feature point is multiplied by the feature score to obtain a feature value corresponding to each target feature point in the target feature map.

[0065] In one embodiment, a feature score threshold is set, and feature point screening processing is performed on the target feature map based on the feature score threshold to obtain a target feature point index that meets a preset number of feature points, and based on a preset step size, the target feature map is segmented to obtain multiple target feature boxes.

[0066] Specifically, the feature score corresponding to each target feature point in the target feature map is compared with the feature score threshold; if the feature score is greater than the feature score threshold, the target feature points corresponding to the feature scores greater than the feature score threshold are retained; if the feature score is not greater than the feature score threshold, the target feature points corresponding to the feature scores not greater than the feature score threshold are deleted; and the corresponding target feature point indexes are obtained based on the feature point positions of all retained target feature points; it is determined whether the number of target feature point indexes is greater than the preset number of feature points, and if so, the feature score threshold is adjusted, and the target feature map is re-screened based on the adjusted feature score threshold. This process is iterative, that is, the number of target feature point indexes is recalculated each time the threshold is adjusted; if the number of target feature point indexes is still greater than the preset number of feature points, the threshold will continue to be halved until the number of target feature point indexes is less than the preset number of feature points.

[0067] Specifically, the preset number of feature points is 1000, and adjusting the feature score threshold means adjusting the feature score threshold to half of the original value. This is because the retained target feature point index is greater than 1000, indicating that there are too many important feature points, which may affect the efficiency of subsequent processing; therefore, the threshold needs to be adjusted to make it more stringent in order to reduce the number of important feature points.

[0068] Preferably, when the number of target feature point indexes corresponding to the retained target feature points is not greater than the preset number of feature points, it is considered that the preset number of feature points is met.

[0069] Specifically, the preset step size is 16; based on the preset step size, the target feature map 0 is divided into boxes with equal length and width to obtain a plurality of target feature boxes.

[0070] In one embodiment, the feature scores, the feature values ​​and the target image boxes corresponding to the target feature point indexes in each incremental target image are spliced ​​to obtain the spliced ​​total features corresponding to each incremental target image.

[0071] In one embodiment, based on the feature score and the target image box corresponding to each incremental target image, non-maximum suppression processing is performed on the spliced ​​total feature of each incremental target image to obtain multiple retained feature point indexes.

[0072] Specifically, all target feature points are sorted based on the feature scores corresponding to each target feature point in each incremental target image to obtain the target feature point corresponding to the highest feature score, and the intersection-over-union (IoU) of the bounding boxes of the target feature point corresponding to the highest feature score and other target feature points is calculated based on the non-maximum suppression algorithm. If the IoU is greater than a preset IoU threshold, the two target feature points are considered to overlap, and the target feature point with a lower feature score will be suppressed. Based on iterative processing, all target feature points that are not suppressed are obtained, and the first target number of target feature points among all target feature points that are not suppressed are selected to obtain multiple retained feature point indexes; wherein, the target number is 500.

[0073] In one embodiment, the target image local features of the target image are obtained and determined based on the first feature scores, the first feature values ​​and the first target image boxes corresponding to the plurality of retained feature point indexes in each incremental target image.

[0074] Specifically, the center point of the first target image box corresponding to each retained feature point index is used as the local center point of the first local feature; the first eigenvalue is used as the local feature intensity corresponding to the first local feature, and the first feature score is used as the local feature score corresponding to the first local feature, and all first local features are integrated to obtain the target image local feature of the target image.

[0075] Step 103: Based on the signboard text, the non-signboard text and the target image vector, a pre-constructed database is searched to obtain a candidate image set.

[0076] In one embodiment, by executing the above-mentioned steps 101 and 102 respectively on multiple sample images containing stores, sample signboard text, sample non-signboard text information, sample signboard area coordinates, sample image vector, sample sub-image vector, and sample image local features corresponding to each sample image are obtained, and the sample signboard text, sample non-signboard text information, sample signboard area coordinates, sample image vector, sample sub-image vector, and sample image local features corresponding to each sample image are stored to obtain a pre-constructed database.

[0077] In one embodiment, sample signboard characters and sample non-signboard characters corresponding to each sample image in a pre-constructed database are obtained.

[0078] In one embodiment, based on the signboard characters and the non-signboard characters, the sample signboard characters and the sample non-signboard characters corresponding to each sample image in the database are searched and processed, and a first candidate image set is obtained based on the search results.

[0079] Specifically, Elasticsearch (ES) is used to search the database based on the signboard text and the non-signboard text.

[0080] Specifically, when using Elasticsearch (ES) for retrieval, the multi-field matching query and weight control functions it supports can be utilized; based on the signboard text, a signboard text retrieval instruction is generated, and a first weight value is set for the signboard text retrieval instruction; based on the non-signboard text, a non-signboard text retrieval instruction is generated, and a second weight value is set for the non-signboard text retrieval instruction; based on the signboard text retrieval instruction and the non-signboard text retrieval instruction, the database is searched to obtain a first candidate image set, wherein the first weight value is greater than the second weight value.

[0081] In one embodiment, a sample image vector corresponding to each sample image in the database is obtained.

[0082] In one embodiment, the image similarities between the target image vector and the sample image vector corresponding to each sample image in the database are calculated respectively, and based on the image similarities, the second candidate image set is determined.

[0083] Specifically, after calculating the image similarity between the target image vector and each sample image vector respectively, the target sample image corresponding to the sample image vector whose image similarity is greater than the preset image similarity is obtained, and based on the size of the image similarity corresponding to each target sample image, each target sample image is sorted, and the first target number of target sample images are selected, and the first target number of target sample images are used as the first candidate image set.

[0084] Preferably, the target number is 5.

[0085] In one embodiment, the first candidate image set and the second candidate image set are integrated to obtain a candidate image set.

[0086] Step 104: Based on the target image vector, the target sub-image vector and the signboard text, respectively calculate the similarity scores between the target image and each candidate image in the candidate image set, and based on the local features of the target image, respectively perform local feature matching processing on the target image and each candidate image in the candidate image set to obtain the number of matching features.

[0087] In one embodiment, the candidate image vector, the candidate sub-image vector and the candidate signboard text corresponding to each candidate image in the candidate image set are obtained.

[0088] In one embodiment, a first similarity is calculated between the image vector of the target image and the selected image vector corresponding to each candidate image in the candidate image set, a second similarity is calculated between the signboard text of the target image and the candidate signboard text corresponding to each candidate image in the candidate image set, and a third similarity is calculated between the target sub-image vector of the target image and the candidate sub-image vector corresponding to each candidate image in the candidate image set.

[0089] Preferably, if there are other signs in addition to the main sign in the target image and the candidate image, other target sub-image vectors corresponding to the other signs in the target image and other candidate sub-image vectors corresponding to the other signs in the candidate image are obtained,

[0090] Calculate the fourth similarity between the other target sub-image vector of the target image and the other candidate sub-image vector corresponding to each candidate image in the candidate image set; if the third similarity is not less than the fourth similarity, retain the third similarity; if the third similarity is less than the fourth similarity, update the fourth similarity to the third similarity.

[0091] In one embodiment, weighted processing is performed on the first similarity, the second similarity, and the third similarity to obtain a similarity score between the target image and each candidate image in the candidate image set.

[0092] Specifically, corresponding weight values ​​are set for the first similarity, the second similarity and the third similarity respectively, and the first similarity, the second similarity and the third similarity are multiplied by their corresponding weight values ​​respectively, and the obtained products are added to obtain the similarity scores between the target image and each candidate image in the candidate image set, wherein the setting of the weight values ​​can be determined according to the needs of the actual application.

[0093] Specifically, based on the target image vector, the target sub-image vector and the sign text, the similarity score between the target image and each candidate image in the candidate image set is calculated, which can be used to filter out falsely recalled images, such as images with similar scenes but dissimilar texts, or images with similar texts but inconsistent backgrounds.

[0094] In one embodiment, local features of the candidate images corresponding to each candidate image in the candidate image set are obtained.

[0095] In one embodiment, feature distances between local features of the target image and each candidate image in the candidate picture set are calculated respectively, and based on the feature distances, feature point screening processing is performed on the target image and each candidate image in the candidate picture set to obtain a first local feature point set corresponding to the target image and a second local feature point set corresponding to each candidate image in the candidate picture set.

[0096] Specifically, the KD tree based on fast nearest neighbor search calculates the feature distance between the target image and the local features of the candidate images corresponding to each candidate image in the candidate picture set. For each local feature point in the target image, the KD tree can quickly find the nearest local feature point in the candidate image and calculate the feature distance between them.

[0097] Specifically, according to the calculated feature distance, a feature distance threshold is set, and local feature points whose special effect distance is less than the feature distance threshold are retained. This step can effectively remove feature points with a long distance and low similarity, thereby improving the matching accuracy.

[0098] Specifically, for the target image, the set of retained local feature points constitutes a first set of local feature points; for each candidate image in the candidate picture set, the set of retained local feature points constitutes a second set of local feature points.

[0099] In one embodiment, the first local feature point set and the second local feature point set are filtered to obtain a first filtered local feature point set and a second filtered local feature point set.

[0100] Specifically, since the first local feature point set and the second local feature point set may contain some outliers and noise points, the RANSAC algorithm is used to filter the first local feature point set and the second local feature point set.

[0101] In one embodiment, the first filtered local feature point set and the second filtered local feature point set are classified and processed respectively to obtain the number of target signboard area feature points and the number of target non-signboard area feature points corresponding to the first filtered local feature point set, as well as the number of candidate signboard area feature points and the number of candidate non-signboard area feature points corresponding to the second filtered local feature point set.

[0102] Specifically, the coordinates of the signboard area in the target image are obtained, and it is determined whether the coordinates corresponding to each first filtered local feature point in the first filtered local feature point set fall within the signboard area coordinates; if so, the first filtered local feature point whose coordinates fall within the signboard area coordinates is used as the target signboard area feature point; otherwise, the first filtered local feature point whose coordinates do not fall within the signboard area coordinates is used as the target non-signboard area feature point; all target signboard area feature points are counted to obtain the number of target signboard area feature points, and all target non-signboard area feature points are counted to obtain the number of target non-signboard area feature points.

[0103] Specifically, the coordinates of the candidate signboard area in the candidate image are obtained, and it is determined whether the coordinates corresponding to each second filtered local feature point in the second filtered local feature point set fall within the candidate signboard area coordinates; if so, the second filtered local feature point whose coordinates fall within the candidate signboard area coordinates is used as the candidate signboard area feature point; otherwise, the second filtered local feature point whose coordinates do not fall within the candidate signboard area coordinates is used as the candidate non-signboard area feature point; all the candidate signboard area feature points are counted to obtain the number of candidate signboard area feature points; and all the candidate non-signboard area feature points are counted to obtain the number of candidate non-signboard area feature points.

[0104] In one embodiment, the number of target signboard area feature points and the number of target non-signboard area feature points corresponding to the target image, as well as the number of candidate signboard area feature points and the number of candidate non-signboard area feature points corresponding to each candidate image in the candidate picture set are used as the matching feature number.

[0105] Step 105: Based on the similarity scores and the number of matching features, the candidate image set is screened to obtain similar images of the target image.

[0106] In one embodiment, when it is detected that there is a target candidate image in the candidate image set whose similarity score with the target image is greater than a preset similarity score threshold, the target candidate image is used as a similar image to the target image.

[0107] In one embodiment, when it is detected that there is no target candidate image in the candidate image set and the similarity score between the target image and the target image is greater than a preset similarity score threshold, it is determined whether the number of matching features satisfies the preset matching feature point rules. If so, the candidate image corresponding to the number of matching features is used as a similar image of the target image.

[0108] Specifically, when judging whether the number of matching features satisfies the preset matching feature point rules, by setting a threshold for the number of signboard area feature points and a threshold for the number of non-signboard area feature points, if the number of target signboard area feature points corresponding to the target image and the number of candidate signboard area feature points corresponding to the target candidate images in the candidate image set are both greater than the threshold for the number of signboard area feature points, and the number of target non-signboard area feature points corresponding to the target image and the number of candidate non-signboard area feature points corresponding to the target candidate images in the candidate image set are both greater than the threshold for the number of non-signboard area feature points, it is determined that the number of matching features satisfies the preset matching feature point rules.

[0109] Example 2, see Figure 2 , Figure 2 : is a schematic diagram of the structure of an embodiment of a store image retrieval device provided by the present application. Corresponding to the above store image retrieval method, the present application also provides a store image retrieval device. The store image retrieval device includes a unit for executing the above store image retrieval method, and the store image retrieval device can be configured in a desktop computer, a tablet computer, a laptop computer, and other terminals. Specifically, the store image retrieval device includes a target image acquisition module 201, a target image feature acquisition module 202, a candidate image retrieval module 203, an image comparison module 204, and a similar image determination module 205.

[0110] The target image acquisition module 201 is used to acquire a target image, perform target recognition on the target image, and obtain the signboard text and non-signboard text on the target image.

[0111] The target image feature acquisition module 202 is used to perform image cropping on the target image to obtain a target sub-image, perform vector extraction on the target sub-image and the target image respectively to obtain a target sub-image vector and a target image vector, and perform local feature extraction on the target image to obtain local features of the target image.

[0112] The candidate image retrieval module 203 is used to search a pre-constructed database based on the signboard text, the non-signboard text and the target image vector to obtain a candidate image set.

[0113] The image comparison module 204 is used to calculate the similarity scores between the target image and each candidate image in the candidate image set based on the target image vector, the target sub-image vector and the sign text, and to perform local feature matching processing on the target image and each candidate image in the candidate image set based on the local features of the target image to obtain the number of matching features.

[0114] The similar image determination module 205 is used to screen the candidate image set based on the similarity score and the number of matching features to obtain similar images of the target image.

[0115] In one embodiment, the target image acquisition module 201 is used to perform target recognition on the target image to obtain the signboard text and non-signboard text on the target image, specifically including: based on a pre-trained signboard recognition model, performing signboard recognition on the target image to obtain the signboard area coordinates; based on a pre-trained text detection model, performing text detection on the target image to obtain the text area, and the text area coordinates corresponding to the text area; based on a pre-trained text recognition model, performing text recognition on the text area to obtain text data, and the text coordinates corresponding to the text data; based on the signboard area coordinates and the text coordinates, classifying the text data to obtain the signboard text and non-signboard text on the target image.

[0116] In one embodiment, the target image feature acquisition module 202 is used to perform vector extraction on the target sub-image and the target image, respectively, to obtain a target sub-image vector and a target image vector, specifically including: performing image increment processing on the target sub-image to obtain an incremental target sub-image set, and performing vector extraction on each incremental target sub-image in the target sub-image set to obtain an incremental target sub-image vector corresponding to each incremental target sub-image; performing averaging processing on all incremental target sub-image vectors to obtain an incremental target sub-image mean vector, and using the incremental target sub-image mean vector as the target sub-image vector corresponding to the target sub-image; performing image increment processing on the target image to obtain an incremental target image set, and performing vector extraction on each incremental target image in the target image set to obtain an incremental target image vector corresponding to each incremental target image; performing averaging processing on all incremental target image vectors to obtain an incremental target image mean vector, and using the incremental target image mean vector as the target image vector corresponding to the target image.

[0117] In one embodiment, the target image feature acquisition module 202 is used to extract local features of the target image to obtain local features of the target image, specifically including: obtaining a pre-trained DELG model, wherein the DELG model includes a local network and a Resnet model, and the output end of the third layer feature map of the Resnet model is connected to the input end of the local network; performing image increment processing on the target image to obtain multiple incremental target images, inputting each incremental target image into the DELG model respectively, so that the Resnet model in the DELG model outputs a target feature map, and inputting the target feature map into the local network to obtain a feature score corresponding to each target feature point in the target feature map; normalizing the target feature map to obtain a normalized original feature value, and multiplying the normalized original feature value and the feature score to obtain the target feature. The target feature point index in the target image is obtained by performing a feature point screening process on the target feature image based on the feature score threshold value to obtain a target feature point index that satisfies a preset number of feature points, and the target feature image is segmented based on a preset step size to obtain a plurality of target feature boxes; the feature score, the feature value and the target image box corresponding to the target feature point index in each incremental target image are spliced ​​to obtain a spliced ​​total feature corresponding to each incremental target image; based on the feature score and the target image box corresponding to each incremental target image, non-maximum suppression is performed on the spliced ​​total feature of each incremental target image to obtain a plurality of retained feature point indexes; and the target image local features of the target image are determined based on the first feature score, the first feature value and the first target image box corresponding to each of the plurality of retained feature point indexes in each incremental target image.

[0118] In one embodiment, the candidate image retrieval module 203 is used to search a pre-constructed database based on the signboard text, the non-signboard text and the target image vector to obtain a candidate image set, specifically including: obtaining sample signboard text and sample non-signboard text corresponding to each sample image in the pre-constructed database, and performing retrieval processing on the sample signboard text and the sample non-signboard text corresponding to each sample image in the database based on the signboard text and the non-signboard text, and obtaining a first candidate image set based on the retrieval result; obtaining a sample image vector corresponding to each sample image in the database, respectively calculating the image similarity between the target image vector and the sample image vector corresponding to each sample image in the database, and determining a second candidate image set based on the image similarity; integrating the first candidate image set and the second candidate image set to obtain a candidate image set.

[0119] In one embodiment, the image comparison module 204 is used to calculate the similarity scores between the target image and each candidate image in the candidate image set based on the target image vector, the target sub-image vector and the signboard text, specifically including: obtaining the candidate image vector, the candidate sub-image vector and the candidate signboard text corresponding to each candidate image in the candidate image set; calculating the first similarity between the image vector of the target image and the selected image vector corresponding to each candidate image in the candidate image set, and calculating the second similarity between the signboard text of the target image and the candidate signboard text corresponding to each candidate image in the candidate image set, and calculating the third similarity between the target sub-image vector of the target image and the candidate sub-image vector corresponding to each candidate image in the candidate image set; weighting the first similarity, the second similarity and the third similarity to obtain the similarity score between the target image and each candidate image in the candidate image set.

[0120] In one embodiment, the image comparison module 204 is used to perform local feature matching processing on the target image and each candidate image in the candidate image set based on the local features of the target image to obtain the number of matching features, specifically including: obtaining the candidate image local features corresponding to each candidate image in the candidate image set; calculating the feature distance between the target image and the candidate image local features corresponding to each candidate image in the candidate image set, and performing feature point screening processing on the target image and each candidate image in the candidate image set based on the feature distance to obtain a first local feature point set corresponding to the target image and a second local feature point set corresponding to each candidate image in the candidate image set; and filtering the first local feature point set. and the second local feature point set are filtered to obtain a first filtered local feature point set and a second filtered local feature point set; the first filtered local feature point set and the second filtered local feature point set are classified to obtain the number of target signboard area feature points and the number of target non-signboard area feature points corresponding to the first filtered local feature point set, as well as the number of candidate signboard area feature points and the number of candidate non-signboard area feature points corresponding to the second filtered local feature point set; the number of target signboard area feature points and the number of target non-signboard area feature points corresponding to the target image, as well as the number of candidate signboard area feature points and the number of candidate non-signboard area feature points corresponding to each candidate image in the candidate picture set are taken as the matching feature number.

[0121] The above-mentioned store image retrieval device can implement the store image retrieval method of the above-mentioned method embodiment. The optional items in the above-mentioned method embodiment are also applicable to this embodiment and will not be described in detail here.

[0122] like Figure 3 As shown, Figure 3 It is a structural diagram of an electronic device provided by the present application; it includes a processor 111, a communication interface 112, a memory 113 and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114, and the memory 113 is used to store computer programs.

[0123] In one embodiment of the present application, the processor 111 is used to implement the store image retrieval method provided by any one of the aforementioned method embodiments when executing the program stored in the memory 113.

[0124] It is understood by those skilled in the art that all or part of the processes in the method for implementing the above embodiment can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a storage medium, which is a computer-readable storage medium. The computer program is executed by at least one processor in the computer system to implement the process steps of the embodiment of the above method.

[0125] Therefore, an embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the store image retrieval method provided in any of the aforementioned method embodiments are implemented.

[0126] The storage medium is a physical, non-transient storage medium, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk, etc., which can store program codes. The computer-readable storage medium can be non-volatile or volatile.

[0127] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0128] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of each unit is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0129] The steps in the method of the embodiment of the present application can be adjusted in order, combined and deleted according to actual needs. The units in the device of the embodiment of the present application can be combined, divided and deleted according to actual needs. In addition, the functional units in the various embodiments of the present application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0130] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, terminal, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application.

[0131] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0132] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

[0133] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.

Claims

1. A store image retrieval method, characterized in that: include: Acquire a target image, perform target recognition on the target image, and obtain signboard text and non-signboard text on the target image; Performing image cropping on the target image to obtain a target sub-image, performing vector extraction on the target sub-image and the target image respectively to obtain a target sub-image vector and a target image vector, and performing local feature extraction on the target image to obtain local features of the target image; Based on the signboard text, the non-signboard text and the target image vector, searching a pre-built database to obtain a candidate image set; Based on the target image vector, the target sub-image vector and the signboard text, respectively calculate the similarity scores between the target image and each candidate image in the candidate image set, and based on the local features of the target image, respectively perform local feature matching processing on the target image and each candidate image in the candidate image set to obtain the number of matching features; Based on the similarity score and the number of matching features, the candidate image set is screened to obtain similar images of the target image.

2. The method according to claim 1, characterized in that Performing target recognition on the target image to obtain the signboard text and non-signboard text on the target image specifically includes: Based on the pre-trained signboard recognition model, perform signboard recognition on the target image to obtain the coordinates of the signboard area; Based on a pre-trained text detection model, perform text detection on the target image to obtain a text area and text area coordinates corresponding to the text area; Based on a pre-trained text recognition model, perform text recognition on the text area to obtain text data and text coordinates corresponding to the text data; Based on the sign area coordinates and the text coordinates, the text data is classified and processed to obtain the sign text and non-sign text on the target image.

3. The method according to claim 1, characterized in that: Performing vector extraction on the target sub-image and the target image respectively to obtain a target sub-image vector and a target image vector specifically includes: Performing image increment processing on the target sub-image to obtain an incremental target sub-image set, and performing vector extraction on each incremental target sub-image in the target sub-image set to obtain an incremental target sub-image vector corresponding to each incremental target sub-image; Performing averaging processing on all incremental target sub-image vectors to obtain an incremental target sub-image mean vector, and using the incremental target sub-image mean vector as the target sub-image vector corresponding to the target sub-image; Performing image increment processing on the target image to obtain an incremental target image set, and performing vector extraction on each incremental target image in the target image set to obtain an incremental target image vector corresponding to each incremental target image; All incremental target image vectors are averaged to obtain an incremental target image mean vector, and the incremental target image mean vector is used as the target image vector corresponding to the target image.

4. The method according to claim 1, characterized in that: Extracting local features of the target image to obtain local features of the target image specifically includes: Obtain a pre-trained DELG model, wherein the DELG model includes a local network and a Resnet model, and an output end of a third-layer feature map of the Resnet model is connected to an input end of the local network; Performing image increment processing on the target image to obtain a plurality of incremental target images, inputting each incremental target image into the DELG model respectively, so that the Resnet model in the DELG model outputs a target feature map, and inputting the target feature map into the local network to obtain a feature score corresponding to each target feature point in the target feature map; Normalizing the target feature map to obtain a normalized original feature value, and multiplying the normalized original feature value and the feature score to obtain a feature value corresponding to each target feature point in the target feature map; Setting a feature score threshold, performing feature point screening processing on the target feature map based on the feature score threshold, obtaining a target feature point index that meets a preset number of feature points, and performing segmentation processing on the target feature map based on a preset step size to obtain a plurality of target feature boxes; Performing splicing processing on the feature score, the feature value and the target image box corresponding to the target feature point index in each incremental target image to obtain a spliced ​​total feature corresponding to each incremental target image; Based on the feature score and the target image box corresponding to each incremental target image, non-maximum suppression processing is performed on the spliced ​​total feature of each incremental target image to obtain multiple retained feature point indexes; A target image local feature of the target image is obtained and determined based on first feature scores, first feature values ​​and first target image boxes corresponding to each of the plurality of retained feature point indexes in each incremental target image.

5. The method according to claim 1, characterized in that: Based on the signboard text, the non-signboard text and the target image vector, a pre-built database is searched to obtain a candidate image set, which specifically includes: Acquire sample signboard characters and sample non-signboard characters corresponding to each sample image in a pre-constructed database, perform retrieval processing on the sample signboard characters and sample non-signboard characters corresponding to each sample image in the database based on the signboard characters and the non-signboard characters, and obtain a first candidate image set based on the retrieval result; Obtaining a sample image vector corresponding to each sample image in the database, respectively calculating image similarities between the target image vector and the sample image vector corresponding to each sample image in the database, and determining a second candidate image set based on the image similarities; The first candidate image set and the second candidate image set are integrated to obtain a candidate image set.

6. The method according to claim 1, characterized in that: The step of calculating the similarity scores between the target image and each candidate image in the candidate image set based on the target image vector, the target sub-image vector and the signboard text comprises: Obtaining a candidate image vector, a candidate sub-image vector and the candidate signboard text corresponding to each candidate image in the candidate image set; respectively calculating a first similarity between the image vector of the target image and the selected image vector corresponding to each candidate image in the candidate image set, respectively calculating a second similarity between the signboard text of the target image and the candidate signboard text corresponding to each candidate image in the candidate image set, and respectively calculating a third similarity between the target sub-image vector of the target image and the candidate sub-image vector corresponding to each candidate image in the candidate image set; The first similarity, the second similarity and the third similarity are weighted to obtain a similarity score between the target image and each candidate image in the candidate image set.

7. The method according to claim 1, characterized in that: The performing local feature matching processing on the target image and each candidate image in the candidate image set based on the local features of the target image to obtain the number of matching features specifically includes: Obtaining local features of candidate images corresponding to each candidate image in the candidate image set; Calculating feature distances between local features of the target image and each candidate image in the candidate picture set respectively, and performing feature point screening processing on the target image and each candidate image in the candidate picture set based on the feature distances to obtain a first local feature point set corresponding to the target image and a second local feature point set corresponding to each candidate image in the candidate picture set; Performing filtering processing on the first local feature point set and the second local feature point set respectively to obtain a first filtered local feature point set and a second filtered local feature point set; Classifying the first filtered local feature point set and the second filtered local feature point set respectively, to obtain the number of target signboard area feature points and the number of target non-signboard area feature points corresponding to the first filtered local feature point set, and the number of candidate signboard area feature points and the number of candidate non-signboard area feature points corresponding to the second filtered local feature point set; The number of target signboard area feature points and the number of target non-signboard area feature points corresponding to the target image, and the number of candidate signboard area feature points and the number of candidate non-signboard area feature points corresponding to each candidate image in the candidate picture set are taken as the matching feature number.

8. A store image retrieval device, characterized in that: include: A target image acquisition module, a target image feature acquisition module, a candidate image retrieval module, an image comparison module and a similar image determination module; The target image acquisition module is used to acquire a target image, perform target recognition on the target image, and obtain the signboard text and non-signboard text on the target image; The target image feature acquisition module is used to perform image cropping on the target image to obtain a target sub-image, perform vector extraction on the target sub-image and the target image respectively to obtain a target sub-image vector and a target image vector, and perform local feature extraction on the target image to obtain a local feature of the target image; The candidate image retrieval module is used to search a pre-built database based on the signboard text, the non-signboard text and the target image vector to obtain a candidate image set; The image comparison module is used to calculate the similarity scores of the target image and each candidate image in the candidate image set based on the target image vector, the target sub-image vector and the signboard text, and to perform local feature matching processing on the target image and each candidate image in the candidate image set based on the local features of the target image to obtain the number of matching features; The similar image determination module is used to screen the candidate image set based on the similarity score and the number of matching features to obtain similar images of the target image.

9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 can be implemented.