Image-based processing method and device, equipment, storage medium and product
By performing text recognition on images, dynamically selecting similarity judgment categories based on the recognition results, and combining multiple algorithms for judgment, the problem of accuracy in image similarity calculation is solved, and the flexibility and accuracy of the judgment results are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2024-11-22
- Publication Date
- 2026-05-22
AI Technical Summary
Existing image similarity calculation methods cannot guarantee the accuracy of the judgment results, especially when there are large differences between different image pairs.
By performing text recognition on target image pairs, determining the target similarity judgment category based on the text recognition results, and using the similarity judgment methods corresponding to different preset similarity judgment categories, the judgment is made by combining at least two similarity algorithms, including text similarity, image color similarity based on histogram, feature similarity based on scale-invariant feature transformation, and hash similarity based on hash value.
It improves the flexibility, robustness, and accuracy of image similarity determination, and adapts to accurate determination in different image scenarios.
Smart Images

Figure CN122073053A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and more particularly to image-based processing methods, apparatus, devices, storage media, and products. Background Technology
[0002] Image similarity calculation is an important image algorithm capability with wide applications in many fields, such as copyright protection, image search, and image deduplication.
[0003] Currently, commonly used image similarity calculation methods include histogram-based color similarity calculation and perceptual hashing algorithms. However, different image pairs requiring similarity assessment may have significant differences, and traditional methods that use a single, fixed image similarity calculation algorithm for similarity determination are difficult to guarantee the accuracy of the results. Summary of the Invention
[0004] This disclosure provides image-based processing methods, apparatus, devices, storage media, and products that can optimize existing image similarity calculation-based processing schemes.
[0005] In a first aspect, embodiments of this disclosure provide an image-based processing method, including:
[0006] Text recognition is performed on a pair of target images to obtain text recognition results, wherein the pair of target images includes a first image and a second image;
[0007] The target similarity judgment category of the target image pair is determined based on the text recognition result. The target similarity judgment category is a preset similarity judgment category in a preset similarity judgment category set. Different preset similarity judgment categories correspond to different similarity judgment methods. The similarity judgment method is a judgment method based on at least two similarity algorithms.
[0008] The similarity determination method corresponding to the target similarity determination category is used to determine whether the first image and the second image are similar images.
[0009] Secondly, embodiments of this disclosure also provide an image-based processing apparatus, including:
[0010] The text recognition module is used to perform text recognition on a target image pair to obtain text recognition results, wherein the target image pair includes a first image and a second image;
[0011] The category determination module is used to determine the target similarity determination category of the target image pair based on the text recognition result. The target similarity determination category is a preset similarity determination category in a preset similarity determination category set. Different preset similarity determination categories correspond to different similarity determination methods. The similarity determination method is a determination method based on at least two similarity algorithms.
[0012] The similarity determination module is used to determine whether the first image and the second image are similar images by adopting the similarity determination method corresponding to the target similarity determination category.
[0013] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:
[0014] One or more processors;
[0015] Storage device for storing one or more programs.
[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the image-based processing method provided in the embodiments of this disclosure.
[0017] Fourthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the image-based processing method provided in embodiments of this disclosure.
[0018] Fifthly, embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements the image-based processing method provided in embodiments of this disclosure.
[0019] The image-based processing scheme provided in this disclosure first performs text recognition on target image pairs requiring similarity determination. Based on the text recognition results, it determines the target similarity determination category for the target image pairs. Different preset similarity determination categories correspond to different similarity determination methods. Using the similarity determination method corresponding to the target similarity determination category, it determines whether the first image and the second image are similar images. The similarity determination method is based on at least two similarity algorithms. By adopting the above technical solution, the flexibility, robustness, and accuracy of image similarity determination can be improved. Attached Figure Description
[0020] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0021] Figure 1 This is a schematic flowchart illustrating an image-based processing method provided in an embodiment of the present disclosure.
[0022] Figure 2 A schematic flowchart illustrating another image-based processing method provided in an embodiment of this disclosure;
[0023] Figure 3 This is a schematic flowchart illustrating yet another image-based processing method provided in an embodiment of the present disclosure;
[0024] Figure 4 A schematic diagram of the structure of an image-based processing apparatus provided in an embodiment of this disclosure;
[0025] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0026] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0027] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0028] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0029] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0030] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0031] Figure 1 This is a flowchart illustrating an image-based processing method provided in an embodiment of the present disclosure. This embodiment is applicable to the processing of image similarity determination. The method can be executed by an image-based processing device, which can be implemented in software and / or hardware. Optionally, it can be implemented by an electronic device, such as a mobile terminal like a mobile phone, smartwatch, tablet computer, or personal digital assistant, or a personal computer (PC) or server.
[0032] like Figure 1 As shown, the method includes:
[0033] Step 101: Perform text recognition on the target image pair to obtain the text recognition result, wherein the target image pair includes a first image and a second image.
[0034] For example, an image pair may contain two images, and a target image pair may be an image pair consisting of two images for which image similarity determination is currently required. The two images in the target image pair are referred to as the first image and the second image, respectively.
[0035] For example, text recognition algorithms can be used to recognize text in any one or two images in a target image pair. Such algorithms can be, for example, Optical Character Recognition (OCR).
[0036] Step 102: Determine the target similarity judgment category of the target image pair based on the text recognition result. The target similarity judgment category is a preset similarity judgment category in a preset similarity judgment category set. Different preset similarity judgment categories correspond to different similarity judgment methods. The similarity judgment method is a judgment method based on at least two similarity algorithms.
[0037] In this embodiment of the disclosure, at least two similarity determination categories (denoted as preset similarity determination categories) can be preset to classify different image similarity determination scenarios. The preset similarity determination category set includes at least two preset similarity determination categories. The specific classification method is not limited. For example, it can be classified according to whether the text recognition result contains text, or it can be classified according to the size of the text region in the image determined by the text recognition result, etc.
[0038] For example, for each preset similarity judgment category, a corresponding similarity judgment method can be preset. Different preset similarity judgment categories correspond to different similarity judgment methods, so that more accurate judgment results can be obtained for the specific judgment scenarios of each preset similarity judgment category.
[0039] For each preset similarity category, a corresponding similarity determination method can be determined based on at least two similarity algorithms. This allows for the use of at least two similarity algorithms to compensate for the shortcomings of a single algorithm, further increasing the accuracy of similarity determination. Differences in similarity determination methods can be reflected in at least one of the following aspects: at least one of the at least two similarity algorithms is different; the order in which the at least two similarity algorithms determine similarity is different; and the determination threshold of at least one of the at least two similarity algorithms is different.
[0040] For example, similarity algorithms may include text similarity algorithms, histogram-based image color similarity algorithms, scale-invariant feature transformation-based feature similarity algorithms, and hash similarity algorithms based on hash values.
[0041] Step 103: Use the similarity determination method corresponding to the target similarity determination category to determine whether the first image and the second image are similar images.
[0042] For example, prior to this step, the target image pair can be preprocessed. Specific preprocessing methods may include resizing and grayscale conversion. Resizing may include adjusting the sizes of the first and second images to be the same; for example, adjusting the larger image to the size of the smaller image. Grayscale conversion may include converting both the first and second images into their corresponding grayscale images.
[0043] For example, after determining the target similarity judgment category in the aforementioned steps, the similarity judgment method corresponding to the target similarity judgment category can be obtained, and the obtained similarity judgment method can be used to determine whether the first image and the second image are similar images, that is, to determine whether the target image pair is a similar image pair, thereby facilitating subsequent related applications, such as copyright protection, image search, and image deduplication.
[0044] The image-based processing method provided in this disclosure first performs text recognition on target image pairs requiring similarity determination. Based on the text recognition results, it determines the target similarity determination category for the target image pairs. Different preset similarity determination categories correspond to different similarity determination methods. Using the similarity determination method corresponding to the target similarity determination category, it determines whether the first image and the second image are similar images. The similarity determination method is based on at least two similarity algorithms. By adopting the above technical solution, the determination categories of image pairs requiring similarity determination are divided according to the text recognition results. For the current target image pair, different determination methods based on at least two similarity algorithms can be dynamically determined to determine image similarity, thereby improving the flexibility, robustness, and accuracy of image similarity determination.
[0045] In some embodiments, the proportion of text regions in an image can be determined based on the text recognition results for classification. Specifically, the step of performing text recognition on a target image pair to obtain text recognition results includes: performing text recognition on a first image in the target image pair to obtain text recognition results. The step of determining the target similarity judgment category of the target image pair based on the text recognition results includes: determining the area of the text regions contained in the first image based on the text recognition results; calculating the proportion of the area of the text regions relative to the image area of the first image; determining the target similarity judgment category of the target image pair as a first judgment category in response to the proportion being greater than or equal to a first preset proportion threshold; determining the target similarity judgment category of the target image pair as a second judgment category in response to the proportion being greater than or equal to a second preset proportion threshold and less than the first preset proportion threshold; and determining the target similarity judgment category of the target image pair as a third judgment category in response to the proportion being less than the second preset proportion threshold. Therefore, classifying different judgment scenarios based on the proportion of text areas in an image can effectively improve the accuracy of similarity judgment results under three different judgment scenarios.
[0046] For example, the text region contained in the first image can be the region where the text is located in the first image. Specifically, it can be the sum of the smallest regions containing complete text. For example, if the first image contains 10 characters, the smallest region containing each character can be calculated, and then the sum of the smallest regions corresponding to the 10 characters can be calculated to obtain the text region contained in the first image. Alternatively, it can be the smallest region containing a text paragraph. For example, if the first image contains a sentence, the smallest region containing the sentence can be determined as the text region contained in the first image. The first preset percentage threshold and the second preset percentage threshold can be set according to the actual situation. For example, the first preset percentage threshold is 70%, and the second preset percentage threshold is 30%. For the first judgment category, the text region has a high percentage, which can be considered a judgment scenario where text is the main focus, and can also be called the text category. For the second judgment category, the text region has a moderate percentage, which can be considered a judgment scenario where both text and image content are relatively important, and can also be called the image-text category. For the third judgment category, the text region has a low percentage, which can be considered a judgment scenario where image content is the main focus, and can also be called the image category.
[0047] Figure 2 This is a flowchart illustrating another image-based processing method provided by an embodiment of the present disclosure. This embodiment optimizes the various optional solutions described above. Specifically, the method includes the following steps:
[0048] Step 201: Perform text recognition on the first image in the target image pair to obtain the text recognition result. The target image pair includes the first image and the second image.
[0049] Step 202: Determine the area of the text region contained in the first image based on the text recognition results.
[0050] Step 203: Calculate the ratio of the region area to the image area of the first image.
[0051] Step 204: Determine whether the percentage is greater than or equal to the first preset percentage threshold. If yes, proceed to step 206; otherwise, proceed to step 205.
[0052] Step 205: Determine whether the percentage is greater than or equal to the second preset percentage threshold. If yes, proceed to step 207; otherwise, proceed to step 208.
[0053] Step 206: Determine the target similarity judgment category of the target image pair as the first judgment category, and determine whether the first image and the second image are similar images based on the text similarity and the image color similarity based on the histogram of the target image pair.
[0054] For example, the calculation steps of text similarity of target image pairs and histogram-based image color similarity of target image pairs used in this step can be performed before this step. That is, if there is sufficient computing power, the calculation can be performed in advance before determining the specific judgment category and used directly when performing this step. Alternatively, the calculation can be performed after determining the first judgment category. There is no specific limitation.
[0055] For example, the text similarity of a target image pair can be calculated as follows: convert the text contained in the first image into a first word vector, convert the text contained in the second image into a second word vector, calculate the cosine similarity between the first and second word vectors, and use this cosine similarity to represent the text similarity; or, calculate the ratio of the intersection to the union of the character sets in the text contained in the first image and the character sets in the text contained in the second image, and use this ratio to represent the text similarity. Other methods for calculating text similarity can also be used, and there are no specific limitations.
[0056] For example, histogram-based image color similarity can be calculated as follows: For the histogram corresponding to the first image (which can be a color histogram or a grayscale histogram) and the histogram corresponding to the second image, calculate the similarity of color distribution or brightness distribution to obtain the image color similarity. Histogram-based image color similarity algorithms have low computational complexity. For the first classification category, since text accounts for a large proportion, a comprehensive judgment using both text similarity and histogram-based image color similarity can effectively balance judgment efficiency and accuracy.
[0057] Optionally, determining whether the first image and the second image are similar images based on the text similarity and histogram-based image color similarity of the target image pair includes: calculating the product of the proportion and the text similarity of the target image pair to obtain a first product; calculating the product of the target value and the histogram-based image color similarity of the target image pair to obtain a second product, wherein the target value is equal to 1 minus the proportion; and determining that the first image and the second image are similar images in response to the sum of the first product and the second product being greater than a first preset threshold. Therefore, dynamically determining the weights of text similarity and image color similarity based on the area proportion of the text region can further improve the accuracy of the determination.
[0058] Optionally, it further includes: determining that the first image and the second image are dissimilar images in response to the sum of the first product and the second product being less than or equal to a first preset threshold.
[0059] For example, let the text similarity be denoted as sim_text, the histogram-based image color similarity as sim_hist, the proportion as x, and the first preset threshold as T1. If x*sim_text+(1-x)*sim_hist>T1, then the first image and the second image can be determined to be similar images. Optionally, the first preset threshold can be 0.9.
[0060] Step 207: Determine the target similarity judgment category of the target image pair as the second judgment category, and determine whether the first image and the second image are similar images based on the text similarity of the target image pair, the image color similarity based on histogram, the feature similarity based on scale-invariant feature transform, and the hash similarity based on hash value.
[0061] For example, the calculation steps of text similarity, histogram-based image color similarity, scale-invariant feature transformation-based feature similarity, and hash value-based hash similarity of target image pairs used in this step can be performed before this step. That is, if there is sufficient computing power, the calculation can be performed in advance before determining the specific judgment category and used directly when executing this step. Alternatively, the calculation can be performed after determining the second judgment category. There is no specific limitation.
[0062] For example, the feature similarity of a target image pair based on Scale-Invariant Feature Transform (SIFT) can be calculated as follows: the image features of the first and second images are analyzed using the SIFT algorithm to extract feature points and feature directions, and the feature points and feature directions are matched to obtain the feature similarity based on SIFT.
[0063] For example, the hash similarity of a target image pair based on hash values can be calculated as follows: feature extraction is performed on the first image and the second image respectively using a perceptual hash algorithm to obtain two hash arrays, and the similarity between the two hash arrays is calculated to obtain the hash similarity based on hash values.
[0064] Optionally, determining whether the first image and the second image are similar images based on the text similarity, histogram-based image color similarity, scale-invariant feature transform-based feature similarity, and hash value-based hash similarity of the target image pair includes: determining the first image and the second image as similar images in response to the target image pair satisfying multiple of the following (1)-(4) (optionally, all four conditions may be satisfied):
[0065] (1) The feature similarity of the target image pair based on scale-invariant feature transformation is greater than the product of the first preset coefficient and the target value, wherein the target value is equal to 1 minus the proportion.
[0066] For example, let sim_sift be the feature similarity of SIFT, K1 be the first preset coefficient, and x be the proportion. If sim_sift > K1*(1-x), then the condition is satisfied. Optionally, the first preset coefficient can be 0.8.
[0067] (2) The hash similarity of the target image pair based on the hash value is greater than the product of the second preset coefficient and the target value, wherein the target value is equal to 1 minus the proportion, and the second preset coefficient is greater than the first preset coefficient.
[0068] For example, let the hash similarity be sim_phash, and the second preset coefficient be K2. If sim_phash > K2*(1-x), then the condition is satisfied. Optionally, the second preset coefficient can be 0.9.
[0069] (3) The proportion and the text similarity of the target image pair are greater than the second preset threshold, wherein the second preset threshold is less than the first preset coefficient.
[0070] For example, let the text similarity be denoted as sim_text, and the second preset threshold be denoted as T2. If sim_text > T2, then the condition can be determined to be met. Optionally, the second preset threshold can be 0.7.
[0071] (4) The image color similarity of the target image pair based on the histogram is greater than a third preset threshold, wherein the third preset threshold is greater than the second preset threshold.
[0072] Let sim_hist be the image color similarity based on the histogram, x be the proportion, and T3 be the third preset threshold. If sim_hist > T3, the condition can be determined to be satisfied. Optionally, the third preset threshold can be 0.8.
[0073] Optionally, in response to the target image pair not satisfying any of (1)-(4) above, the first image and the second image are determined to be dissimilar images.
[0074] Step 208: Determine the target similarity judgment category of the target image pair as the third judgment category. Based on the image color similarity based on histogram, the feature similarity based on scale-invariant feature transformation, and the hash similarity based on hash value of the target image pair, determine whether the first image and the second image are similar images.
[0075] For example, the calculation steps of image color similarity based on histogram, feature similarity based on scale-invariant feature transform, and hash similarity based on hash value used in this step can be performed before this step. That is, if there is sufficient computing power, the calculation can be performed in advance before determining the specific judgment category and used directly when executing this step. Alternatively, the calculation can be performed after determining the third judgment category. There is no specific limitation.
[0076] For example, for the third category of judgment, text accounts for a relatively low proportion, and similarity judgment can be mainly based on image content.
[0077] Optionally, determining whether the first image and the second image are similar images based on histogram-based image color similarity, scale-invariant feature transform-based feature similarity, and hash value-based hash similarity of the target image pair includes: determining that the first image and the second image are similar images in response to the target image pair satisfying at least one of the following (5)-(7):
[0078] (5) The feature similarity of the target image pair based on scale-invariant feature transformation is greater than the fourth preset threshold.
[0079] For example, let the feature similarity of SIFT be denoted as sim_sift, and the fourth preset threshold be denoted as T4. If sim_sift > T4, it can be determined that the condition is met. Optionally, the fourth preset threshold can be 0.65.
[0080] (6) The hash similarity of the target image pair based on the hash value is greater than the fifth preset threshold.
[0081] For example, let the hash similarity be sim_phash, and the fifth preset threshold be T5. If sim_phash > T5, the condition can be determined to be met. Optionally, the fifth preset threshold can be 0.95.
[0082] (7) The feature similarity of the target image pair based on scale-invariant feature transformation is greater than the sixth preset threshold, and the image color similarity of the target image pair based on histogram is greater than the seventh preset threshold, wherein the sixth preset threshold is less than the fourth preset threshold.
[0083] For example, let SIFT feature similarity be denoted as sim_sift, histogram-based image color similarity as sim_hist, the sixth preset threshold as T6, and the seventh preset threshold as T7. If sim_sift > T6 && sim_hist > T7, then... Optionally, the sixth preset threshold can be 0.3 or 0.9.
[0084] Optionally, in response to the target image pair not satisfying all of the items in (5)-(7) above, the first image and the second image are determined to be dissimilar images.
[0085] The image-based processing method provided in this embodiment calculates the ratio of the area of the text region contained in the first image to the image area of the first image based on the text recognition result, and divides the image into three judgment scenarios based on the ratio. The method sets corresponding similarity judgment methods for the three judgment categories, which can effectively improve the accuracy of similarity judgment results under the three different judgment scenarios and further enhance the robustness of the similarity judgment scheme.
[0086] In some embodiments, the method further includes: in response to the target similarity determination category being the third determination category, determining the target color richness type of the first image, wherein the target color richness type is one color richness type in a preset color richness type set, the preset color richness type set including single types and rich types; determining a seventh preset threshold based on the target color richness type, wherein the seventh preset threshold corresponding to the single type is greater than the seventh preset threshold corresponding to the rich type. Thus, for the third determination category where image content is primary, different seventh preset thresholds are further selected based on the different color richness of the image to perform histogram-based image color similarity correlation determination, further improving the accuracy of the determination results.
[0087] For example, hash similarity based on hash values and feature similarity based on scale-invariant feature transform are less sensitive to color richness than image color similarity based on histograms. Therefore, it is not necessary to dynamically determine the corresponding threshold. However, image color similarity based on histograms is relatively more sensitive to color richness. Dynamically determining the corresponding seventh preset threshold can effectively improve the accuracy of the judgment results.
[0088] In some embodiments, determining the target color richness type of the first image includes: performing quantization processing on each channel of each pixel in the first image based on a preset numerical interval to obtain a quantized image corresponding to the first image; counting the number of pixels containing the current pixel value in the quantized image for each pixel value; sorting the pixel values in the quantized image in descending order according to the number of pixels to obtain a sorting result; calculating the sum of the number of pixels for the first N pixel values in the sorting result; and determining that the target color richness type of the first image is a single type in response to the ratio of the sum of the number of pixels to the total number of pixels in the quantized image being greater than an eighth preset threshold. This allows for quick and accurate determination of the color richness type.
[0089] For example, the first image is an RGB (red, green, and blue) image. Each color channel (red, green, and blue) is quantized based on a preset numerical interval (e.g., 20). Specifically, the channel value is divided by the preset numerical interval, rounded down, and then multiplied by the preset numerical interval. This allows pixels with similar colors to be approximated as the same pixel, facilitating subsequent pixel count. For instance, for a pixel with original RGB values of (10, 30, 70) and quantized RGB values of (0, 20, 60), and another pixel with original RGB values of (11, 31, 71) and quantized RGB values of (0, 20, 60), the number of pixels corresponding to each quantized pixel value is counted. The number of pixels with the highest frequency of the N (N can be 5, for example) values is accumulated. If the ratio of the accumulated number to the total number of pixels is greater than a certain value, such as an eighth preset threshold (e.g., 70%), the target color richness type of the first image can be determined to be a single type; otherwise, it is a rich type.
[0090] In some embodiments, determining the target similarity category of the target image pair based on the text recognition result includes: in response to the text recognition result indicating that the target image pair contains text, determining the target similarity category of the target image pair as a fourth category; and in response to the text recognition result indicating that the target image pair does not contain text, determining the target similarity category of the target image pair as a fifth category. Thus, fast and accurate classification can be performed based on whether the text recognition result contains text.
[0091] Optionally, the similarity determination method corresponding to the target similarity determination category is used to determine whether the first image and the second image are similar images, including: in response to the target similarity determination category of the target image pair being a fourth determination category, determining whether the text similarity of the target image pair is greater than a first value; if so, determining whether the hash similarity of the target image pair based on hash value is greater than a second value; if so, determining whether the feature similarity of the target image pair based on scale-invariant feature transform is greater than a third value; if so, determining that the first image and the second image are similar images in response to the target image pair satisfying at least one of the preset determination conditions; in response to the target similarity determination category of the target image pair being a fifth determination category, determining whether the hash similarity of the target image pair based on hash value is greater than a second value; if so, determining whether the feature similarity of the target image pair based on scale-invariant feature transform is greater than a third value; if so, determining that the first image and the second image are similar images in response to the target image pair satisfying at least one of the preset determination conditions. Therefore, if the target image pair contains text, text similarity can be determined first. If the text similarity does not meet the requirements, the determination result can be obtained quickly, thus improving the determination efficiency.
[0092] Figure 3 This is a flowchart illustrating another image-based processing method provided in this disclosure. This disclosure optimizes the various optional solutions described in the above embodiments. Specifically, the method includes the following steps:
[0093] Step 301: Perform text recognition on the target image pair to obtain the text recognition result, wherein the target image pair includes a first image and a second image.
[0094] Optionally, text recognition can be performed on any one of the images in the target image pair.
[0095] Step 302: Determine whether the text recognition result contains text. If yes, proceed to step 303; otherwise, proceed to step 307.
[0096] For example, if the target image pair contains text, the target similarity determination category is determined to be the fourth determination category, and step 303 is executed in response to the target image pair being determined to be the fourth determination category; if the target image pair does not contain text, the target similarity determination category is determined to be the fifth determination category, and step 307 is executed in response to the target image pair being determined to be the fifth determination category.
[0097] Step 303: Determine whether the text similarity of the target image pair is greater than the first value. If yes, proceed to step 304; otherwise, proceed to step 311.
[0098] For example, the first value can be preset, such as 0.65.
[0099] Step 304: Determine whether the hash similarity of the target image pair based on the hash value is greater than the second value. If yes, proceed to step 305; otherwise, proceed to step 311.
[0100] For example, the second value can be preset, such as 0.7.
[0101] Step 305: Determine whether the feature similarity of the target image pair based on scale-invariant feature transformation is greater than the third value. If yes, proceed to step 306; otherwise, proceed to step 311.
[0102] For example, the third value can be preset, such as 0.3.
[0103] Step 306: Determine if the target image satisfies at least one of the preset judgment conditions. If so, proceed to step 310; otherwise, proceed to step 311.
[0104] For example, the target image pair satisfying the preset set of judgment conditions includes: the histogram-based image color similarity of the target image pair is greater than a fourth value, wherein the fourth value is greater than the second value and the third value; the feature similarity of the target image pair is greater than a fifth value, wherein the fifth value is greater than the third value; and the hash similarity of the target image pair is greater than a sixth value, wherein the sixth value is greater than the second value. Therefore, after performing a relatively low-requirement serial judgment on hash similarity and feature similarity in the aforementioned steps, ensuring that the first image and the second image have relatively comprehensive similarity, a higher-requirement parallel judgment is then performed, reducing false judgments and improving judgment accuracy.
[0105] For example, the fourth, fifth, and sixth values can be preset, such as the fourth value being 0.9, the fifth value being 0.65, and the sixth value being 0.9.
[0106] Step 307: Determine whether the hash similarity of the target image pair is greater than the second value. If yes, proceed to step 308; otherwise, proceed to step 311.
[0107] Step 308: Determine whether the feature similarity of the target image pair is greater than the third value. If yes, proceed to step 309; otherwise, proceed to step 311.
[0108] Step 309: Determine if the target image satisfies at least one of the preset judgment conditions. If so, proceed to step 310; otherwise, proceed to step 311.
[0109] Step 310: Determine that the first image and the second image are similar images.
[0110] Step 311: Determine that the first image and the second image are dissimilar images.
[0111] The image-based processing method provided in this disclosure first performs text recognition on the target image pair for which similarity determination is required. Based on the text recognition results, it is determined whether the target image pair contains text. If it contains text, a serial determination based on text similarity, hash similarity, and feature similarity is performed first, followed by a parallel determination based on image color similarity, feature similarity, and hash similarity. If it does not contain text, a serial determination based on hash similarity and feature similarity is performed first, followed by a parallel determination based on image color similarity, feature similarity, and hash similarity, ensuring the accuracy of the determination results.
[0112] Figure 4 This is a schematic diagram of the structure of an image-based processing apparatus provided in an embodiment of this disclosure, such as... Figure 4 As shown, the device includes:
[0113] The text recognition module 401 is used to perform text recognition on a target image pair to obtain a text recognition result, wherein the target image pair includes a first image and a second image;
[0114] The category determination module 402 is used to determine the target similarity determination category of the target image pair based on the text recognition result. The target similarity determination category is a preset similarity determination category in a preset similarity determination category set. Different preset similarity determination categories correspond to different similarity determination methods. The similarity determination method is a determination method based on at least two similarity algorithms.
[0115] The similarity determination module 403 is used to determine whether the first image and the second image are similar images by using the similarity determination method corresponding to the target similarity determination category.
[0116] The image-based processing apparatus provided in this embodiment first performs text recognition on target image pairs requiring similarity determination. Based on the text recognition results, it determines the target similarity determination category for the target image pairs. Different preset similarity determination categories correspond to different similarity determination methods. Using the similarity determination method corresponding to the target similarity determination category, it determines whether the first image and the second image are similar images. The similarity determination method is based on at least two similarity algorithms. By adopting the above technical solution, the flexibility, robustness, and accuracy of image similarity determination can be improved.
[0117] Optionally, the text recognition module is used to: perform text recognition on the first image in the target image pair to obtain the text recognition result;
[0118] The category determination module includes:
[0119] The region area determination unit is used to determine the region area of the text region contained in the first image based on the text recognition result;
[0120] A proportion calculation unit is used to calculate the proportion of the area of the region relative to the image area of the first image;
[0121] The first category determination unit is used to determine the target similarity determination category of the target image pair as the first determination category in response to the proportion being greater than or equal to the first preset proportion threshold.
[0122] The second category determination unit is used to determine the target similarity determination category of the target image pair as the second determination category in response to the proportion being greater than or equal to the second preset proportion threshold and less than the first preset proportion threshold.
[0123] The third category determination unit is used to determine the target similarity judgment category of the target image pair as the third judgment category in response to the proportion being less than the second preset proportion threshold.
[0124] Optionally, the similarity determination module includes at least one of the following:
[0125] The first determination unit is configured to, in response to the target similarity determination category being the first determination category, determine whether the first image and the second image are similar images based on the text similarity of the target image pair and the image color similarity based on the histogram;
[0126] The second determination unit is used to determine whether the first image and the second image are similar images based on the text similarity of the target image pair, the image color similarity based on the histogram, the feature similarity based on the scale-invariant feature transform, and the hash similarity based on the hash value in response to the target similarity determination category being the second determination category.
[0127] The third determination unit is used to determine whether the first image and the second image are similar images based on the histogram-based image color similarity, the scale-invariant feature transformation-based feature similarity, and the hash value-based hash similarity of the target image pair in response to the target similarity determination category being the third determination category.
[0128] Optionally, determining whether the first image and the second image are similar images based on the text similarity and histogram-based image color similarity of the target image pair includes:
[0129] Calculate the product of the percentage and the text similarity of the target image pair to obtain the first product;
[0130] Calculate the product of the target value and the histogram-based image color similarity of the target image pair to obtain a second product, wherein the target value is equal to 1 minus the proportion;
[0131] In response to the fact that the sum of the first product and the second product is greater than a first preset threshold, the first image and the second image are determined to be similar images;
[0132] Optionally, determining whether the first image and the second image are similar images based on the text similarity, histogram-based image color similarity, scale-invariant feature transform-based feature similarity, and hash value-based hash similarity of the target image pair includes:
[0133] The first image and the second image are determined to be similar images in response to the target image pair satisfying multiple of the following conditions:
[0134] The feature similarity of the target image pair based on scale-invariant feature transformation is greater than the product of a first preset coefficient and a target value, wherein the target value is equal to 1 minus the proportion;
[0135] The hash similarity of the target image pair based on the hash value is greater than the product of the second preset coefficient and the target value, wherein the target value is equal to 1 minus the proportion, and the second preset coefficient is greater than the first preset coefficient;
[0136] The text similarity of the target image pair is greater than a second preset threshold, wherein the second preset threshold is less than the first preset coefficient;
[0137] The histogram-based color similarity of the target image pair is greater than a third preset threshold, wherein the third preset threshold is greater than the second preset threshold;
[0138] Optionally, determining whether the first image and the second image are similar images based on histogram-based image color similarity, scale-invariant feature transform-based feature similarity, and hash value-based hash similarity of the target image pair includes:
[0139] The first image and the second image are determined to be similar images in response to the target image pair satisfying at least one of the following:
[0140] The feature similarity of the target image pair based on scale-invariant feature transformation is greater than a fourth preset threshold;
[0141] The hash similarity of the target image pair based on hash value is greater than a fifth preset threshold;
[0142] The feature similarity of the target image pair based on scale-invariant feature transform is greater than a sixth preset threshold, and the image color similarity of the target image pair based on histogram is greater than a seventh preset threshold, wherein the sixth preset threshold is less than the fourth preset threshold.
[0143] Optionally, the device may also include:
[0144] A color richness determination module is used to determine the target color richness type of the first image in response to the target similarity determination category being the third determination category, wherein the target color richness type is a color richness type in a preset color richness type set, and the preset color richness type set includes single types and rich types;
[0145] A threshold determination module is used to determine a seventh preset threshold based on the target color richness type, wherein the seventh preset threshold corresponding to the single type is greater than the seventh preset threshold corresponding to the richness type;
[0146] Optionally, determining the target color richness type of the first image includes:
[0147] For each channel of each pixel in the first image, quantization processing is performed based on a preset numerical interval to obtain the quantized image corresponding to the first image;
[0148] For each pixel value in the quantized image, count the number of pixels containing the current pixel value in the quantized image;
[0149] The pixel values in the quantized image are sorted in descending order according to the number of pixels to obtain the sorting result;
[0150] Calculate the sum of the number of pixels for the first N pixel values in the sorting result;
[0151] In response to the ratio of the quantity to the total number of pixels in the quantized image being greater than an eighth preset threshold, the target color richness type of the first image is determined to be a single type.
[0152] Optionally, the category determination module includes:
[0153] The fourth category determination unit is used to determine the target similarity determination category of the target image pair as the fourth determination category in response to the text recognition result being text-containing;
[0154] The fifth category determination unit is used to determine the target similarity determination category of the target image pair as the fifth determination category in response to the text recognition result being that it does not contain text;
[0155] Specifically, determining whether the first image and the second image are similar images using the similarity determination method corresponding to the target similarity determination category includes:
[0156] In response to the target image pair being classified as a fourth determination category, it is determined whether the text similarity of the target image pair is greater than a first value. If so, it is determined whether the hash similarity of the target image pair based on hash value is greater than a second value. If so, it is determined whether the feature similarity of the target image pair based on scale-invariant feature transformation is greater than a third value. If so, in response to the target image pair satisfying at least one of the preset determination conditions, the first image and the second image are determined to be similar images.
[0157] In response to the target image pair being classified as the fifth classification category, it is determined whether the hash similarity of the target image pair based on the hash value is greater than the second value. If so, it is determined whether the feature similarity of the target image pair based on scale-invariant feature transformation is greater than the third value. If so, in response to the target image pair satisfying at least one of the preset judgment conditions, the first image and the second image are determined to be similar images.
[0158] The set of preset judgment conditions for the target image pairs includes:
[0159] The histogram-based color similarity of the target image pair is greater than a fourth value, and the fourth value is greater than the second value and the third value.
[0160] The feature similarity of the target image pair is greater than a fifth value, wherein the fifth value is greater than the third value;
[0161] The hash similarity of the target image pair is greater than a sixth value, wherein the sixth value is greater than the second value.
[0162] The image-based processing apparatus provided in this disclosure can execute the image-based processing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method.
[0163] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0164] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Reference is made below. Figure 5 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 5 The diagram below shows the structure of the terminal device or server 500. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0165] like Figure 5 As shown, electronic device 500 may include a processing unit (e.g., central processing unit, graphics processor, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from storage device 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. An edit / output (I / O) interface 505 is also connected to bus 504.
[0166] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0167] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.
[0168] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0169] The electronic device provided in this embodiment and the image-based processing method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0170] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the image-based processing method provided in the above embodiments.
[0171] This disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the image-based processing method provided in the above embodiments.
[0172] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0173] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0174] The aforementioned computer-readable medium carries one or more programs. When the electronic device executes the aforementioned one or more programs, the electronic device causes the following actions: to perform text recognition on a pair of target images to obtain a text recognition result, wherein the pair of target images includes a first image and a second image; to determine a target similarity judgment category for the pair of target images based on the text recognition result, wherein the target similarity judgment category is a preset similarity judgment category from a preset set of similarity judgment categories, different preset similarity judgment categories correspond to different similarity judgment methods, and the similarity judgment method is a judgment method based on at least two similarity algorithms; and to determine whether the first image and the second image are similar images using the similarity judgment method corresponding to the target similarity judgment category.
[0175] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0176] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0177] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The names of modules do not necessarily limit the module itself; for example, a similarity determination module can also be described as "a module that determines whether the first image and the second image are similar images by using the similarity determination method corresponding to the target similarity determination category".
[0178] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0179] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0180] According to one or more embodiments of this disclosure, an image-based processing method is provided, comprising:
[0181] Text recognition is performed on a pair of target images to obtain text recognition results, wherein the pair of target images includes a first image and a second image;
[0182] The target similarity judgment category of the target image pair is determined based on the text recognition result. The target similarity judgment category is a preset similarity judgment category in a preset similarity judgment category set. Different preset similarity judgment categories correspond to different similarity judgment methods. The similarity judgment method is a judgment method based on at least two similarity algorithms.
[0183] The similarity determination method corresponding to the target similarity determination category is used to determine whether the first image and the second image are similar images.
[0184] According to one or more embodiments of this disclosure, the step of performing text recognition on a target image pair to obtain text recognition results includes:
[0185] Perform text recognition on the first image in the target image pair to obtain the text recognition result;
[0186] The step of determining the target similarity category of the target image pair based on the text recognition result includes:
[0187] The area of the text region contained in the first image is determined based on the text recognition results;
[0188] Calculate the ratio of the area of the region to the area of the first image;
[0189] In response to the proportion being greater than or equal to a first preset proportion threshold, the target similarity judgment category of the target image pair is determined to be the first judgment category;
[0190] In response to the proportion being greater than or equal to a second preset proportion threshold and less than a first preset proportion threshold, the target similarity judgment category of the target image pair is determined to be the second judgment category;
[0191] In response to the proportion being less than the second preset proportion threshold, the target similarity judgment category of the target image pair is determined to be the third judgment category.
[0192] According to one or more embodiments of this disclosure, determining whether the first image and the second image are similar images using a similarity determination method corresponding to the target similarity determination category includes at least one of the following:
[0193] In response to the target similarity determination category being the first determination category, the first image and the second image are determined to be similar images based on the text similarity of the target image pair and the image color similarity based on the histogram;
[0194] In response to the target similarity determination category being the second determination category, the first image and the second image are determined to be similar images based on the text similarity of the target image pair, the image color similarity based on histogram, the feature similarity based on scale-invariant feature transform, and the hash similarity based on hash value.
[0195] In response to the target similarity determination category being the third determination category, the first image and the second image are determined to be similar images based on the histogram-based image color similarity, the scale-invariant feature transformation-based feature similarity, and the hash value-based hash similarity of the target image pair.
[0196] According to one or more embodiments of this disclosure, determining whether the first image and the second image are similar images based on the text similarity and histogram-based image color similarity of the target image pair includes:
[0197] Calculate the product of the percentage and the text similarity of the target image pair to obtain the first product;
[0198] Calculate the product of the target value and the histogram-based image color similarity of the target image pair to obtain a second product, wherein the target value is equal to 1 minus the proportion;
[0199] In response to the fact that the sum of the first product and the second product is greater than a first preset threshold, the first image and the second image are determined to be similar images;
[0200] The step of determining whether the first image and the second image are similar images based on the text similarity, histogram-based image color similarity, scale-invariant feature transform-based feature similarity, and hash value-based hash similarity of the target image pair includes:
[0201] The first image and the second image are determined to be similar images in response to the target image pair satisfying multiple of the following conditions:
[0202] The feature similarity of the target image pair based on scale-invariant feature transformation is greater than the product of a first preset coefficient and a target value, wherein the target value is equal to 1 minus the proportion;
[0203] The hash similarity of the target image pair based on the hash value is greater than the product of the second preset coefficient and the target value, wherein the target value is equal to 1 minus the proportion, and the second preset coefficient is greater than the first preset coefficient;
[0204] The text similarity of the target image pair is greater than a second preset threshold, wherein the second preset threshold is less than the first preset coefficient;
[0205] The histogram-based color similarity of the target image pair is greater than a third preset threshold, wherein the third preset threshold is greater than the second preset threshold;
[0206] The step of determining whether the first image and the second image are similar images based on histogram-based image color similarity, scale-invariant feature transform-based feature similarity, and hash value-based hash similarity of the target image pair includes:
[0207] The first image and the second image are determined to be similar images in response to the target image pair satisfying at least one of the following:
[0208] The feature similarity of the target image pair based on scale-invariant feature transformation is greater than a fourth preset threshold;
[0209] The hash similarity of the target image pair based on hash value is greater than a fifth preset threshold;
[0210] The feature similarity of the target image pair based on scale-invariant feature transform is greater than a sixth preset threshold, and the image color similarity of the target image pair based on histogram is greater than a seventh preset threshold, wherein the sixth preset threshold is less than the fourth preset threshold.
[0211] According to one or more embodiments of this disclosure, it further includes:
[0212] In response to the target similarity determination category being the third determination category, the target color richness type of the first image is determined, wherein the target color richness type is a color richness type in a preset color richness type set, and the preset color richness type set includes single type and rich type;
[0213] A seventh preset threshold is determined based on the target color richness type, wherein the seventh preset threshold corresponding to the single type is greater than the seventh preset threshold corresponding to the richness type;
[0214] The step of determining the target color richness type of the first image includes:
[0215] For each channel of each pixel in the first image, quantization processing is performed based on a preset numerical interval to obtain the quantized image corresponding to the first image;
[0216] For each pixel value in the quantized image, count the number of pixels containing the current pixel value in the quantized image;
[0217] The pixel values in the quantized image are sorted in descending order according to the number of pixels to obtain the sorting result;
[0218] Calculate the sum of the number of pixels for the first N pixel values in the sorting result;
[0219] In response to the ratio of the quantity to the total number of pixels in the quantized image being greater than an eighth preset threshold, the target color richness type of the first image is determined to be a single type.
[0220] According to one or more embodiments of this disclosure, determining the target similarity category of the target image pair based on the text recognition result includes:
[0221] In response to the text recognition result indicating that the text is contained, the target similarity judgment category of the target image pair is determined to be the fourth judgment category;
[0222] In response to the text recognition result being that it does not contain text, the target similarity judgment category of the target image pair is determined to be the fifth judgment category;
[0223] Specifically, determining whether the first image and the second image are similar images using the similarity determination method corresponding to the target similarity determination category includes:
[0224] In response to the target image pair being classified as a fourth determination category, it is determined whether the text similarity of the target image pair is greater than a first value. If so, it is determined whether the hash similarity of the target image pair based on hash value is greater than a second value. If so, it is determined whether the feature similarity of the target image pair based on scale-invariant feature transformation is greater than a third value. If so, in response to the target image pair satisfying at least one of the preset determination conditions, the first image and the second image are determined to be similar images.
[0225] In response to the target image pair being classified as the fifth classification category, it is determined whether the hash similarity of the target image pair based on the hash value is greater than the second value. If so, it is determined whether the feature similarity of the target image pair based on scale-invariant feature transformation is greater than the third value. If so, in response to the target image pair satisfying at least one of the preset judgment conditions, the first image and the second image are determined to be similar images.
[0226] The set of preset judgment conditions for the target image pairs includes:
[0227] The histogram-based color similarity of the target image pair is greater than a fourth value, and the fourth value is greater than the second value and the third value.
[0228] The feature similarity of the target image pair is greater than a fifth value, wherein the fifth value is greater than the third value;
[0229] The hash similarity of the target image pair is greater than a sixth value, wherein the sixth value is greater than the second value.
[0230] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0231] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0232] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. An image-based processing method, characterized in that, include: Text recognition is performed on a pair of target images to obtain text recognition results, wherein the pair of target images includes a first image and a second image; The target similarity judgment category of the target image pair is determined based on the text recognition result. The target similarity judgment category is a preset similarity judgment category in a preset similarity judgment category set. Different preset similarity judgment categories correspond to different similarity judgment methods. The similarity judgment method is a judgment method based on at least two similarity algorithms. The similarity determination method corresponding to the target similarity determination category is used to determine whether the first image and the second image are similar images.
2. The method according to claim 1, characterized in that, The process of performing text recognition on the target image pair to obtain the text recognition result includes: Perform text recognition on the first image in the target image pair to obtain the text recognition result; The step of determining the target similarity category of the target image pair based on the text recognition result includes: The area of the text region contained in the first image is determined based on the text recognition results; Calculate the ratio of the area of the region to the area of the first image; In response to the proportion being greater than or equal to a first preset proportion threshold, the target similarity judgment category of the target image pair is determined to be the first judgment category; In response to the proportion being greater than or equal to a second preset proportion threshold and less than a first preset proportion threshold, the target similarity judgment category of the target image pair is determined to be the second judgment category; In response to the proportion being less than the second preset proportion threshold, the target similarity judgment category of the target image pair is determined to be the third judgment category.
3. The method according to claim 2, characterized in that, The similarity determination method corresponding to the target similarity determination category is used to determine whether the first image and the second image are similar images, including at least one of the following: In response to the target similarity determination category being the first determination category, the first image and the second image are determined to be similar images based on the text similarity of the target image pair and the image color similarity based on the histogram; In response to the target similarity determination category being the second determination category, the first image and the second image are determined to be similar images based on the text similarity of the target image pair, the image color similarity based on histogram, the feature similarity based on scale-invariant feature transform, and the hash similarity based on hash value. In response to the target similarity determination category being the third determination category, the first image and the second image are determined to be similar images based on the histogram-based image color similarity, the scale-invariant feature transformation-based feature similarity, and the hash value-based hash similarity of the target image pair.
4. The method according to claim 3, Its features are, in, Determining whether the first image and the second image are similar images based on the text similarity and histogram-based image color similarity of the target image pair includes: Calculate the product of the percentage and the text similarity of the target image pair to obtain the first product; Calculate the product of the target value and the histogram-based image color similarity of the target image pair to obtain a second product, wherein the target value is equal to 1 minus the proportion; In response to the fact that the sum of the first product and the second product is greater than a first preset threshold, the first image and the second image are determined to be similar images; The step of determining whether the first image and the second image are similar images based on the text similarity, histogram-based image color similarity, scale-invariant feature transform-based feature similarity, and hash value-based hash similarity of the target image pair includes: The first image and the second image are determined to be similar images in response to the target image pair satisfying multiple of the following conditions: The feature similarity of the target image pair based on scale-invariant feature transformation is greater than the product of a first preset coefficient and a target value, wherein the target value is equal to 1 minus the proportion; The hash similarity of the target image pair based on the hash value is greater than the product of the second preset coefficient and the target value, wherein the target value is equal to 1 minus the proportion, and the second preset coefficient is greater than the first preset coefficient; The text similarity of the target image pair is greater than a second preset threshold, wherein the second preset threshold is less than the first preset coefficient; The histogram-based color similarity of the target image pair is greater than a third preset threshold, wherein the third preset threshold is greater than the second preset threshold; The step of determining whether the first image and the second image are similar images based on histogram-based image color similarity, scale-invariant feature transform-based feature similarity, and hash value-based hash similarity of the target image pair includes: The first image and the second image are determined to be similar images in response to the target image pair satisfying at least one of the following: The feature similarity of the target image pair based on scale-invariant feature transformation is greater than a fourth preset threshold; The hash similarity of the target image pair based on hash value is greater than a fifth preset threshold; The feature similarity of the target image pair based on scale-invariant feature transform is greater than a sixth preset threshold, and the image color similarity of the target image pair based on histogram is greater than a seventh preset threshold, wherein the sixth preset threshold is less than the fourth preset threshold.
5. The method according to claim 4, characterized in that, Also includes: In response to the target similarity determination category being the third determination category, the target color richness type of the first image is determined, wherein the target color richness type is a color richness type in a preset color richness type set, and the preset color richness type set includes single type and rich type; A seventh preset threshold is determined based on the target color richness type, wherein the seventh preset threshold corresponding to the single type is greater than the seventh preset threshold corresponding to the richness type; The step of determining the target color richness type of the first image includes: For each channel of each pixel in the first image, quantization processing is performed based on a preset numerical interval to obtain the quantized image corresponding to the first image; For each pixel value in the quantized image, count the number of pixels containing the current pixel value in the quantized image; The pixel values in the quantized image are sorted in descending order according to the number of pixels to obtain the sorting result; Calculate the sum of the number of pixels for the first N pixel values in the sorting result; In response to the ratio of the quantity to the total number of pixels in the quantized image being greater than an eighth preset threshold, the target color richness type of the first image is determined to be a single type.
6. The method according to claim 1, characterized in that, The step of determining the target similarity category of the target image pair based on the text recognition result includes: In response to the text recognition result indicating that the text is contained, the target similarity judgment category of the target image pair is determined to be the fourth judgment category; In response to the text recognition result being that it does not contain text, the target similarity judgment category of the target image pair is determined to be the fifth judgment category; Specifically, determining whether the first image and the second image are similar images using the similarity determination method corresponding to the target similarity determination category includes: In response to the target image pair being classified as a fourth determination category, it is determined whether the text similarity of the target image pair is greater than a first value. If so, it is determined whether the hash similarity of the target image pair based on hash value is greater than a second value. If so, it is determined whether the feature similarity of the target image pair based on scale-invariant feature transformation is greater than a third value. If so, in response to the target image pair satisfying at least one of the preset determination conditions, the first image and the second image are determined to be similar images. In response to the target image pair being classified as the fifth classification category, it is determined whether the hash similarity of the target image pair based on the hash value is greater than the second value. If so, it is determined whether the feature similarity of the target image pair based on scale-invariant feature transformation is greater than the third value. If so, in response to the target image pair satisfying at least one of the preset judgment conditions, the first image and the second image are determined to be similar images. The set of preset judgment conditions for the target image pairs includes: The histogram-based color similarity of the target image pair is greater than a fourth value, and the fourth value is greater than the second value and the third value. The feature similarity of the target image pair is greater than a fifth value, wherein the fifth value is greater than the third value; The hash similarity of the target image pair is greater than a sixth value, wherein the sixth value is greater than the second value.
7. An image-based processing apparatus, characterized in that, include: The text recognition module is used to perform text recognition on a target image pair to obtain text recognition results, wherein the target image pair includes a first image and a second image; The category determination module is used to determine the target similarity determination category of the target image pair based on the text recognition result. The target similarity determination category is a preset similarity determination category in a preset similarity determination category set. Different preset similarity determination categories correspond to different similarity determination methods. The similarity determination method is a determination method based on at least two similarity algorithms. The similarity determination module is used to determine whether the first image and the second image are similar images by adopting the similarity determination method corresponding to the target similarity determination category.
8. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the image-based processing method as described in any one of claims 1-6.
9. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the image-based processing method as described in any one of claims 1-6.
10. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the image-based processing method as described in any one of claims 1-6.