Image retrieval method and device, equipment and storage medium

By controlling the character area ratio and performing appropriate geometric and color transformation during the image enhancement process, training sample pairs are constructed and image retrieval model is optimized, which solves the problem of data enhancement destroying structured features and improves the accuracy and reliability of image retrieval.

CN120336573AActive Publication Date: 2025-07-18HEFEI IFLYTEK TOYCLOUD TECH
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510812302.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-18
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

Data enhancement schemes in the prior art are prone to destroy structured features, making it difficult for image retrieval models to effectively capture important structural semantics, and reduce the accuracy and reliability of search results.

Method used

By enhancing the original structured image, it is ensured that the area ratio of the first character area area of the image sample to the second character area of the original structured image is not lower than the area ratio threshold. Combined with geometric transformation, color transformation and noise interference, a training sample pair is constructed and the image retrieval model is optimized using the contrast loss function.

Benefits of technology

While image enhancement, structured features are effectively protected, the image retrieval model's generalization ability of image retrieval model to image changes is enhanced, and the accuracy and reliability of search results are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336573A_ABST
    Figure CN120336573A_ABST
Patent Text Reader

Abstract

The invention provides an image retrieval method and device, equipment and a storage medium, and relates to the technical field of image processing, and the method comprises the steps: obtaining a to-be-retrieved image; inputting the to-be-retrieved image into the image retrieval model to obtain an image retrieval result output by the image retrieval model; the image retrieval model is obtained by training an initial image retrieval model based on at least one training sample pair, and image samples in the training sample pair are obtained by performing image enhancement on an original structured image; the area ratio corresponding to the area of the first character region of the image sample and the area of the second character region of the original structured image matched with the image sample is not smaller than an area ratio threshold value. According to the method, structural features are effectively protected while image enhancement is achieved, the generalization ability of the image retrieval model to image changes is improved, and then the accuracy and reliability of image retrieval based on the image retrieval model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to an image retrieval method, apparatus, device, and storage medium. Background Art

[0002] In the technical field of image processing, especially when processing image retrieval with strong structured features (such as textbook image retrieval, literature image retrieval, etc.), the core information of such images not only includes text content, but also includes structured features such as text layout, formula structure, table layout, and chart elements.

[0003] In related technologies, although the contrast learning method effectively improves the generalization ability of the model to image changes through some data augmentation schemes, when processing such images with strong structured features, the existing data augmentation schemes are extremely likely to destroy these structured features. Due to the destruction of structured features, problems such as loss of key information and distortion of feature representation will occur, which will cause the model to be difficult to effectively capture important structural semantics, thereby reducing the accuracy and reliability of the retrieval results. Summary of the Invention

[0004] The present invention provides an image retrieval method, apparatus, device, and storage medium, which are used to solve the problem that the data augmentation scheme in the prior art is extremely likely to destroy structured features, resulting in the model being difficult to effectively capture important structural semantics, thereby reducing the accuracy and reliability of the retrieval results, and realizing effective protection of structured features while performing image enhancement, improving the generalization ability of the model to image changes, and the accuracy and reliability of image enhancement retrieval.

[0005] The present invention provides an image retrieval method, and the method includes: Obtain an image to be retrieved; Input the image to be retrieved into an image retrieval model to obtain an image retrieval result output by the image retrieval model; the image retrieval model is trained based on at least one pair of training samples, and the image samples in the at least one pair of training samples are obtained by performing image enhancement on an original structured image; Wherein, the area ratio of the area of the first character region of the image sample to the area of the second character region of the original structured image matching the image sample is not less than an area ratio threshold.

[0006] According to an image retrieval method provided by the present invention, the image samples in the at least one pair of training samples are obtained through the following method: Crop at least one candidate region image from the original structured image; Determine that the target candidate region image in the at least one candidate region image is the first image sample corresponding to the original structured image; the area ratio of the area of the first character region of the target candidate region image to the area of the second character region of the original structured image is not less than the area ratio threshold. Based on the first image sample and / or the second image sample corresponding to the first image sample, obtain the image samples in the at least one training sample pair; the second image sample is obtained by performing at least one of geometric transformation, color transformation, and noise interference on the first image sample for image enhancement.

[0007] According to an image retrieval method provided by the present invention, the cropping of at least one candidate region image from the original structured image includes: Based on the structured layout features of the original structured image, determine the minimum cropping ratio corresponding to the original structured image. Crop at least one candidate region image not less than the minimum cropping ratio from the original structured image.

[0008] According to an image retrieval method provided by the present invention, the cropping of at least one candidate region image from the original structured image includes: Determine the aspect ratio of the original structured image. Based on the aspect ratio of the original structured image, crop at least one candidate region image from the original structured image.

[0009] According to an image retrieval method provided by the present invention, it further includes: Perform a geometric transformation on the first image sample to obtain a second image sample. Wherein, when performing an affine transformation on the first image sample, control the parameter values of the affine transformation parameters not to exceed the corresponding preset parameter value ranges; the affine transformation parameters include rotation, scaling, translation, and shear.

[0010] According to an image retrieval method provided by the present invention, it further includes: Perform a geometric transformation on the first image sample to obtain a geometrically transformed image sample. Perform a color transformation on the geometrically transformed image sample to obtain a second image sample.

[0011] According to an image retrieval method provided by the present invention, it further includes: Perform a tensor conversion on the first image sample and / or the second image sample to obtain a standardized tensor image sample.

[0012] The present invention also provides an image retrieval device, and the device includes: A first image retrieval module for obtaining an image to be retrieved; A second image retrieval module for inputting the image to be retrieved into an image retrieval model to obtain an image retrieval result output by the image retrieval model; the image retrieval model is obtained by training an initial image retrieval model based on at least one pair of training samples, and the image samples in the at least one pair of training samples are obtained by performing image enhancement on an original structured image; Wherein, the area ratio of the area of the first character region of the image sample to the area of the second character region of the original structured image matching the image sample is not less than an area ratio threshold.

[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the image retrieval method as described in any one of the above.

[0014] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the image retrieval method as described in any one of the above.

[0015] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the image retrieval method as described in any one of the above.

[0016] The image retrieval method provided by the present invention, by ensuring that the area ratio of the area of the first character region of the image sample after image enhancement to the area of the second character region of the original structured image is not lower than the area ratio threshold, reduces the damage to the structured features during the image enhancement process, thereby effectively protecting the structured features while performing image enhancement, improving the generalization ability of the image retrieval model to image changes, and further improving the accuracy and reliability of image retrieval based on the image retrieval model. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 is one of the flow diagrams of the image retrieval method provided by the present invention.

[0019] Figure 2 is another flow diagram of the image retrieval method provided by the present invention.

[0020] Figure 3 It is the third schematic flow chart of the image retrieval method provided by the present invention.

[0021] Figure 4 It is the schematic structural diagram of the image retrieval device provided by the present invention.

[0022] Figure 5 It is the schematic structural diagram of the electronic device provided by the present invention. Specific embodiments

[0023] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0024] Contrastive learning is a self-supervised learning method. It constructs positive sample pairs (similar samples) and negative sample pairs (dissimilar samples) through data augmentation schemes (such as transformation operations like cropping, rotating, and color adjustment of images), and enables the model to come into contact with multimodal presentation forms of the same semantic content during the learning process through positive sample pairs and negative sample pairs.

[0025] However, when existing contrastive learning methods process such images with strong structural features, these structural features carrying key semantic information are extremely easily damaged through simple image augmentation schemes. For example, random cropping or rotation may cut off the relationship between the formula and the context, making the formula difficult to recognize. Such operations will cause the enhanced image samples to deviate from the semantic expression of the original structural features, thereby destroying the semantic consistency of the positive sample pairs, causing the model to misclassify the structurally damaged positive sample pairs as dissimilar, and then learning incorrect feature representations. And the negative sample pairs may generate false similarities due to structural perturbations, causing the model to confuse semantic differences, resulting in problems such as a decline in matching accuracy and semantic understanding deviation in the image retrieval task, and it is difficult to meet the strict requirements of strong structured scenarios for feature integrity and semantic accuracy.

[0026] Based on this, the embodiments of the present invention provide an image retrieval method to reduce the damage to structural features during the image augmentation process, effectively protect the structural features while performing image augmentation, improve the generalization ability of the image retrieval model to image changes, and further improve the accuracy and reliability of image retrieval based on the image retrieval model.

[0027] Figure 1 It is one of the schematic flow charts of the image retrieval method provided by the present invention. As Figure 1 shown, the method includes the following steps 110 and 120.

[0028] Step 110, obtain the image to be retrieved.

[0029] Here, the image to be retrieved refers to the image submitted by the user that needs to retrieve the same or relevant content from the database or knowledge base. For example, the user submits an image of a mathematical formula in a textbook and hopes to retrieve the specific page of the textbook where the mathematical formula in this image of the mathematical formula comes from. Another example is that the user submits a screenshot of a table in a document and hopes to retrieve the page of the document where the table in this screenshot of the table comes from.

[0030] In one example, the image to be retrieved can be obtained passively by an image acquisition device. For example, the user takes an image with a mobile device installed with an image retrieval device. After taking the picture, the user manually uploads the taken image to the image retrieval device. Another example is that the user obtains an image from a third-party website. These images can be generated through the scanning function provided by the website or directly downloaded electronic images. Then the user manually uploads the obtained image to the image retrieval device.

[0031] In one example, the image to be retrieved can also be obtained actively by the image retrieval device. For example, the image retrieval device captures the writing content of the user on the image acquisition interface in real time through the configured image acquisition interface. These writing contents can be text, formulas, charts, etc. After the image retrieval device captures these writing contents through the image acquisition interface, it converts these writing contents into images. Another example is that the image retrieval device uses its built-in camera to scan the paper materials provided by the user into images in real time. The user only needs to place the paper materials within the shooting range of the camera, and the image retrieval device can automatically perform scanning to obtain images.

[0032] Step 120, input the image to be retrieved into the image retrieval model to obtain the image retrieval result output by the image retrieval model; the image retrieval model is obtained by training an initial image retrieval model based on at least one pair of training samples, and the image samples in the at least one pair of training samples are obtained by performing image enhancement on the original structured image.

[0033] Wherein, the area ratio of the area of the first character region of the image sample to the area of the second character region of the original structured image matching the image sample is not less than the area ratio threshold.

[0034] Here, the image retrieval result refers to one or more images that the image retrieval model finds in the database or knowledge base and are highly similar to the image to be retrieved in terms of text content, text layout, formula structure, table layout and other features.

[0035] When the image retrieval results are multiple images, the multiple images can usually be presented in a sorted form on the display interface of the image retrieval results, and can also be presented in descending order of similarity. In addition, each image can also be accompanied by a similarity score, which is used to measure the degree of matching with the image to be retrieved.

[0036] In one example, the base model of the image retrieval model can adopt a model with the ability to extract image visual features, such as a convolutional neural network model, a Transformer model, a multimodal model, etc.

[0037] In practical applications, multiple training sample pairs are pre-constructed, and the multiple training sample pairs include multiple positive sample pairs and multiple negative sample pairs. A positive sample pair means that two images look different visually but have the same content. For example, the same textbook page is scanned using different scanners or scanning devices (such as scanners or scanning devices with different resolutions, contrasts, brightness, etc.) to construct positive sample pairs. A negative sample pair means that the contents of two images are different (such as different textbook pages). After constructing multiple positive sample pairs and multiple negative sample pairs, the multiple positive sample pairs and multiple negative sample pairs are used, and combined with a contrast loss function, the parameters of the base model are updated through backpropagation to obtain a trained image retrieval model.

[0038] In this step, image samples can be constructed by performing image enhancement processing on the original structured image, such as cropping, scaling, grayscaling, etc. Then, two images that look different visually but have the same content are selected from the image samples or the original structured image to construct positive sample pairs. Two images with different contents are selected from the image samples or the original structured image to construct negative sample pairs.

[0039] Here, the original structured image can refer to an image that has not undergone any enhancement or processing and has clear structured semantic features.

[0040] In one example, when cropping a specific proportion of the image from the original structured image as an image sample, a ratio threshold of the area can be set to measure whether the structured features are damaged in the cropped image sample.

[0041] Specifically, the area ratio is a quantization index used to measure the change degree of the area of a specific region (i.e., the character region) in the image sample relative to the area of the corresponding region in the original image. By setting a ratio threshold of the area, it can be ensured that during the image enhancement process, the area change of the character region will not exceed a certain range, thereby avoiding excessive damage to the structured features.

[0042] In one example, the area ratio threshold can be determined based on the semantic recognition critical damage rate of each type of structured feature, that is, the maximum acceptable area loss ratio when the semantic recognition accuracy of the structured feature significantly decreases during the cropping process. For example, for an original image that only contains the same type of structured feature, cropped samples with different area ratios are generated, and then a semantic recognition model is used to evaluate the semantic accuracy of the cropped samples with different area ratios. The area ratio when the semantic accuracy drops to P% of the original level is used as the threshold for this type of feature. Here, P can be flexibly set according to actual needs.

[0043] Finally, for each original structured image, the area ratio threshold corresponding to this original structured image can be determined based on the types of structured features existing in this original structured image, the area ratio of the character regions of each type of structured feature, and the semantic recognition critical damage rate of each type of structured feature. For example, the area ratio threshold can be calculated according to the formula where T , n is the number of types of structured features, is the area ratio of the character region of the i th type of feature, is the i th type of feature's semantic recognition critical damage rate.

[0044] For example, during the image enhancement process, for different original structured images 1, original structured image 2, and original structured image 3, based on the types of structured features existing in these three images, the proportion of each type of structured feature, and the semantic recognition critical damage rate of each type of structured feature, combined with the formula , the corresponding area ratio thresholds are determined to be A%, B%, and C% respectively.

[0045] After cropping multiple candidate region images from the original structured image 1, according to the set area ratio threshold A%, images whose area ratio of the first character region of the image to the second character region of the original structured image 1 is not less than A% are selected from the multiple candidate region images as image samples. After cropping multiple candidate region images from the original structured image 2, according to the set area ratio threshold B%, images whose area ratio of the first character region of the image to the second character region of the original structured image 2 is not less than B% are selected from the multiple candidate region images as image samples. After cropping multiple candidate region images from the original structured image 3, according to the set area ratio threshold C%, images whose area ratio of the first character region of the image to the second character region of the original structured image 3 is not less than C% are selected from the multiple candidate region images as image samples.

[0046] In one example, the area ratio threshold can also be determined through experimental values, that is, designing experiments to verify the influence of different area ratios on the model performance. For example, by setting multiple different area ratios, and respectively creating datasets corresponding to each area ratio for training the model. Finally, using performance metrics (such as accuracy, recall rate, F1-score, etc.) to evaluate the performance of the model trained under the datasets corresponding to different area ratios, and selecting the area ratio with the best performance on the validation set as the area ratio threshold.

[0047] For example, during the image enhancement process, if the area ratio threshold is determined to be D% according to the experiment, then for all the original structured images, after cropping out the candidate region images from the original structured images, according to the set area ratio threshold D%, select as image samples those images from all the candidate region images whose area ratio of the first character region area to the second character region area of the original structured image is not less than D%.

[0048] In addition, the area ratio threshold can also be flexibly set by the user according to their needs or experience, and there is no limitation on this.

[0049] It should be noted here that the character region area can refer to the pixel area occupied by the character in the image, or the contour area occupied by the character contour in the image, and there is no limitation on this. However, it should be understood that the calculation method of the character region area for the candidate region images cropped from the same original structured image or the same batch of original structured images should be the same as that of the character region area of this original structured image.

[0050] In one example, when performing geometric transformation on the original structured image, such as affine transformation, a preset parameter value range can also be set to avoid over-transformation from destroying the structured features.

[0051] It should be noted that affine transformation is a linear geometric transformation, which can perform operations such as translation, rotation, scaling, and shearing on the image through matrix operations, and maintain the parallelism of straight lines. Structured features refer to the features in the image with clear geometric shapes, regular arrangements, or fixed patterns. Therefore, for structured features, affine transformation will affect their original structures (such as formula symbol misalignment or incomplete formulas, etc.). In this embodiment, by limiting the range of its parameter values, over-transformation from destroying the structured features is avoided.

[0052] Here, the setting of the range of preset parameter values can be determined according to the actual application scenario and image features. For example, if the structured features in the image are not sensitive to position changes, the range of parameter values of the translation parameter can be set larger. If the structured features in the image are not sensitive to direction changes, the range of parameter values of the rotation parameter can be set larger. If the structured features in the image are not sensitive to size changes, the range of parameter values of the scaling parameter can be set larger. If the structured features in the image are not sensitive to inclination changes, the range of parameter values of the shear parameter can be set larger. On the contrary, the corresponding range of parameter values should be set smaller to ensure that during the image enhancement process, the change of the structured features is within an acceptable range, thus avoiding the destruction of the structured features.

[0053] In some examples, when performing image enhancement processing on the original structured image, such as cropping, scaling, grayscale conversion and other steps, other constraints can also be introduced to protect the structured features. For example, a regularization term can be used to limit the intensity of the enhancement operation to construct an image sample. For another example, when cropping the original structured image, the minimum cropping ratio can be limited.

[0054] It should be understood that the minimum cropping ratio refers to the ratio of the area of the entire image of the image sample to the area of the entire image of the original structured image. For example, a minimum cropping ratio (such as 0.5) can be set according to experience, and the minimum cropping ratio corresponding to the original structured image can also be determined according to the structured layout features in the original structured image (such as text layout structure features, formula and symbol layout features, chart and graphic relationship features) to ensure that when cropping based on this minimum cropping ratio subsequently, the destruction of the structured features during the cropping process is reduced.

[0055] Based on the above image enhancement steps, it can be ensured that under the condition of small sample training, both efficient model convergence and optimization are achieved, and the structured features of the image can be effectively protected, the destruction of the structured features is reduced, the generalization ability of the model to image changes is improved, and thus the accuracy and reliability of image enhancement retrieval are improved.

[0056] The image retrieval method provided by the embodiment of the present invention reduces the destruction of the structured features during the image enhancement process by ensuring that the area ratio of the area of the first character region of the image sample after image enhancement to the area of the second character region of the original structured image is not lower than the area ratio threshold, thereby effectively protecting the structured features while enhancing the image, improving the generalization ability of the image retrieval model to image changes, and further improving the accuracy and reliability of image retrieval based on the image retrieval model.

[0057] In some embodiments, the image samples in the training sample pairs used by the above image retrieval model can be obtained through the following method, refer to Figure 2 , Figure 2This is the second schematic flow diagram of the image retrieval method provided by the present invention. As Figure 2 shown, the method includes the following steps 210, step 220, and step 230.

[0058] Step 210: Crop at least one candidate region image from the original structured image.

[0059] In this step, a ratio range (for example, 0.5 to 1.0) can be preset in advance. This ratio range represents the range of the area ratio of the region allowed to be cropped to the original image.

[0060] Then, randomly crop at least one candidate region image from the original structured image according to the preset ratio range. The area ratio of this candidate region image to the original structured image is within the preset ratio range.

[0061] Here, the width-to-height ratio of the cropped candidate region image is not restricted and can be flexibly adjusted according to the actual situation. For example, it is the same as the width-to-height ratio of the original structured image, or different from the width-to-height ratio of the original structured image.

[0062] Step 220: Determine the target candidate region image among the at least one candidate region image as the first image sample corresponding to the original structured image; the area ratio of the first character region area of the target candidate region image to the second character region area of the original structured image is not less than the area ratio threshold.

[0063] After obtaining multiple candidate region images, based on the preset area ratio threshold corresponding to the character region, select the image that meets the requirements from the multiple candidate region images as the first image sample.

[0064] Here, the character region area can refer to the pixel area occupied by the character in the image, or the contour area occupied by the character contour in the image, etc., and there is no restriction on this. However, it should be understood that the calculation method of the character region area of the candidate region image cropped from the same original structured image should be the same as that of the character region area of this original structured image.

[0065] Here, the setting of the area ratio threshold can be the same as above, and will not be specifically described here.

[0066] In one example, the second text region mask in the original structured image and the first text region mask in the candidate region image can be obtained respectively through a text line detection algorithm. Here, both the first text region mask and the second text region mask are binary images, where the pixel value of the text line region is 1 (or 255), and the pixel value of the non-text line region is 0. Based on this, the text line region in the image can be determined.

[0067] Next, count the number of pixels with a value of 1 (or 255) in the first text area mask to determine the area of the first character area of the first text area mask. Similarly, count the number of pixels with a value of 1 (or 255) in the second text area mask to determine the area of the second character area of the second text area mask. Finally, calculate the area ratio of the first character area to the area corresponding to the second character area of the original structured image.

[0068] Step 230, obtain the image samples in the at least one training sample pair based on the first image sample and / or the second image sample corresponding to the first image sample; the second image sample is obtained by performing at least one of geometric transformation, color transformation, and noise interference on the first image sample for image enhancement.

[0069] In one example, the first image sample after the above cropping can be used as the image sample for model training, and the second image sample obtained by performing image transformation on the first image sample can also be used as the image sample for model training. The first image sample and the second image sample can also be used together as the image sample for model training, and there is no limitation on this.

[0070] Here, the second image sample is obtained by performing one or more image enhancements on the first image sample. For example, performing geometric transformation (such as translation, rotation, scaling, shearing, and affine transformation, etc.) on the first image sample. Color transformation (such as brightness adjustment, contrast adjustment, color space conversion, etc.) can also be performed on the first image sample. In addition, random noise (such as Gaussian blur) can be added to the first image sample to simulate the noise situation in actual image acquisition.

[0071] Based on this, after using the image samples obtained by the above image enhancements, multiple training sample pairs are constructed, and the multiple training sample pairs include multiple positive sample pairs and multiple negative sample pairs. Among them, the positive sample pair can be the first image sample and the second image sample from the same original structured image, or two second image samples from the same original structured image but with different image enhancements. The negative sample pair can be the first image sample and / or the second image sample from different original structured images.

[0072] After constructing multiple positive sample pairs and multiple negative sample pairs, use the multiple positive sample pairs and multiple negative sample pairs, and combine with the contrast loss function to update the base model parameters through backpropagation to obtain a trained image retrieval model. In this way, the accuracy and reliability of image retrieval can be improved through the trained image retrieval model.

[0073] The image retrieval method provided by the embodiments of the present invention reduces the damage to the structured features during the cropping process by ensuring that the area ratio of the first character area of the image sample after image enhancement to the second character area of the original structured image is not lower than the area ratio threshold, thereby effectively protecting the structured features while enhancing the image. Then, an image enhancement transformation is continued on the basis of the first image sample to obtain a second image sample, improving the generalization ability of the image retrieval model to image changes, and further improving the accuracy and reliability of image retrieval based on the image retrieval model.

[0074] In some embodiments, the above candidate region images can be obtained in the following manner. Refer to Figure 3 , Figure 3 FIG. 3 is a schematic flowchart of the image retrieval method provided by the present invention. As Figure 3 shown, the method includes the following steps 3101, step 3102, step 320, and step 330.

[0075] Step 3101: Based on the structured layout features of the original structured image, determine the minimum cropping ratio corresponding to the original structured image.

[0076] Here, the structured layout features refer to the spatial arrangement, organizational form, and mutual relationship of various structured elements (such as text, formulas, symbols, charts, etc.) in the image.

[0077] For example, text layout structure features (row and column distribution of text, paragraph spacing, hierarchical relationship between headings and text, etc.), formula and symbol layout features (formula typesetting, arrangement of special symbols such as superscripts and subscripts), chart and graphic relationship features (chart position, graphic relationship), etc.

[0078] In one example, through the identified structured layout features of the original structured image, determine the minimum region boundary corresponding to the structured features, and then determine the cropping boundary according to the minimum region boundary corresponding to the structured features. Here, the minimum region boundary corresponding to the structured features refers to the minimum region range in the image that can completely contain a certain structured feature. By identifying the structured layout features, determining the boundaries of each structured feature, and integrating all the boundaries, the minimum region boundary is determined.

[0079] Finally, according to the determined cropping boundary, calculate the ratio of the cropped image to the original image, and this ratio can be used as the minimum cropping ratio corresponding to the original structured image.

[0080] Here, the minimum cropping ratio corresponding to each original structured image can be determined according to its own structured layout features. It is also possible to uniformly determine the minimum cropping ratio according to all the structured layout features after determining the structured layout features corresponding to all the structured images, and there is no limitation on this.

[0081] Step 3102, crop at least one candidate region image not less than the minimum cropping ratio from the original structured image.

[0082] After determining the minimum cropping ratio, based on the minimum cropping ratio, crop multiple candidate region images from the original structured image.

[0083] Step 320, determine that the target candidate region image among the at least one candidate region image is the first image sample corresponding to the original structured image; the area ratio of the first character region area of the target candidate region image to the area of the second character region area of the original structured image is not less than the area ratio threshold.

[0084] After obtaining multiple candidate region images, based on the preset area ratio threshold corresponding to the character region, select the images that meet the requirements from the multiple candidate region images as the first image samples.

[0085] Step 330, obtain the image samples in the at least one training sample pair based on the first image sample and / or the second image sample corresponding to the first image sample; the second image sample is obtained by performing at least one of geometric transformation, color transformation, and noise interference on the first image sample for image enhancement.

[0086] In one example, the first image sample after the above cropping can be used as the image sample for model training, and the second image sample obtained by further performing image transformation on the first image sample can also be used as the image sample for model training. The first image sample and the second image sample can also be used together as the image sample for model training, and there is no limitation on this.

[0087] Here, the second image sample is obtained by performing one or more image enhancements on the first image sample. For example, performing geometric transformation (such as translation, rotation, scaling, shearing, and affine transformation, etc.) on the first image sample. Also, for example, color transformation (such as brightness adjustment, contrast adjustment, color space conversion, etc.) can be performed on the first image sample. In addition, random noise (such as Gaussian blur) can be added to the first image sample to simulate the noise situation in actual image acquisition.

[0088] Based on this, after using the image samples obtained by the above image enhancement, multiple training sample pairs are constructed, and the multiple training sample pairs include multiple positive sample pairs and multiple negative sample pairs. Among them, the positive sample pair can be the first image sample and the second image sample from the same original structured image, or two second image samples from the same original structured image but with different image enhancements. The negative sample pair can be the first image sample and / or the second image sample from different original structured images.

[0089] After constructing multiple positive sample pairs and multiple negative sample pairs, use the multiple positive sample pairs and multiple negative sample pairs, and combine with the contrast loss function to update the base model parameters through backpropagation to obtain a trained image retrieval model. In this way, through the trained image retrieval model, the accuracy and reliability of image retrieval are improved.

[0090] The image retrieval method provided by the embodiments of the present invention reduces the damage to the structured features during the cropping process through the dual cropping constraints of the minimum cropping ratio and the area ratio threshold, thereby effectively protecting the structured features while enhancing the image, improving the generalization ability of the image retrieval model to image changes, and further improving the accuracy and reliability of image retrieval based on the image retrieval model.

[0091] Based on the above embodiments, the cropping of at least one candidate region image from the original structured image may further include: Determine the aspect ratio of the original structured image; Based on the aspect ratio of the original structured image, crop at least one candidate region image from the original structured image.

[0092] In this embodiment, the aspect ratio of the original structured image is used as a reference for cropping. For example, if the aspect ratio of the original image is 0.65, then a candidate region image within a specific ratio range (e.g., 0.5 - 1.0) is randomly cropped from the original structured image according to the aspect ratio of 0.65. This ratio range represents the range of the area ratio of the cropped region to the original image.

[0093] The image retrieval method provided by the embodiments of the present invention crops at least one candidate region image from the original structured image based on the aspect ratio of the original structured image, which can ensure that the cropped candidate region image can completely retain the key structured features and avoid the risk of damage to the structured features caused by blind aspect ratio cropping.

[0094] Based on the above embodiments, it further includes: Perform a geometric transformation on the first image sample to obtain a second image sample; Wherein, when performing an affine transformation on the first image sample, control the parameter values of the affine transformation parameters not to exceed the corresponding preset parameter value ranges; the affine transformation parameters include rotation, scaling, translation, and shearing.

[0095] In one example, the geometric transformation includes but is not limited to: rigid body transformation, affine transformation, and projection transformation.

[0096] Here, a rigid body transformation refers to a geometric transformation that preserves the shape and size of a figure, only changing its position and orientation, such as translation and rotation. An affine transformation adds linear transformations of scaling and shearing on the basis of a rigid body transformation, maintaining the parallelism of straight lines, but not guaranteeing the invariance of length and angle. A projective transformation includes translation, rotation, scaling, shearing, and perspective transformation.

[0097] Here, the setting of the preset parameter value range can be determined according to the actual application scenario and image features. For example, if the structured features in the image are not sensitive to position changes, then the parameter value range of the translation parameter can be set larger. If the structured features in the image are not sensitive to direction changes, then the parameter value range of the rotation parameter can be set larger. If the structured features in the image are not sensitive to size changes, then the parameter value range of the scaling parameter can be set larger. If the structured features in the image are not sensitive to inclination changes, then the parameter value range of the shearing parameter can be set larger. Conversely, the corresponding parameter value range should be set smaller to ensure that during the image enhancement process, the changes in the structured features are within an acceptable range, thereby avoiding the destruction of the structured features.

[0098] In one example, an affine transformation can be performed on the first image based on the following affine transformation parameter matrix T : ; where represents that the rotation angle of the image around the origin is within ; represents that the translation ratios of the image in the horizontal and vertical directions are both within ; represents that the scaling ratios of the image in the horizontal and vertical directions are both within ; represents that the shearing angles of the image in the horizontal and vertical directions are both within .

[0099] The image retrieval method provided by the embodiments of the present invention increases data diversity by performing multi-type geometric transformations on the first image sample. Additionally, when performing an affine transformation, the parameter values of the affine transformation parameters are restricted within the preset parameter value range to avoid damage to the structured features caused by excessive transformation operations, thereby effectively protecting the structured features while enhancing the image, improving the generalization ability of the image retrieval model to image changes, and further improving the accuracy and reliability of image retrieval based on the image retrieval model.

[0100] Based on the above embodiments, it further includes: Performing a geometric transformation on the first image sample to obtain a geometrically transformed image sample; Performing a color transformation on the geometrically transformed image sample to obtain a second image sample.

[0101] In this embodiment, first, geometric transformations are performed on the first image sample, such as at least one of rigid transformation, affine transformation, and projection transformation. Then, color transformation is performed on the geometrically transformed image sample, such as color jittering (simulating illumination changes and shooting condition differences in real scenes by randomly adjusting the brightness, contrast, saturation, and hue of the image), to obtain the transformed second image sample.

[0102] In one example, when performing color jittering processing, RGB values can be randomly added or subtracted to simulate image acquisition under different illumination conditions. The distribution range of pixel values can also be adjusted to simulate low-contrast scanning or faded printed paper materials. By adjusting the vividness of colors, the color distortion problem during image acquisition of color charts or illustrations in the image can be simulated. By randomly offsetting the camera in the HSV space, the white balance difference during image acquisition can be simulated.

[0103] It should be noted that in this embodiment, by first performing geometric transformation to change the spatial structure of the image and then performing color transformation to change the color attributes of the pixels, the integrity of the structural features of the image in space can be ensured, and the complexity of color transformation can be reduced, thereby improving the efficiency and accuracy of the entire processing flow.

[0104] Based on the above embodiments, it further includes: Performing tensor conversion on the first image sample and / or the second image sample to obtain an image sample of a standardized tensor.

[0105] It should be understood that images are stored in the form of matrices in a computer. For example, a color image is a three-dimensional matrix, and a grayscale image is a two-dimensional matrix. A tensor is a high-dimensional extension of a matrix and is used to uniformly represent data in deep learning.

[0106] In this embodiment, after obtaining multiple constructed first image samples and multiple second image samples, the dimensions of these image samples are adjusted to unify the original dimensions, and then linear normalization and zero-mean normalization are performed on the images, so as to obtain an image sample of a standardized tensor.

[0107] The image retrieval method provided by the embodiment of the present invention can ensure that the format and distribution of the input data meet the expectations of the model by performing tensor conversion to obtain an image sample of a standardized tensor, thereby improving the training speed and performance of the model.

[0108] Based on any of the above embodiments, the present invention further provides an image retrieval device, refer to Figure 4 , and the device includes: A first image retrieval module 410, configured to obtain an image to be retrieved; A second image retrieval module 420, configured to input the image to be retrieved into an image retrieval model to obtain an image retrieval result output by the image retrieval model; the image retrieval model is obtained by training an initial image retrieval model based on at least one pair of training samples, and the image samples in the at least one pair of training samples are obtained by performing image enhancement on an original structured image; Wherein, the area ratio of the area of the first character region of the image sample to the area of the second character region of the original structured image matching the image sample is not less than an area ratio threshold.

[0109] The image retrieval device provided by the present invention reduces the damage to the structured features during the image enhancement process by ensuring that the area ratio of the area of the first character region of the image sample after image enhancement to the area of the second character region of the original structured image is not lower than the area ratio threshold, thereby effectively protecting the structured features while performing image enhancement, improving the generalization ability of the image retrieval model to image changes, and further improving the accuracy and reliability of image retrieval based on the image retrieval model.

[0110] Figure 5 An example of a schematic physical structure diagram of an electronic device is shown as Figure 5 shown. The electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call logic instructions in the memory 530 to execute an image retrieval method, and the method includes: Obtain an image to be retrieved; Input the image to be retrieved into an image retrieval model to obtain an image retrieval result output by the image retrieval model; the image retrieval model is obtained by training an initial image retrieval model based on at least one pair of training samples, and the image samples in the at least one pair of training samples are obtained by performing image enhancement on an original structured image; Wherein, the area ratio of the area of the first character region of the image sample to the area of the second character region of the original structured image matching the image sample is not less than an area ratio threshold.

[0111] In addition, when the logical instructions in the above-mentioned memory 530 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0112] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the image retrieval method provided by the above-mentioned various methods. The method includes: Obtain the image to be retrieved; Input the image to be retrieved into the image retrieval model to obtain the image retrieval result output by the image retrieval model. The image retrieval model is trained from an initial image retrieval model based on at least one pair of training samples. The image samples in the at least one pair of training samples are obtained by performing image enhancement on the original structured image; Wherein, the area ratio of the area of the first character region of the image sample to the area of the second character region of the original structured image matched by the image sample is not less than the area ratio threshold.

[0113] In yet another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the image retrieval method provided by the above-mentioned various methods. The method includes: Obtain the image to be retrieved; Input the image to be retrieved into the image retrieval model to obtain the image retrieval result output by the image retrieval model. The image retrieval model is trained from an initial image retrieval model based on at least one pair of training samples. The image samples in the at least one pair of training samples are obtained by performing image enhancement on the original structured image; Wherein, the area ratio of the area of the first character region of the image sample to the area of the second character region of the original structured image matched by the image sample is not less than the area ratio threshold.

[0114] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.

[0115] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An image retrieval method, characterized in that, The method includes: Obtaining an image to be retrieved; Inputting the image to be retrieved into an image retrieval model to obtain an image retrieval result output by the image retrieval model; the image retrieval model is obtained by training an initial image retrieval model based on at least one pair of training samples, and the image samples in the at least one pair of training samples are obtained by performing image enhancement on the original structured image; Wherein, the area ratio of the area of the first character region of the image sample to the area of the second character region of the original structured image matching the image sample is not less than the area ratio threshold.

2. The image retrieval method according to claim 1, wherein The image samples in the at least one pair of training samples are obtained by the following method: Cropping at least one candidate region image from the original structured image; Determining the target candidate region image in the at least one candidate region image as the first image sample corresponding to the original structured image; The area ratio of the area of the first character region of the target candidate region image to the area of the second character region of the original structured image is not less than the area ratio threshold; Based on the first image sample and / or the second image sample corresponding to the first image sample, obtaining the image samples in the at least one pair of training samples; the second image sample is obtained by performing at least one of geometric transformation, color transformation, and noise interference on the first image sample for image enhancement.

3. The image retrieval method according to claim 2, wherein The cropping of at least one candidate region image from the original structured image includes: Based on the structured layout features of the original structured image, determining the minimum cropping ratio corresponding to the original structured image; Cropping at least one candidate region image not less than the minimum cropping ratio from the original structured image.

4. The image retrieval method according to claim 2, characterized in that The cropping of at least one candidate region image from the original structured image includes: Determining the aspect ratio of the original structured image; Based on the aspect ratio of the original structured image, cropping at least one candidate region image from the original structured image.

5. The image retrieval method according to claim 2, wherein It further includes: Performing geometric transformation on the first image sample to obtain a second image sample; Wherein, in the case of performing affine transformation on the first image sample, controlling the parameter values of the affine transformation parameters not to exceed the corresponding preset parameter value ranges; the affine transformation parameters include rotation, scaling, translation, and shear.

6. The image retrieval method according to claim 2, wherein It further includes: Performing geometric transformation on the first image sample to obtain a geometrically transformed image sample; Performing color transformation on the geometrically transformed image sample to obtain a second image sample.

7. The image retrieval method according to any one of claims 2 to 6, characterized in that, It further includes: Performing tensor conversion on the first image sample and / or the second image sample to obtain a standardized tensor image sample.

8. An image retrieval device, characterized in that, The apparatus includes: A first image retrieval module for obtaining an image to be retrieved; A second image retrieval module for inputting the image to be retrieved into an image retrieval model to obtain an image retrieval result output by the image retrieval model; the image retrieval model is obtained by training an initial image retrieval model based on at least one pair of training samples, and the image samples in the at least one pair of training samples are obtained by performing image enhancement on the original structured image; Wherein, the area ratio of the area of the first character region of the image sample to the area of the second character region of the original structured image matching the image sample is not less than the area ratio threshold.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the image retrieval method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the image retrieval method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Image similarity retrieval method

    CN108595596A

  • Effective fine image classification method based on optimized feature weight

    CN111639206A

  • Image retrieval method and device based on multi-task learning, equipment and medium

    CN114282037A

  • Model training method, image retrieval method, equipment and computer readable medium

    CN117523330A

  • Deep learning-based qualification image classification method and system

    CN117788957A