Image retrieval method, device, equipment and storage medium

By controlling the character area ratio and performing appropriate geometric and color transformations during the image enhancement process to construct training sample pairs, the problem of data enhancement schemes destroying structured features is solved, and the accuracy and reliability of image retrieval are improved.

CN120336573BActive Publication Date: 2025-09-12HEFEI IFLYTEK TOYCLOUD TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510812302.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-09-12
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

In existing technologies, data augmentation schemes easily destroy structural features, making it difficult for the model to effectively capture important structural semantics, thereby reducing the accuracy and reliability of image retrieval results.

Method used

By ensuring that the character area ratio of the image sample is not lower than the area ratio threshold during the image enhancement process, and combining geometric transformation, color transformation and noise interference, training sample pairs are constructed to reduce the destruction of structured features and protect the integrity of structured features.

Benefits of technology

It improves the generalization ability of the image retrieval model to image changes, enhances the accuracy and reliability of image retrieval, and ensures the integrity of structured features during the enhancement process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336573B_ABST
    Figure CN120336573B_ABST
Patent Text Reader

Abstract

The present invention provides an image retrieval method, apparatus, device, and storage medium, relating to the field of image processing technology. The method comprises obtaining an image to be retrieved; inputting the image to be retrieved into an image retrieval model to obtain an image retrieval result output by the image retrieval model; the image retrieval model is obtained by training an initial image retrieval model based on at least one training sample pair, wherein the image samples in the training sample pair are obtained by performing image enhancement on the original structured image; and the area ratio of the area of ​​the first character region of the image sample to the area of ​​the second character region of the original structured image that matches the image sample is not less than an area ratio threshold. The present invention effectively protects structural features while enhancing the image, improves the generalization ability of the image retrieval model to image changes, and thereby improves the accuracy and reliability of image retrieval based on the image retrieval model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to an image retrieval method, apparatus, device and storage medium. Background Art

[0002] In the field of image processing technology, especially when processing image retrieval with strong structured features (such as textbook image retrieval, document image retrieval, etc.), the core information of such images not only includes text content, but also includes structured features such as text typesetting, formula structure, table layout, and chart elements.

[0003] In related technologies, contrastive learning methods effectively improve the model's generalization ability to image changes through certain data augmentation schemes. However, when processing images with strong structural features, existing data augmentation schemes are prone to destroying these structural features. This destruction of structural features can lead to the loss of key information and distortion of feature representations, making it difficult for the model to effectively capture important structural semantics, thereby reducing the accuracy and reliability of retrieval results. Summary of the Invention

[0004] The present invention provides an image retrieval method, apparatus, device and storage medium, which are used to solve the problem that data enhancement schemes in the prior art are very likely to destroy structural features, making it difficult for the model to effectively capture important structural semantics, thereby reducing the accuracy and reliability of the retrieval results. The method effectively protects structural features while enhancing the image, improves the model's generalization ability to image changes, and enhances the accuracy and reliability of image enhancement retrieval.

[0005] The present invention provides an image retrieval method, comprising:

[0006] Get the image to be retrieved;

[0007] Inputting the image to be retrieved into an image retrieval model to obtain an image retrieval result output by the image retrieval model; the image retrieval model is obtained by training an initial image retrieval model based on at least one training sample pair, and the image samples in the at least one training sample pair are obtained by performing image enhancement on the original structured image;

[0008] The area ratio of the first character region of the image sample to the second character region of the original structured image that matches the image sample is not less than an area ratio threshold.

[0009] According to an image retrieval method provided by the present invention, the image samples in the at least one training sample pair are obtained by:

[0010] Crop at least one candidate region image from the original structured image;

[0011] Determining that a target candidate region image in the at least one candidate region image is a first image sample corresponding to the original structured image; an area ratio of an area of ​​a first character region of the target candidate region image to an area corresponding to a second character region of the original structured image is not less than an area ratio threshold;

[0012] Based on the first image sample and / or the second image sample corresponding to the first image sample, an image sample in the at least one training sample pair is obtained; the second image sample is obtained by performing at least one image enhancement of geometric transformation, color transformation and noise interference on the first image sample.

[0013] According to an image retrieval method provided by the present invention, cropping at least one candidate region image from an original structured image includes:

[0014] Determining a minimum cropping ratio corresponding to the original structured image based on the structured layout features of the original structured image;

[0015] At least one candidate region image having a cropping ratio not less than the minimum cropping ratio is cropped from the original structured image.

[0016] According to an image retrieval method provided by the present invention, cropping at least one candidate region image from an original structured image includes:

[0017] determining the aspect ratio of the original structured image;

[0018] At least one candidate region image is cropped from the original structured image based on the aspect ratio of the original structured image.

[0019] An image retrieval method according to the present invention further includes:

[0020] Performing a geometric transformation on the first image sample to obtain a second image sample;

[0021] Wherein, when performing affine transformation on the first image sample, parameter values ​​of the parameters for controlling the affine transformation do not exceed corresponding preset parameter value ranges; the affine transformation parameters include rotation, scaling, translation and shearing.

[0022] An image retrieval method according to the present invention further includes:

[0023] Performing a geometric transformation on the first image sample to obtain a geometrically transformed image sample;

[0024] Performing color transformation on the geometrically transformed image sample to obtain a second image sample.

[0025] An image retrieval method according to the present invention further includes:

[0026] Perform tensor conversion on the first image sample and / or the second image sample to obtain normalized tensor image samples.

[0027] The present invention also provides an image retrieval device, comprising:

[0028] A first image retrieval module, used to obtain an image to be retrieved;

[0029] a second image retrieval module, configured to input the image to be retrieved into an image retrieval model to obtain an image retrieval result output by the image retrieval model; the image retrieval model is obtained by training an initial image retrieval model based on at least one training sample pair, wherein the image samples in the at least one training sample pair are obtained by performing image enhancement on the original structured image;

[0030] The area ratio of the first character region of the image sample to the second character region of the original structured image that matches the image sample is not less than an area ratio threshold.

[0031] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, any of the above-described image retrieval methods is implemented.

[0032] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which implements any of the above-mentioned image retrieval methods when executed by a processor.

[0033] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any one of the above-mentioned image retrieval methods.

[0034] The image retrieval method provided by the present invention reduces the damage to structured features during the image enhancement process by ensuring that the area ratio of the first character area of ​​the image sample after image enhancement to the area of ​​the second character area of ​​the original structured image is not less than an area ratio threshold, thereby effectively protecting the structured features while enhancing the image, improving the generalization ability of the image retrieval model to image changes, and further improving the accuracy and reliability of image retrieval based on the image retrieval model. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0036] Figure 1 This is one of the flow charts of the image retrieval method provided by the present invention.

[0037] Figure 2 This is the second flowchart of the image retrieval method provided by the present invention.

[0038] Figure 3 This is the third flow chart of the image retrieval method provided by the present invention.

[0039] Figure 4 It is a structural diagram of the image retrieval device provided by the present invention.

[0040] Figure 5 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0041] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0042] Contrastive learning is a self-supervised learning method that constructs positive sample pairs (similar samples) and negative sample pairs (dissimilar samples) through data augmentation schemes (such as cropping, rotating, color adjustment, and other transformation operations on images). Through positive and negative sample pairs, the model is exposed to multimodal presentations of the same semantic content during the learning process.

[0043] However, when processing images with strong structured features, existing contrastive learning methods can easily destroy these structured features that carry key semantic information through simple image enhancement schemes. For example, random cropping or rotation may sever the relationship between a formula and its context, making the formula difficult to recognize. Such operations can cause the enhanced image samples to deviate from the semantic expression of the original structured features, thereby destroying the semantic consistency of positive sample pairs. This can cause the model to mistakenly treat structurally damaged positive sample pairs as dissimilar, leading to incorrect feature representations. Negative sample pairs may produce false similarities due to structural perturbations, causing the model to confuse semantic differences, resulting in problems such as reduced matching accuracy and semantic understanding deviations in image retrieval tasks. This makes it difficult to meet the strict requirements for feature integrity and semantic accuracy in strongly structured scenarios.

[0044] Based on this, an embodiment of the present invention provides an image retrieval method to reduce the damage to structured features during image enhancement, effectively protect structured features while enhancing the image, improve the generalization ability of the image retrieval model to image changes, and thereby improve the accuracy and reliability of image retrieval based on the image retrieval model.

[0045] Figure 1 This is one of the flow charts of the image retrieval method provided by the present invention, such as Figure 1 As shown, the method includes the following steps 110 and 120.

[0046] Step 110: Obtain the image to be retrieved.

[0047] Here, a "searchable image" refers to an image submitted by a user for which the same or related content needs to be retrieved from a database or knowledge base. For example, a user may submit an image of a math formula from a textbook and want to retrieve the specific textbook and page of the textbook from which the math formula in the image originates. Another example is a user submitting a screenshot of a table from a document and wanting to retrieve the specific document and page of the document from which the table in the screenshot originates.

[0048] In one example, the image to be retrieved may be passively acquired by an image acquisition device. For example, a user may capture an image using a mobile device equipped with an image retrieval device. After capturing the image, the user manually uploads the captured image to the image retrieval device. Alternatively, a user may obtain an image from a third-party website. These images may be generated using a scanning function provided by the website or directly downloaded as electronic images. The user then manually uploads the captured image to the image retrieval device.

[0049] In one example, the image to be retrieved can also be actively acquired by the image retrieval device. For example, the image retrieval device uses a configured image acquisition interface to capture the user's written content on the image acquisition interface in real time. This written content can be text, formulas, charts, etc. After the image retrieval device captures this written content through the image acquisition interface, it converts it into an image. For example, the image retrieval device uses its built-in camera to scan paper materials provided by the user into images in real time. The user only needs to place the paper material within the camera's shooting range, and the image retrieval device will automatically scan and acquire the image.

[0050] Step 120: input the image to be retrieved into the image retrieval model to obtain the image retrieval result output by the image retrieval model; the image retrieval model is obtained by training the initial image retrieval model based on at least one training sample pair, and the image samples in the at least one training sample pair are obtained by image enhancement of the original structured image.

[0051] The area ratio of the first character region of the image sample to the second character region of the original structured image that matches the image sample is not less than an area ratio threshold.

[0052] Here, the image retrieval result refers to the image retrieval model finding one or more images in the database or knowledge base based on the input image to be retrieved that are highly similar to the image to be retrieved in terms of text content, text layout, formula structure, table layout and other features.

[0053] When the image search results contain multiple images, they are typically displayed in a sorted order on the image search result display interface, or in descending order of similarity. In addition, each image can be accompanied by a similarity score, which measures the degree of match with the image being searched.

[0054] In one example, the base model of the image retrieval model can adopt a convolutional neural network model, a Transformer model, a multimodal model, or other models capable of extracting image visual features.

[0055] In practical applications, multiple training sample pairs are pre-constructed, including multiple positive and negative pairs. A positive pair refers to two images that appear visually different but have the same content. For example, a positive pair is constructed by scanning the same textbook page using different scanners or scanning devices (e.g., scanners or scanning devices with different resolutions, contrast, brightness, etc.). A negative pair refers to two images with different content (e.g., different textbook pages). After constructing these multiple positive and negative pairs, the base model parameters are updated through backpropagation using these pairs and a contrastive loss function to obtain a trained image retrieval model.

[0056] In this step, image samples can be constructed by performing image enhancement processing on the original structured image, such as cropping, scaling, and grayscale conversion. Then, two images that appear visually different but have the same content are selected from the image samples or the original structured image to construct a positive sample pair. Two images with different content are selected from the image samples or the original structured image to construct a negative sample pair.

[0057] Here, the original structured image may refer to an image that has not been enhanced or processed in any way and has clear structured semantic features.

[0058] In one example, when an image of a specific ratio is cropped from an original structured image as an image sample, an area ratio threshold may be set to measure whether the structural features of the cropped image sample are destroyed.

[0059] Specifically, the area ratio is a quantitative metric used to measure the extent to which the area of ​​a specific region (i.e., the character region) in an image sample changes relative to the corresponding area in the original image. By setting an area ratio threshold, we can ensure that the area change of the character region does not exceed a certain range during the image enhancement process, thereby preventing excessive destruction of structural features.

[0060] In one example, the area ratio threshold can be determined based on the critical destruction rate of semantic recognition of each type of structured feature, that is, the maximum acceptable area loss ratio of the structured feature that causes a significant decrease in its semantic recognition accuracy during the cropping process. For example, for an original image containing only structured features of the same type, cropped samples with different area ratios are generated, and then the semantic recognition model is used to evaluate the semantic accuracy of the cropped samples with each area ratio, and the area ratio when the semantic accuracy drops to P% of the original level is used as the threshold for this type of feature. Here, P can be flexibly set according to actual needs.

[0061] Finally, for each original structured image, the area ratio threshold corresponding to the original structured image can be determined based on the various types of structured features in the original structured image, the character area ratio of each type of structured feature, and the semantic recognition critical destruction rate of each type of structured feature. For example, the area ratio threshold corresponding to the original structured image can be determined based on the formula Calculate the area ratio threshold T ,in, n is the number of types of structured features, It is i The character area ratio of the class feature, It is i Critical destruction rate of semantic recognition of class features.

[0062] For example, in the image enhancement process, for different original structured images 1, original structured image 2, and original structured image 3, according to the various types of structured features existing in these three images, the proportion of each type of structured features, and the semantic recognition critical destruction rate of each type of structured features, combined with the formula , and determine the corresponding area ratio thresholds as A%, B% and C% respectively.

[0063] After cropping multiple candidate area images from the original structured image 1, images whose area ratio between the first character area of ​​the image and the area corresponding to the second character area of ​​the original structured image 1 is not less than A% are screened out from the multiple candidate area images as image samples based on the set area ratio threshold value A%. After cropping multiple candidate area images from the original structured image 2, images whose area ratio between the first character area of ​​the image and the area corresponding to the second character area of ​​the original structured image 2 is not less than B% are screened out from the multiple candidate area images as image samples based on the set area ratio threshold value B%. After cropping multiple candidate area images from the original structured image 3, images whose area ratio between the first character area of ​​the image and the area corresponding to the second character area of ​​the original structured image 3 is not less than C% are screened out from the multiple candidate area images as image samples based on the set area ratio threshold value C%.

[0064] In one example, the area ratio threshold can be determined experimentally. This involves designing an experiment to verify the impact of different area ratios on model performance. For example, by setting multiple area ratios and creating datasets corresponding to each area ratio for model training, performance metrics (such as precision, recall, and F1 score) are used to evaluate the performance of the trained model under the datasets corresponding to different area ratios. The area ratio that achieves the best performance on the validation set is selected as the area ratio threshold.

[0065] For example, in the image enhancement process, the area ratio threshold is determined to be D% based on experiments. Then, for all original structured images, after cropping the candidate area images from the original structured images, images in which the area ratio of the first character area of ​​the image to the area of ​​the second character area of ​​the original structured image is not less than D% are screened out from all candidate area images according to the set area ratio threshold D%. These images are used as image samples.

[0066] In addition, the area ratio threshold can be flexibly set according to user needs or user experience, and there is no restriction on this.

[0067] It should be noted that the character region area can refer to the pixel area occupied by the character in the image or the contour area occupied by the character outline in the image, without limitation. However, it should be understood that the character region area calculation method for candidate region images cropped from the same original structured image or the same batch of original structured images should be the same as the calculation method for the character region area of ​​the original structured image.

[0068] In one example, when performing a geometric transformation, such as an affine transformation, on the original structured image, a preset parameter value range may be set to avoid excessive transformation that would destroy the structured features.

[0069] It should be noted that an affine transformation is a linear geometric transformation that can translate, rotate, scale, and shear an image through matrix operations, maintaining the parallelism of straight lines. Structural features refer to features in an image that have a clear geometric shape, regular arrangement, or fixed pattern. Therefore, affine transformations can affect the original structure of structural features (such as misplaced symbols or incomplete formulas). In this embodiment, the parameter value range is limited to prevent excessive transformations from destroying structural features.

[0070] Here, the setting of the preset parameter value range can be determined according to the actual application scenario and image characteristics. For example, if the structured features in the image are not sensitive to position changes, then the parameter value range of the translation parameter can be set larger. If the structured features in the image are not sensitive to direction changes, then the parameter value range of the rotation parameter can be set larger. If the structured features in the image are not sensitive to size changes, then the parameter value range of the scaling parameter can be set larger. If the structured features in the image are not sensitive to tilt changes, then the parameter value range of the shearing parameter can be set larger. On the contrary, the corresponding parameter value range should be set smaller to ensure that the changes in the structured features are within an acceptable range during the image enhancement process, thereby avoiding the destruction of the structured features.

[0071] In some examples, during image enhancement processing on the original structured image, such as cropping, scaling, and grayscaling, other constraints can be introduced to protect structural features, such as using regularization terms to limit the intensity of the enhancement operation to construct image samples. Another example is when cropping the original structured image, a minimum cropping ratio can be specified.

[0072] It should be understood that the minimum cropping ratio refers to the ratio of the entire image area of ​​the image sample to the entire image area of ​​the original structured image. For example, a minimum cropping ratio (e.g., 0.5) can be set based on experience. Alternatively, the minimum cropping ratio corresponding to the original structured image can be determined based on the structural layout features of the original structured image (e.g., text layout features, formula and symbol layout features, and chart and text relationship features). This ensures that subsequent cropping based on this minimum cropping ratio minimizes damage to structural features during the cropping process.

[0073] Based on the above image enhancement steps, we can ensure that under the condition of small sample training, we can not only achieve efficient model convergence and optimization, but also effectively protect the structural features of the image, reduce the damage to the structural features, and improve the model's generalization ability to image changes, thereby enhancing the accuracy and reliability of image enhancement retrieval.

[0074] The image retrieval method provided by an embodiment of the present invention reduces the damage to structured features during the image enhancement process by ensuring that the area ratio of the first character area of ​​the image sample after image enhancement to the area of ​​the second character area of ​​the original structured image is not less than an area ratio threshold, thereby effectively protecting the structured features while enhancing the image, improving the generalization ability of the image retrieval model to image changes, and further improving the accuracy and reliability of image retrieval based on the image retrieval model.

[0075] In some embodiments, the image samples in the training sample pairs used by the above image retrieval model can be obtained by the following method: Figure 2 , Figure 2 This is the second flow chart of the image retrieval method provided by the present invention, such as Figure 2 As shown, the method includes the following steps 210, 220 and 230.

[0076] Step 210: crop at least one candidate region image from the original structured image.

[0077] In this step, a ratio range (eg, 0.5 to 1.0) may be preset, and the ratio range indicates the ratio range of the area of ​​the cropped region allowed to occupy in the original image.

[0078] Next, at least one candidate region image is randomly cropped from the original structured image according to a preset ratio range, wherein the area ratio of the candidate region image to the original structured image is within the preset ratio range.

[0079] Here, the aspect ratio of the cropped candidate region image is not restricted and can be flexibly adjusted according to actual conditions, for example, the aspect ratio can be the same as that of the original structured image, or different from that of the original structured image.

[0080] Step 220: Determine that the target candidate area image in the at least one candidate area image is the first image sample corresponding to the original structured image; the area ratio of the first character area area of ​​the target candidate area image to the second character area area of ​​the original structured image is not less than an area ratio threshold.

[0081] After obtaining a plurality of candidate region images, an image that meets the requirements is selected from the plurality of candidate region images as a first image sample based on a preset area ratio threshold corresponding to the character region.

[0082] Here, the character region area may refer to the pixel area occupied by the character in the image, or the outline area occupied by the character outline in the image, etc., without limitation. However, it should be understood that the character region area calculation method for the candidate region image cropped from the same original structured image should be the same as the calculation method for the character region area of ​​the original structured image.

[0083] Here, the setting of the area ratio threshold may be the same as above and will not be described in detail here.

[0084] In one example, a text line detection algorithm can be used to obtain a second text region mask from the original structured image and a first text region mask from the candidate region image. Both the first and second text region masks are binary images, where the pixel values ​​of text line regions are 1 (or 255) and the pixel values ​​of non-text line regions are 0. Based on this, the text line regions in the image can be determined.

[0085] Next, the number of pixels in the first text region mask with a median value of 1 (or 255) is counted to determine the area of ​​the first character region in the first text region mask. Similarly, the number of pixels in the second text region mask with a median value of 1 (or 255) is counted to determine the area of ​​the second character region in the second text region mask. Finally, the area ratio of the first character region area to the second character region area in the original structured image is calculated.

[0086] Step 230: Obtain an image sample in the at least one training sample pair based on the first image sample and / or the second image sample corresponding to the first image sample; the second image sample is obtained by performing at least one image enhancement of geometric transformation, color transformation, and noise interference on the first image sample.

[0087] In one example, the first image sample after the above cropping can be used as the image sample for model training, and the second image sample after the first image sample undergoes image transformation can also be used as the image sample for model training. Alternatively, both the first image sample and the second image sample can be used as the image sample for model training, without limitation.

[0088] Here, the second image sample is obtained by performing one or more image enhancements on the first image sample. For example, the first image sample may be subjected to a geometric transformation (such as translation, rotation, scaling, shearing, and affine transformation). The first image sample may also be subjected to a color transformation (such as brightness adjustment, contrast adjustment, color space conversion, etc.). Furthermore, random noise (such as Gaussian blur) may be added to the first image sample to simulate the noise conditions encountered in actual image acquisition.

[0089] Based on this, after using the image samples obtained by the above image enhancement, multiple training sample pairs are constructed, including multiple positive sample pairs and multiple negative sample pairs. A positive sample pair can be a first image sample and a second image sample derived from the same original structured image, or two second image samples derived from the same original structured image but subjected to different image enhancements. A negative sample pair can be a first image sample and / or a second image sample derived from different original structured images.

[0090] After constructing multiple positive and negative sample pairs, we use these pairs, combined with a contrastive loss function, to update the base model parameters through backpropagation to obtain a trained image retrieval model. This improves the accuracy and reliability of image retrieval through the trained image retrieval model.

[0091] The image retrieval method provided by an embodiment of the present invention reduces the damage to structured features during the cropping process by ensuring that the area ratio of the first character area of ​​the image sample after image enhancement to the area of ​​the second character area of ​​the original structured image is not less than an area ratio threshold, thereby effectively protecting the structured features while enhancing the image, and continuing to perform image enhancement transformation on the basis of the first image sample to obtain the second image sample, thereby improving the generalization ability of the image retrieval model to image changes, and thereby improving the accuracy and reliability of image retrieval based on the image retrieval model.

[0092] In some embodiments, the above candidate region images can be obtained by the following methods: Figure 3 , Figure 3 This is the third flow chart of the image retrieval method provided by the present invention, as shown in FIG. Figure 3 As shown, the method includes the following steps 3101, 3102, 320 and 330.

[0093] Step 3101: Determine a minimum cropping ratio corresponding to the original structured image based on the structured layout features of the original structured image.

[0094] Here, the structural layout features refer to the spatial arrangement, organizational form and mutual relationship of various structural elements (such as text, formulas, symbols, charts, etc.) in the image.

[0095] For example, text layout structure features (row and column distribution of text, paragraph spacing, hierarchical relationship between title and text, etc.), formula and symbol layout features (formula layout, arrangement of special symbols such as superscripts and subscripts), chart and image-text relationship features (chart position, image-text relationship), etc.

[0096] In one example, the structured layout features of the original structured image are identified to determine the minimum region boundary corresponding to the structured features, and then the cropping boundary is determined based on the minimum region boundary corresponding to the structured features. Here, the minimum region boundary corresponding to the structured features refers to the smallest region in the image that can completely contain a structured feature. The minimum region boundary is determined by identifying the structured layout features, determining the boundaries of each structured feature, and combining all the boundaries.

[0097] Finally, based on the determined cropping boundary, the ratio of the cropped image to the original image is calculated, and this ratio can be used as the minimum cropping ratio corresponding to the original structured image.

[0098] Here, the minimum cropping ratio corresponding to each original structured image can be determined based on its own structured layout features. Alternatively, after determining the structured layout features corresponding to all structured images, the minimum cropping ratio can be uniformly determined based on all structured layout features. There is no limitation to this.

[0099] Step 3102: crop from the original structured image at least one candidate region image having a ratio not less than the minimum cropping ratio.

[0100] After determining the minimum cropping ratio, multiple candidate region images are cropped from the original structured image based on the minimum cropping ratio.

[0101] Step 320: Determine that the target candidate area image in the at least one candidate area image is the first image sample corresponding to the original structured image; the area ratio of the first character area area of ​​the target candidate area image to the second character area area of ​​the original structured image is not less than an area ratio threshold.

[0102] After obtaining a plurality of candidate region images, an image that meets the requirements is selected from the plurality of candidate region images as a first image sample based on a preset area ratio threshold corresponding to the character region.

[0103] Step 330: Obtain an image sample in the at least one training sample pair based on the first image sample and / or the second image sample corresponding to the first image sample; the second image sample is obtained by performing at least one image enhancement of geometric transformation, color transformation, and noise interference on the first image sample.

[0104] In one example, the first image sample after the above cropping can be used as the image sample for model training, and the second image sample after further image transformation of the first image sample can also be used as the image sample for model training. Alternatively, both the first image sample and the second image sample can be used as the image sample for model training, without limitation.

[0105] Here, the second image sample is obtained by performing one or more image enhancements on the first image sample. For example, geometric transformations (such as translation, rotation, scaling, shearing, and affine transformations) may be performed on the first image sample. For example, color transformations (such as brightness adjustment, contrast adjustment, and color space conversion) may also be performed on the first image sample. Furthermore, random noise (such as Gaussian blur) may be added to the first image sample to simulate the noise experienced in actual image acquisition.

[0106] Based on this, after using the image samples obtained by the above image enhancement, multiple training sample pairs are constructed, including multiple positive sample pairs and multiple negative sample pairs. A positive sample pair can be a first image sample and a second image sample derived from the same original structured image, or two second image samples derived from the same original structured image but subjected to different image enhancements. A negative sample pair can be a first image sample and / or a second image sample derived from different original structured images.

[0107] After constructing multiple positive and negative sample pairs, we use these pairs, combined with a contrastive loss function, to update the base model parameters through backpropagation to obtain a trained image retrieval model. This improves the accuracy and reliability of image retrieval through the trained image retrieval model.

[0108] The image retrieval method provided by the embodiment of the present invention reduces the damage to structured features during the cropping process through the dual cropping constraints of the minimum cropping ratio and the area ratio threshold, thereby effectively protecting the structured features while enhancing the image, improving the generalization ability of the image retrieval model to image changes, and further improving the accuracy and reliability of image retrieval based on the image retrieval model.

[0109] Based on the above embodiment, the step of cropping at least one candidate region image from the original structured image may further include:

[0110] determining the aspect ratio of the original structured image;

[0111] At least one candidate region image is cropped from the original structured image based on the aspect ratio of the original structured image.

[0112] In this embodiment, the aspect ratio of the original structured image is used as a reference for cropping. For example, if the aspect ratio of the original image is 0.65, candidate area images with a specific ratio range (for example, 0.5~1.0) are randomly cropped from the original structured image according to the aspect ratio of 0.65. The ratio range indicates the area ratio range of the allowed cropped area to the original image.

[0113] The image retrieval method provided by an embodiment of the present invention crops at least one candidate region image from the original structured image based on the aspect ratio of the original structured image, thereby ensuring that the cropped candidate region image can fully retain key structural features and avoiding the risk of structural feature damage caused by blind aspect ratio cropping.

[0114] Based on the above embodiment, the following further comprises:

[0115] Performing a geometric transformation on the first image sample to obtain a second image sample;

[0116] Wherein, when performing affine transformation on the first image sample, parameter values ​​of the parameters for controlling the affine transformation do not exceed corresponding preset parameter value ranges; the affine transformation parameters include rotation, scaling, translation and shearing.

[0117] In one example, the geometric transformation includes, but is not limited to, rigid body transformation, affine transformation, and projective transformation.

[0118] Here, rigid transformations refer to geometric transformations that preserve the shape and size of an object, changing only its position and orientation, such as translation and rotation. Affine transformations add linear transformations such as scaling and shearing to rigid transformations, preserving the parallelism of lines but not their lengths or angles. Projective transformations include translation, rotation, scaling, shearing, and perspective transformations.

[0119] Here, the setting of the preset parameter value range can be determined according to the actual application scenario and image characteristics. For example, if the structured features in the image are not sensitive to position changes, then the parameter value range of the translation parameter can be set larger. If the structured features in the image are not sensitive to direction changes, then the parameter value range of the rotation parameter can be set larger. If the structured features in the image are not sensitive to size changes, then the parameter value range of the scaling parameter can be set larger. If the structured features in the image are not sensitive to tilt changes, then the parameter value range of the shearing parameter can be set larger. Conversely, the corresponding parameter value range should be set smaller to ensure that the changes in the structured features are within an acceptable range during the image enhancement process, thereby avoiding the destruction of the structured features.

[0120] In one example, the affine transformation parameter matrix can be based on the following T Perform an affine transformation on the first image:

[0121] ;

[0122] in, Indicates that the image is rotated around the origin at an angle of ; Indicates that the image's horizontal and vertical translation ratios are both ; Indicates that the image is scaled horizontally and vertically. ; Indicates that the image is sheared at both horizontal and vertical angles. .

[0123] The image retrieval method provided by an embodiment of the present invention increases data diversity by performing multiple types of geometric transformations on a first image sample. In addition, when performing an affine transformation, the parameter values ​​of the affine transformation parameters are limited within a preset parameter value range to avoid excessive transformation operations from damaging structured features. This effectively protects structured features while enhancing the image, improves the generalization ability of the image retrieval model to image changes, and thereby improves the accuracy and reliability of image retrieval based on the image retrieval model.

[0124] Based on the above embodiment, the following further comprises:

[0125] Performing a geometric transformation on the first image sample to obtain a geometrically transformed image sample;

[0126] Performing color transformation on the geometrically transformed image sample to obtain a second image sample.

[0127] In this embodiment, a geometric transformation, such as at least one of a rigid body transformation, an affine transformation, and a projective transformation, is first performed on the first image sample. A color transformation, such as color dithering (randomly adjusting the brightness, contrast, saturation, and hue of the image to simulate lighting changes and shooting conditions in a real scene), is then performed on the geometrically transformed image sample to obtain a transformed second image sample.

[0128] In one example, color dithering can be performed by randomly adding or subtracting RGB values ​​to simulate image acquisition under varying lighting conditions. Pixel value distribution can also be adjusted to simulate low-contrast scans or faded printed paper. Color vividness can be adjusted to simulate color distortion during image acquisition, such as in color charts or illustrations. Randomly offsetting the camera in HSV space can simulate white balance differences during image acquisition.

[0129] It should be noted that in this embodiment, by first performing a geometric transformation to change the spatial structure of the image, and then performing a color transformation to change the color attributes of the pixels, this ensures the spatial integrity of the image's structural features while reducing the complexity of the color transformation, thereby improving the efficiency and accuracy of the entire processing flow.

[0130] Based on the above embodiment, the following further comprises:

[0131] Perform tensor conversion on the first image sample and / or the second image sample to obtain normalized tensor image samples.

[0132] It's important to understand that images are stored in computers as matrices. For example, color images are three-dimensional matrices, and grayscale images are two-dimensional matrices. Tensors, on the other hand, are a high-dimensional extension of matrices and are used in deep learning to uniformly represent data.

[0133] In this embodiment, after obtaining multiple constructed first image samples and multiple second image samples, these image samples are dimensionally adjusted to unify the original dimensions, and then the images are linearly normalized and zero-mean standardized to obtain image samples of the standardized tensor.

[0134] The image retrieval method provided by the embodiment of the present invention obtains image samples of standardized tensors by performing tensor conversion, which can ensure that the format and distribution of input data meet the expectations of the model, thereby improving the training speed and performance of the model.

[0135] Based on any of the above embodiments, the present invention further provides an image retrieval device, referring to Figure 4 , the device comprises:

[0136] A first image retrieval module 410 is used to obtain an image to be retrieved;

[0137] A second image retrieval module 420 is configured to input the image to be retrieved into an image retrieval model to obtain an image retrieval result output by the image retrieval model; the image retrieval model is obtained by training an initial image retrieval model based on at least one training sample pair, wherein the image samples in the at least one training sample pair are obtained by performing image enhancement on the original structured image;

[0138] The area ratio of the first character region of the image sample to the second character region of the original structured image that matches the image sample is not less than an area ratio threshold.

[0139] The image retrieval device provided by the present invention reduces the damage to structured features during the image enhancement process by ensuring that the area ratio of the first character area of ​​the image sample after image enhancement to the area of ​​the second character area of ​​the original structured image is not less than an area ratio threshold, thereby effectively protecting the structured features while enhancing the image, improving the generalization ability of the image retrieval model to image changes, and further improving the accuracy and reliability of image retrieval based on the image retrieval model.

[0140] Figure 5 An example of a physical structure diagram of an electronic device is shown below. Figure 5 As shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communications interface 520, and the memory 530 communicate with each other via the communication bus 540. The processor 510 may call logic instructions in the memory 530 to execute an image retrieval method, which includes:

[0141] Get the image to be retrieved;

[0142] Inputting the image to be retrieved into an image retrieval model to obtain an image retrieval result output by the image retrieval model; the image retrieval model is obtained by training an initial image retrieval model based on at least one training sample pair, and the image samples in the at least one training sample pair are obtained by performing image enhancement on the original structured image;

[0143] The area ratio of the first character region of the image sample to the second character region of the original structured image that matches the image sample is not less than an area ratio threshold.

[0144] Furthermore, the logic instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0145] In another aspect, the present invention further provides a computer program product, comprising a computer program, which may be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the image retrieval method provided by each of the above methods, wherein the method comprises:

[0146] Get the image to be retrieved;

[0147] Inputting the image to be retrieved into an image retrieval model to obtain an image retrieval result output by the image retrieval model; the image retrieval model is obtained by training an initial image retrieval model based on at least one training sample pair, and the image samples in the at least one training sample pair are obtained by performing image enhancement on the original structured image;

[0148] The area ratio of the first character region of the image sample to the second character region of the original structured image that matches the image sample is not less than an area ratio threshold.

[0149] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the image retrieval method provided by the above methods, the method comprising:

[0150] Get the image to be retrieved;

[0151] Inputting the image to be retrieved into an image retrieval model to obtain an image retrieval result output by the image retrieval model; the image retrieval model is obtained by training an initial image retrieval model based on at least one training sample pair, and the image samples in the at least one training sample pair are obtained by performing image enhancement on the original structured image;

[0152] The area ratio of the first character region of the image sample to the second character region of the original structured image that matches the image sample is not less than an area ratio threshold.

[0153] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0154] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0155] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. An image retrieval method, characterized in that: The method comprises: Get the image to be retrieved; Inputting the image to be retrieved into an image retrieval model to obtain an image retrieval result output by the image retrieval model; the image retrieval model is obtained by training an initial image retrieval model based on at least one training sample pair, and the image samples in the at least one training sample pair are obtained by performing image enhancement on the original structured image; Among them, the area ratio of the first character area of ​​the image sample to the area of ​​the second character area of ​​the original structured image matching the image sample is not less than an area ratio threshold, and the area ratio threshold of the original structured image is determined based on the character area ratio of each type of structured feature existing in the original structured image, and the semantic recognition critical destruction rate of each type of structured feature; the semantic recognition critical destruction rate includes the maximum acceptable area loss ratio that causes the semantic recognition accuracy of the structured feature to decrease during the cropping process.

2. The image retrieval method according to claim 1, wherein: The image samples in the at least one training sample pair are obtained in the following manner: Crop at least one candidate region image from the original structured image; Determining that a target candidate region image in the at least one candidate region image is a first image sample corresponding to the original structured image; The area ratio of the first character region of the target candidate region image to the second character region of the original structured image is not less than an area ratio threshold; Based on the first image sample and / or the second image sample corresponding to the first image sample, an image sample in the at least one training sample pair is obtained; the second image sample is obtained by performing at least one image enhancement of geometric transformation, color transformation and noise interference on the first image sample.

3. The image retrieval method according to claim 2, wherein: The step of cropping at least one candidate region image from the original structured image includes: Determining a minimum cropping ratio corresponding to the original structured image based on the structured layout features of the original structured image; At least one candidate region image having a cropping ratio not less than the minimum cropping ratio is cropped from the original structured image.

4. The image retrieval method according to claim 2, wherein: The step of cropping at least one candidate region image from the original structured image includes: determining the aspect ratio of the original structured image; At least one candidate region image is cropped from the original structured image based on the aspect ratio of the original structured image.

5. The image retrieval method according to claim 2, wherein: Also includes: Performing a geometric transformation on the first image sample to obtain a second image sample; Wherein, when performing affine transformation on the first image sample, parameter values ​​of the parameters for controlling the affine transformation do not exceed corresponding preset parameter value ranges; the affine transformation parameters include rotation, scaling, translation and shearing.

6. The image retrieval method according to claim 2, wherein: Also includes: Performing a geometric transformation on the first image sample to obtain a geometrically transformed image sample; Performing color transformation on the geometrically transformed image sample to obtain a second image sample.

7. The image retrieval method according to any one of claims 2 to 6, characterized in that: Also includes: Perform tensor conversion on the first image sample and / or the second image sample to obtain normalized tensor image samples.

8. An image retrieval device, characterized in that: The device comprises: A first image retrieval module, used to obtain an image to be retrieved; a second image retrieval module, configured to input the image to be retrieved into an image retrieval model to obtain an image retrieval result output by the image retrieval model; the image retrieval model is obtained by training an initial image retrieval model based on at least one training sample pair, wherein the image samples in the at least one training sample pair are obtained by performing image enhancement on the original structured image; Among them, the area ratio of the first character area of ​​the image sample to the area of ​​the second character area of ​​the original structured image matching the image sample is not less than an area ratio threshold, and the area ratio threshold of the original structured image is determined based on the character area ratio of each type of structured feature existing in the original structured image, and the semantic recognition critical destruction rate of each type of structured feature; the semantic recognition critical destruction rate includes the maximum acceptable area loss ratio that causes the semantic recognition accuracy of the structured feature to decrease during the cropping process.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the image retrieval method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the image retrieval method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Effective fine image classification method based on optimized feature weight

    CN111639206A

  • Model training method, image retrieval method, equipment and computer readable medium

    CN117523330A

  • Deep learning-based qualification image classification method and system

    CN117788957A

  • Method and apparatus for training image processing model, and image classifying method and apparatus

    US20240203097A1