Comparison learning model training method, image retrieval method and device
By comparing the feature alignment and training of the learning model, the problem of inaccurate feature vector extraction in image retrieval is solved, and the accuracy and efficiency of image retrieval is improved, especially in semantic content discrimination tasks.
Patent Information
- Application Number
- CN202510812303.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-07-18
AI Technical Summary
In the existing image retrieval methods, the accuracy of directly performing image similarity calculation is not high, and how to effectively extract image feature vectors has become an urgent problem.
Through the training method of the comparative learning model, the real-time image samples of the document are aligned with the electronic document image samples, positive and negative sample image pairs are constructed, and the image comparison learning model is trained to focus on the essential semantic features of the image and reduce the interference feature differences that are irrelevant to the semantic content.
It improves the accuracy and efficiency of image retrieval, can more effectively reduce the semantic content differences of the same image, expand the semantic content differences of different images, and improves the accuracy of similarity discrimination between the two images.
Smart Images

Figure CN120339650A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image retrieval, and in particular to a method for training a contrast learning model, an image retrieval method and device. Background Art
[0002] In the process of image retrieval, in order to meet the requirement of quickly locating a document real-shot image, it is necessary to retrieve in an electronic document image library based on the document real-shot image to determine the most similar document real-shot image thereto.
[0003] Existing image retrieval usually calculates image similarity and retrieves based on the calculated similarity. The method of directly calculating image similarity is not as accurate as calculating similarity through image feature vectors. And how to effectively extract the image feature vectors to be retrieved has become an urgent problem to be solved in the industry. Summary of the Invention
[0004] The present invention provides a method for training a contrast learning model, an image retrieval method and device, which are used to effectively extract the image feature vectors to be retrieved.
[0005] The present invention provides a method for training a contrast learning model, including the following steps: Align the features of a document real-shot image sample and an electronic document image sample with the same page content as the document real-shot image sample to obtain an electronic document image sample with features aligned with the document real-shot image sample; Take an electronically aligned document image sample and a document real-shot image sample with the same page content as a positive sample image pair, and take an electronically aligned document image sample and a document real-shot image sample with different page content as a negative sample image pair; Train an initial image contrast learning model based on the positive sample image pair and the negative sample image pair until a preset training condition is met to obtain an image contrast learning model; Wherein, the image contrast learning model is used to output an electronic document image feature vector and a document real-shot image feature vector according to the input electronic document image and document real-shot image.
[0006] According to a method for training a contrast learning model provided by the present invention, the training of the initial image contrast learning model based on the positive sample image pair and the negative sample image pair until a preset training condition is met to obtain an image contrast learning model includes: Input the positive sample image pair into the initial image contrast learning model to obtain a positive sample image feature vector pair output by the initial image contrast learning model; Input the negative sample image pairs into the initial image contrast learning model to obtain the negative sample image feature vector pairs output by the initial image contrast learning model; Determine the contrast loss based on the feature similarity of the positive sample image feature vector pairs and the feature similarity of the negative sample image feature vector pairs; Based on the contrast loss, adjust the parameters of the initial image contrast learning model until the preset training conditions are met to obtain the image contrast learning model.
[0007] According to a training method of a contrast learning model provided by the present invention, the feature alignment of the document real-shot image sample and the electronic document image sample with the same page content as the document real-shot image sample to obtain the electronic document image sample feature-aligned with the document real-shot image sample includes: Perform noise addition processing on the electronic document image sample according to the noise features in the document real-shot image sample with the same page content as the electronic document image sample, so that the electronic document image sample learns the noise features of the document real-shot image sample to obtain the electronic document image sample feature-aligned with the document real-shot image sample; Wherein, the noise addition processing includes at least one of the following: noise injection processing, illumination simulation processing.
[0008] According to a training method of a contrast learning model provided by the present invention, the illumination simulation processing of the electronic document image sample includes: Perform illumination analysis on the document real-shot image sample with the same page content as the electronic document image sample to obtain an illumination analysis result; Determine the illumination simulation method of the electronic document image sample according to the illumination analysis result; Perform illumination simulation processing on the electronic document image sample based on the illumination simulation method; Wherein, the illumination simulation method includes one or more of color temperature offset adjustment, non-uniform illumination synthesis, surface specular reflection synthesis, and dynamic range adjustment; the color temperature offset adjustment is used to adjust the RGB channel weights of the electronic document image sample; the non-uniform illumination synthesis is used to generate a gradient mask and superimpose it on the electronic document image sample; the surface specular reflection synthesis is used to add an elliptical highlight area at random positions in the electronic document image sample; the dynamic range adjustment is used to perform gamma correction and histogram clipping on the electronic document image sample.
[0009] According to a training method of a contrast learning model provided by the present invention, the noise injection processing of the electronic document image sample includes: Perform noise analysis on the actual-shot image samples of documents with the same content as the electronic document image samples, and determine the noise type; the noise type includes Gaussian noise and salt-and-pepper noise; Superimpose the noise of the noise type on the electronic document image samples.
[0010] The present invention also provides an image retrieval method, including: Construct an image pair based on any electronic document image in the document image library and the actual-shot image to be retrieved, traverse each electronic document image in the document image library, and obtain multiple image pairs; Input each image pair into the contrast learning model respectively, and output the feature vectors of the electronic document image and the feature vectors of the actual-shot document image corresponding to each image pair; Use the electronic document image in the target image pair as the image retrieval result of the actual-shot image to be retrieved; wherein, the target image pair is the image pair with the highest feature similarity between the feature vectors of the electronic document image and the feature vectors of the actual-shot document image among all the image pairs; The contrast learning model is trained based on the training method of the contrast learning model described in any one of the above.
[0011] The present invention also provides an image retrieval device, including the following modules: An image pair construction module, configured to construct an image pair based on any electronic document image in the document image library and the actual-shot image to be retrieved, traverse each electronic document image in the document image library, and obtain multiple image pairs; A model processing module, configured to input each image pair into the contrast learning model respectively, and output the feature vectors of the electronic document image and the feature vectors of the actual-shot document image corresponding to each image pair; A retrieval module, configured to use the electronic document image in the target image pair as the image retrieval result of the actual-shot image to be retrieved; wherein, the target image pair is the image pair with the highest feature similarity between the feature vectors of the electronic document image and the feature vectors of the actual-shot document image among all the image pairs; The contrast learning model is trained based on the training method of the contrast learning model described in any one of the above.
[0012] The present invention also provides an image retrieval system, including an image acquisition device and the image retrieval device described above; The image acquisition device is used to obtain the actual-shot image to be retrieved.
[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor, and when the processor executes the program, it implements the image retrieval method described above.
[0014] The present invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the image retrieval method as described above is implemented.
[0015] The training method, image retrieval method and device of the contrast learning model provided by the present invention align the features of the document real-shot image sample with the electronic document image sample with the same page content of the document real-shot image sample, so that the electronic document image sample learns the interference features in the document real-shot image sample to reduce the non-semantic feature differences between the electronic document image sample and the document real-shot image sample with the same content. Based on the samples after feature alignment, positive and negative sample image pairs are constructed to train the image contrast learning model, so that the trained image contrast learning model will not pay attention to the interference features irrelevant to the semantic content, and will more centrally understand the core semantic content of the image, avoiding invalid learning on irrelevant features, and guiding the model to focus on the essential semantic feature content of the learning image. As a result, in the process of feature extraction of two images by the trained image contrast learning model, the semantic content difference of the same image can be more effectively reduced, the semantic content difference of different images can be expanded, and the accuracy of subsequent similarity judgment of the two images can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0017] Figure 1 It is a flow chart of the training method of the contrastive learning model provided by the present invention.
[0018] Figure 2 It is a flowchart of the image retrieval method provided by the present invention.
[0019] Figure 3 It is a structural schematic diagram of the image retrieval device provided by the present invention.
[0020] Figure 4 It is a structural schematic diagram of a training device for a contrastive learning model provided by the present invention.
[0021] Figure 5 It is a structural schematic diagram of the image retrieval system provided by the present invention.
[0022] Figure 6 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0023] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0024] Figure 1 is a schematic flowchart of a method for training a contrastive learning model provided by the present invention. As Figure 1 shown, the method includes the following: In step 110, the feature alignment is performed between the physical document image sample and the electronic document image sample with the same page content as the physical document image sample, so as to obtain the electronic document image sample with the feature alignment with the physical document image sample.
[0025] In the present invention, the electronic document image sample refers to the document page image directly derived from the digital source file or generated by the digital system. The electronic document image sample can be specifically converted from the electronic document: for example, the document pages in formats such as Portable Document Format (PDF), Electronic Publication (EPUB), Word format, slide format, Hyper Text Markup Language (HTML), etc. are directly rendered or exported as image files (such as Portable Network Graphics (PNG), Joint Photographic Experts Group (JPEG)).
[0026] Optionally, the electronic document image sample can be obtained based on the parsing method of the rendering engine. By writing Python code and using an open-source library to render the electronic image page into a bitmap image, the electronic document image sample is obtained. The rendering process supports adjusting the resolution (Dots Per Inch, DPI) and color space. The resolution of the obtained electronic document image sample can meet the standard of being greater than or equal to 300 DPI.
[0027] In the present invention, the physical document image sample is the digital image data obtained by photographing or scanning a real paper document (such as textbooks, printed materials, books, newspapers, magazine pages, handwritten manuscripts, etc.) through physical imaging devices such as digital cameras, mobile phone cameras or scanners.
[0028] It should be noted that the actual-shot image samples of documents often contain various interference factors introduced during the shooting or scanning process, such as uneven ambient lighting (too bright, too dark, shadows, color cast), lens distortion (especially the curvature at the page edges), the texture or folds of the paper itself, ink penetration, out-of-focus blur, camera shake blur, sensor noise (graininess, color noise), compression artifacts, and perspective distortion caused by the shooting angle, etc. While the electronic document image samples, being directly generated electronic documents, are high-quality images with a clean background, uniform color, and brightness.
[0029] In the present invention, the electronic document image sample with the same page content as the actual-shot image sample of the document specifically refers to a digitalized document image that is completely consistent with the actual-shot image at the semantic level. It accurately reproduces all the valid information carried in the actual-shot image, including core elements such as text, formulas, tables, charts, layout structure, and logical order.
[0030] Align the features of the actual-shot image sample of the document with the electronic document image sample that has the same page content as the actual-shot image sample of the document, so that the electronic document image sample can learn the interference factors in the actual-shot image sample of the document, making the obtained electronic document image sample able to simulate the interference factors in the real shooting environment.
[0031] Optionally, the electronic document image sample can accurately reproduce the noise information of the actual-shot image sample of the document with the same page content as the electronic document image sample through noise injection. Specifically, it can include: random noise distribution and lighting characteristics; the random noise distribution includes Gaussian noise and salt-and-pepper noise, and the lighting characteristics can be achieved through lighting simulation techniques, which can include temperature offset adjustment, non-uniform lighting synthesis, surface reflection synthesis, and dynamic range adjustment, for effectively restoring physical interferences such as uneven lighting, shadow transition, and highlight reflection in the actual-shot image. This process transforms the electronic document image from an idealized digital form into a real-shot version with controllable artificial distortion, making it converge with the corresponding actual-shot image sample of the document in terms of the underlying feature distribution (such as lighting, texture, local contrast, noise pattern), thereby significantly reducing the inter-domain difference between the two and providing a pair of positive sample images with consistent semantic content and aligned visual features for the subsequent contrast learning model.
[0032] It should be noted that in the positive sample image pair, the electronic document image samples are feature aligned (specifically, noise can be added, lighting simulation can be performed), so that the electronic document image samples input to the first branch and the real-shot images input to the second branch are closer in terms of underlying interference features. This greatly reduces the difficulty of learning to bring the features of the two image samples closer based on the subsequent image contrast learning model, allowing the image contrast learning model to focus more on learning matching at the semantic content level rather than overcoming huge underlying feature differences. Without feature alignment, the difference between the relatively pure electronic document image samples and the noisy real-shot document image samples is too large, and the noise will cause greater interference, making it difficult for the image contrast learning model to learn effectively, or the learned features mainly reflect the underlying feature differences rather than the semantic content differences, resulting in low discrimination accuracy in the process of semantic difference discrimination.
[0033] In step 120, an electronic document image sample after feature alignment and a real document image sample with the same page content are taken as a positive sample image pair, and an electronic document image sample and a real document image sample with different page contents are taken as a negative sample image pair.
[0034] After the feature alignment of the document real-shot image sample and the electronic document image sample with the same page content as the document real-shot image sample is performed based on step 110, the electronic document image sample with the feature alignment with the document real-shot image sample is obtained.
[0035] After feature alignment, interference features can be superimposed on the electronic document image sample to eliminate the influence of interference features between the actual document image sample and the electronic document image sample, thereby obtaining an electronic document image sample that is feature aligned with the actual document image sample.
[0036] An electronic document image sample with the same page content and an image sample of a real document are combined into a positive sample image pair. The positive sample image pair reflects the corresponding relationship between image features of the same document content under different acquisition methods, which helps the model learn the core semantic features of the document content. An electronic document image sample with different page content and an image sample of a real document are combined into a negative sample image pair. The negative sample image pair is used to guide the model to distinguish different document contents and improve the model's sensitivity to content differences.
[0037] In step 130, the initial image contrast learning model is trained based on the positive sample image pair and the negative sample image pair until the preset training conditions are met to obtain the image contrast learning model. The image contrast learning model is used to output the electronic document image feature vector and the document real-shot image feature vector according to the input electronic document image and the document real-shot image.
[0038] In the present invention, the image contrast learning model is a self-supervised deep learning architecture that can process image pairs through a two-branch neural network. Among them, the goal of the image contrast learning model is to learn the feature representation of images, so that the two image feature vectors generated from positive sample image pairs are as similar as possible, while the two image feature vectors generated from negative sample image pairs are as dissimilar as possible.
[0039] It should be noted that the initial image contrast learning model of the present invention is an untrained contrast learning model. The image contrast learning model is the model obtained after being trained with positive and negative sample image pairs. For the training process of the initial image contrast learning model, it can be specifically achieved by defining a loss function.
[0040] In the present invention, the image contrast learning model can adopt a contrast learning model with a two-branch encoder sharing weights. The two branches of the contrast learning model with a two-branch encoder can be composed based on the same backbone network. These two branches share parameters to ensure that the way of extracting features for the same type of input is consistent.
[0041] A training sample of the image contrast learning model in the present invention is an image pair. This image pair can be a positive sample image pair (an electronically-documented image sample and a document real-shot image sample with the same page content after feature alignment) or a negative sample image pair (an electronically-documented image sample and a document real-shot image sample with different page content). For a positive sample image pair, after inputting it into the initial image contrast learning model of the two-branch encoder, a pair of feature vectors corresponding to the positive sample image pair is output.
[0042] For a negative sample image pair, after inputting it into the initial image contrast learning model of the two-branch encoder, a pair of feature vectors corresponding to the negative sample image pair is output.
[0043] By designing the loss function of the training process, during the training of the initial image contrast learning model, the feature distance between the two feature vectors corresponding to the positive sample image pair is minimized and the feature distance between the two feature vectors corresponding to the negative sample image pair is maximized.
[0044] Optionally, the specific formula of the loss function can be: ; where is the loss value, is the similarity between an electronically-documented image sample with the same page content after feature alignment and a document real-shot image sample , is an electronically-documented image sample with different page content and a document real-shot image sample similarity is the temperature coefficient (usually set to 0.1 - 0.5), which is used to control the sharpness of the distribution is the number of negative sample image pairs is index of
[0045] Through repeated training, the initial image contrast learning model will gradually learn to map the features of samples with the same page content (an electronic document image sample and a document real - shot image sample after aligning one feature of the same page content) to nearby positions in the feature space, while dispersing the features of samples with different page content (an electronic document image sample and a document real - shot image sample with different page content) to distant positions
[0046] The trained image contrast learning model can shield the interference of non - semantic features and guide the model to focus its attention on learning the essential semantic feature content of the images. Thus, the trained image contrast learning model is more accurate and efficient in the discrimination task of two images based on semantic content
[0047] Based on the trained image contrast learning model, the image retrieval process can be realized. Specifically, the image to be retrieved is obtained in advance and a document image library is constructed. The document image library can be generated based on a large number of electronic documents that may be related to the content to be retrieved. A large number of electronic documents are converted to obtain multiple electronic document images. Based on the obtained multiple electronic document images, a document image library is constructed. Among them, the constructed document image library contains the clear electronic document images corresponding to the real - shot images to be retrieved
[0048] Based on the real - shot image to be retrieved and all the electronic document images in the document image library, multiple image pairs are constructed. Specifically, for each electronic document image in the document image library, it is paired with the real - shot image to be retrieved to form an image pair. In this way, if there are N electronic document images in the document image library, then N image pairs will be constructed
[0049] Each constructed image pair is respectively input into the pre - trained image contrast learning model. The image contrast learning model can extract representative feature vectors in the images by learning the feature differences and similarities between a large number of image pairs. In this process, the image contrast learning model will extract the features of the two images in the image pair and output two corresponding feature vectors, forming a feature vector pair, specifically the feature vector of the electronic document image corresponding to the image pair and the feature vector of the document real - shot image. Therefore, for N image pairs, the image contrast learning model will output N feature vector pairs
[0050] Calculate the similarity of each pair of feature vectors. There are various methods for similarity calculation, such as cosine similarity, Euclidean distance, etc. By calculating the similarity, the similarity degree between each pair of images can be quantified, providing a basis for determining the subsequent retrieval results.
[0051] Based on the similarity of each pair of feature vectors among the multiple pairs of feature vectors obtained by calculation, from all the image pairs in the document image library, determine the target image pair with the highest feature similarity between the feature vector of the electronic document image and the feature vector of the actual document image, and return the electronic document image in the target image pair with the highest similarity as the retrieval result, thereby realizing the retrieval process of the image to be retrieved from the document image library.
[0052] The training method of the contrast learning model provided by the present invention aligns the features of the actual document image sample and the electronic document image sample with the same page content as the actual document image sample, so that the electronic document image sample learns the interference features in the actual document image sample to reduce the non-semantic feature differences between the electronic document image samples with the same content and the actual document image samples. And based on the samples after feature alignment, positive and negative sample image pairs are constructed to train the image contrast learning model, so that the trained image contrast learning model does not pay attention to the interference feature differences irrelevant to the semantic content, and more focuses on understanding and matching the core semantic content of the images, avoiding ineffective learning on irrelevant features, and guiding the model to focus its attention on learning the essential semantic feature content of the images. Thus, when the trained image contrast learning model extracts the features of two images, it can more effectively reduce the semantic content differences of the same images, expand the semantic content differences of different images, and improve the accuracy of subsequent discrimination of the differences between two images based on semantic content.
[0053] Based on the above embodiments, based on the positive sample image pair and the negative sample image pair, train the initial image contrast learning model until the preset training conditions are met to obtain the image contrast learning model, including: input the positive sample image pair into the initial image contrast learning model to obtain the positive sample image feature vector pair output by the initial image contrast learning model; input the negative sample image pair into the initial image contrast learning model to obtain the negative sample image feature vector pair output by the initial image contrast learning model; determine the contrast loss based on the feature similarity of the positive sample image feature vector pair and the feature similarity of the negative sample image feature vector pair; based on the contrast loss, adjust the parameters of the initial image contrast learning model until the preset training conditions are met to obtain the image contrast learning model.
[0054] For the model training process, the initial image contrast learning model is iteratively trained based on positive sample image pairs and negative sample image pairs until the preset training conditions are met, thereby obtaining a trained image contrast learning model.
[0055] Specifically, after determining the feature similarity between the positive sample image feature vectors and the negative sample image feature vectors, the contrast loss is determined through these similarities. The contrast loss is a quantitative metric that reflects the accuracy of the model in distinguishing between positive and negative sample image pairs.
[0056] To minimize the contrast loss, techniques such as backpropagation can be used to adjust the parameters of the initial image contrast learning model. In the continuously repeated process, the feature vectors and their similarities of the positive and negative sample image pairs are recalculated according to the new parameter settings in each iteration, and then the contrast loss is updated. As the training progresses, the model gradually learns how to more accurately capture the similarities and differences between images, and the contrast loss decreases accordingly.
[0057] Finally, when the contrast loss drops below the preset threshold, the training process ends. The obtained image contrast learning model at this time already has the capabilities of feature extraction and contrast, and can output high-quality electronic document image feature vectors and document real-shot image feature vectors based on the input electronic document images and document real-shot images, providing strong support for subsequent tasks such as image retrieval and classification.
[0058] Based on the above embodiments, aligning the features of the document real-shot image sample with the electronic document image sample having the same page content as the document real-shot image sample to obtain an electronic document image sample with features aligned with the document real-shot image sample includes: adding noise to the electronic document image sample according to the noise features in the document real-shot image sample having the same page content as the electronic document image sample, so that the electronic document image sample learns the noise features of the document real-shot image sample to obtain an electronic document image sample with features aligned with the document real-shot image sample; wherein, the noise addition process includes at least one of the following: noise injection process, illumination simulation process.
[0059] During the training process of the image contrast learning model, in order to align the features of the electronic document image samples with the actual captured images of the documents with the same corresponding page content. Among them, feature alignment is to enable the electronic document image samples to learn the interference features in the actual captured image samples of the documents, so as to reduce the non-semantic feature differences between the electronic document image samples with the same content and the actual captured image samples of the documents. It is necessary to add noise to the electronic document image samples so that they can learn and simulate the noise features in the actual captured image samples of the documents. This processing process aims to enhance the similarity of the electronic document image samples and the actual captured image samples in terms of interference features, thereby improving the robustness and accuracy of the contrast learning model.
[0060] The noise addition process mainly includes two methods: noise injection processing and lighting simulation processing. According to the purpose of achieving the best effect of feature alignment, either of these two methods can be selected and used alone, or they can be used simultaneously.
[0061] For the noise injection processing process, it can be to perform noise analysis on the actual captured image samples of the documents, determine the noise type of the actual captured image samples of the documents, and superimpose the noise of the noise type on the electronic document image samples, so that the noise distribution of the electronic document image samples after noise injection is the same as that of the actual captured image samples of the documents.
[0062] For the lighting simulation processing, it can be to perform lighting analysis on the actual captured image samples of the documents, and determine the lighting analysis results of the actual captured image samples of the documents. Based on the lighting analysis results, perform lighting simulation processing on the electronic document image samples, so that the lighting distribution of the electronic document image samples after lighting simulation processing is the same as that of the actual captured image samples of the documents.
[0063] Optionally, in order to further improve the accuracy of feature alignment, geometric transformation can also be performed on the actual captured image samples of the documents, and / or distortion correction processing can be performed on the distorted areas in the actual captured image samples of the documents.
[0064] The geometric transformation can specifically include rotation processing or cropping processing. Make the geometric features of the actual captured image samples after geometric transformation the same as those of the electronic document image samples. The distortion correction processing of the distorted area can be to correct the bending distortion generated during the document shooting process.
[0065] Optionally, during the distortion correction process, for slightly bent documents, perspective transformation can be used for correction. Perspective transformation adjusts the shape of the image by mapping four corner points to make it look flatter. For severely bent documents, thin plate spline interpolation method can be used. The thin plate spline interpolation method fits a set of control points by minimizing the bending energy, thereby achieving smooth correction of the image.
[0066] Based on geometric transformation and distortion correction processing, it is possible to simulate the deformation and angle change of document content caused by factors such as the user's hand-held angle and distance in the actual shooting environment, further improving the accuracy of feature alignment.
[0067] Based on the above embodiments, the illumination simulation processing of the electronic document image sample includes: performing illumination analysis on a document real-shot image sample with the same page content as the electronic document image sample to obtain an illumination analysis result; determining the illumination simulation method for the electronic document image sample according to the illumination analysis result; performing illumination simulation processing on the electronic document image sample based on the illumination simulation method; wherein, the illumination simulation method includes one or more of color temperature offset adjustment, non-uniform illumination synthesis, surface reflection synthesis, and dynamic range adjustment; the color temperature offset adjustment is used to adjust the RGB channel weights of the electronic document image sample; the non-uniform illumination synthesis is used to generate a gradient mask and superimpose it on the electronic document image sample; the surface reflection synthesis is used to add elliptical highlight regions at random positions in the electronic document image sample; the dynamic range adjustment is used to perform gamma correction and histogram clipping on the electronic document image sample.
[0068] Illumination analysis extracts illumination features from a document real-shot image sample with the same page content as the electronic document image sample, providing a basis for subsequent illumination simulation. The extracted features include the direction, intensity, color temperature (CT) of the illumination, and the uniformity of the illumination, etc. Through illumination analysis, the illumination conditions in the real-shot environment can be understood, such as whether it is natural light or artificial light, whether the illumination is uniform, whether there are obvious shadow or highlight regions, etc. These illumination features are used in the subsequent illumination simulation process.
[0069] Based on the result of the illumination analysis, an illumination simulation method can be adopted to perform illumination simulation processing on the electronic document image sample. The methods of illumination simulation processing include: color temperature offset adjustment, non-uniform illumination synthesis, surface reflection synthesis, and dynamic range adjustment. These four methods can be used alone or in combination, so that the illumination distribution of the electronic document image sample after illumination simulation processing is the same as that of the document real-shot image sample.
[0070] Specifically, if the illumination analysis result shows that there is an obvious color temperature deviation in the document real-shot image sample, such as the overall tone being yellowish or bluish, the illumination simulation processing method of color temperature offset can be selected for adjustment.
[0071] When the illumination analysis results show that the document image samples have local uneven illumination, you can select the illumination simulation processing method of non-uniform illumination synthesis for adjustment. If the illumination analysis results show that the document image samples have obvious surface reflections, similar to the highlights on the surface of paper, you can select the illumination simulation processing method of surface reflection synthesis for adjustment. If the illumination analysis results show that the document image samples have a wide dynamic range, you can select the illumination simulation processing method of dynamic range adjustment for adjustment. If multiple defects occur at the same time, you can combine multiple illumination simulation processing methods according to the multiple defects.
[0072] It should be noted that for color temperature offset adjustment: color temperature is a physical quantity that describes the color of the light source. Different color temperatures will cause the image to present different tones. By adjusting the RGB channel weights of the electronic document image sample, the lighting effects under different color temperatures can be simulated, such as simulating cold light (blue tone) and warm light (yellow tone) lighting effects, so that the electronic document image sample is closer to the actual document image sample in color.
[0073] For non-uniform lighting synthesis: In actual shooting, due to the uneven position and intensity distribution of the light source, non-uniform lighting often occurs in the image. In order to simulate this effect, a gradient mask or shadow mask can be generated and superimposed on the electronic document image sample, and the non-uniform lighting synthesis can be achieved by adjusting the transparency and gradient mode of the mask or mask.
[0074] For surface reflection synthesis: In real-life images, the document surface may reflect light due to the lighting during the shooting process, and these reflective areas appear as highlight areas in the image. In order to simulate this reflective effect, elliptical highlight areas can be added at random positions in the electronic document image sample to simulate the mirror reflection process of a smooth document.
[0075] Dynamic range adjustment: Dynamic range refers to the brightness difference between the brightest and darkest areas that can be represented in an image. Actual images may have limited dynamic range due to strong or weak lighting, resulting in highlight overflow or loss of shadow details. To simulate this effect, gamma correction and histogram cropping can be performed on the electronic document image sample to adjust the image brightness and contrast so that its dynamic range is closer to the actual document image sample.
[0076] By performing illumination simulation on electronic document image samples, the robustness of the image contrast learning model to the influence of illumination in the actual shooting environment can be significantly improved. During the training process of the image contrast learning model, the electronic document image samples that have been processed with illumination simulation are used for comparative learning with the document image samples that are actually shot, so that the image contrast learning model can learn to accurately extract image features under illumination interference, thereby improving the performance of the image contrast learning model in real application scenarios.
[0077] Based on the above embodiments, the noise injection processing of the electronic document image sample includes: performing noise analysis on the actual document image sample with the same page content as the electronic document image sample to determine the noise type; the noise type includes Gaussian noise and salt-and-pepper noise; superimposing the noise of the noise type on the electronic document image sample.
[0078] In the present invention, Gaussian noise is a kind of random noise, and its probability density function follows a Gaussian distribution (normal distribution). In an image, Gaussian noise appears as tiny random fluctuations in pixel values, usually affecting the overall brightness and contrast of the image, but not significantly changing the structural information of the image.
[0079] In the present invention, salt-and-pepper noise, also known as impulse noise, appears as randomly occurring black and white pixel points in the image. Salt-and-pepper noise will significantly damage the local structural information of the image and have a greater impact on the image quality.
[0080] After determining the noise type, these noises are superimposed on the electronic document image sample. Simulate the noise environment in the actual document image sample on the electronic document image sample.
[0081] Specifically, for Gaussian noise, random numbers conforming to the Gaussian distribution can be generated and added to each pixel value of the electronic document image sample according to a certain proportion. During the superimposing process, the intensity of the noise needs to be controlled to ensure that the superimposed image still remains recognizable while simulating the noise effect in the actual document image sample. For salt-and-pepper noise, pixel points in the image can be randomly selected and their values can be set to the maximum value (white) or the minimum value (black) for simulation. When superimposing salt-and-pepper noise, the density of the noise also needs to be controlled to avoid overly damaging the structural information of the image.
[0082] Superimposing the noise of the noise type on the electronic document image sample makes the noise distribution of the electronic document image sample after noise injection the same as that of the actual document image sample.
[0083] By performing noise injection processing on the electronic document image sample, the robustness of the image contrast learning model to noise in the actual shooting environment can be significantly improved. During the training process of the image contrast learning model, using the electronic document image sample after noise injection processing and the actual document image sample for contrast learning can enable the image contrast learning model to learn to accurately extract image semantic features under noise interference, thereby improving the performance of the image contrast learning model in real application scenarios.
[0084] The present invention also provides an image retrieval method, Figure 2 which is a schematic flowchart of the image retrieval method provided by the present invention, as Figure 2As shown in the figure, the method includes: Step 210: Construct an image pair based on any electronic document image in the document image library and the real-shot image to be retrieved, traverse each electronic document image in the document image library, and obtain multiple image pairs; Step 220: Input each image pair into the contrastive learning model respectively, and output the feature vector of the electronic document image and the feature vector of the real-shot document image corresponding to each image pair; Step 230: Use the electronic document image in the target image pair as the image retrieval result of the real-shot image to be retrieved; wherein, the target image pair is the image pair with the highest feature similarity between the feature vector of the electronic document image and the feature vector of the real-shot document image among all the image pairs; The contrastive learning model is trained based on the training method of the contrastive learning model described above.
[0085] In the present invention, the document image library can be pre-generated based on a large number of electronic documents that may be related to the content to be retrieved. Convert a large number of electronic documents to obtain multiple electronic document images. Among them, the constructed document image library contains the clear electronic document image corresponding to the real-shot image to be retrieved.
[0086] Furthermore, in the process of constructing the document image library, since an electronic document may contain many pages. For example, for a textbook electronic document, a single page of the textbook can be converted into an image. And label information is marked in the document image corresponding to each page. The label information can reflect the name of the document and the page information. Specifically, it can be constructed based on the International Standard Book Number (ISBN) code of the book and the page number. The label information can be specifically used for verification in the subsequent retrieval process. Verify whether the electronic document image in the retrieved document image library is consistent with the book name and page information of the real-shot image to be retrieved. If they are consistent, it is regarded as a successful retrieval.
[0087] For the training process of the initial image contrastive learning model, positive sample image pairs and negative sample image pairs can be constructed to implement the training process.
[0088] Specifically, align the features of the real-shot document image sample and the electronic document image sample with the same page content as the real-shot document image sample, so that the electronic document image sample learns the interference factors in the real-shot document image sample, and enables the electronic document image sample to simulate the interference factors in the real shooting environment.
[0089] Specifically, the electronic document image sample can accurately reproduce the random noise distribution (such as Gaussian noise, salt-and-pepper noise) of the real-shot image through noise injection, and restore physical interferences such as uneven illumination, shadow transition, and specular reflection of the real-shot image through illumination simulation techniques (including color temperature shift, non-uniform illumination synthesis, specular area addition, and dynamic range adjustment). This process transforms the electronic document image from an idealized digital form into a version with controllable artificial distortion, making it converge with the corresponding real-shot image in terms of the underlying feature distribution (such as texture, local contrast, noise pattern), thereby significantly reducing the inter-domain difference between the two and providing a basis for positive sample image pairs with consistent semantic content and aligned visual features for subsequent image contrast learning models.
[0090] Take an electronically-documented image sample with aligned features of the same page content and a real-shot image sample of the document as a positive sample image pair, and take an electronically-documented image sample and a real-shot image sample of the document with different page content as a negative sample image pair.
[0091] Among them, the image contrast learning model can adopt a contrast learning model with a shared-weight double-branch encoder. The two branches of the contrast learning model with a double-branch encoder can be composed based on the same backbone network. These two branches share parameters to ensure the same way of extracting features for the same type of input.
[0092] The first image in the image pair (such as the electronically-documented image sample with aligned features) is input into the first branch encoder in the double-branch encoder. The second image in the image pair (such as the real-shot image sample of the document) is input into the second branch encoder. Each branch independently extracts features from the input image and outputs the corresponding feature vectors.
[0093] For a positive sample image pair, after being input into the initial image contrast learning model of the double-branch encoder, two feature vectors corresponding to the positive sample image pair are output. For a negative sample image pair, after being input into the initial image contrast learning model of the double-branch encoder, two feature vectors corresponding to the negative sample image pair are output.
[0094] By designing the loss function of the training process, during the training of the initial image contrast learning model, minimize the feature distance between the two feature vectors corresponding to the positive sample image pair and maximize the feature distance between the two feature vectors corresponding to the negative sample image pair. After the training is completed, the image contrast learning model is obtained.
[0095] It is understandable that by actively simulating typical interferences of real-shot images (such as noise, blur, uneven illumination) on electronic document images, the differences in interference characteristics between the two are reduced. This enables the trained image contrast learning model not to focus on the differences in interference characteristics unrelated to semantic content (such as the purity of electronic images and the noise or illumination effects of real-shot images), and can be more concentrated on understanding and matching the core semantic content of images (such as text information, layout structure, key graphics), avoiding ineffective learning on irrelevant features. The trained image contrast learning model can shield the interference of non-semantic features and guide the model to focus its attention on learning the essential semantic feature content of images. Thus, the trained image contrast learning model is more accurate and efficient in the discrimination task of two images based on semantic content.
[0096] After obtaining the image contrast learning model and the document image library, they can be used in the retrieval process of the real-shot image to be retrieved.
[0097] In the present invention, the real-shot image to be retrieved can be any document image obtained by shooting that needs to be retrieved.
[0098] Based on the real-shot image to be retrieved and all the electronic document images in the document image library, multiple image pairs are constructed. Specifically, for each electronic document image in the document image library, it is paired with the real-shot image to be retrieved to form an image pair. Thus, if there are N electronic document images in the document image library, then N image pairs will be constructed.
[0099] Each constructed image pair is respectively input into the pre-trained image contrast learning model. The image contrast learning model can extract representative feature vectors in the images by learning the feature differences and similarities between a large number of image pairs. In this process, the image contrast learning model will extract features from the two images in the image pair and output two corresponding feature vectors, forming a pair of feature vectors, specifically the feature vector of the electronic document image corresponding to the image pair and the feature vector of the real-shot document image. Therefore, for N image pairs, the image contrast learning model will output N pairs of feature vectors.
[0100] Calculate the similarity of each pair of feature vectors. There are various methods for calculating similarity, such as cosine similarity, Euclidean distance, etc. In practical applications, an appropriate similarity calculation method can be selected according to specific requirements and scenarios. By calculating the similarity, the similarity degree between each pair of images can be quantified, providing a basis for determining the subsequent retrieval results.
[0101] Based on the similarity of each pair of feature vectors among multiple calculated feature vector pairs, from all image pairs in the document image library, determine the target image pair with the highest feature similarity between the feature vectors of the electronic document image and the feature vectors of the actual captured document image, and return the electronic document image in the target image pair with the highest similarity as the retrieval result.
[0102] The image retrieval method provided by the present invention constructs positive and negative sample image pairs through sample construction after feature alignment for training the image contrast learning model. The trained image contrast learning model will not focus on the interference feature differences irrelevant to the semantic content, but more concentratedly understand and match the core semantic content of the image, avoiding ineffective learning on irrelevant features, and guiding the model to focus its attention on learning the essential semantic feature content of the image. Therefore, the trained image contrast learning model is more accurate in discrimination in the discrimination task based on semantic content during the process of image retrieval, improving the accuracy of retrieval.
[0103] The image retrieval device provided by the present invention will be described below. The image retrieval device described below can be mutually corresponding and referred to the image retrieval method described above.
[0104] As Figure 3 As shown in the structural schematic diagram of the image retrieval device provided by the present invention, the device includes: An image pair construction module 310, configured to construct an image pair according to any electronic document image in the document image library and the actual captured image to be retrieved, traverse each electronic document image in the document image library, and obtain multiple image pairs; A model processing module 320, configured to input each image pair into the contrast learning model respectively, and output the feature vectors of the electronic document image and the feature vectors of the actual captured document image corresponding to each image pair; A retrieval module 330, configured to use the electronic document image in the target image pair as the image retrieval result of the actual captured image to be retrieved; wherein, the target image pair is the image pair with the highest feature similarity between the feature vectors of the electronic document image and the feature vectors of the actual captured document image among all the image pairs; The contrast learning model is trained by the training method of the contrast learning model described above.
[0105] The image retrieval device provided by the present invention constructs positive and negative sample image pairs through feature-aligned samples for training an image contrast learning model. The trained image contrast learning model will not focus on the interference feature differences unrelated to semantic content, but more concentratedly understand and match the core semantic content of the images, avoiding ineffective learning on irrelevant features and guiding the model to focus its attention on learning the essential semantic feature content of the images. Thus, the trained image contrast learning model is more accurate in discrimination on the discrimination task based on semantic content during the process of image retrieval, improving the accuracy of retrieval.
[0106] The present invention also provides a training device for a contrast learning model. Figure 4 As shown in the structural schematic diagram of the training device for the contrast learning model provided by the present invention, Figure 4 the device includes: A feature alignment module 410, configured to align the features of a document real-shot image sample with an electronic document image sample having the same page content as the document real-shot image sample, to obtain an electronic document image sample whose features are aligned with those of the document real-shot image sample; A sample pair construction module 420, configured to use an electronic document image sample with aligned features and a document real-shot image sample having the same page content as a positive sample image pair, and use an electronic document image sample and a document real-shot image sample having different page contents as a negative sample image pair; A training module 430, configured to train an initial image contrast learning model based on the positive sample image pair and the negative sample image pair until a preset training condition is satisfied, to obtain an image contrast learning model; wherein, the image contrast learning model is configured to output an electronic document image feature vector and a document real-shot image feature vector according to the input electronic document image and the document real-shot image.
[0107] The training device for the contrastive learning model provided by the present invention aligns the features of a document real-shot image sample with an electronic document image sample having the same page content as the document real-shot image sample, so that the electronic document image sample learns the interference features in the document real-shot image sample, thereby reducing the non-semantic feature differences between the electronic document image samples with the same content and the document real-shot image samples. And based on the samples after feature alignment, positive and negative sample image pairs are constructed to train the image contrastive learning model, so that the trained image contrastive learning model does not pay attention to the interference feature differences irrelevant to the semantic content, but more centrally understands and matches the core semantic content of the images, avoiding ineffective learning on irrelevant features and guiding the model to focus its attention on learning the essential semantic feature content of the images. Thus, when the trained image contrastive learning model extracts the features of two images, it can more effectively reduce the semantic content differences of the same images, expand the semantic content differences of different images, and improve the accuracy of subsequent discrimination of the differences between two images based on semantic content.
[0108] Based on the above embodiments, the training module 430 is specifically configured to: Based on the positive sample image pair and the negative sample image pair, train the initial image contrastive learning model until a preset training condition is met, and obtain the image contrastive learning model, including: Input the positive sample image pair into the initial image contrastive learning model to obtain a positive sample image feature vector pair output by the initial image contrastive learning model; Input the negative sample image pair into the initial image contrastive learning model to obtain a negative sample image feature vector pair output by the initial image contrastive learning model; Determine the contrast loss based on the feature similarity of the positive sample image feature vector pair and the feature similarity of the negative sample image feature vector pair; Based on the contrast loss, adjust the parameters of the initial image contrastive learning model until a preset training condition is met, and obtain the image contrastive learning model.
[0109] Based on the above embodiments, the feature alignment module 410 is specifically configured to: The process of aligning the features of the document real-shot image sample with the electronic document image sample having the same page content as the document real-shot image sample to obtain an electronic document image sample with features aligned with the document real-shot image sample includes: According to the noise features in the document real-shot image sample having the same page content as the electronic document image sample, perform noise addition processing on the electronic document image sample, so that the electronic document image sample learns the noise features of the document real-shot image sample, and obtain an electronic document image sample with features aligned with the document real-shot image sample; Among them, the noise addition process includes at least one of the following: noise injection process, illumination simulation process.
[0110] Based on the above embodiments, the feature alignment module 410 is further specifically configured to: The illumination simulation process of the electronic document image sample includes: Performing illumination analysis on a real-shot document image sample with the same page content as the electronic document image sample to obtain an illumination analysis result; Determining the illumination simulation method of the electronic document image sample according to the illumination analysis result; Performing illumination simulation processing on the electronic document image sample based on the illumination simulation method; Among them, the illumination simulation method includes one or more of color temperature offset adjustment, non-uniform illumination synthesis, surface specular reflection synthesis, and dynamic range adjustment; the color temperature offset adjustment is used to adjust the RGB channel weights of the electronic document image sample; the non-uniform illumination synthesis is used to generate a gradient mask and superimpose it on the electronic document image sample; the surface specular reflection synthesis is used to add elliptical highlight areas at random positions in the electronic document image sample; the dynamic range adjustment is used to perform gamma correction and histogram clipping on the electronic document image sample.
[0111] Based on the above embodiments, the feature alignment module 410 is further specifically configured to: The noise injection process of the electronic document image sample includes: Performing noise analysis on a real-shot document image sample with the same page content as the electronic document image sample to determine the noise type; the noise type includes Gaussian noise and salt-and-pepper noise; Superimposing the noise of the noise type on the electronic document image sample.
[0112] The present invention also provides an image retrieval system, Figure 5 As shown in the structural schematic diagram of the image retrieval system provided by the present invention, Figure 5 As shown, the system includes an image acquisition device 510 and an image retrieval device 520.
[0113] The image acquisition device 510 is used to acquire a real-shot image to be retrieved.
[0114] An image retrieval device 520 is configured to construct an image pair based on any electronic document image in a document image library and a real-shot image to be retrieved, traverse each electronic document image in the document image library to obtain a plurality of image pairs; input each image pair into a contrastive learning model, and output an electronic document image feature vector and a document real-shot image feature vector corresponding to each image pair; use the electronic document image in the target image pair as the image retrieval result of the real-shot image to be retrieved; wherein, the target image pair is the image pair with the highest feature similarity between the electronic document image feature vector and the document real-shot image feature vector among all the image pairs; the contrastive learning model is trained by the training method of the contrastive learning model described above.
[0115] Figure 6 An example of the physical structure diagram of an electronic device is shown as Figure 6 As shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communication interface 620, and the memory 630 complete mutual communication through the communication bus 640. The processor 610 can call the logical instructions in the memory 630 to execute the training method of the contrastive learning model, and the method includes: aligning the features of a document real-shot image sample and an electronic document image sample with the same page content as the document real-shot image sample to obtain an electronic document image sample with features aligned with the document real-shot image sample; Taking an electronically-document-image sample with aligned features and a document real-shot image sample with the same page content as a positive sample image pair, and taking an electronically-document-image sample and a document real-shot image sample with different page contents as a negative sample image pair; Based on the positive sample image pair and the negative sample image pair, training an initial image contrastive learning model until a preset training condition is met to obtain an image contrastive learning model; Wherein, the image contrastive learning model is configured to output an electronic document image feature vector and a document real-shot image feature vector according to the input electronic document image and the document real-shot image.
[0116] In addition, when the logical instructions in the above-mentioned memory 630 can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0117] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the training method of the contrast learning model provided by the above-mentioned various methods. The method includes: aligning the features of a document actual shooting image sample with an electronic document image sample having the same page content as the document actual shooting image sample to obtain an electronic document image sample with features aligned with the document actual shooting image sample; Taking an electronically aligned document image sample and a document actual shooting image sample with the same page content as a positive sample image pair, and taking an electronically aligned document image sample and a document actual shooting image sample with different page content as a negative sample image pair; Based on the positive sample image pair and the negative sample image pair, training an initial image contrast learning model until a preset training condition is met to obtain an image contrast learning model; Wherein, the image contrast learning model is used to output an electronic document image feature vector and a document actual shooting image feature vector according to the input electronic document image and the document actual shooting image.
[0118] In yet another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is used to execute the training method of the contrast learning model provided by the above-mentioned various methods. The method includes: aligning the features of a document actual shooting image sample with an electronic document image sample having the same page content as the document actual shooting image sample to obtain an electronic document image sample with features aligned with the document actual shooting image sample; An electronic document image sample and a document actual photo image sample with one feature of the page content aligned are used as a positive sample image pair, and an electronic document image sample and a document actual photo image sample with different page content are used as a negative sample image pair; Based on the positive sample image pair and the negative sample image pair, the initial image contrast learning model is trained until the preset training conditions are met, and an image contrast learning model is obtained; Among them, the image contrast learning model is used to output an electronic document image feature vector and a document actual photo image feature vector according to the input electronic document image and the document actual photo image.
[0119] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.
[0120] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solutions, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0121] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A training method for a contrastive learning model, characterized in that Including: Aligning the features of the actual photographed image sample of the document with the electronic document image sample having the same page content as that of the actual photographed image sample of the document, to obtain an electronic document image sample with features aligned with the actual photographed image sample of the document; Taking an electronic document image sample with features aligned and an actual photographed image sample of the document having the same page content as a positive sample image pair, and taking an electronic document image sample and an actual photographed image sample having different page contents as a negative sample image pair; Training an initial image contrast learning model based on the positive sample image pair and the negative sample image pair until a preset training condition is satisfied, to obtain an image contrast learning model; Wherein, the image contrast learning model is used to output an electronic document image feature vector and an actual photographed image feature vector according to the input electronic document image and the actual photographed image of the document.
2. The training method of the contrastive learning model according to claim 1, wherein The training of the initial image contrast learning model based on the positive sample image pair and the negative sample image pair until a preset training condition is satisfied to obtain an image contrast learning model includes: Inputting the positive sample image pair into the initial image contrast learning model to obtain a positive sample image feature vector pair output by the initial image contrast learning model; Inputting the negative sample image pair into the initial image contrast learning model to obtain a negative sample image feature vector pair output by the initial image contrast learning model; Determining a contrast loss based on the feature similarity of the positive sample image feature vector pair and the feature similarity of the negative sample image feature vector pair; Adjusting the parameters of the initial image contrast learning model based on the contrast loss until a preset training condition is satisfied, to obtain an image contrast learning model.
3. The training method of the contrastive learning model according to claim 1, wherein The aligning the features of the actual photographed image sample of the document with the electronic document image sample having the same page content as that of the actual photographed image sample of the document to obtain an electronic document image sample with features aligned with the actual photographed image sample of the document includes: Performing noise addition processing on the electronic document image sample according to the noise features in the actual photographed image sample of the document having the same page content as the electronic document image sample, so that the electronic document image sample learns the noise features of the actual photographed image sample, to obtain an electronic document image sample with features aligned with the actual photographed image sample of the document; Wherein, the noise addition processing includes at least one of the following: noise injection processing, illumination simulation processing.
4. The training method of the contrastive learning model according to claim 3, characterized in that The illumination simulation processing of the electronic document image sample includes: Performing illumination analysis on the actual photographed image sample of the document having the same page content as the electronic document image sample to obtain an illumination analysis result; Determining the illumination simulation method of the electronic document image sample according to the illumination analysis result; Performing illumination simulation processing on the electronic document image sample based on the illumination simulation method; Among them, the light simulation method includes one or more of color temperature offset adjustment, non-uniform light synthesis, surface specular reflection synthesis, and dynamic range adjustment; the color temperature offset adjustment is used to adjust the RGB channel weights of the electronic document image sample; the non-uniform light synthesis is used to generate a gradient mask and superimpose it on the electronic document image sample; the surface specular reflection synthesis is used to add elliptical high-light areas at random positions in the electronic document image sample; the dynamic range adjustment is used to perform gamma correction and histogram clipping on the electronic document image sample.
5. The training method of the contrastive learning model according to claim 3, wherein The noise injection process for the electronic document image sample includes: Performing noise analysis on a document real-shot image sample with the same page content as the electronic document image sample to determine the noise type; the noise type includes Gaussian noise and salt-and-pepper noise; Superimposing the noise of the noise type on the electronic document image sample.
6. An image retrieval method, characterized in that, It includes: Constructing an image pair based on any electronic document image in the document image library and the real-shot image to be retrieved, traversing each electronic document image in the document image library, and obtaining multiple image pairs; Inputting each image pair into the contrast learning model respectively, and outputting the feature vectors of the electronic document image and the document real-shot image corresponding to each image pair; Taking the electronic document image in the target image pair as the image retrieval result of the real-shot image to be retrieved; where the target image pair is the image pair with the highest feature similarity between the feature vectors of the electronic document image and the document real-shot image among all the image pairs; The contrast learning model is trained based on the training method of the contrast learning model according to any one of claims 1 to 5.
7. An image retrieval device, characterized in that, It includes: An image pair construction module, configured to construct an image pair based on any electronic document image in the document image library and the real-shot image to be retrieved, traverse each electronic document image in the document image library, and obtain multiple image pairs; A model processing module, configured to input each image pair into the contrast learning model respectively, and output the feature vectors of the electronic document image and the document real-shot image corresponding to each image pair; A retrieval module, configured to take the electronic document image in the target image pair as the image retrieval result of the real-shot image to be retrieved; where the target image pair is the image pair with the highest feature similarity between the feature vectors of the electronic document image and the document real-shot image among all the image pairs; The contrast learning model is trained based on the training method of the contrast learning model according to any one of claims 1 to 5.
8. An image retrieval system, characterized in that, It includes an image acquisition device and the image retrieval device according to claim 7; The image acquisition device is used to obtain the real-shot image to be retrieved.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the image retrieval method according to claim 6.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the image retrieval method according to claim 6.
Citation Information
Patent Citations
Document / image searching method and program, and document / image recording and searching device
CN101133429A
Cultural relic image retrieval system based on comparative learning
CN114610941A
Chest X-ray quality evaluation method fusing medical field knowledge
CN116245828A
Document image registration data synthesis method, system and device and medium
CN116452641A
Binary code image retrieval method and system based on comparative learning
CN117573915A
Cited By
Document retrieval method and device, equipment, storage medium and computer program product
CN121658634A