Image comparison method and device, computer equipment and storage medium

Through dense semantic matching and pixel-level alignment, combined with image quality evaluation, the problem of low image comparison accuracy in the prior art is solved, and higher contrast accuracy and reliability are achieved.

CN119992132APending Publication Date: 2025-05-13SHENZHEN YISHIHUOLALA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510086385.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When using images with semantic complexity, the existing image comparison algorithm has low accuracy and is prone to misjudgment when comparing images with different view angles or position offsets of the same content. The uneven input image quality also leads to a decrease in the accuracy of the comparison.

Method used

Through dense semantic matching, the image to be detected is pixel-level aligned with the reference image, a stitching image is generated, and the image to be detected in the stitching image is consistently compared with the reference image, and the image quality of the image to be detected is evaluated, and the final comparison result is determined based on the consistency comparison results and quality evaluation results.

Benefits of technology

It improves the accuracy of image comparison, can comprehensively consider the overall information of the image, compare it from multiple aspects, which is more comprehensive and accurate than traditional methods, and effectively recognizes and processes comparison errors caused by image quality issues, improving the reliability of the final comparison results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992132A_ABST
    Figure CN119992132A_ABST
Patent Text Reader

Abstract

The invention discloses an image comparison method and device, computer equipment and a storage medium, and the method comprises the steps: carrying out the dense semantic matching of a to-be-detected image and a reference image, and obtaining an aligned to-be-detected image; splicing the reference image and the aligned to-be-detected image to obtain a spliced image; performing consistency comparison on the to-be-detected image in the spliced image and the reference image to obtain a consistency comparison result; performing quality evaluation on the image quality of the to-be-detected image to obtain a quality evaluation result; and determining a final comparison result based on the consistency comparison result and the quality evaluation result. Moreover, by evaluating the quality of the to-be-detected image, a comparison error caused by an image quality problem can be avoided, and the reliability of a comparison result is improved. The overall information of the image can be comprehensively considered, and the comparison is more comprehensive and accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence and computer vision technology, and in particular to an image comparison method, device, computer equipment and storage medium. Background Art

[0002] Image comparison is one of the important tasks in the field of computer vision. The input of the comparison algorithm is a pair of images, and the output is the consistency prediction score of the content on the image. The higher the score, the higher the consistency of the content predicted by the algorithm. The most common application scenario of the image comparison algorithm is one-to-one face verification. Other application scenarios include card verification, liveness recognition, human and vehicle re-identification, etc.

[0003] Traditional image comparison algorithms mainly rely on handcrafted features, such as SIFT (Scale Invariant Feature Transform), SURF (Speeded Up Robust Features), and HOG (Histogram of Oriented Gradients). These methods first extract representative local features from the image, and then calculate the similarity between images by feature matching. With the development of deep learning, some simple convolutional neural network (CNN) architectures are used for image comparison. These networks map images to a feature space by training on large-scale datasets, and use the distance of feature vectors to measure the similarity of images.

[0004] However, manual features mainly focus on local features of images, such as textures, edges, and key points, and have weak processing capabilities for complex semantic information and high-level semantic concepts. When processing images with semantic complexity, the accuracy is low. Although deep learning methods can extract more advanced semantic features, when performing image comparison, images of the same content with different perspectives or position offsets may be misjudged as dissimilar only through distance calculation of feature vectors. In addition, the input images may come from different devices and environments, and the image quality varies, which also leads to a decrease in comparison accuracy. Summary of the invention

[0005] Based on this, it is necessary to provide an image comparison method, device, computer equipment and storage medium to address the above technical issues, so as to solve at least one problem existing in the above-mentioned prior art.

[0006] In a first aspect, an image comparison method is provided, comprising:

[0007] Dense semantic matching is performed on the image to be detected and the reference image to obtain an aligned image to be detected;

[0008] Splicing the reference image with the aligned image to be detected to obtain a spliced ​​image;

[0009] Performing consistency comparison between the image to be detected in the spliced ​​image and the reference image to obtain a consistency comparison result;

[0010] Performing quality assessment on the image quality of the image to be detected to obtain a quality assessment result;

[0011] Based on the consistency comparison result and the quality assessment result, a final comparison result is determined.

[0012] In one embodiment, the step of performing dense semantic matching on the image to be detected and the reference image to obtain an aligned image to be detected includes:

[0013] Performing feature extraction on the image to be detected and the reference image respectively to obtain a feature map to be detected and a reference feature map;

[0014] Matching the feature map to be detected with the reference feature map to obtain feature similarity;

[0015] Based on the feature similarity, generating an optical flow field;

[0016] The image to be detected and the reference image are aligned at the pixel level based on the optical flow field to obtain the image to be detected aligned with the reference.

[0017] In one embodiment, the pixel-level alignment of the image to be detected and the reference image based on the optical flow field includes:

[0018] Each pixel point in the feature map to be detected is moved to a new position according to the displacement vector of the corresponding position in the optical flow field.

[0019] In one embodiment, the performing quality assessment on the image quality of the image to be detected to obtain a quality assessment result includes:

[0020] Construct quality assessment prompt text based on preset quality assessment standards;

[0021] The quality assessment prompt text and the image to be detected are input into the visual language model to obtain the quality assessment result.

[0022] In one embodiment, the quality assessment prompt text includes:

[0023] Determine whether the image to be detected has traces of being copied; and / or

[0024] Determine whether the clarity of the image to be detected is lower than a preset clarity condition; and / or

[0025] Determine whether the image to be detected has traces of tampering.

[0026] In one embodiment, performing consistency comparison between the image to be detected in the spliced ​​image and the reference image to obtain a consistency comparison result includes:

[0027] According to the preset consistency comparison standard, construct image consistency comparison prompt text;

[0028] Inputting the image consistency comparison prompt text and the spliced ​​image into the visual language large model to obtain a consistency evaluation result between the image to be detected and the reference image;

[0029] In one embodiment, the image to be detected is a car sticker image to be detected, the reference image is a reference car sticker image, and the image consistency comparison prompt text includes:

[0030] Determining whether the trademark sticker in the reference vehicle sticker image exists in the vehicle sticker image to be detected; and / or

[0031] Determine whether the vehicle sticker image to be detected is damaged compared to the reference vehicle sticker image; and / or

[0032] Determine whether the position of the trademark sticker in the to-be-detected vehicle sticker image is different from the position of the trademark sticker in the reference vehicle sticker image; and / or

[0033] Determine whether the text of the trademark sticker in the to-be-detected vehicle sticker image is different from the text of the trademark sticker in the reference vehicle sticker image.

[0034] In a second aspect, an image comparison device is provided, comprising:

[0035] A dense semantic matching unit, used for performing dense semantic matching between the image to be detected and the reference image to obtain an aligned image to be detected;

[0036] A stitched image acquisition unit, used for stitching the reference image with the aligned image to be detected to obtain a stitched image;

[0037] A comparison unit, used for performing consistency comparison between the image to be detected in the spliced ​​image and the reference image to obtain a consistency comparison result;

[0038] An image quality assessment unit, used to perform quality assessment on the image to be detected to obtain a quality assessment result;

[0039] A comparison result determination unit is used to determine a final comparison result based on the consistency comparison result and the quality assessment result.

[0040] In a third aspect, a computer device is provided, comprising a memory, a processor, and computer-readable instructions stored in the memory and running on the processor, wherein the image comparison method as described above is implemented when the processor executes the computer-readable instructions.

[0041] In a fourth aspect, a readable storage medium is provided, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by a processor, the image comparison method as described above is implemented.

[0042] The above-mentioned image comparison method, device, computer equipment and storage medium, the method of which is implemented, includes: dense semantic matching of the image to be detected and the reference image to obtain an aligned image to be detected; splicing the reference image and the aligned image to be detected to obtain a spliced ​​image; performing consistency comparison between the image to be detected and the reference image in the spliced ​​image to obtain a consistency comparison result; performing quality assessment on the image quality of the image to be detected to obtain a quality assessment result; and determining the final comparison result based on the consistency comparison result and the quality assessment result. In the embodiment of the present application, the image to be detected and the reference image are aligned at the pixel level by dense semantic matching, which not only considers the semantic information of the image, but also solves the problem of image position offset. The accuracy of image comparison can be improved to a certain extent. By splicing the reference image and the aligned image to be detected, and then performing consistency comparison, the overall information of the image can be comprehensively considered, and comparison can be performed from multiple aspects such as overall layout, semantic information and details, which is more comprehensive and accurate than the traditional method of only comparing from local or feature vectors. In addition, by evaluating the quality of the image to be tested, it is possible to effectively identify the comparison errors that may be caused by the quality problems of the image itself (such as traces of reshoots, low clarity, traces of tampering, etc.). Determining the final comparison result based on the quality evaluation results and the consistency comparison results can avoid drawing wrong conclusions due to poor image quality and improve the reliability of the final comparison result. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.

[0044] Figure 1 is a schematic diagram of a process of an image comparison method in one embodiment of the present invention;

[0045] Figure 2 is a schematic diagram of an application environment of an image comparison method in one embodiment of the present invention;

[0046] Figure 3 is a schematic diagram of an application environment of an image alignment method according to an embodiment of the present invention;

[0047] Figure 4 is a schematic diagram of an application environment of an image stitching method according to an embodiment of the present invention;

[0048] Figure 5 is a structural schematic diagram of an image contrast device in one embodiment of the present invention;

[0049] Figure 6 is a schematic diagram of a computer device in one embodiment of the present invention. DETAILED DESCRIPTION

[0050] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0051] In one embodiment, if Figure 1 , Figure 2 As shown, an image comparison method is provided, comprising the following steps:

[0052] In step S110, dense semantic matching is performed on the image to be detected and the reference image to obtain an aligned image to be detected;

[0053] In the embodiment of the present application, an image to be detected can be collected, and the image to be detected may include information such as car stickers, trademarks, and clothing, so as to facilitate subsequent review of the car stickers, trademarks, and clothing. The car sticker refers to a sticker pasted on the surface of a vehicle, which may include a variety of styles according to its purpose, and is usually provided with information such as brand logos, product information, slogans, and patterns.

[0054] It can be understood that the reference image refers to a pre-photographed image including a standard target object, such as an image of a standard car sticker style and pasting position, a standard trademark pattern, a standard clothing style, and the like.

[0055] Optionally, after obtaining the image to be detected, the image to be detected and the reference image can be input into a dense semantic matching neural network model, such as the PWC-Net (Pyramid-Warping-CostVolume-Network) model, FlowNet, Mask R-CNN, etc. for dense matching. Taking PWC-Net as an example, it can effectively estimate the optical flow field between images by constructing an image pyramid, performing feature distortion and cost volume calculation, and then realize dense semantic matching of images to obtain pixel-level aligned images to be detected. The image to be detected and the reference image are aligned at the pixel level through dense semantic matching, which not only takes into account the semantic information of the image, but also solves the problem of image position offset, which can improve the accuracy of image comparison to a certain extent.

[0056] like Figure 3 As shown, taking a car sticker image as an example, after dense semantic alignment between the car sticker to be detected and the reference car sticker, the car sticker to be detected aligned with the reference car sticker can be obtained.

[0057] In step S120, the reference image is spliced ​​with the aligned image to be detected to obtain a spliced ​​image;

[0058] Optionally, the reference image and the aligned image to be detected may be spliced ​​to obtain a spliced ​​image. Specifically, the reference image and the aligned image to be detected may be spliced ​​horizontally, horizontally, vertically, faded, or based on a specific area. Taking horizontal splicing as an example, the hconcat function of OpenCV may be used to splice the images together.

[0059] Taking horizontal stitching as an example, the stitched image can be Figure 4 As shown, the reference image can be placed on the left and the aligned image to be detected can be placed on the right.

[0060] In step S130, a consistency comparison is performed between the image to be detected in the spliced ​​image and the reference image to obtain a consistency comparison result.

[0061] Optionally, the color features, texture features, shape features, text semantic features and other features of the image to be detected and the reference image can be compared for consistency. Exemplarily, for color feature comparison, the color histograms of the reference image and the image to be detected can be calculated respectively. The color histogram counts the frequency of occurrence of different colors in the image and can reflect the overall color distribution of the image. By comparing the similarity of the color histograms of the two, the image consistency can be preliminarily judged. For example, the difference between the histograms can be quantified using metrics such as Bhattacharyya distance and chi-square distance. If the difference value is small, it means that the two images are similar in color distribution and have high consistency. For texture features, the image to be detected and the reference image can be divided into multiple small areas respectively. In each area, the grayscale value of its neighborhood pixels and the center pixel is compared with the central pixel based on the center pixel. Binary codes are generated according to the comparison results, and these codes constitute LBP features. The LBP features of each area of ​​the reference image and the image to be detected are calculated. The higher the similarity, the more consistent the texture. For shape features, edge detection algorithms (such as Canny algorithm) can be used to extract the contours of objects in the image. For the reference image and the image to be detected, their contour information is obtained respectively. Then, the shape is described by calculating the geometric features of the contour, such as the perimeter, area, and shape complexity. These geometric features of the corresponding shapes in the two images are compared. If the feature values ​​are close, it means that the shape consistency is high. For text semantic features, the text in the reference image and the image to be detected can be extracted and segmented to obtain multiple independent characters. They are then compared to determine whether the characters are consistent. If they are consistent, it means that the reference image is consistent with the image to be detected.

[0062] In addition to the above methods, instruction prompt text can also be constructed according to the consistency comparison requirements, such as determining whether the trademark sticker in the reference car sticker image exists in the car sticker image to be detected; determining whether the car sticker image to be detected is damaged compared to the reference car sticker image; determining whether the position of the trademark sticker in the car sticker image to be detected is different from the position of the trademark sticker in the reference car sticker image; determining whether the text of the trademark sticker in the car sticker image to be detected is different from the text of the trademark sticker in the reference car sticker image, etc. Then the spliced ​​image and the instruction prompt text can be input into the visual language large model. If the model outputs "yes", it means that the image to be detected is consistent with the reference image, and the subsequent operations can be continued. If the model outputs "no", it means that the image to be detected is inconsistent with the reference image.

[0063] It should be noted that if the image to be detected is inconsistent with the reference image, a prompt message can be output to inform the user to make adjustments in time. For example, if the car sticker is pasted incorrectly, it needs to be corrected in time.

[0064] Among them, the visual language model adds image input to the large language model, is trained on extremely large-scale image-text data pairs, has strong generalization capabilities, and performs well in tasks such as zero-sample image-text question answering. The visual language model can be used to judge noise and interference such as moiré, reflection, blur, etc. in images. It can be Qwen-VL, CLIP (Contrastive Language-Image Pretraining), ALIGN (A Large-scale In-languageImage-Text Pretraining), VisualBERT, etc.

[0065] In step S140, the image quality of the image to be detected is evaluated to obtain a quality evaluation result;

[0066] Optionally, it can be determined whether the image clarity of the image to be detected meets the preset image clarity requirement; determine whether there is color difference in the image color quality; determine whether the image has problems such as copying, tampering, dirtiness, damage, etc.

[0067] Exemplarily, for image clarity: the gradient amplitude of the image can be calculated, and a clear image usually has a higher gradient amplitude. For example, the gradient of the image in the horizontal and vertical directions is calculated by the Sobel operator, and then the gradient amplitude is calculated. If the average gradient amplitude is lower than the preset threshold, it indicates that the image clarity is poor. For image color: the color gamut range covered by the color in the image can be calculated and compared with the standard color gamut. For example, in the RGB color space, the distribution range of different color components can be counted. If the color gamut coverage is lower than the preset value, it means that the image color is not rich enough. Or the deviation of the color of the corresponding pixel between the image to be detected and the reference image can be calculated. This deviation can be measured in the color space using a metric such as the Euclidean distance. If the deviation exceeds the preset standard, the color quality does not meet the requirements. Alternatively, a pre-trained image classification model, such as VGG16, ResNet, etc., can be used. The model pre-trained on a large-scale image dataset can extract high-level semantic features of the image. The image to be detected is input into the model to obtain the feature representation of a specific layer. By comparing these features with the feature distribution of known high-quality images (for example, using a metric such as Kullback-Leibler divergence), the quality of the image to be detected is judged. If the feature distribution difference exceeds a preset threshold, the image quality does not meet the standard.

[0068] In addition to the above methods, instruction prompt text can also be generated according to the image quality assessment requirements, such as determining whether the image to be detected has traces of reshoots, whether the clarity is lower than the preset clarity condition, whether there are traces of tampering, etc. Then, the image to be detected and the instruction prompt text can be input into the pre-built visual language large model. If the model outputs "yes", it means that the quality of the image to be detected meets the preset quality assessment standard, and subsequent operations can be continued. If the model outputs "no", it means that the quality of the image to be detected does not meet the preset quality assessment standard.

[0069] Among them, the visual language large model adds image input on the basis of the large language model, is trained on extremely large-scale image-text data pairs, has strong generalization ability, and performs well in tasks such as zero-sample image-text question answering. The visual language large model can be used to judge noise and interference such as moiré, reflection, blur, etc. in the image. It can be Qwen-VL, CLIP (Contrastive Language-Image Pretraining), ALIGN (A Large-scale In-languageImage-Text Pretraining), VisualBERT, etc.

[0070] It should be noted that in order to ensure that the image format to be detected is correct (such as common JPEG, PNG, etc.), and the resolution and other parameters are suitable for processing by the visual language model. If the image resolution is too high or too low, appropriate scaling preprocessing may be required to meet the input requirements of the visual language model.

[0071] In an embodiment of the present application, an image comparison method is provided, including: performing dense semantic matching on an image to be detected and a reference image to obtain an aligned image to be detected; splicing the reference image and the aligned image to be detected to obtain a spliced ​​image; performing consistency comparison on the image to be detected in the spliced ​​image and the reference image to obtain a consistency comparison result; performing quality assessment on the image quality of the image to be detected to obtain a quality assessment result; and determining a final comparison result based on the consistency comparison result and the quality assessment result. In an embodiment of the present application, the image to be detected and the reference image are aligned at the pixel level by dense semantic matching, which not only considers the semantic information of the image, but also solves the problem of image position offset. The accuracy of image comparison can be improved to a certain extent. By splicing the reference image and the aligned image to be detected, and then performing consistency comparison, the overall information of the image can be comprehensively considered, and comparison can be performed from multiple aspects such as overall layout, semantic information and details, which is more comprehensive and accurate than the traditional method of only comparing from local or feature vectors. In addition, the quality of the image to be detected is evaluated, and the comparison error that may be caused by the quality problem of the image itself (such as remake traces, low clarity, tampering traces, etc.) can be effectively identified. Determining the final comparison result based on the quality assessment result and the consistency comparison result can avoid drawing erroneous conclusions due to poor image quality and improve the reliability of the final comparison result.

[0072] In one embodiment of the present application, the step of performing dense semantic matching on the image to be detected and the reference image to obtain an aligned image to be detected includes:

[0073] Performing feature extraction on the image to be detected and the reference image respectively to obtain a feature map to be detected and a reference feature map;

[0074] Matching the feature map to be detected with the reference feature map to obtain feature similarity;

[0075] Based on the feature similarity, generating an optical flow field;

[0076] The image to be detected and the reference image are aligned at the pixel level based on the optical flow field to obtain the image to be detected aligned with the reference.

[0077] Optionally, the image to be detected and the reference image can be uniformly scaled to a suitable size first, and then the image to be detected and the reference image are input into the constructed feature extraction network for feature extraction. After operations such as convolution and pooling, the feature extraction network can extract the corresponding feature map, which may include features such as edges, textures, and shapes. Image pyramids can be constructed for the input image to be detected and the reference image respectively. The image pyramid is a series of images of different resolutions obtained by continuously downsampling the original image. After the image pyramid is constructed, feature distortion operations can be performed at different scales. It can be understood that according to the optical flow estimated at the previous scale, the feature map of the reference image is distorted to a position corresponding to the feature map of the image to be detected, so that the two images are more aligned in space, which facilitates subsequent optical flow estimation and pixel-level alignment.

[0078] After feature distortion, the cost volume is obtained by comparing the distorted reference image feature map and the feature map of the image to be detected. The cost volume can reflect the degree of feature difference between the image to be detected and the reference image at each position. Each element in the cost volume represents the similarity measure between the features of the two images at a specific position. By analyzing the cost volume, the corresponding relationship between the pixels between the images, i.e., the optical flow, can be found.

[0079] The difference between the predicted optical flow and the real optical flow obtained by manual annotation is calculated based on the loss function, such as the L1 or L2 loss function. Then, the loss value can be back-propagated to each layer of the network through the back-propagation algorithm, and the parameters of the network can be adjusted so that the predicted optical flow field can more accurately reflect the real motion relationship between images, and the loss value is continuously minimized, thereby optimizing the result of optical flow estimation and obtaining the optical flow field.

[0080] After the optical flow field is obtained, the original image to be detected is transformed using methods such as bilinear interpolation to generate an aligned image to be detected, thereby achieving pixel-level alignment.

[0081] In an embodiment of the present application, the pixel-level alignment of the image to be detected and the reference image based on the optical flow field includes:

[0082] Each pixel point in the feature map to be detected is moved to a new position according to the displacement vector of the corresponding position in the optical flow field.

[0083] Optionally, find the corresponding position of each pixel in the feature map to be detected in the optical flow field. Then calculate the new position (X, Y) of the pixel based on the displacement vector in the optical flow field. Since the new position X, Y may be a decimal, bilinear interpolation can be used to determine the pixel value of the position. Each pixel is moved to the new position in the above manner, and the pixel value of the new position is calculated according to bilinear interpolation, and finally a new feature map after the move is generated, that is, the aligned image to be detected.

[0084] In an embodiment of the present application, the step of performing quality assessment on the image quality of the image to be detected to obtain a quality assessment result includes:

[0085] Construct quality assessment prompt text based on preset quality assessment standards;

[0086] The quality assessment prompt text and the image to be detected are input into the visual language model to obtain the quality assessment result.

[0087] Optionally, instruction prompt text can be generated according to the image quality assessment requirements, and then the image to be detected and the instruction prompt text can be input into a pre-built visual language model. If the model outputs "yes", it means that the quality of the image to be detected meets the preset quality assessment standards, and subsequent operations can continue. If the model outputs "no", the image to be detected is reselected.

[0088] In one embodiment of the present application, the quality assessment prompt text includes:

[0089] Determine whether the image to be detected has traces of being copied; and / or

[0090] Determine whether the clarity of the image to be detected is lower than a preset clarity condition; and / or

[0091] Determine whether the image to be detected has traces of tampering.

[0092] Optionally, if the instruction prompt text is "Please determine whether the following image has any traces of copying. If so, please answer 'yes', if not, answer 'no'; at the same time, determine whether the image clarity is lower than the preset clarity condition. If so, please answer 'yes', if not, answer 'no'; finally, determine whether the image has any signs of tampering. If so, answer 'yes', if not, answer 'no'." Then, the image to be detected and the instruction prompt text can be input into the pre-built visual language large model. If the model outputs yes, it means that the quality of the image to be detected meets the preset quality assessment standard, and subsequent operations can continue. If the model outputs "no", it means that the quality of the image to be detected does not meet the preset quality assessment standard.

[0093] It can be understood that the prompt text may include one of the above prompt texts or any combination thereof.

[0094] Among them, the large visual language model can be Qwen-VL, CLIP (Contrastive Language-Image Pretraining), ALIGN (A Large-scale In-language Image-Text Pretraining), VisualBERT, etc.

[0095] In an embodiment of the present application, performing consistency comparison between the image to be detected in the spliced ​​image and the reference image to obtain a consistency comparison result includes:

[0096] According to the preset consistency comparison standard, construct image consistency comparison prompt text;

[0097] Inputting the image consistency comparison prompt text and the spliced ​​image into the visual language large model to obtain a consistency evaluation result between the image to be detected and the reference image;

[0098] Optionally, an instruction prompt text can be constructed according to the consistency comparison requirements, and then the stitched image and the instruction prompt text can be input into the visual language model. If the model outputs "yes", it means that the image to be detected is consistent with the reference image, and subsequent operations can continue. If the model outputs "no", it means that the image to be detected is inconsistent with the reference image.

[0099] In one embodiment of the present application, the image to be detected is a car sticker image to be detected, the reference image is a reference car sticker image, and the image consistency comparison prompt text includes:

[0100] Determining whether the trademark sticker in the reference vehicle sticker image exists in the vehicle sticker image to be detected; and / or

[0101] Determine whether the vehicle sticker image to be detected is damaged compared to the reference vehicle sticker image; and / or

[0102] Determine whether the position of the trademark sticker in the to-be-detected vehicle sticker image is different from the position of the trademark sticker in the reference vehicle sticker image; and / or

[0103] Determine whether the text of the trademark sticker in the to-be-detected vehicle sticker image is different from the text of the trademark sticker in the reference vehicle sticker image.

[0104] Optionally, if the prompt text is "determining whether the trademark sticker in the reference car sticker image exists in the car sticker image to be detected", if yes, answer 'yes', if no, answer 'no'; if the prompt text is determining whether the car sticker image to be detected is damaged compared to the reference car sticker image, if yes, answer

[0105] If the prompt text is to determine whether the position of the trademark sticker in the to-be-detected car sticker image is different from that in the reference car sticker image, answer 'yes' if yes, otherwise answer 'no'; if the prompt text is to determine whether the text of the trademark sticker in the to-be-detected car sticker image is different from that of the trademark sticker in the reference car sticker image, answer 'yes' if yes, otherwise answer 'no'. Then the spliced ​​image and the instruction prompt text can be input into the visual language model. If the model outputs "yes", it means that the to-be-detected image is consistent with the reference image, and the subsequent operations can be continued. If the model outputs "no", it means that the to-be-detected image is inconsistent with the reference image.

[0106] It can be understood that the prompt text may include one of the above prompt texts or any combination thereof.

[0107] Among them, the large visual language model can be Qwen-VL, CLIP (Contrastive Language-Image Pretraining), ALIGN (A Large-scale In-language Image-Text Pretraining), VisualBERT, etc.

[0108] In the embodiment of the present application, the image to be detected is aligned with the reference image at the pixel level by dense semantic matching, which not only considers the semantic information of the image, but also solves the problem of image position offset. The accuracy of image comparison can be improved to a certain extent. By splicing the reference image and the aligned image to be detected, and then performing consistency comparison, the overall information of the image can be comprehensively considered, and the comparison is performed from multiple aspects such as overall layout, semantic information and details, which is more comprehensive and accurate than the traditional method of only comparing from local or feature vectors. In addition, the quality of the image to be detected is evaluated, and the comparison error that may be caused by the quality problem of the image itself (such as remake traces, low clarity, tampering traces, etc.) can be effectively identified. The final comparison result is determined jointly according to the quality assessment result and the consistency comparison result, which can avoid drawing wrong conclusions due to poor image quality and improve the reliability of the final comparison result. The comparison process does not require human participation, and can detect whether the image is damaged or part of the object is displaced, and other minor changes, and there is no need to retrain the model when the object style increases, such as the increase in the style of car stickers.

[0109] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.

[0110] In one embodiment, an image comparison device is provided, which corresponds one-to-one to the image comparison method in the above embodiment. Figure 5 As shown, the image comparison device includes a dense semantic matching unit 10, a spliced ​​image acquisition unit 20, a comparison unit 30, an image quality assessment unit 40 and a comparison result determination unit 50. Each functional module is described in detail as follows:

[0111] A dense semantic matching unit 10 is used to perform dense semantic matching on the image to be detected and the reference image to obtain an aligned image to be detected;

[0112] A stitched image acquisition unit 20 is used to stitch the reference image with the aligned image to be detected to obtain a stitched image;

[0113] A comparison unit 30 is used to perform consistency comparison between the image to be detected in the spliced ​​image and the reference image to obtain a consistency comparison result;

[0114] An image quality assessment unit 40 is used to perform quality assessment on the image to be detected to obtain a quality assessment result;

[0115] The comparison result determination unit 50 is used to determine a final comparison result based on the consistency comparison result and the quality assessment result.

[0116] In one embodiment of the present application, the dense semantic matching unit 10 is further used for:

[0117] Performing feature extraction on the image to be detected and the reference image respectively to obtain a feature map to be detected and a reference feature map;

[0118] Matching the feature map to be detected with the reference feature map to obtain feature similarity;

[0119] Based on the feature similarity, generating an optical flow field;

[0120] The image to be detected and the reference image are aligned at the pixel level based on the optical flow field to obtain the image to be detected aligned with the reference.

[0121] In one embodiment of the present application, the dense semantic matching unit 10 is further used for:

[0122] Each pixel point in the feature map to be detected is moved to a new position according to the displacement vector of the corresponding position in the optical flow field.

[0123] In an embodiment of the present application, the image quality determination unit 20 is further configured to:

[0124] Construct quality assessment prompt text based on preset quality assessment standards;

[0125] The quality assessment prompt text and the image to be detected are input into the visual language model to obtain the quality assessment result.

[0126] In an embodiment of the present application, the image quality determination unit 20 is further configured to:

[0127] Determine whether the image to be detected has traces of being copied; and / or

[0128] Determine whether the clarity of the image to be detected is lower than a preset clarity condition; and / or

[0129] Determine whether the image to be detected has traces of tampering.

[0130] In one embodiment of the present application, the comparison unit 40 is further used for:

[0131] According to the preset consistency comparison standard, construct image consistency comparison prompt text;

[0132] Inputting the image consistency comparison prompt text and the spliced ​​image into the visual language large model to obtain a consistency evaluation result between the image to be detected and the reference image;

[0133] In one embodiment of the present application, the image to be detected is a vehicle sticker image to be detected, the reference image is a reference vehicle sticker image, and the comparison unit 40 is further used to:

[0134] Determining whether the trademark sticker in the reference vehicle sticker image exists in the vehicle sticker image to be detected; and / or

[0135] Determine whether the vehicle sticker image to be detected is damaged compared to the reference vehicle sticker image; and / or

[0136] Determine whether the position of the trademark sticker in the to-be-detected vehicle sticker image is different from the position of the trademark sticker in the reference vehicle sticker image; and / or

[0137] Determine whether the text of the trademark sticker in the to-be-detected vehicle sticker image is different from the text of the trademark sticker in the reference vehicle sticker image.

[0138] In the embodiment of the present application, the image to be detected is aligned with the reference image at the pixel level by dense semantic matching, which not only considers the semantic information of the image, but also solves the problem of image position offset. The accuracy of image comparison can be improved to a certain extent. By splicing the reference image and the aligned image to be detected, and then performing consistency comparison, the overall information of the image can be comprehensively considered, and the comparison is performed from multiple aspects such as overall layout, semantic information and details, which is more comprehensive and accurate than the traditional method of only comparing from local or feature vectors. In addition, the quality of the image to be detected is evaluated, and the comparison error that may be caused by the quality problem of the image itself (such as remake traces, low clarity, tampering traces, etc.) can be effectively identified. The final comparison result is determined jointly according to the quality assessment result and the consistency comparison result, which can avoid drawing wrong conclusions due to poor image quality and improve the reliability of the final comparison result. The comparison process does not require human participation, and can detect whether the image is damaged or part of the object is displaced, and other minor changes, and there is no need to retrain the model when the object style increases, such as the increase in the style of car stickers.

[0139] For the specific definition of the image comparison device, please refer to the definition of the image comparison method above, which will not be repeated here. Each module in the above-mentioned image comparison device can be implemented in whole or in part by software, hardware and a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.

[0140] In one embodiment, a computer device is provided. The computer device may be a terminal device, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a readable storage medium. The readable storage medium stores computer-readable instructions. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer-readable instructions are executed by the processor, an image comparison method is implemented. The readable storage medium provided in this embodiment includes a non-volatile readable storage medium and a volatile readable storage medium.

[0141] In an embodiment of the present application, a computer device is provided, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, the steps of the above-mentioned image comparison method are implemented.

[0142] In an embodiment of the application, a readable storage medium is provided, which stores computer-readable instructions. When the computer-readable instructions are executed by a processor, the steps of the above-mentioned image comparison method are implemented.

[0143] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing related hardware through computer-readable instructions, and the computer-readable instructions can be stored in a non-volatile readable storage medium or a volatile readable storage medium. When the computer-readable instructions are executed, they may include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0144] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0145] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. An image comparison method, characterized in that: The method comprises: Dense semantic matching is performed on the image to be detected and the reference image to obtain an aligned image to be detected; Splicing the reference image with the aligned image to be detected to obtain a spliced ​​image; Performing consistency comparison between the image to be detected in the spliced ​​image and the reference image to obtain a consistency comparison result; Performing quality assessment on the image quality of the image to be detected to obtain a quality assessment result; Based on the consistency comparison result and the quality assessment result, a final comparison result is determined.

2. The image comparison method according to claim 1, characterized in that: The step of performing dense semantic matching on the image to be detected and the reference image to obtain an aligned image to be detected includes: Performing feature extraction on the image to be detected and the reference image respectively to obtain a feature map to be detected and a reference feature map; Matching the feature map to be detected with the reference feature map to obtain feature similarity; Based on the feature similarity, generating an optical flow field; The image to be detected and the reference image are aligned at the pixel level based on the optical flow field to obtain the image to be detected aligned with the reference.

3. The image comparison method according to claim 2, characterized in that: The performing pixel-level alignment on the image to be detected and the reference image based on the optical flow field comprises: Each pixel point in the feature map to be detected is moved to a new position according to the displacement vector of the corresponding position in the optical flow field.

4. The image comparison method according to claim 1, characterized in that: The step of performing quality assessment on the image quality of the image to be detected to obtain a quality assessment result includes: Construct quality assessment prompt text based on preset quality assessment standards; The quality assessment prompt text and the image to be detected are input into the visual language model to obtain the quality assessment result.

5. The image comparison method according to claim 4, characterized in that: The quality assessment prompt text includes: Determine whether the image to be detected has traces of being copied; and / or Determine whether the clarity of the image to be detected is lower than a preset clarity condition; and / or Determine whether the image to be detected has traces of tampering.

6. The image comparison method according to claim 1, characterized in that: The performing consistency comparison between the image to be detected in the spliced ​​image and the reference image to obtain a consistency comparison result includes: According to the preset consistency comparison standard, construct image consistency comparison prompt text; The image consistency comparison prompt text and the spliced ​​image are input into the visual language large model to obtain the consistency evaluation result of the image to be detected and the reference image.

7. The image comparison method according to claim 6, characterized in that: The image to be detected is a car sticker image to be detected, the reference image is a reference car sticker image, and the image consistency comparison prompt text includes: Determining whether the trademark sticker in the reference vehicle sticker image exists in the vehicle sticker image to be detected; and / or Determine whether the vehicle sticker image to be detected is damaged compared to the reference vehicle sticker image; and / or Determine whether the position of the trademark sticker in the to-be-detected vehicle sticker image is different from the position of the trademark sticker in the reference vehicle sticker image; and / or Determine whether the text of the trademark sticker in the to-be-detected vehicle sticker image is different from the text of the trademark sticker in the reference vehicle sticker image.

8. An image contrast device, characterized in that: The device comprises: A dense semantic matching unit, used for performing dense semantic matching between the image to be detected and the reference image to obtain an aligned image to be detected; A stitched image acquisition unit, used for stitching the reference image with the aligned image to be detected to obtain a stitched image; A comparison unit, used for performing consistency comparison between the image to be detected in the spliced ​​image and the reference image to obtain a consistency comparison result; An image quality assessment unit, used to perform quality assessment on the image to be detected to obtain a quality assessment result; A comparison result determination unit is used to determine a final comparison result based on the consistency comparison result and the quality assessment result.

9. A computer device comprising a memory, a processor, and computer-readable instructions stored in the memory and executed on the processor, characterized in that: When the processor executes the computer-readable instructions, the image comparison method according to any one of claims 1 to 7 is implemented.

10. A readable storage medium having computer readable instructions stored thereon, characterized in that: When the computer-readable instructions are executed by a processor, the image comparison method according to any one of claims 1 to 7 is implemented.