Image processing method and device, computer equipment, storage medium and program product
By identifying similar regions and calculating structural and information similarity in image processing, the problem of inaccurate image content similarity in existing technologies is solved, achieving higher accuracy in judgment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-04-07
AI Technical Summary
Existing average hashing and perceptual hashing algorithms cannot accurately reflect the similarity in image content when judging image similarity.
By obtaining similar regions in the first and second images, structural similarity and information similarity are determined, and the content similarity of the images is judged by combining the two.
It improves the accuracy of image content similarity, reduces the interference of redundant information on similarity calculation, and can more accurately judge the visual and semantic similarity of images.
Smart Images

Figure CN121811074A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to an image processing method, apparatus, computer equipment, storage medium, and program product. Background Technology
[0002] With the development of artificial intelligence and internet technology, there is a growing demand for image comparison analysis in many scenarios. For example, in autonomous driving, captured images can be matched with maps or pre-collected images for subsequent location analysis. Another example is comparing user-captured images with pre-stored images to determine if the screenshot contains similar content for further processing.
[0003] In traditional techniques, the similarity between images can be determined using average hashing algorithms, perceptual hashing algorithms, or large models.
[0004] However, average hashing, perceptual hashing, or large models primarily judge the similarity between images as a whole, which does not accurately reflect the similarity of images in terms of content. Summary of the Invention
[0005] Therefore, it is necessary to provide an image processing method, apparatus, computer device, computer-readable storage medium, and computer program product that can improve the accuracy of image content similarity in response to the above-mentioned technical problems.
[0006] On one hand, this application provides an image processing method, comprising: acquiring a first image and a second image; determining a region in the first image similar to the second image to obtain a third image; determining a structural similarity between the second image and the third image, wherein the structural similarity reflects the similarity between the features presented by pixels in the second image and the features presented by pixels in the third image; determining an information similarity between the second image and the third image, wherein the information similarity reflects the similarity between the information contained in the second image and the information contained in the third image; and determining a content similarity between the first image and the second image based on the structural similarity and the information similarity.
[0007] On the other hand, this application also provides an image processing apparatus, comprising: a region detection module, configured to acquire a first image and a second image, and determine a region similar to the second image from the first image to obtain a third image; a first similarity determination module, configured to determine the structural similarity between the second image and the third image, wherein the structural similarity reflects the similarity between the features presented by pixels in the second image and the features presented by pixels in the third image; a second similarity determination module, configured to determine the information similarity between the second image and the third image, wherein the information similarity reflects the similarity between the information contained in the second image and the information contained in the third image; and a third similarity determination module, configured to determine the content similarity between the first image and the second image based on the structural similarity and the information similarity.
[0008] On the other hand, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described image processing method.
[0009] On the other hand, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-described image processing method.
[0010] On the other hand, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps in the above-described image processing method.
[0011] The aforementioned image processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product acquire a first image and a second image, determine a region in the first image that is similar to the second image to obtain a third image, determine the structural similarity between the second and third images, determine the informational similarity between the second and third images, and determine the content similarity between the first and second images based on the structural and informational similarities. By determining the region in the first image that is similar to the second image, the interference of redundant information in the first image that is unrelated to the content in the second image on the subsequent similarity calculation can be reduced, thus helping to improve the accuracy of the similarity. Since structural similarity is used to reflect the similarity between the features presented by pixels in the second image and the features presented by pixels in the third image, it can reflect visual similarity. Since informational similarity is used to reflect the similarity between the information contained in the second image and the information contained in the third image, it can reflect semantic similarity. Since both visual and semantic similarity can reflect content similarity, combining structural and informational similarity can better determine similarity, thereby improving the accuracy of content similarity. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is an application environment diagram of an image processing method in one embodiment;
[0014] Figure 2 This is a flowchart illustrating an image processing method in one embodiment;
[0015] Figure 3 This is a flowchart illustrating the image processing method in another embodiment;
[0016] Figure 4 This is a structural block diagram of an image processing device in one embodiment;
[0017] Figure 5 This is an internal structural diagram of a computer device in one embodiment;
[0018] Figure 6 This is a diagram of the internal structure of a computer device in another embodiment. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0020] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.
[0021] The image processing method provided in this application embodiment can be applied to, for example... Figure 1The application environment shown includes computer device 102 and terminal 104. Computer device 102 communicates with terminal 104 via a network. A data storage system can store the data that computer device 102 needs to process. The data storage system can be integrated onto computer device 102 or located in the cloud or on other network servers.
[0022] Specifically, computer device 102 acquires a first image and a second image, determines a region in the first image that is similar to the second image to obtain a third image, determines the structural similarity between the second and third images, determines the informational similarity between the second and third images, and determines the content similarity between the first and second images based on the structural and informational similarities. The structural similarity reflects the similarity between the features presented by pixels in the second and third images, and the informational similarity reflects the similarity between the information contained in the second and third images. The second image may be pre-stored in a data storage system. The first image may be acquired from terminal 104, for example, it may be an image captured by the user and sent by terminal 104. Alternatively, the first image may also be pre-stored in a data storage system.
[0023] The computer equipment can be a terminal or a server. Terminals can be, but are not limited to, various personal computers, laptops, smartphones, tablets, drones, low-altitude aircraft, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle systems, and projection equipment. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted displays. Head-mounted displays can be virtual reality (VR) devices, augmented reality (AR) devices, and smart glasses. Servers can be independent physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers providing cloud computing services.
[0024] Multimodal AI (Artificial Intelligence) large-scale models are poised to trigger a new industrial revolution. Tracking current industry developments, the multimodal development of large-scale models is deepening and is expected to become the mainstream of AI large-scale models. Following the rapid embedding of text-to-image capabilities into various large-scale models, text-to-video is the next important direction for multimodal applications of large-scale models. Recently, several manufacturers have successively released related products or updates, significantly improving the quality of text-to-video, achieving higher definition, smoother playback, and arbitrary video modification. It can be said that multimodality is an essential path to achieving general artificial intelligence and will inevitably become a cutting-edge direction in the development of large-scale models. However, there are still many problems when using large-scale models for screenshot analysis. For example, analyzing whether two website screenshots are completely identical, where one screenshot is manually taken by the user and includes not only the website page but also the browser title bar, Windows 10 desktop taskbar, and other extraneous information, while the other is a complete screenshot of the website without the browser title bar, Windows 10 desktop taskbar, etc., cannot accurately determine content consistency by relying solely on large-scale models. Based on this, the image processing method of this application is proposed.
[0025] In one exemplary embodiment, such as Figure 2 As shown, an image processing method is provided, which can be executed by a terminal or a server, or by both a terminal and a server, and can be applied to... Figure 1 Taking computer device 102 as an example, the following steps are included:
[0026] Step 202: Obtain the first image and the second image, and determine the region in the first image that is similar to the second image to obtain the third image.
[0027] The first image and the second image can be any two images. The first image and the second image can be the same size or different. The first image can be from the user, such as a screenshot, and the second image can be a reference image, such as a pre-captured image.
[0028] For example, the first image can be a sub-image of the second image. For instance, in the field of autonomous driving, the first image can be an image captured by an in-vehicle terminal, and the second image can be a pre-captured map image. The image captured by the in-vehicle terminal is a part of the map image.
[0029] For example, the second image can be a sub-image of the first image. For instance, the first image is a user-captured image of a website containing distracting information. The second image is a pre-captured image of a website that does not contain distracting information. Distracting information includes, for example, redundant information such as a browser title bar or desktop taskbar.
[0030] For example, a first image can be converted to grayscale to obtain a first grayscale image, and a second image can be converted to grayscale to obtain a second grayscale image. Based on the first and second grayscale images, a coordinate transformation relationship between the first and second images can be determined. Based on this coordinate transformation relationship, a third image can be obtained by identifying regions in the first image similar to the second image. Thus, compared to the first image, the third image has reduced redundant information, retaining only information similar to or identical to the second image.
[0031] Step 204: Determine the structural similarity between the second image and the third image. The structural similarity is used to reflect the similarity between the features presented by pixels in the second image and the features presented by pixels in the third image.
[0032] The features presented by a pixel include, but are not limited to, the features presented by the pixel's brightness or pixel value.
[0033] For example, the coordinate transformation relationship between the first image and the second image can be determined, and the second image and the third image can be aligned using the coordinate transformation relationship. After alignment, the structural similarity between the second image and the third image can be calculated.
[0034] For example, the second image can be compared with the third image using coordinate transformation to obtain an aligned second image. The structural similarity can be determined based on the features of the pixels in the aligned second image and the features of the pixels in the third image.
[0035] Step 206: Determine the information similarity between the second image and the third image. The information similarity is used to reflect the similarity between the information contained in the second image and the information contained in the third image.
[0036] The information can be any kind of information, such as, but not limited to, text or icons. Each element in the image can be understood as a type of information. Various information can be obtained by labeling the elements in the image.
[0037] For example, a computer device can use a multimodal model to determine the information similarity between a second image and a third image. Since information in an image is also part of its content, information similarity reflects the similarity of the images in terms of content.
[0038] For example, a computer device can determine information similarity based on the elements contained in the second image and the elements contained in the third image.
[0039] Step 208: Determine the content similarity between the first image and the second image based on structural similarity and information similarity.
[0040] For example, the smaller or larger of the structural similarity and information similarity can be used as the content similarity between the first image and the second image.
[0041] For example, the average of the calculated structural similarity and information similarity can be used as the content similarity between the first image and the second image.
[0042] For example, structural similarity and information similarity can be weighted and the result of the weighted calculation can be used as the content similarity between the first image and the second image.
[0043] In the image processing method described above, a first image and a second image are acquired. A third image is obtained by identifying regions in the first image that are similar to the second image. The structural similarity between the second and third images is determined, as well as the informational similarity between them. Based on the structural and informational similarities, the content similarity between the first and second images is determined. By identifying regions in the first image that are similar to the second image, the interference of redundant information in the first image that is unrelated to the content in the second image on subsequent similarity calculations can be reduced, thus improving the accuracy of similarity. Since structural similarity reflects the similarity between the features presented by pixels in the second and third images, it can reflect visual similarity. Since informational similarity reflects the similarity between the information contained in the second and third images, it can reflect semantic similarity. Because both visual and semantic similarity can reflect content similarity, combining structural and informational similarity can better determine similarity, thereby improving the accuracy of content similarity.
[0044] In an exemplary embodiment, information similarity includes a first information similarity and / or a second information similarity. Determining the information similarity between a second image and a third image includes: annotating information in the second image to obtain a first annotation set, and annotating information in the third image to obtain a second annotation set; determining the first information similarity between the second image and the third image based on the similarity between the first annotation set and the second annotation set; and / or, determining the second information similarity between the second image and the third image using a multimodal model.
[0045] The second and third images may contain one or more types of information, including but not limited to text or icons. Each element in the first annotation set represents a piece of information annotated from the second image, and each element in the second annotation set represents a piece of information annotated from the third image. A multimodal model is an artificial intelligence model capable of recognizing multimodal information; a multimodal model can be a large-scale multimodal model.
[0046] For example, an image annotation tool can be used to annotate the second and third images respectively, obtaining a first annotation set and a second annotation set. The image annotation tool can be any tool capable of annotating images. It can be based on artificial intelligence, such as large models, or on technologies other than artificial intelligence.
[0047] For example, the intersection of the first annotation set and the second annotation set can be determined, the number of elements in the intersection can be calculated, and the first information similarity can be determined based on the number of intersection elements. For instance, the first information similarity is positively correlated with the number of intersection elements.
[0048] For example, the second and third images can be input into a multimodal model, which calculates the similarity between information in the second and third images to obtain a second information similarity. For instance, a task prompt can be generated to instruct the multimodal model to analyze the consistency of information or content contained in the images and output a similarity score. Then, the task prompt, the second image, and the third image are input into the multimodal model to obtain the similarity score output by the multimodal model, i.e., the second information similarity.
[0049] For example, the smaller or larger of the structural similarity, the first information similarity, and the second information similarity can be used as the content similarity between the first image and the second image.
[0050] For example, the average of the calculated structural similarity, the first information similarity, and the second information similarity can be used as the content similarity between the first image and the second image.
[0051] For example, structural similarity, first information similarity, and second information similarity can be weighted and calculated, and the result of the weighted calculation can be used as the content similarity between the first image and the second image.
[0052] In this embodiment, the similarity between the first and second annotation sets reflects the similarity of the content or information contained in the second and third images. Therefore, determining the first information similarity based on the similarity between the first and second annotation sets ensures that the first information similarity accurately reflects the content similarity. Furthermore, since the multimodal model has the ability to analyze multimodal information, the second information similarity obtained using the multimodal model can also accurately reflect the content similarity.
[0053] In an exemplary embodiment, determining the first information similarity between the second image and the third image based on the similarity between the first annotation set and the second annotation set includes: determining the intersection and union between the first annotation set and the second annotation set; and determining the first information similarity between the second image and the third image based on the ratio between the number of elements in the intersection and the number of elements in the union.
[0054] For example, the number of elements contained in the union can be determined to obtain the number of union elements, and the ratio of the number of intersection elements to the number of union elements can be used as the first information similarity.
[0055] For example, if the first labeled set is list1, the second labeled set is list2, and their intersection is list3, then list1 = [list3, list-1], list2 = [list3, list-2], where list-1 refers to the set of elements that are in list1 but not in list2, and list-2 refers to the set of elements that are in list2 but not in list1. The similarity of the first information can then be calculated as: result = len(list3) / (len(list3) + len(list-1) + len(list-2)), where result represents the similarity of the first information.
[0056] In this embodiment, since the ratio between the number of elements in the intersection and the number of elements in the union reflects the proportion of identical elements in the two images, the first information similarity can be determined based on the ratio, which can improve the reliability of the first information similarity.
[0057] In an exemplary embodiment, determining the structural similarity between the second image and the third image includes: determining the coordinate transformation relationship between the first image and the second image; aligning the third image to the second image using the coordinate transformation relationship to obtain the aligned third image; and determining the structural similarity based on the features of pixels in the aligned third image and the features of pixels in the second image.
[0058] The coordinate transformation relationship can transform the first image from its own coordinate system to the coordinate system of the second image, so that the two images are in the same plane.
[0059] For example, coordinate transformation relationships can be used to map the third image from its own coordinate system to the coordinate system of the second image. For instance, perspective transformation can be performed on the third image using coordinate transformation relationships to obtain an aligned third image.
[0060] For example, if the aligned third image differs in size or shape from the second image, the aligned third image and / or the second image can be cropped to make their sizes and shapes consistent before structural similarity calculation. Taking cropping the second image as an example, after obtaining the cropped second image, multiple first local regions can be determined from the cropped second image, and multiple second local regions can be determined from the aligned third image. The first local regions and the second local regions correspond one-to-one, and the first local region and its corresponding second local region are two regions in the same position. That is, the position of the first local region in the cropped second image is consistent with the position of its corresponding second local region in the aligned third image. Structural similarity can be calculated based on these multiple first local regions and multiple second local regions.
[0061] For example, if the aligned third image has the same size and shape as the second image, then multiple first local regions can be determined from the second image, and multiple second local regions can be determined from the aligned third image, with the first local regions corresponding one-to-one with the second local regions.
[0062] For example, for each pair of local regions (a first local region and a second local region with a corresponding relationship), the brightness similarity between the first local region and the second local region can be calculated, the contrast similarity between the first local region and the second local region can be calculated, and the cosine similarity between the first local region and the second local region can be calculated. Based on at least one of the brightness similarity, contrast similarity, and cosine similarity, the comprehensive similarity between the first local region and the second local region can be determined. Since there are multiple pairs of local regions, each pair of local regions can correspond to a comprehensive similarity. The structural similarity can be determined for the comprehensive similarity corresponding to each pair of local regions. For example, the average of the comprehensive similarities corresponding to each pair of local regions can be calculated, and the result of the average calculation can be used as the structural similarity.
[0063] For example, for each pair of local regions, a regional brightness characterization value for a first local region and a regional brightness characterization value for a second local region can be determined. The regional brightness characterization value reflects the overall brightness of the region. The regional brightness characterization value can be represented using the average grayscale value of the pixels. The brightness similarity between the first and second local regions can be determined based on the regional brightness characterization values of the first and second local regions.
[0064] For example, the formula for calculating brightness similarity can be:
[0065]
[0066] in, Representing brightness similarity, x and y represent the first and second local regions, respectively. and The first local region represents the region brightness characterization value, and the second local region represents the region brightness characterization value, respectively. C1 is a preset value.
[0067] For example, the structural similarity index (SSIM) between the aligned third image and the second image can be calculated, and this structural similarity index can be used as the structural similarity.
[0068] For example, the difference image between the second and third images can be used, for instance, morphological operations can be used to remove smaller differences in the difference image, the difference regions can be found, and difference regions with areas smaller than an area threshold can be filtered out. Significant differences can then be marked in the third image to obtain a labeled image corresponding to the third image, which can then be saved. The standard image can serve as visual evidence for verifying, interpreting, and auditing the similarity of the content.
[0069] In this embodiment, the reliability of the structural similarity can be guaranteed by aligning the images before calculating the structural similarity.
[0070] In an exemplary embodiment, determining a region similar to a second image from a first image to obtain a third image includes: determining a coordinate transformation relationship between the first image and the second image; using the coordinate transformation relationship to determine a similar region mask, the similar region mask being used to reflect the position of similar regions in the first image and the second image; and using the similar region mask to obtain the third image from the first image.
[0071] The similar region mask is a binary mask image. The pixel value in the similar region mask is 1 or 0. The position of the pixel value 1 in the similar region mask represents the position of the similar region.
[0072] For example, the projection position of the second image onto the first image can be determined using coordinate transformation relationships; a binary mask image is generated based on the projection position, which is a similarity region mask. The size of the binary mask image is the same as that of the first image. The pixel value of the pixel corresponding to the projection position in the binary mask image is a first preset value, and the pixel values of the remaining pixels in the mask image are second preset values; the image region corresponding to the projection position is obtained from the first image using the mask image to obtain the third image. The first preset value is, for example, 1, and the second preset value is 0.
[0073] In this embodiment, since the similar region mask is used to reflect the position of similar regions in the first image and the second image, the similar regions, i.e. the third image, can be accurately obtained through the similar region mask.
[0074] In an exemplary embodiment, determining the coordinate transformation relationship between a first image and a second image includes: converting the first image to grayscale to obtain a first grayscale image, and converting the second image to grayscale to obtain a second grayscale image; performing feature extraction on the first grayscale image to obtain feature vectors of multiple first feature points, and performing feature extraction on the second grayscale image to obtain feature vectors of multiple second feature points; matching the multiple first feature points with the multiple second feature points based on the feature vectors to obtain multiple matching pairs, each matching pair containing one first feature point and one second feature point; and determining the coordinate transformation relationship based on the pixel coordinates of the first feature point and the pixel coordinates of the second feature point in each matching pair.
[0075] For example, a SIFT (Scale-Invariant Feature Transform) feature detector can be used to extract a first feature point and its feature vector from a first grayscale image, and to extract a second feature point and its feature vector from a second grayscale image.
[0076] For example, the FLANN (Fast Library for Approximate Nearest Neighbors) feature matcher can be used to perform nearest neighbor matching between multiple first feature points and multiple second feature points based on feature vectors, and reliable matching pairs can be selected by combining the distance ratio criterion.
[0077] For example, the homography matrix between two images can be estimated using the RANSAC (Random Sample Consensus) algorithm based on each matching pair.
[0078] In this embodiment, by extracting feature points and feature vectors and matching feature points, the coordinate transformation relationship can be determined based on the pixel coordinates of the feature points, thus ensuring the reliability of the coordinate transformation relationship.
[0079] In one exemplary embodiment, such as Figure 3 As shown, an image processing method is provided, including:
[0080] Step 302: Obtain the first image and the second image.
[0081] Step 304: Convert the first image to grayscale to obtain a first grayscale image, convert the second image to grayscale to obtain a second grayscale image, extract features from the first grayscale image to obtain feature vectors of multiple first feature points, and extract features from the second grayscale image to obtain feature vectors of multiple second feature points.
[0082] Step 306: Match multiple first feature points with multiple second feature points based on feature vectors to obtain multiple matching pairs. Determine the coordinate transformation relationship based on the pixel coordinates of the first feature point and the pixel coordinates of the second feature point in each matching pair.
[0083] Each matching pair contains a first feature point and a second feature point.
[0084] Step 308: Determine the similar region mask using coordinate transformation relationship, and obtain the third image from the first image using the similar region mask.
[0085] The similarity region mask is used to reflect the location of similar regions in the first image and the second image.
[0086] Step 310: Label the information in the second image to obtain a first label set, label the information in the third image to obtain a second label set, determine the intersection and union between the first and second label sets, and determine the first information similarity between the second and third images based on the ratio between the number of elements in the intersection and the number of elements in the union.
[0087] Step 312: Use a multimodal model to determine the second information similarity between the second image and the third image.
[0088] Step 314: Align the third image with the second image using coordinate transformation to obtain the aligned third image. Determine the structural similarity based on the features of the pixels in the aligned third image and the features of the pixels in the second image.
[0089] Step 316: Determine the content similarity between the first image and the second image based on structural similarity, first information similarity, and second information similarity.
[0090] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.
[0091] Based on the same inventive concept, this application also provides an image processing apparatus for implementing the image processing method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more image processing apparatus embodiments provided below can be found in the limitations of the image processing method described above, and will not be repeated here.
[0092] In one exemplary embodiment, such as Figure 4 As shown, an image processing apparatus is provided, including: a region detection module 402, a first similarity determination module 404, a second similarity determination module 406, and a third similarity determination module 408, wherein:
[0093] The region detection module 402 is used to acquire a first image and a second image, and to determine a region in the first image that is similar to the second image to obtain a third image.
[0094] The first similarity determination module 404 is used to determine the structural similarity between the second image and the third image. The structural similarity is used to reflect the similarity between the features presented by pixels in the second image and the features presented by pixels in the third image.
[0095] The second similarity determination module 406 is used to determine the information similarity between the second image and the third image. The information similarity is used to reflect the similarity between the information contained in the second image and the information contained in the third image.
[0096] The third similarity determination module 408 is used to determine the content similarity between the first image and the second image based on structural similarity and information similarity.
[0097] In some embodiments, information similarity includes a first information similarity and / or a second information similarity. The second similarity determination module 406 is further configured to annotate information in the second image to obtain a first annotation set, annotate information in the third image to obtain a second annotation set, determine the first information similarity between the second image and the third image based on the similarity between the first annotation set and the second annotation set, and / or determine the second information similarity between the second image and the third image using a multimodal model.
[0098] In some embodiments, the second similarity determination module 406 is further configured to determine the intersection and union between the first annotation set and the second annotation set; and to determine the first information similarity between the second image and the third image based on the ratio between the number of elements in the intersection and the number of elements in the union.
[0099] In some embodiments, the first similarity determination module 404 is further configured to determine the coordinate transformation relationship between the first image and the second image; align the third image to the second image using the coordinate transformation relationship to obtain the aligned third image; and determine the structural similarity based on the features of pixels in the aligned third image and the features of pixels in the second image.
[0100] In some embodiments, the region detection module 402 is further configured to determine the coordinate transformation relationship between the first image and the second image; determine a similar region mask using the coordinate transformation relationship, the similar region mask being used to reflect the position of similar regions in the first image and the second image; and obtain a third image from the first image using the similar region mask.
[0101] In some embodiments, the region detection module 402 is further configured to convert the first image to grayscale to obtain a first grayscale image, and convert the second image to grayscale to obtain a second grayscale image; perform feature extraction on the first grayscale image to obtain feature vectors of multiple first feature points, and perform feature extraction on the second grayscale image to obtain feature vectors of multiple second feature points; match the multiple first feature points with the multiple second feature points based on the feature vectors to obtain multiple matching pairs, each matching pair containing one first feature point and one second feature point; and determine the coordinate transformation relationship based on the pixel coordinates of the first feature point and the pixel coordinates of the second feature point in each matching pair.
[0102] Each module in the aforementioned image processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0103] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 5As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores at least a portion of the data involved in the image processing method. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements an image processing method.
[0104] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 6 As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements an image processing method. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.
[0105] Those skilled in the art will understand that Figure 5 and Figure 6The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0106] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0107] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0108] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0109] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0110] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0111] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. An image processing method, characterized in that, The method includes: Acquire a first image and a second image, and determine a region from the first image that is similar to the second image to obtain a third image; Determine the structural similarity between the second image and the third image, wherein the structural similarity is used to reflect the similarity between the features presented by pixels in the second image and the features presented by pixels in the third image; Determine the information similarity between the second image and the third image, wherein the information similarity is used to reflect the similarity between the information contained in the second image and the information contained in the third image; Based on the structural similarity and the information similarity, the content similarity between the first image and the second image is determined.
2. The method according to claim 1, characterized in that, The information similarity includes a first information similarity and / or a second information similarity, and determining the information similarity between the second image and the third image includes: The information in the second image is annotated to obtain a first annotation set, and the information in the third image is annotated to obtain a second annotation set; Based on the similarity between the first annotation set and the second annotation set, a first information similarity between the second image and the third image is determined; and / or, A second information similarity between the second image and the third image is determined using a multimodal model.
3. The method according to claim 2, characterized in that, The step of determining the first information similarity between the second image and the third image based on the similarity between the first annotation set and the second annotation set includes: Determine the intersection and union of the first annotation set and the second annotation set; The first information similarity between the second image and the third image is determined based on the ratio between the number of elements in the intersection and the number of elements in the union.
4. The method according to any one of claims 1 to 3, characterized in that, Determining the structural similarity between the second image and the third image includes: Determine the coordinate transformation relationship between the first image and the second image; The third image is aligned with the second image using the coordinate transformation relationship to obtain the aligned third image; The structural similarity is determined based on the features of pixels in the aligned third image and the features of pixels in the second image.
5. The method according to any one of claims 1 to 3, characterized in that, The step of determining a region in the first image that is similar to the second image to obtain a third image includes: Determine the coordinate transformation relationship between the first image and the second image; The coordinate transformation relationship is used to determine a similar region mask, which is used to reflect the position of similar regions in the first image and the second image; The third image is obtained from the first image using the similar region mask.
6. The method according to claim 5, characterized in that, Determining the coordinate transformation relationship between the first image and the second image includes: The first image is converted to grayscale to obtain a first grayscale image, and the second image is converted to grayscale to obtain a second grayscale image; Feature extraction is performed on the first grayscale image to obtain feature vectors of multiple first feature points, and feature extraction is performed on the second grayscale image to obtain feature vectors of multiple second feature points. Based on the feature vector, the plurality of first feature points are matched with the plurality of second feature points to obtain a plurality of matching pairs, each matching pair containing a first feature point and a second feature point; Based on the pixel coordinates of the first feature point and the pixel coordinates of the second feature point in each matching pair, the coordinate transformation relationship is determined.
7. An image processing apparatus, characterized in that, The device includes: A region detection module is used to acquire a first image and a second image, and to determine a region from the first image that is similar to the second image to obtain a third image; A first similarity determination module is used to determine the structural similarity between the second image and the third image, wherein the structural similarity is used to reflect the similarity between the features presented by pixels in the second image and the features presented by pixels in the third image; The second similarity determination module is used to determine the information similarity between the second image and the third image, wherein the information similarity is used to reflect the similarity between the information contained in the second image and the information contained in the third image; The third similarity determination module is used to determine the content similarity between the first image and the second image based on the structural similarity and the information similarity.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.