Image background similarity analysis method and device, equipment and medium

By segmenting the image background into objects and scene areas and constructing dynamic weight calculation similarity scores, the problem of inaccurate image background similarity analysis in the prior art is solved, and a more efficient and accurate image similarity evaluation is achieved.

CN120339653APending Publication Date: 2025-07-18WELAB INFORMATION TECH SHENZHEN LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510333740.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When analyzing the similarity of image backgrounds, relying solely on background features can easily lead to misjudgment, and it is impossible to accurately distinguish the differences between different objects or regions, resulting in inaccurate similarity scores.

Method used

By segmenting the image background into object areas and scene areas, a distribution vector is constructed and dynamic weights are assigned, the weighting coefficients and weights of each area are calculated, and the similarity score of the image background is calculated.

Benefits of technology

It improves the accuracy and reliability of image background similarity analysis, reduces noise interference, shortens processing time, ensures in-depth processing of important areas, and obtains more accurate similarity comparison results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339653A_ABST
    Figure CN120339653A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and discloses an image background similarity analysis method and device, equipment and a medium, and the method comprises the steps: extracting an image background of a first image obtained in advance, and segmenting the extracted image background into an object region and a scene region; constructing a distribution vector according to the object area of the image, and distributing an object area weight according to the distribution vector; calculating a weighting coefficient of a scene area in the image, and constructing a scene area weight by using the weighting coefficient of the scene area; and calculating a similarity score of the image background of the first image and the image background of the second image according to the first object area weight, the first scene area weight, the second object area weight and the second scene area weight. By calculating the weights of different areas, the similarity score between the two image backgrounds can be calculated more accurately, so that the similarity scores between the two image backgrounds can be compared more accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular, to an image background similarity analysis method, apparatus, device, and medium. Background Art

[0002] Image background similarity is a method for evaluating image similarity and plays an important role in applications such as image matching, object detection, image classification, and image retrieval.

[0003] Currently, analyzing image background similarity mainly involves extracting the background features of an image and calculating the similarity between the backgrounds of two images. However, in some scenarios, the backgrounds are often relatively similar. Therefore, only extracting the background for similarity analysis may lead to misjudgment, unable to distinguish the differences between different objects or regions, resulting in inaccurate similarity scores calculated. Summary of the Invention

[0004] The present invention provides an image background similarity analysis method, apparatus, device, and medium. By calculating the weights of different regions, the similarity score between the backgrounds of two images can be calculated more accurately, thereby more accurately comparing their similarity scores.

[0005] In a first aspect, an image background similarity analysis method is provided, including:

[0006] Extracting the image background of a pre-acquired first image and segmenting the extracted image background into an object region and a scene region;

[0007] Constructing a first distribution vector based on the object region of the first image and assigning a first object region weight according to the first distribution vector;

[0008] Calculating the weighted coefficient of the scene region in the first image and constructing a first scene region weight using the weighted coefficient of the first scene region;

[0009] Extracting the image background of a pre-acquired second image and segmenting the extracted image background into an object region and a scene region;

[0010] Constructing a second distribution vector based on the object region of the second image and assigning a second object region weight according to the second distribution vector;

[0011] Calculating the weighted coefficient of the scene region in the second image and constructing a second scene region weight using the weighted coefficient of the second scene region;

[0012] Calculate the similarity score of the image background of the first image and the image background of the second image according to the first object region weight, the first scene region weight, the second object region weight, and the second scene region weight.

[0013] In a second aspect, an apparatus for analyzing image background similarity is provided, including:

[0014] A first image processing module, configured to extract the image background of a pre-acquired first image, segment the extracted image background into an object region and a scene region, construct a first distribution vector according to the object region of the first image, assign a first object region weight according to the first distribution vector, calculate a weighting coefficient of the scene region in the first image, and construct a first scene region weight by using the weighting coefficient of the first scene region;

[0015] A second image processing module, configured to extract the image background of a pre-acquired second image, segment the extracted image background into an object region and a scene region, construct a second distribution vector according to the object region of the second image, assign a second object region weight according to the second distribution vector, calculate a weighting coefficient of the scene region in the second image, and construct a second scene region weight by using the weighting coefficient of the second scene region;

[0016] A similarity analysis module, configured to calculate the similarity score of the image background of the first image and the image background of the second image according to the first object region weight, the first scene region weight, the second object region weight, and the second scene region weight.

[0017] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned method for analyzing image background similarity are implemented.

[0018] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned method for analyzing image background similarity are implemented.

[0019] In the solutions implemented by the above-mentioned method, device, computer equipment and storage medium for analyzing the similarity of image backgrounds, through background segmentation, the pixels in the non-target areas of the image can be removed, reducing noise interference and making the signal-to-noise ratio of the target area higher; after extracting the object area and the scene area, object recognition and scene understanding can be carried out more accurately, thereby improving the accuracy and reliability of the recognition and understanding results and providing a more accurate basis for subsequent image analysis and processing; by calculating the dynamic weight of the object area according to the distribution vector, the importance of the object area in the image can be more reasonably reflected, thus better guiding subsequent image processing tasks; at the same time, by dynamically assigning weights to the object area and different parts of the scene area, the processing time can be shortened; at the same time, it is also possible to ensure that more in-depth and detailed processing is carried out on more important areas, so that the processing results are more accurate and of high quality; by calculating the weights of different areas, the similarity score between two segmented images can be calculated more accurately, which can well help us judge the difference between two images and thus more accurately compare their similarity scores. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0021] Figure 1 is a schematic diagram of an application environment of a method for analyzing the similarity of image backgrounds in an embodiment of the present invention;

[0022] Figure 2 is a schematic flowchart of a method for analyzing the similarity of image backgrounds in an embodiment of the present invention;

[0023] Figure 3 is a schematic structural diagram of a device for analyzing the similarity of image backgrounds in an embodiment of the present invention;

[0024] Figure 4 is a schematic structural diagram of a computer device in an embodiment of the present invention;

[0025] Figure 5 is another schematic structural diagram of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0027] An image background similarity analysis method provided by an embodiment of the present invention can be applied, for example, in Figure 1In the application environment, the client communicates with the server through the network. The server can extract the image background of the pre-acquired first image, and divide the extracted image background into an object area and a scene area; construct a first distribution vector according to the object area of the first image, and allocate a first object area weight according to the first distribution vector; calculate the weighting coefficient of the scene area in the first image, and construct a first scene area weight by using the weighting coefficient of the first scene area; extract the image background of the pre-acquired second image, and divide the extracted image background into an object area and a scene area; construct a second distribution vector according to the object area of the second image, and allocate a second object area weight according to the second distribution vector; calculate the weighting coefficient of the scene area in the second image, and construct a second scene area weight by using the weighting coefficient of the second scene area; calculate the similarity score of the image background of the first image and the image background of the second image according to the first object area weight, the first scene area weight, the second object area weight, and the second scene area weight, and feedback the similarity score back to the client. The present invention provides an image background similarity analysis device. For the similarity analysis service, the image background of the pre-acquired first image is extracted, and the extracted image background is divided into an object area and a scene area; a first distribution vector is constructed according to the object area of the first image, and a first object area weight is allocated according to the first distribution vector; the weighting coefficient of the scene area in the first image is calculated, and a first scene area weight is constructed by using the weighting coefficient of the first scene area; the image background of the pre-acquired second image is extracted, and the extracted image background is divided into an object area and a scene area; a second distribution vector is constructed according to the object area of the second image, and a second object area weight is allocated according to the second distribution vector; the weighting coefficient of the scene area in the second image is calculated, and a second scene area weight is constructed by using the weighting coefficient of the second scene area; the similarity score of the image background of the first image and the image background of the second image is calculated according to the first object area weight, the first scene area weight, the second object area weight, and the second scene area weight. By calculating the weights of different areas, the similarity score between the two image backgrounds can be calculated more accurately, so as to compare the similarity scores between them more accurately. Among them, the client can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.

[0028] Please refer to Figure 2 as shown in Figure 2 which is a schematic flowchart of an image background similarity analysis method provided by an embodiment of the present invention, and includes the following steps:

[0029] S1. Extract the image background of the pre-acquired first image, and segment the extracted image background into an object region and a scene region.

[0030] In the embodiments of the present invention, the extraction refers to separately segmenting the background part in the image to form a new background image or background mask, and the segmentation refers to respectively extracting the pixel points included in the object region and the pixel points included in the scene region according to the image background of the first image.

[0031] Specifically, in the preprocessing process, the background of the first image can be segmented by threshold segmentation or image segmentation, and the background segmentation of the first image obtained can be used as the input for image similarity calculation; when extracting the object region and the scene region, it depends on the specific application scenario and requirements, select appropriate segmentation algorithms and parameters to achieve the optimal segmentation effect, and perform reasonable post-processing operations according to the specific target and background features to improve the accuracy and reliability of image processing and analysis.

[0032] In the embodiments of the present invention, the extraction of the image background of the pre-acquired first image includes:

[0033] Perform enhancement processing on the pre-acquired first image to obtain a first enhanced image;

[0034] Extract the background pixels of the first enhanced image;

[0035] Perform background segmentation on the first enhanced image according to the background pixels to obtain the image background of the first image.

[0036] In the embodiments of the present invention, the enhancement processing refers to performing a series of processes on the image to improve the quality of the image, increase the performance in aspects such as contrast, sharpness, and brightness of the image, the extraction refers to extracting the pixels belonging to the background part from the image, and the background segmentation refers to separately segmenting the background part in the image to form a new background image or background mask.

[0037] Specifically, perform processing such as noise removal and contrast enhancement on the first image. Noise removal can use filters such as median filtering and Gaussian filtering; contrast enhancement can use methods such as histogram stretching and histogram equalization; further enhance the contour and texture of the image by highlighting edges and details. A commonly used method is Laplacian operator sharpening, and Gaussian filters and Laplacian filters can be used for image sharpening.

[0038] Further, according to the features of the enhanced image, background pixels can be selected. It can be done manually or automatically. If automatic selection is used, methods such as threshold segmentation or clustering can be employed. Based on the extraction of background pixels, a fixed threshold method can be used to segment the image. For each pixel, if it belongs to the background, it is set to black; otherwise, it is set to white. This method is simple and easy to understand, but a suitable threshold needs to be set manually. Too high or too low a threshold will result in poor segmentation effects.

[0039] In addition, the Gaussian Mixture Model (GMM) can be used to model the distribution of background pixels, and pixels below a certain random segmentation point in the distribution are determined as background pixels. This method can adapt to the complex distribution of background pixels and can automatically estimate the threshold, but the segmentation effect may be unstable in complex image scenarios.

[0040] In the embodiments of the present invention, the segmentation of the extracted image background into an object region and a scene region includes:

[0041] Performing object mask segmentation on the image background of the first image to obtain the object region of the first image;

[0042] Performing scene analysis on the image background of the first image to obtain an analysis result;

[0043] Performing scene mask segmentation on the image background of the first image according to the analysis result to obtain the scene region of the first image.

[0044] In the embodiments of the present invention, the object mask segmentation refers to precise pixel-level segmentation of the object region. The scene analysis refers to semantic analysis and reasoning of the scene based on the category information of the object and the context information of the scene to identify elements such as themes, activities, and environments existing in the scene and make a judgment on the scene. The scene mask segmentation refers to precise pixel-level segmentation of the scene region.

[0045] Specifically, using an object mask generation algorithm, by performing operations such as iterative segmentation or region growing on the foreground object, the first image is segmented into an object and a background. The object mask is usually a binary image, where the pixel value of the object region is 1 and the pixel value of the background region is 0. For the generated object mask, algorithms such as morphological processing and connected component analysis can be used for further processing and analysis. For example, a morphological processing algorithm can be used to perform morphological transformation on the object mask to eliminate noise and small objects. In addition, a connected component analysis algorithm can be used to filter or merge isolated regions or small regions in the object mask. After obtaining the object mask, post-processing is also required to further optimize the boundary and shape of the object region to obtain the object region.

[0046] Specifically, the object region in the first image is recognized to obtain the category information of the object. Based on the category information of the object and the context information of the scene, semantic analysis and reasoning of the scene are performed to identify elements such as the theme, activity, and environment existing in the scene, understand and interpret the scene, and infer which category of scene in the real world the object in the scene belongs to, such as indoor, outdoor, urban street, natural scenery, etc. According to the above analysis results, the image background of the first image is segmented by a scene mask to obtain the scene region of the first image.

[0047] In the embodiment of the present invention, for the image after enhancement processing, noise and background features are better separated, which is conducive to using a more accurate method to extract background pixels, thereby effectively reducing the interference of segmentation noise. At the same time, through background segmentation, the target to be sought can be better separated from the background, eliminating interference information and improving the accuracy of target segmentation. The image after enhancement and segmentation processing is the image background of the first image, which can be directly used during image processing, saving the time and computational effort for extracting background pixels and segmentation again, and facilitating subsequent image processing.

[0048] In the embodiment of the present invention, through background segmentation, the pixels in the non-target region of the image can be removed, reducing noise interference and making the signal-to-noise ratio of the target region higher. After extracting the object region and the scene region, object recognition and scene understanding can be performed more accurately, thereby improving the accuracy and reliability of the recognition and understanding results, and providing a more accurate basis for subsequent image analysis and processing.

[0049] S2. Construct a first distribution vector according to the object region of the first image, and assign a first object region weight according to the first distribution vector.

[0050] In the embodiment of the present invention, the construction refers to grouping the object regions detected in the first image according to their respective object categories, counting the frequency of each category appearing in the image, and constructing these frequencies into a vector. The assignment refers to the process of dynamically balancing and adjusting the importance of object regions by assigning different weight values to different object regions in the image according to the weights of the object regions in the first distribution vector.

[0051] Specifically, the process of constructing the first distribution vector has a very wide range of applications. For example, it can be used in fields such as object detection, image processing, and computer vision. By calculating and analyzing the frequency of object appearance, the distribution of objects can be better understood, which is conducive to subsequent object recognition and processing.

[0052] Furthermore, according to the frequency or weight of the object regions, the weight of the object regions can be evenly distributed according to their importance. If a certain group of object categories appears frequently, it indicates higher saliency in the entire image, and thus a higher weight should be given.

[0053] In an embodiment of the present invention, constructing the first distribution vector according to the object regions of the first image includes:

[0054] Identifying the object categories of the object regions of the first image;

[0055] Grouping the same categories according to the object categories to obtain the corresponding object categories for each group;

[0056] Performing frequency statistics on the corresponding object categories of each group to obtain the frequency of each group of object categories;

[0057] Constructing a frequency distribution vector for the object regions of the first image according to the frequency of each group of object categories to obtain the first distribution vector.

[0058] In an embodiment of the present invention, the identification refers to performing target identification on all object regions in the image to determine their respective object categories. The same-category grouping refers to grouping all the objects detected in the image according to their categories, and each group corresponds to a different object category. The frequency statistics refer to counting the frequency of each category group in the image. The construction refers to constructing a frequency distribution vector from the corresponding frequency sequence for each category group.

[0059] Specifically, first, detect all the objects from the object regions in the first image and mark them; this step can be completed using technologies such as deep learning-based object detection models and traditional object detection algorithms. Then, group the detected objects according to their categories, and each group corresponds to a different object category. Data structures such as dictionaries and hash tables can be used to achieve the grouping of categories. For each category group, count the frequency of its appearance in the background image. Specifically, the entire image can be traversed. For each pixel, determine whether it belongs to a certain object category group, and then increment the frequency of the corresponding category in that group by 1. For each category group, construct a frequency distribution vector from its corresponding frequency sequence; the dimension size of this frequency distribution vector should be the same as the number of object categories, each dimension corresponding to a different object category, and its value corresponding to the number of times that category appears in the frequency distribution.

[0060] In addition, the object area can be divided into several grid areas according to preset rules. For each grid area, all pixels inside it are traversed. For each pixel, its color information is extracted and counted, that is, the occurrence times of each color are accumulated. After counting the occurrence times of all pixels' colors, the color frequency of this grid area is obtained. The color frequency can be put into a vector. For each element in the vector, it represents the number of times this color appears in this grid area. After counting the color frequencies of all grid areas, the color frequency distribution of the entire object area can be obtained, and this distribution can be used for the description and extraction of the object color feature.

[0061] Further, for each grid area, the color frequency vector can be normalized into a color distribution feature vector by the following method: divide the color frequency vector by the sum of all elements so that the sum of all elements is 1; divide each element of the color distribution vector by the square root of the sum of the squares of all elements of this vector. The color distribution feature vectors obtained for each grid area are combined into the color distribution feature vector of the grid area. For example, the feature vectors of all grid areas are concatenated together to obtain a total color distribution feature vector.

[0062] In addition, a prepared training dataset can be used to train a deep learning model, such as a convolutional neural network (CNN), for extracting the depth feature vector of each object. During the training process, loss functions such as cross-entropy can be used for optimization, and the weights are updated iteratively through the backpropagation algorithm. Using the trained deep learning model, the depth of the object in the image is extracted to obtain the depth feature vector of this object. Using the above method, the depth feature vector of the object in the image is extracted and the frequency distribution vector is constructed.

[0063] Specifically, for the input image background object, the features of each pixel in the picture are extracted through a convolutional neural network (CNN) and other means to obtain a feature map. Using the Self-Attention mechanism, by constructing Query, Key, and Value vectors, the features of each pixel in the picture are encoded. In this process, the features of each pixel are regarded as Key and Value, so as to help the network better remember and capture the picture features, while the Query vector is constructed as the category vector for each object category. Based on the above constructed Query, Key, and Value vectors, the similarity distribution between each pixel in the picture can be obtained by performing corresponding operations, so as to calculate the importance weights that each pixel needs to pay attention to.

[0064] In the embodiment of the present invention, the allocating the first object area weight according to the first distribution vector includes:

[0065] Extract features from the object region of the first image to obtain a first feature image;

[0066] Construct a dynamic weight vector corresponding to the object region of the first image according to the first distribution vector;

[0067] Perform weighted processing on the first feature image and the dynamic weight vector to obtain a first object region weight.

[0068] In the embodiments of the present invention, the feature extraction refers to extracting representative features from the object region for subsequent object recognition, classification, positioning and other tasks. The construction refers to calculating the frequency distribution vector and then converting it into a dynamic weight vector. The weighted processing refers to the process of multiplying each pixel of the constructed dynamic weight vector and the first feature image.

[0069] Specifically, extract the features of the input image. For example, use a convolutional neural network to extract features from the image to obtain a first feature image. Subsequently, construct a corresponding dynamic weight vector according to each detected object category. This vector usually consists of two parts: one is to assign higher weights to the pixels near the position of the object of the corresponding category, and the other is to assign lower weights to the pixels farther away from the object. The specific form of the dynamic weight vector can be adjusted according to the needs of the task.

[0070] In detail, multiply each pixel of the constructed dynamic weight vector and the feature map to obtain a weighted synthesized feature map, that is, the processing result. In this process, the object regions that are focused on will be assigned higher weights, while other regions will have their weights reduced, so as to retain the most important object information and reduce interference. The dynamic allocation of the dynamic weights of each object category utilizes the object position and category information. By performing the process of dynamic allocation, feature extraction and weighted processing are performed on the object information, so as to obtain a more accurate and effective classification result.

[0071] In the embodiments of the present invention, the dynamic weight vector is calculated according to the first distribution vector, which can reflect the importance degree of each object region in the image. By combining the dynamic weight vector with other image processing algorithms, more intelligent and adaptive image processing can be realized. At the same time, it is convenient to construct different dynamic weight vectors according to different tasks or application scenarios to realize a more flexible image processing process. By performing weighted processing on the first feature image and the dynamic weight vector, the weight of each object region can be obtained more accurately, so as to better reflect the importance of the object region.

[0072] In the embodiments of the present invention, after extracting the object region and the scene region, object recognition and scene understanding can be performed more accurately, thereby improving the accuracy and reliability of the recognition and understanding results, and providing a more accurate basis for subsequent image analysis and processing; by constructing a frequency distribution vector, the frequency of each object appearing in the image can be understood, so as to better understand the distribution of objects. At the same time, the first distribution vector can represent the color distribution of the object region, and this representation method is more intuitive, facilitating comparison and contrast analysis between different object regions; the first distribution vector can also be used to improve the performance of machine learning algorithms.

[0073] In the embodiments of the present invention, by calculating the dynamic weight of the object region according to the first distribution vector, the importance of the object region in the image can be more reasonably reflected, thereby better guiding subsequent image processing tasks; according to the first distribution vector, the importance degree of different object regions in the image can be obtained, so as to provide appropriate processing strategies for different image processing tasks.

[0074] S3. Calculate the weighting coefficient of the scene region in the first image, and construct the weight of the first scene region by using the weighting coefficient of the first scene region.

[0075] In the embodiments of the present invention, the calculation refers to calculating a weighting coefficient for each sub-region in all sub-regions of the entire scene region according to the characteristics of each sub-region and the determination result of the scene category to which the sub-region belongs, to represent the contribution degree of the sub-region to the entire scene, and the construction refers to constructing the processing weight of each sub-region according to the weighting coefficient of each sub-region and the current state of the entire scene region.

[0076] Specifically, for different scenes and target objects, different information features of the target region can be obtained through an image analysis algorithm, and then combined with the weighting coefficient, weights are assigned to each scene region.

[0077] In the embodiments of the present invention, the calculation of the weighting coefficient of the scene region in the first image includes:

[0078] Divide the scene region of the first image into multiple groups of partial scene regions;

[0079] Perform feature analysis on each of the partial scene regions one by one to obtain partial scene features;

[0080] Identify the partial scene types according to the partial scene features;

[0081] Calculate the weighting coefficients of the partial scene regions one by one according to the partial scene types to obtain the partial scene weighting coefficients corresponding to the partial scene regions;

[0082] Generate the weighting factor of the scene area in the first image according to the combination of the partial scene weighting factors.

[0083] In the embodiments of the present invention, the division refers to dividing the entire scene area into multiple sub-areas, the feature analysis refers to analyzing different parts of the scene to understand the feature information such as color, texture, shape, etc. of each part, the recognition refers to identifying the scene type to which the sub-area belongs by performing feature extraction and processing on each sub-area, and the weighting factor calculation refers to calculating a weighting factor for each sub-area in all sub-areas of the entire scene area according to the features of each sub-area and the determination result of the scene category to which the sub-area belongs to represent the contribution degree of the sub-area to the entire scene.

[0084] Specifically, divide the entire scene area into several areas with obvious boundaries. For example, in a picture, the part belonging to the sky, the part of the building, the part of the vegetation, the part of the road, etc. can be divided into different sub-areas respectively, and then extract the features of each sub-area, such as color distribution, texture features, object layout, etc., and match these features with predefined scene categories to determine the scene type to which each sub-area belongs. Taking an indoor picture as an example, it can be divided into several sub-areas such as the area containing desks and chairs, the area of the window and the interior, etc., and then extract their respective features for these sub-areas, and determine the scene types to which they belong according to these features, such as office, bedroom, living room, etc.

[0085] Furthermore, perform the weighting factor calculation in the weighting factor calculation for each of the partial scene areas one by one. When calculating the weighting factor, different algorithms can be used, such as the weighting factor calculation based on feature similarity, the weighting factor calculation based on pixel depth, etc.

[0086] In the embodiments of the present invention, the construction of the first scene area weight by using the weighting factor of the first scene area includes:

[0087] Perform dynamic weight allocation on the partial scene area according to the partial scene weighting factor to obtain the partial scene dynamic weight;

[0088] Generate the first scene area weight according to the combination of the partial scene dynamic weights.

[0089] In the embodiments of the present invention, the dynamic weight allocation refers to performing weighting processing on different parts of the scene area according to the image analysis result and the distribution vector to obtain the corresponding weight value of each scene area.

[0090] Specifically, the obtained weighting coefficients can be used to determine the importance of each sub-region in the entire scene, and each sub-region can be weighted according to these coefficients. For example, for a sub-region with a larger weighting coefficient, the processing weight should also be larger to pay more attention to these important regions and perform more in-depth processing and improvement. On the contrary, for a sub-region with a smaller weighting coefficient, the processing weight can be correspondingly reduced to reduce the processing of these sub-regions so as not to affect the processing efficiency and results. Finally, dynamic weight allocation can be implemented by various algorithms according to the actual situation, such as dynamic weight allocation based on weighting coefficients, dynamic weight allocation based on states, etc.

[0091] In the embodiments of the present invention, after the entire scene area is divided into multiple sub-regions, the feature analysis of each sub-region is more detailed, and the scene category to which each sub-region belongs can be determined more accurately, thereby improving the accuracy of the entire scene classification; the recognition of some scene features realizes the local analysis of the scene area, reduces the processing scale of the entire image, improves the efficiency and speed of image processing, and makes image analysis more efficient and fast; by calculating the weighting coefficients for different parts of the scene area and performing dynamic weight allocation for some scene areas according to this, the processing time can be shortened; calculating the weighting coefficients for different parts of the scene and performing dynamic weight allocation for some scene areas according to this can ensure more in-depth and detailed processing of the more important regions, thereby making the processing results more accurate and of high quality.

[0092] In the embodiments of the present invention, dynamic weight allocation can make the processing time more efficient. When performing image processing, only those parts of the scene with higher weights need to be processed, while those parts of the scene with lower weights do not need to be processed too much, which can improve the processing efficiency; at the same time, it can enhance the flexibility of image processing and make the algorithm more adaptable to different scenes and applications.

[0093] S4. Extract the image background of the pre-acquired second image, and segment the extracted image background into an object region and a scene region.

[0094] In the embodiments of the present invention, extracting the image background of the pre-acquired second image and segmenting the extracted image background into an object region and a scene region are the same as the steps of S1 above, and will not be elaborated here.

[0095] Specifically, the second image is segmented by a threshold segmentation method or an image segmentation method; when extracting the object region and the scene region of the image background of the second image, it depends on the specific application scenario and requirements, and a suitable segmentation algorithm and parameters are selected to achieve the optimal segmentation effect, and reasonable post-processing operations are performed according to the specific target and background features.

[0096] S5. Construct a second distribution vector based on the object regions of the second image, and assign second object region weights according to the second distribution vector.

[0097] In the embodiments of the present invention, constructing a second distribution vector based on the object regions of the second image and assigning second object region weights according to the second distribution vector are the same as the steps of S2 above, and will not be elaborated here.

[0098] Specifically, the process of constructing the second distribution vector can adopt object detection, image processing, and computer vision. By calculating and analyzing the frequency of object appearance, the distribution of objects can be better understood. A vector is constructed based on the frequency of object region appearance, and then weights are assigned to the object regions according to the vector. If a certain group of object categories appears frequently, it indicates that its significance in the entire image is higher, and then higher weights should be given.

[0099] S6. Calculate the weighting coefficient of the scene region in the second image, and construct a second scene region weight using the weighting coefficient of the second scene region.

[0100] In the embodiments of the present invention, calculating the weighting coefficient of the scene region in the second image and constructing a second scene region weight using the weighting coefficient of the second scene region are the same as the steps of S3 above, and will not be elaborated here.

[0101] Specifically, for different scenes and target objects in the scene region of the second image, different information features of the target region can be obtained through image analysis algorithms, and then combined with the weighting coefficient to assign weights to each scene region.

[0102] S7. Calculate the similarity score between the image background of the first image and the image background of the second image according to the first object region weight, the first scene region weight, the second object region weight, and the second scene region weight.

[0103] In the embodiments of the present invention, the calculation refers to extracting and calculating the image background regions of the two images, and calculating the similarity score between the two image backgrounds according to the weight values of different regions.

[0104] Specifically, calculating the similarity score between the image background of the first image and the image background of the second image according to the first object region weight, the first scene region weight, the second object region weight, and the second scene region weight includes region division, similarity calculation, weight assignment, and weighting, which can more accurately evaluate the similarity between different images and the contributions of different scenes and objects to the image similarity.

[0105] In an embodiment of the present invention, calculating the similarity score of the image backgrounds of the first image and the second image based on the first object region weight, the first scene region weight, the second object region weight, and the second scene region weight includes:

[0106] Calculating the similarity score of the object regions of the first image and the second image according to the first object region weight and the second object region weight;

[0107] Calculating the similarity score of the scene regions of the first image and the second image according to the first scene region weight and the second scene region weight;

[0108] Calculating the initial similarity score of the image backgrounds of the first image and the second image;

[0109] Performing weighted fusion on the object region similarity score, the scene region similarity score, and the initial similarity score to obtain the similarity score of the image backgrounds of the first image and the second image.

[0110] In an embodiment of the present invention, the calculation refers to comparing the object regions and scene regions in two segmented images and the similarity scores between the two segmented images, and the weighted fusion refers to performing weighted summation on different similarity score metrics to obtain a comprehensive similarity score.

[0111] Specifically, for the object regions, scene regions, and the image backgrounds of two images, various image similarity algorithms (such as SSIM, PSNR, etc.) can be used to calculate the similarity scores. The weight values are calculated according to the first object region weight, the first scene region weight, the second object region weight, and the second scene region weight, and the scores of all regions are weighted to obtain the similarity score of the entire image.

[0112] In addition, a multi-task learning framework can be used to optimize the weight factors of the scene region and object region features. A commonly used optimization method is to gradually update the weights through the task loss function. In the initial stage, the weight factors can be initialized in a uniform distribution manner; for example, the weight factors can be initialized as random values with a uniform distribution, and the weights are continuously adjusted during the subsequent training process to improve the performance of the model. During the training process, the task loss function can be used to update the weights, and commonly used loss functions include mean square error (MSE), cross-entropy (Cross-Entropy), etc.

[0113] Specifically, first define the objective function, then use the objective function to calculate the gradient of the model, calculate with the weight factor as the independent variable, and use the calculated gradient to update the weight factor. Generally, optimization algorithms such as Stochastic Gradient Descent (SGD) or Newton's method can be used to update the weights. Repeat the first three steps until the model converges, and finally obtain the similarity score between the two images.

[0114] In the embodiments of the present invention, the similarity scores of the scene region and the object region in the image background can be used as supplementary information to help the model calculate the similarity of the images more accurately and improve the accuracy of the calculation results. At the same time, by calculating the similarity scores of the object regions, the similarities and differences between the objects can be better captured, thereby improving the effect of object tracking. By calculating the similarity scores of the scene regions, the robustness of the algorithm can be increased, making it more adaptable to image processing tasks in different scenarios.

[0115] In the embodiments of the present invention, by calculating the weights of different regions, the similarity scores between two segmented images can be calculated more accurately, which can well help us judge the differences between the two image backgrounds, and thus compare the similarity scores between them more accurately. At the same time, the similarity scores between the two image backgrounds can be presented intuitively, which can well help us compare the similarities of the two image backgrounds, thereby improving the quality of image matching.

[0116] It can be seen that in the above solution, for the similarity analysis service, the image background of the first image obtained in advance is extracted, and the extracted image background is segmented into an object region and a scene region; a first distribution vector is constructed according to the object region of the first image, and a first object region weight is assigned according to the first distribution vector; the weighted coefficient of the scene region in the first image is calculated, and a first scene region weight is constructed using the weighted coefficient of the first scene region; the image background of the second image obtained in advance is extracted, and the extracted image background is segmented into an object region and a scene region; a second distribution vector is constructed according to the object region of the second image, and a second object region weight is assigned according to the second distribution vector; the weighted coefficient of the scene region in the second image is calculated, and a second scene region weight is constructed using the weighted coefficient of the second scene region; the similarity score of the image background of the first image and the image background of the second image is calculated according to the first object region weight, the first scene region weight, the second object region weight, and the second scene region weight. By calculating the weights of different regions, the similarity scores between two image backgrounds can be calculated more accurately, and thus the similarity scores between them can be compared more accurately.

[0117] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0118] In one embodiment, an image background similarity analysis device is provided. This image background similarity analysis device corresponds one-to-one with the image background similarity analysis method in the above embodiment. As Figure 3 shown, this image background similarity analysis device includes a first image processing module 101, a second image processing module 102, and a similarity analysis module 103. The detailed description of each functional module is as follows:

[0119] The first image processing module 101 is used to extract the image background of the first image obtained in advance, segment the extracted image background into an object area and a scene area, construct a first distribution vector according to the object area of the first image, allocate a first object area weight according to the first distribution vector, calculate the weighting coefficient of the scene area in the first image, and construct a first scene area weight using the weighting coefficient of the first scene area;

[0120] The second image processing module 102 is used to extract the image background of the second image obtained in advance, segment the extracted image background into an object area and a scene area, construct a second distribution vector according to the object area of the second image, allocate a second object area weight according to the second distribution vector, calculate the weighting coefficient of the scene area in the second image, and construct a second scene area weight using the weighting coefficient of the second scene area;

[0121] The similarity analysis module 103 is used to calculate the similarity score between the image background of the first image and the image background of the second image according to the first object area weight, the first scene area weight, the second object area weight, and the second scene area weight.

[0122] In one embodiment, the first image processing module 101:

[0123] When used to extract the image background of the first image obtained in advance, it is used to:

[0124] Perform enhancement processing on the first image obtained in advance to obtain a first enhanced image;

[0125] Extract the background pixels of the first enhanced image;

[0126] Perform background segmentation on the first enhanced image according to the background pixels to obtain the image background of the first image;

[0127] When used to segment the extracted image background into an object region and a scene region, it is used for:

[0128] Perform object mask segmentation on the image background of the first image to obtain the object region of the first image;

[0129] Perform scene analysis on the image background of the first image to obtain an analysis result;

[0130] Perform scene mask segmentation on the image background of the first image according to the analysis result to obtain the scene region of the first image;

[0131] When used to construct a first distribution vector according to the object region of the first image, it is used for:

[0132] Identify the object categories of the object region of the first image;

[0133] Perform grouping of the same kind according to the object categories to obtain the corresponding object categories for each group;

[0134] Perform frequency statistics on the corresponding object categories of each group to obtain the frequency of each group of object categories;

[0135] Construct a frequency distribution vector for the object region of the first image according to the frequency of each group of object categories to obtain a first distribution vector;

[0136] When used to assign a first object region weight according to the first distribution vector, it is used for:

[0137] Perform feature extraction on the object region of the first image to obtain a first feature image;

[0138] Construct a dynamic weight vector corresponding to the object region of the first image according to the first distribution vector;

[0139] Perform weighted processing on the first feature image and the dynamic weight vector to obtain a first object region weight;

[0140] When used to calculate the weighted coefficient of the scene region in the first image, it is used for:

[0141] Divide the scene region of the first image into multiple groups of partial scene regions;

[0142] Perform feature analysis on each of the partial scene regions one by one to obtain partial scene features;

[0143] Identify partial scene types according to the partial scene features;

[0144] Calculate the weighted coefficient for each of the partial scene regions one by one according to the partial scene types to obtain the partial scene weighted coefficients corresponding to the partial scene regions;

[0145] Generate the weighting coefficient of the scene area in the first image according to the combination of the partial scene weighting coefficients.

[0146] In one embodiment, the steps of the second image processing module 102 are the same as those of the first image processing module 101 described above, and will not be elaborated here.

[0147] In one embodiment, when calculating the similarity score of the image background of the first image and the image background of the second image according to the first object area weight, the first scene area weight, the second object area weight, and the second scene area weight, the similarity analysis module 103 is configured to:

[0148] Calculate the similarity score of the object area of the first image and the object area of the second image according to the first object area weight and the second object area weight;

[0149] Calculate the similarity score of the scene area of the first image and the scene area of the second image according to the first scene area weight and the second scene area weight;

[0150] Calculate the initial similarity score of the image background of the first image and the image background of the second image;

[0151] Perform weighted fusion on the object area similarity score, the scene area similarity score, and the initial similarity score to obtain the similarity score of the image background of the first image and the image background of the second image.

[0152] The present invention provides an image background similarity analysis device. For the service of analyzing similarity, the image background of a pre-acquired first image is extracted, and the extracted image background is segmented into an object area and a scene area; a first distribution vector is constructed according to the object area of the first image, and a first object area weight is assigned according to the first distribution vector; a weighted coefficient of the scene area in the first image is calculated, and a first scene area weight is constructed by using the weighted coefficient of the first scene area; the image background of a pre-acquired second image is extracted, and the extracted image background is segmented into an object area and a scene area; a second distribution vector is constructed according to the object area of the second image, and a second object area weight is assigned according to the second distribution vector; a weighted coefficient of the scene area in the second image is calculated, and a second scene area weight is constructed by using the weighted coefficient of the second scene area; a similarity score of the image background of the first image and the image background of the second image is calculated according to the first object area weight, the first scene area weight, the second object area weight, and the second scene area weight. By calculating the weights of different areas, the similarity score between the two image backgrounds can be calculated more accurately, so as to compare the similarity scores between them more accurately.

[0153] For the specific limitations on an image background similarity analysis device, reference can be made to the limitations on an image background similarity analysis method in the foregoing text, which will not be elaborated here. Each module in the above image background similarity analysis device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in the form of hardware or be independent of it, or can be stored in the memory in the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the above modules.

[0154] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 4 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of an image background similarity analysis method.

[0155] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be as Figure 5As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of an image background similarity analysis method.

[0156] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:

[0157] Extract the image background of a pre-acquired first image, and segment the extracted image background into an object region and a scene region;

[0158] Construct a first distribution vector according to the object region of the first image, and assign a first object region weight according to the first distribution vector;

[0159] Calculate the weighting coefficient of the scene region in the first image, and construct a first scene region weight using the weighting coefficient of the first scene region;

[0160] Extract the image background of a pre-acquired second image, and segment the extracted image background into an object region and a scene region;

[0161] Construct a second distribution vector according to the object region of the second image, and assign a second object region weight according to the second distribution vector;

[0162] Calculate the weighting coefficient of the scene region in the second image, and construct a second scene region weight using the weighting coefficient of the second scene region;

[0163] Calculate the similarity score between the image background of the first image and the image background of the second image according to the first object region weight, the first scene region weight, the second object region weight, and the second scene region weight.

[0164] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the following steps are implemented:

[0165] Extract the image background of a pre-acquired first image, and segment the extracted image background into an object region and a scene region;

[0166] Construct a first distribution vector based on the object region of the first image, and assign a first object region weight according to the first distribution vector;

[0167] Calculate the weighting coefficient of the scene region in the first image, and construct a first scene region weight using the weighting coefficient of the first scene region;

[0168] Extract the image background of the pre-acquired second image, and segment the extracted image background into an object region and a scene region;

[0169] Construct a second distribution vector based on the object region of the second image, and assign a second object region weight according to the second distribution vector;

[0170] Calculate the weighting coefficient of the scene region in the second image, and construct a second scene region weight using the weighting coefficient of the second scene region;

[0171] Calculate the similarity score between the image background of the first image and the image background of the second image according to the first object region weight, the first scene region weight, the second object region weight, and the second scene region weight.

[0172] It should be noted that for the functions or steps that can be realized by the above computer-readable storage medium or computer device, reference can be made to the relevant descriptions on the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described in detail here.

[0173] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0174] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0175] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. If non-company software tools or components appear in the embodiments of the application, they are only used for example introduction and do not represent actual use. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or equivalently replace some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention and should all be included in the protection scope of the present invention.

Claims

1. An image background similarity analysis method, characterized in that, Including: Extract the image background of a pre-acquired first image, and segment the extracted image background into an object region and a scene region; Construct a first distribution vector based on the object region of the first image, and assign a first object region weight according to the first distribution vector; Calculate the weighting coefficient of the scene region in the first image, and construct a first scene region weight using the weighting coefficient of the first scene region; Extract the image background of a pre-acquired second image, and segment the extracted image background into an object region and a scene region; Construct a second distribution vector based on the object region of the second image, and assign a second object region weight according to the second distribution vector; Calculate the weighting coefficient of the scene region in the second image, and construct a second scene region weight using the weighting coefficient of the second scene region; Calculate the similarity score between the image background of the first image and the image background of the second image according to the first object region weight, the first scene region weight, the second object region weight, and the second scene region weight.

2. The method for analyzing the similarity of image backgrounds according to claim 1, characterized in that The extracting the image background of a pre-acquired first image includes: Perform enhancement processing on the pre-acquired first image to obtain a first enhanced image; Extract the background pixels of the first enhanced image; Perform background segmentation on the first enhanced image according to the background pixels to obtain the image background of the first image.

3. The image background similarity analysis method according to claim 1, wherein The segmenting the extracted image background into an object region and a scene region includes: Perform object mask segmentation on the image background of the first image to obtain the object region of the first image; Perform scene analysis on the image background of the first image to obtain an analysis result; Perform scene mask segmentation on the image background of the first image according to the analysis result to obtain the scene region of the first image.

4. The method for analyzing the similarity of an image background according to claim 1, wherein The constructing a first distribution vector based on the object region of the first image includes: Identify the object categories of the object region of the first image; Perform grouping of the same category according to the object categories to obtain the corresponding object categories for each group; Perform frequency statistics on the corresponding object categories of each group to obtain the frequency of each group of object categories; Construct a frequency distribution vector for the object region of the first image according to the frequency of each group of object categories to obtain a first distribution vector.

5. The method for analyzing the similarity of image backgrounds according to claim 1, wherein, The assigning a first object region weight according to the first distribution vector includes: Extract features from the object region of the first image to obtain a first feature image; Construct a dynamic weight vector corresponding to the object region of the first image according to the first distribution vector; Perform weighted processing on the first feature image and the dynamic weight vector to obtain a first object region weight.

6. The method for analyzing the similarity of image backgrounds according to claim 1, wherein, The calculating the weighting coefficient of the scene region in the first image includes: Divide the scene region of the first image into multiple groups of partial scene regions; Perform feature analysis on each of the partial scene regions one by one to obtain partial scene features; Identify partial scene types according to the partial scene features; Calculate the weighting coefficient for each of the partial scene regions one by one according to the partial scene types to obtain the partial scene weighting coefficients corresponding to the partial scene regions; Generate the weighting coefficient of the scene area in the first image according to the combination of the partial scene weighting coefficients.

7. The method for analyzing the similarity of image backgrounds according to claim 1, characterized in that, Calculating the similarity score of the image background of the first image and the image background of the second image according to the first object area weight, the first scene area weight, the second object area weight, and the second scene area weight includes: Calculating the similarity score of the object area of the first image and the object area of the second image according to the first object area weight and the second object area weight; Calculating the similarity score of the scene area of the first image and the scene area of the second image according to the first scene area weight and the second scene area weight; Calculating the initial similarity score of the image background of the first image and the image background of the second image; Performing weighted fusion on the object area similarity score, the scene area similarity score, and the initial similarity score to obtain the similarity score of the image background of the first image and the image background of the second image.

8. An image background similarity analysis device, characterized in that, Including: A first image processing module, configured to extract the image background of a pre-acquired first image, segment the extracted image background into an object area and a scene area, construct a first distribution vector according to the object area of the first image, allocate a first object area weight according to the first distribution vector, calculate the weighting coefficient of the scene area in the first image, and construct a first scene area weight using the weighting coefficient of the first scene area; A second image processing module, configured to extract the image background of a pre-acquired second image, segment the extracted image background into an object area and a scene area, construct a second distribution vector according to the object area of the second image, allocate a second object area weight according to the second distribution vector, calculate the weighting coefficient of the scene area in the second image, and construct a second scene area weight using the weighting coefficient of the second scene area; A similarity analysis module, configured to calculate the similarity score of the image background of the first image and the image background of the second image according to the first object area weight, the first scene area weight, the second object area weight, and the second scene area weight.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the image background similarity analysis method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the image background similarity analysis method according to any one of claims 1 to 7 are implemented.