Image region labeling method and device based on semantic relationship
Through the semantic relationship-based image area labeling method, and the image segmentation and sparse reconstruction technology are used to solve the problem that traditional annotation methods cannot perform area-level annotation, and the accuracy and adaptability of image area labeling are improved.
Patent Information
- Application Number
- CN202311474912.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-07
- Publication Date
- 2025-05-13
AI Technical Summary
Traditional image annotation methods cannot effectively perform regional-level image annotation, and cannot meet the retrieval and understanding needs of massive images.
The image area labeling method based on semantic relationship is adopted, and the target image is segmented, the feature vectors of the test area are obtained, and the symbiotic relationship between the semantics of the marked area and the feature vectors of the unlabeled area is used for sparse reconstruction to obtain the annotation result of the unlabeled area.
The accuracy of image area labeling is improved, so that the annotation results give priority to the constraints of the semantic symbiotic relationship between semantics and marked areas, and adapt to the needs of massive images.
Smart Images

Figure CN119992148A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a method and device for annotating image regions based on semantic relations. Background Art
[0002] With the development of Internet technology and multimedia technology, network images are showing a trend of massive growth. The diversification of images and the trend of big data have brought severe challenges to traditional image retrieval engines. Traditional image annotation methods realize the global annotation of images, that is, image-level label annotation. The above methods are not suitable for the current development of massive network images. Due to the characteristics of large image data volume, rapid growth and many types, they cannot provide good assistance for image retrieval and image understanding.
[0003] Therefore, how to perform accurate and effective region-level image annotation has become a technical issue that needs to be urgently addressed in this field. Summary of the invention
[0004] The object of the present invention is to provide a method and device for annotating image regions based on semantic relations, which can perform more reliable image region annotation.
[0005] To achieve the above object, the present invention provides an image region labeling method based on semantic relationship, comprising:
[0006] Segmenting the target image to obtain a plurality of non-overlapping test areas, and obtaining a feature vector of each of the test areas;
[0007] For each first test area, sparsely reconstruct the first test area according to the feature vector of the first test area, the first semantics of each second test area, and the symbiotic relationship between the semantics of the areas obtained in advance, to obtain a labeling result of the first test area;
[0008] The first test area is an unmarked test area; and the second test area is a marked test area.
[0009] In one embodiment of the present invention, for each first test area, sparsely reconstructing the first test area according to the symbiotic relationship between the feature vector of the first test area, the first semantics of each second test area, and the semantics of the pre-obtained area to obtain the annotation result of the first test area includes:
[0010] Based on the symbiotic relationship between the first semantics and the semantics of the region, an original dictionary is optimized to obtain a dictionary corresponding to the first test region;
[0011] Performing sparse reconstruction on the first test area according to the feature vector of the first test area and a dictionary corresponding to the first test area to obtain a labeling result of the first test area;
[0012] The original dictionary is obtained according to the semantics of all sample images and each sample area in the sample images.
[0013] In one embodiment of the present invention, the optimizing of the original dictionary based on the symbiotic relationship between the first semantics and the semantics of the region to obtain the dictionary corresponding to the first test region includes:
[0014] Determining a second semantics based on the symbiotic relationship between the first semantics and the semantics of the region; the second semantics being a semantics having a symbiotic relationship with the first semantics;
[0015] The portion of the original dictionary that is irrelevant to the second semantics is removed to obtain a dictionary corresponding to the first test area.
[0016] In one embodiment of the present invention, sparsely reconstructing the first test area according to the feature vector of the first test area and the dictionary corresponding to the first test area to obtain the labeling result of the first test area includes:
[0017] Performing sparse reconstruction on the first test area according to the feature vector of the first test area and a dictionary corresponding to the first test area to obtain a sparse code of the first test area;
[0018] Based on the sparse coding of the first test area, semantics of the first test area is obtained as a labeling result of the first test area.
[0019] In one embodiment of the present invention, after performing sparse reconstruction on each first test area according to the symbiotic relationship between the feature vector of the first test area, the first semantics of each second test area, and the semantics of the pre-obtained area, and obtaining the annotation result of the first test area, the method further includes:
[0020] Acquire the actual position relationship between the first test area and each of the second test areas;
[0021] When the labeling result of the first test area does not conform to the actual position relationship with the first semantics, each first test area is re-labeled.
[0022] In one embodiment of the present invention, the obtaining of the actual positional relationship between the first test area and each of the second test areas includes:
[0023] Obtaining the centroid of the first test area and the centroid of each of the second test areas;
[0024] For each of the second test areas, based on the positional relationship between the centroid of the first test area and the centroid of the second test area, an actual positional relationship between the first test area and the second test area is obtained.
[0025] In one embodiment of the present invention, a semantic relationship-based image region labeling device includes:
[0026] An acquisition module, used for segmenting the target image to obtain a plurality of non-overlapping test areas, and obtaining a feature vector of each of the test areas;
[0027] a labeling module, configured to perform sparse reconstruction on each first test area according to the feature vector of the first test area, the first semantics of each second test area, and the symbiotic relationship between the semantics of the areas obtained in advance, and obtain a labeling result of the first test area;
[0028] The first test area is an unmarked test area; and the second test area is a marked test area.
[0029] In one embodiment of the present invention, an electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of any of the above-mentioned methods for labeling image regions based on semantic relationships are implemented.
[0030] In one embodiment of the present invention, a non-transitory computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of any of the above-mentioned methods for labeling image regions based on semantic relationships.
[0031] In one embodiment of the present invention, a computer program product includes a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned methods for labeling image regions based on semantic relationships are implemented.
[0032] Compared with the prior art, the method and device for image region labeling based on semantic relationships of the present invention have the beneficial effect that, in the process of labeling the target image region, based on a sparse representation method and utilizing the symbiotic relationship between the semantics of each test region, sparsely reconstructing the unlabeled test region according to the semantics of the labeled test region, predicting the label of the unlabeled test region, and obtaining the labeling result of the unlabeled test region, the labeling result of the unlabeled test region preferentially satisfies the constraint of the symbiotic relationship between the semantics and the semantics of each labeled test region, thereby improving the accuracy of image region labeling. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 This is one of the flowcharts of the method for labeling image regions based on semantic relations according to one embodiment of the present invention;
[0034] Figure 2 is a schematic diagram of a semantic symbiotic relationship in an image region labeling method based on a semantic relationship according to an embodiment of the present invention;
[0035] Figure 3 is a second flow chart of an image region labeling method based on semantic relationship according to an embodiment of the present invention;
[0036] Figure 4 is a structural schematic diagram of an image region labeling device based on semantic relationship according to an embodiment of the present invention;
[0037] Figure 5 is a schematic structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0038] The specific implementation modes of the present invention are described in detail below in conjunction with the accompanying drawings, but it should be understood that the protection scope of the present invention is not limited by the specific implementation modes.
[0039] Unless explicitly stated otherwise, throughout the specification and claims, the term “comprise” or variations such as “include” or “comprising”, etc., will be understood to include the stated elements or components but not to exclude other elements or components.
[0040] In the related art, compared with image-level image annotation, region-level image annotation is a more detailed image annotation method. Image-level image annotation is to annotate the image with a label of the semantics based on the overall semantics of the image. For example, an image can be labeled as "landscape", "building" or "person", etc. Region-level image annotation is to annotate the region with a label of the semantics based on the semantics of each region in the image. For example, the regions in an image labeled as "landscape" can be labeled as "mountains", "trees" and "rivers", etc.
[0041] Region-level image annotation can better solve the problems that global image annotation cannot handle. Region-level image annotation can find the relationship between the semantic region of the image and the semantic label of the image, realize the correlation between the semantic region of the image and the semantic label of the image, and achieve one-to-one correspondence, which is of great help to image retrieval and image understanding, thereby improving the performance of image retrieval.
[0042] For example, for the same image, image-level image annotation can label the entire image as "building", "grass", "tree", and "sky", while region-level image annotation can further label different regions in the image as "building", "grass", "tree", and "sky" respectively.
[0043] In the multi-instance learning (MIL) problem, the training samples are called bags, each of which contains multiple instances. Instances do not have concept labels, only bags have concept labels. If at least one instance in a bag is a positive example, the bag is labeled as positive; if all the examples in the bag are negative examples, the bag is labeled as negative. The multi-instance learning algorithm is to obtain a classifier that can predict unknown bags or examples by learning the training bag. Due to the existence of noise, the process of training the MIL classifier is prone to errors, which affects the accuracy of the final annotation.
[0044] The SVM and MIL classifiers can be combined to solve the shortcomings of each classifier and improve the accuracy of annotation. For example, the problem of region-based automatic semantic annotation of images can be transformed into a multi-instance learning problem, and an asymmetric SVM can be designed to study multi-instance learning of automatic semantic annotation of images. The classifier-based annotation algorithm mainly processes smaller-scale data sets, and the algorithm model is relatively complex, and the scalability of the algorithm is poor.
[0045] With the progress of sparse coding, the method of combining sparse coding with image region annotation has emerged. On the basis of sparse coding, sparse coding is combined with image region annotation and has achieved certain results.
[0046] For example, the double-layer sparse coding algorithm (BLSC) combined with sparse representation is used for image region annotation: first, the image is divided into regions; then each test region is reconstructed using the set double-layer sparse model; based on the sparse reconstruction coefficient, the label is located in the target region. A semantic label of an image can form a separate local semantic region, and images with common semantic labels also contain the same semantic region. Two constraints are used to solve the sparse representation coefficients during the sparse reconstruction process. The first constraint is that an image or a region in an image is reconstructed from sub-blocks of images labeled with the same label, and the second constraint is to reconstruct from as few images as possible. However, in the process of constructing the dictionary, this method assumes that the base regions are independent of each other and ignores the contextual relationship between semantic regions.
[0047] For example, the unified dictionary learning and region tagging with hierarchical coefficient representation algorithm (Unified Dictionary Learning and Region Tagging with Hierarchical Sparse Representation) that combines the structured nature of image layout can achieve certain results. Using this method, in the process of region sparse representation, in the traditional process of using L1 norm, a hierarchical structure can be added to the sparse coding to form a tree-guided dictionary learning. In this structure, the hierarchical structure between feature points, regions, and images is encoded by forming a tree-guided dictionary learning. In the process of region reconstruction, a semi-hierarchical structure can be used to guide the sparse reconstruction of the sample region.
[0048] like Figures 1 to 5 As shown, the method and device for labeling image regions based on semantic relations according to a preferred embodiment of the present invention can be implemented in the following manner.
[0049] Figure 1 FIG. 1 is a flow chart of a method for labeling image regions based on semantic relationships according to an embodiment of the present invention. Figure 1 As shown, the method may include the following steps: step 101 and step 102.
[0050] Step 101: segment the target image to obtain multiple non-overlapping test areas, and obtain a feature vector for each test area.
[0051] Specifically, the target image can be segmented based on any image segmentation method, and the target image can be segmented into multiple non-overlapping regions, each of which serves as a test region.
[0052] The above-mentioned image segmentation method can adopt any threshold-based segmentation method, any region-based segmentation method, any edge-based segmentation method, any segmentation method based on a specific theory, or a combined segmentation method combining at least two of the above methods.
[0053] The embodiment of the present invention does not specifically limit the specific image segmentation method used.
[0054] The goal of segmenting the target image is to completely segment the target image according to semantic regions.
[0055] For each test area, any feature extraction method may be used to extract the features of the test area, thereby obtaining a feature vector of the test area.
[0056] The above-mentioned feature extraction method may adopt any feature extraction method based on deep learning, or any feature extraction method not based on deep learning, such as SIFT, HOG, SURF, ORB, LBP, HAAR, etc.
[0057] Preferably, a convolutional neural network may be used to extract features of the test area to obtain a feature vector of the test area.
[0058] Optionally, the features of the test area include low-level features and high-level features. The low-level features may mainly include visual features, such as color histogram, texture, and color moment, etc. The high-level features may mainly include semantic features.
[0059] Step 102: for each first test area, sparsely reconstruct the first test area according to the symbiotic relationship between the feature vector of the first test area, the first semantics of each second test area and the semantics of the pre-obtained area, and obtain the labeling result of the first test area; wherein the first test area is an unlabeled test area; and the second test area is a labeled test area.
[0060] Specifically, the innovative point of the embodiment of the present invention is that in the process of image region annotation based on sparse representation in the related art, the annotation work is improved by introducing high-level information of the image. Specifically, in the process of sparse representation, the co-occurrence relationship between region labels is introduced to improve the accuracy of annotation.
[0061] The focus of the embodiment of the present invention is to combine the low-level visual feature information of the image (i.e. the aforementioned low-level features) and the high-level information of the image (i.e. the aforementioned high-level features) in the process of region labeling, increase subjective description, and thus improve the accuracy and robustness of labeling.
[0062] Image feature information plays an important role in image retrieval, image understanding, image recognition, scene recognition and other fields. Region-level image annotation based on sparse representation can take advantage of the characteristics of sparse representation that are easy to solve and convenient to calculate. Region-level image annotation based on sparse representation mainly measures the similarity by inputting the feature information of the extracted image and calculating the difference between the above feature information and the standard data.
[0063] In the related art, the method of using low-level visual information as the feature of the image can no longer solve the problems of a wide variety of objects, diversified images, and similar features of different objects in the current image. In order to better describe different objects and overcome the defects of low-level visual features, high-level semantic information can be introduced into the fields of image recognition and scene classification, and good results have been achieved. The reason is that it not only combines the information contained in the low-level visual features, but also overcomes the erroneous results caused by the similarity or identity of simple visual features in the process of image retrieval, understanding and recognition by adding subjective descriptions.
[0064] In the embodiment of the present invention, based on the basic sparse representation algorithm, image region annotation is performed in combination with high-level semantic information of the image. The information used in the image region annotation process may include low-level features such as color histogram, texture and color moment, and at least one high-level feature such as the semantic symbiosis relationship of the region label and the spatial relationship of the region.
[0065] Sparse representation has the advantages of strong robustness and high accuracy. It has good accuracy and has good applications in the process of image region annotation.
[0066] In region annotation based on image annotation and object recognition, contextual restrictions of sample regions can be used to improve the performance of image understanding and region annotation. The above restrictions mainly include three types: visual restrictions, spatial restrictions and semantic restrictions. In various embodiments of the present invention, the lexical co-occurrence relationship (i.e., semantic symbiosis relationship) in semantic restrictions and the spatial position relationship in spatial restrictions are mainly used.
[0067] The embodiment of the present invention focuses on the sparse representation of the image. In the process of reconstructing the semantic region of the image, the sparse representation is mainly used to perform sparse reconstruction.
[0068] The encoding process in sparse reconstruction can adopt a graph-guided region annotation sparse reconstruction algorithm. In the graph-guided region annotation sparse reconstruction algorithm, the contextual relationship between regions is flexibly modeled. Optionally, a graph-guided fusion penalty term can be defined to support the sparsity of the difference between the two reconstruction coefficients, so that the graph structure obtained from the sample region can be well introduced into the sparse reconstruction process.
[0069] By using sparse representation to annotate image regions, better annotation effects can be achieved by considering the contextual relationship between regions. Therefore, making full use of the semantic symbiosis relationship and spatial position relationship between semantic regions can further improve the annotation effect. The existence of semantic symbiosis and spatial position relationship between semantic regions is determined by the distribution law of objects in nature. Moreover, in images collected by humans, the semantic symbiosis and spatial position relationship between semantic regions can be well displayed. By using the semantic symbiosis relationship between regions and the spatial position information between regions to annotate different regions of the entire image, the contextual relationship between image regions can be further utilized, thereby improving the accuracy of annotation and providing more help for image understanding and image retrieval.
[0070] Sparse learning methods are widely used in the field of computer vision to learn basic visual patterns or to select the most significant data features in images.
[0071] Image set X = (X1, X2, ..., X r ,...,X N ), contains N images in total, of which the rth image is X r Each image has been well segmented into Nr semantic regions, X r =(x r1 ,x r2 ,...,x rNr ), r=1,2,...,N. Each semantic region X rNr Extract d-dimensional features.
[0072] The images in the training set and the images in the test set have M labeled labels, and the standard for labeling each image is known. Define E∈{0,1} M×j Represents the label description matrix of all regions. When the jth region contains the i-th label, E(i,j)=1; otherwise, E(i,j)=0.
[0073] In the embodiment of the present invention, for an unlabeled test area x u , through a learned reconstruction sparse vector α∈R j Reconstruct from the sample area in the image set X. For each reconstruction process, the reconstruction error is used to measure the reconstruction situation. The purpose of reconstruction is to make the reconstruction error as small as possible.
[0074] The reconstruction error can be represented by the Euclidean distance between the test area and the reconstructed area. The specific formula is as follows:
[0075]
[0076] Optionally, an L1 norm of the reconstruction coefficient vector can be added to formula (2) to modify the sparse reconstruction framework and obtain the reconstruction objective function. The specific formula of the objective function is as follows:
[0077]
[0078] λ≥0 in formula (2) is a regularization parameter that controls the number of non-zero terms in α. The optimization problem of formula (2) can be solved using the Lasso algorithm or the like. r Represents the image X r The subvector of reconstruction coefficients for the region in .
[0079] The reconstruction coefficient vector α is divided into N sub-vectors. The sparse reconstruction framework can be rewritten using the Group Lasso penalty term as follows:
[0080]
[0081] Among them, δ ≥ 0 is the regularization parameter of the Group Lasso penalty term.
[0082] In the process of image region annotation, the accuracy of sparse reconstruction can be improved by analyzing the semantic and / or spatial correlations between regions.
[0083] The semantic relationship between semantic tags is mainly reflected in the semantic co-occurrence relationship of words. In the natural world, various objects appear in a certain pattern. By analyzing the entire pattern and finding the co-occurrence relationship between the semantics of words, we can improve the understanding of image content and image area annotation.
[0084] The semantics of a region refers to the semantics of the semantic tags that annotate the region. The symbiotic relationship between semantics is mainly the co-occurrence relationship (also known as the "co-occurrence relationship"), which refers to the tendency of words to appear together.
[0085] Optionally, a 1-N co-occurrence relationship may be used to represent the co-occurrence relationship between the semantics of the regions.
[0086] Define each tag as T-tag. For example, the tag grass is defined as T-grass and the tag cow is defined as T-cow.
[0087] The 1-N co-occurrence relationship means that when Tag-1 appears, there are usually N corresponding tags Tag-N in many images. Tag-1 is defined as the first-level tag and Tag-N is defined as the second-level tag.
[0088] It should be noted that the 1-N co-occurrence relationship does not mean that all corresponding tags appear in the image where Tag-1 appears. In the image where Tag-1 appears, Tag-N may appear once or multiple times. The first-level tag of the co-occurrence relationship is the tag to be determined first. Through the first-level tag, other tags can be predicted.
[0089] Figure 2 It is a schematic diagram of the symbiotic relationship of semantics in the image region labeling method based on semantic relationship according to an embodiment of the present invention. Figure 2 It is an example of a simulation diagram obtained by summarizing numerous images and combining statistics with the database (including multiple sample images) used in the embodiment of the present application.
[0090] Figure 2 In , the direction of the connecting arrows between different semantic labels is from the first-level label to the second-level label. Figure 2 As shown, the first-level tag (Level-1 Tag) is "grassland", and there are 8 second-level tags (Level-2 Tags) with symbiotic relationships, namely "cow", "horse", "sheep", "grass", "dog", "building", "sky" and "tree".
[0091] In step 102, the label of a certain test area can be first predicted using sparse representation to obtain the annotation result of the area; then, based on the area, the label of a test area that has not been annotated is predicted, and the prediction result must give priority to satisfying the constraints of the symbiotic relationship between the semantics and the semantics of the aforementioned annotation results; and so on, the label of each test area that has not been annotated is predicted, and the prediction result must give priority to satisfying the constraints of the symbiotic relationship between the semantics and the semantics of each annotated test area, until each test area in the target image is annotated.
[0092] In some embodiments, for each first test area, sparse reconstruction is performed on the first test area according to the symbiotic relationship between the feature vector of the first test area, the first semantics of each second test area, and the semantics of the pre-obtained area, to obtain a labeling result for the first test area, including: optimizing the original dictionary based on the first semantics and the symbiotic relationship between the semantics of the area to obtain a dictionary corresponding to the first test area; wherein the original dictionary is obtained based on the semantics of all sample images and each sample area in the sample images.
[0093] Specifically, during the sparse reconstruction of the test area, the dictionary needs to be overcomplete to ensure that the required reconstructed items are obtained. It is understandable that an overly complete dictionary will introduce more errors in the sparse reconstruction process, thereby reducing the reconstruction effect. In the real world, the layout of objects is regular, which is intuitively reflected in the obtained pictures. Different semantic areas have a certain correlation, and semantic co-occurrence relationships are very important relationships between objects. In the sparse reconstruction process, semantic co-occurrence relationships can be introduced, and semantic co-occurrence relationships can be used to optimize the dictionary and reduce the introduction of noise, which can improve the effect of sparse reconstruction and thus improve the accuracy of area annotation.
[0094] The original dictionary is a dictionary obtained based on the semantics of all sample images and sample regions in the sample images. The original dictionary can be optimized based on the symbiotic relationship between the first semantics and the semantics of the region, and redundant information in the original dictionary can be removed to obtain a dictionary corresponding to the first test region, thereby reducing noise.
[0095] It can be understood that the dictionary corresponding to the first test area is a proper subset of the original dictionary.
[0096] According to the feature vector of the first test area and the dictionary corresponding to the first test area, the first test area is sparsely reconstructed to obtain a labeling result of the first test area.
[0097] Specifically, based on the feature vector of the first test area and the optimized dictionary (that is, the dictionary corresponding to the first test area), the first test area may be sparsely reconstructed to obtain a labeling result of the first test area.
[0098] The specific process of performing sparse reconstruction based on the dictionary can be found in the above-mentioned embodiment and will not be described in detail here.
[0099] In some embodiments, based on the symbiotic relationship between the first semantics and the semantics of the region, the original dictionary is optimized to obtain a dictionary corresponding to the first test area, including: determining the second semantics based on the symbiotic relationship between the first semantics and the semantics of the region; removing parts of the original dictionary that are not related to the second semantics to obtain a dictionary corresponding to the first test area; the second semantics is a semantics that has a symbiotic relationship with the first semantics.
[0100] Specifically, Figure 2 What is shown is an ideal situation of symbiotic relationship, which means that in the image where the label "grassland" appears, semantic labels such as the label "cow", the label "horse" and the label "building" will appear.
[0101] For 1-N semantic symbiosis, when determining that a test area is labeled Tag-1, the embodiment of the present application will group the corresponding N images of Tag-N that meet the semantic symbiosis relationship 1-N into a new image sample library as the Tag-1 sample library. When the target image is detected to contain the tag Tag-1, the aforementioned reconstructed Tag-1 sample library is used to continue sparse reconstruction for the remaining test areas to annotate the remaining semantic areas in the target image.
[0102] In the above process, the dictionary used for sparse reconstruction is optimized, eliminating some redundant information, thereby improving the accuracy of labeling.
[0103] The following uses the semantics of the label "grass" as the first semantics to illustrate how to optimize the original dictionary. First, statistics can be performed on the training sample library. The training sample library includes multiple sample images. All sample images in the training sample library that have the label "grass" can be found, and all sample images in the training sample library that contain other labels that have a symbiotic relationship with the label "grass" can be unified into a sample library dedicated to the label "grass". It can be understood that the semantics of other labels that have a symbiotic relationship with the label "grass" are the second semantics.
[0104] All sample images with the label "grass" in the training sample library can be expressed as formula (4).
[0105]
[0106] Among them, Image grass represents the set of sample images with the label "grass".
[0107] For the collection Image grass All the tags in are counted, as shown in formula (5).
[0108] Tags grass ={...,tag-k,...} X r ∈Image grass (5).
[0109] Set Tag grass All sample images with the label appearing in are made into a sample library dedicated to the label "grassland", as shown in formula (6).
[0110] Database grass =X r X rNr ∈Tag grass (6).
[0111] It is understandable that based on Database grassThe obtained dictionary is a dictionary corresponding to the first test area. Compared with the original dictionary obtained based on the training sample library, the dictionary corresponding to the first test area removes some redundant information.
[0112] The goal of the embodiment of the present application is to predict each semantic region (i.e., test region) of the target image. When the label of the first test region is predicted to be Tag-1, by summarizing and counting the sample images in the training sample library, combined with the following Figure 2 The model shown can filter out sample images containing Tag-N labels based on our initial training sample library, form a new sample library, and obtain a new dictionary. Through the above process, the size of the sample library and the size of the dictionary can be reduced, some redundant information can be reduced, and the accuracy of region annotation can be improved.
[0113] In some embodiments, based on the feature vector of the first test area and the dictionary corresponding to the first test area, the first test area is sparsely reconstructed to obtain the annotation result of the first test area, including: based on the feature vector of the first test area and the dictionary corresponding to the first test area, the first test area is sparsely reconstructed to obtain the sparse coding of the first test area.
[0114] Specifically, after obtaining the dictionary corresponding to the first test area, a sparse representation method may be executed to sparsely reconstruct the first test area according to the feature vector of the first test area and the dictionary corresponding to the first test area to obtain a sparse code of the first test area.
[0115] Based on the sparse coding of the first test area, the semantics of the first test area is obtained as a labeling result of the first test area.
[0116] Specifically, based on the sparse coding of the first test area and the dictionary corresponding to the first test area, the semantics of the first test area can be obtained, and the semantics is used as a label of the first test area to mark the first test area.
[0117] In some embodiments, for each first test area, sparse reconstruction is performed on the first test area according to the symbiotic relationship between the feature vector of the first test area, the first semantics of each second test area, and the semantics of the pre-obtained area. After obtaining the annotation result of the first test area, the actual position relationship between the first test area and each second test area is obtained.
[0118] Specifically, after annotation based on semantic relationship, the spatial position relationship between image regions can be used to judge the obtained annotation results, and the annotation results can be improved on this basis.
[0119] Spatial position relationship (abbreviated as "position relationship") has important applications in image retrieval and scene recognition. For example, based on the in-depth study of the low-level visual features of the high-level semantics of the image, semantic extraction and retrieval algorithms of high-level semantic layers such as image semantics and spatial relationship semantics can be proposed, thereby effectively extracting high-level semantic information of the image.
[0120] Simply collecting 2D images cannot fully reproduce the spatial layout relationship between objects. In order to improve the efficiency and accuracy of annotation and to ensure the universality of the method of the present invention, the upper, middle and lower positional relationships in the spatial layout between different regions can be used to assist in annotation based on the block division of the target image.
[0121] In an embodiment of the present invention, the spatial structure information between different regions of a single image can be used to perform a secondary judgment on the annotation results obtained based on the semantic co-occurrence relationship: when it does not conform to the rules of the spatial structure relationship of the image region, the annotation results are further improved, thereby improving the accuracy and robustness of the image region annotation method.
[0122] Optionally, based on the feature vector of the first test area, the first semantics of each second test area, and the pre-obtained positional relationship of the areas, the first test area may be sparsely reconstructed to obtain a labeling result of the first test area. The above labeling result is compared with the labeling result obtained in step 102. If the two are inconsistent, the labeling result obtained in step 102 is corrected based on the above labeling result to improve the labeling result.
[0123] Optionally, the positional relationship of regions can be used in sparse reconstruction based on the spatial group sparse coding (SGSC) algorithm. This algorithm fully considers the advantages of group sparse coding in the sparse reconstruction process and adds the spatial correlation between semantic regions into the algorithm. In the sparse coding process, a group-specific spatial kernel function is added to generate a regularization term that is easier to interpret. The joint version of the SGSC model can uniformly encode the interrelated regions within an image.
[0124] When the labeling result of the first test area does not conform to the actual position relationship with the first semantics, each first test area is re-labeled.
[0125] Specifically, in theory, the annotation result of the first test area and the first semantics are consistent with the spatial position layout of objects in the real world. For example, in the spatial position layout of objects in the real world, the sky is above the grass.
[0126] The accuracy of the annotation results can be determined by comparing the actual positional relationships between different test areas obtained based on the annotation results. If the spatial position layout of objects in the real world is consistent, it means that the annotation results based on the sparse representation based on the semantic symbiosis relationship are accurate; if the spatial position layout of objects in the real world is not consistent, it means that the annotation results based on the sparse representation based on the semantic symbiosis relationship are inaccurate and need to be further improved to improve the accuracy of the annotation.
[0127] For example, in the target image, the actual positional relationship between the first test area labeled "sky" and the second test area labeled "ocean" is that the first test area labeled "sky" is below the second test area labeled "ocean", and the "sky" and "ocean" do not conform to the unknown relationship that the "sky" is below the "ocean". Therefore, the first test area originally labeled "sky" needs to be re-labeled to obtain a labeling result that conforms to the spatial position layout of objects in the real world, thereby improving the accuracy of labeling.
[0128] In some embodiments, obtaining the actual positional relationship between the first test area and each second test area includes: obtaining the center of mass of the first test area and the center of mass of each second test area; for each second test area, obtaining the actual positional relationship between the first test area and the second test area based on the positional relationship between the center of mass of the first test area and the center of mass of the second test area.
[0129] Specifically, the centroid of the region can be used to determine the positional relationship of the region in the target image. It should be noted that the use of the centroid of the region to determine the positional relationship of the region in the target image requires that the shooting angle of the target image cannot be located vertically above the object.
[0130] After segmenting the target image, each test area can be labeled, and then the centroid of each test area can be obtained.
[0131] Assume that there are Nr test regions in image Xr. r1 ,X r2 ,...,X rNr}∈X r , extract the Y-axis value of the centroid of each test area.
[0132]
[0133] The positional relationship between different test areas can be determined by comparing the positions of the centroids of different test areas on the Y axis in the two-dimensional coordinate system of the target image.
[0134]
[0135] Figure 3FIG2 is a second flow chart of a method for labeling image regions based on semantic relationships according to an embodiment of the present invention. Figure 3 As shown, the image region annotation method may include the following steps: segmenting (Segmantation) and feature extraction (Feature Extraction) of the sample image; obtaining the semantic symbiosis relationship between the regions; segmenting and feature extraction of the target image; sparse reconstruction (Sparse Reconstruction) under the guidance of semantic symbiosis (Semantic Occurrence Guided); tag propagation (Tag propagation); and improving the annotation results under the spatial relationship assistant (Spatial relationship assistant).
[0136] The process of labeling the target image region may include four steps: segmenting and extracting features of the target image, sparse reconstruction under the guidance of semantic symbiotic relationships, and improving the labeling results under the guidance of labels and spatial relationships. Segmenting and extracting features of the sample image, and obtaining the symbiotic relationship between the semantics of the region are steps that are performed in advance before executing the process of labeling the target image region.
[0137] It can be understood that, for both the sample image and the target image, the feature vector of each region is extracted based on image segmentation.
[0138] For each test area, sparse reconstruction can be performed under the guidance of semantic symbiosis through steps to obtain a reconstruction coefficient vector.
[0139] It is understandable that the above-mentioned image region annotation method including improving the annotation result by using the position relationship may include:
[0140] Step ①: Segment the target image into regions and perform block processing.
[0141] Step ②: Label the segmented image.
[0142] Step ③: For each segmented area, extract the centroid of the area, and compare the positions of the centroids of different test areas on the Y-axis in the two-dimensional coordinate system of the target image to determine the positional relationship between different test areas.
[0143] Step ④: Improve the annotation results by comparing the spatial position relationship between different areas, thereby improving the accuracy of annotation.
[0144] The beneficial effect of the present invention is that, in the process of image region annotation of the target image, based on the sparse representation method and utilizing the symbiotic relationship between the semantics of each test region, the unlabeled test region is sparsely reconstructed according to the semantics of the labeled test region, the label of the unlabeled test region is predicted, and the annotation result of the unlabeled test region is obtained, so that the annotation result of the unlabeled test region preferentially satisfies the constraint of the symbiotic relationship between the semantics and the semantics of each labeled test region, thereby improving the accuracy of image region annotation.
[0145] The present invention proposes a sparse representation algorithm that combines the symbiotic relationship of high-level semantic information of an image to achieve image region labeling. The algorithm integrates the contextual relationship of semantic regions into the image region labeling process. By combining the symbiotic relationship and spatial relationship of semantic regions, the construction of the dictionary is optimized during the sparse reconstruction process. On the basis of the initial labeling, a secondary judgment is made using the spatial position relationship to improve the already labeled labels, thereby improving the accuracy and robustness of the labeling. By experimenting with the algorithm proposed by the present invention on a public dataset for image labeling, the experiment shows that compared with related technologies, the algorithm has higher robustness and accuracy.
[0146] The image region labeling apparatus based on semantic relationship provided by the present invention is described below. The image region labeling apparatus based on semantic relationship described below and the image region labeling method based on semantic relationship described above can be referred to each other.
[0147] Figure 4 is a schematic diagram of the structure of the image region labeling device based on semantic relationship provided by the present invention. Based on the content of any of the above embodiments, Figure 4 As shown, the device includes an acquisition module 401 and a marking module 402, wherein:
[0148] The acquisition module 401 is used to segment the target image to obtain multiple non-overlapping test areas and obtain a feature vector of each test area;
[0149] The annotation module 402 is used to perform sparse reconstruction on each first test area according to the feature vector of the first test area, the first semantics of each second test area and the symbiotic relationship between the semantics of the areas obtained in advance, and obtain the annotation result of the first test area;
[0150] The first test area is an unmarked test area; and the second test area is a marked test area.
[0151] Specifically, the acquisition module 401 and the marking module 402 may be electrically connected.
[0152] Optionally, the marking module 402 may include:
[0153] A dictionary optimization unit, configured to optimize the original dictionary based on the symbiotic relationship between the first semantics and the semantics of the region, to obtain a dictionary corresponding to the first test region;
[0154] A sparse reconstruction unit, configured to perform sparse reconstruction on the first test area according to the feature vector of the first test area and a dictionary corresponding to the first test area, and obtain a labeling result of the first test area;
[0155] The original dictionary is obtained based on the semantics of all sample images and each sample area in the sample images.
[0156] Optionally, the dictionary optimization unit can be specifically used to determine the second semantics based on the symbiotic relationship between the first semantics and the semantics of the region; the second semantics is the semantics that has a symbiotic relationship with the first semantics; remove the part of the original dictionary that is not related to the second semantics to obtain the dictionary corresponding to the first test area.
[0157] Optionally, the sparse reconstruction unit can be specifically used to perform sparse reconstruction on the first test area based on the feature vector of the first test area and the dictionary corresponding to the first test area, so as to obtain the sparse coding of the first test area; based on the sparse coding of the first test area, obtain the semantics of the first test area as the annotation result of the first test area.
[0158] Optionally, the device may further include:
[0159] A position relationship module, used to obtain the actual position relationship between the first test area and each second test area;
[0160] The result improvement module is used to re-annotate each first test area respectively when the annotation result of the first test area does not conform to the position relationship with the first semantics.
[0161] Optionally, the position relationship module can be specifically used to obtain the center of mass of the first test area and the center of mass of each second test area; for each second test area, based on the positional relationship between the center of mass of the first test area and the center of mass of the second test area, obtain the actual positional relationship between the first test area and the second test area.
[0162] The image region labeling device based on semantic relationship provided in an embodiment of the present invention is used to execute the image region labeling method based on semantic relationship provided in the present invention. Its implementation method is consistent with the implementation method of the image region labeling method based on semantic relationship provided in the present invention, and can achieve the same beneficial effects, which will not be repeated here.
[0163] The image region labeling device based on semantic relationship is used in the image region labeling method based on semantic relationship in the above embodiments. Therefore, the description and definition in the image region labeling method based on semantic relationship in the above embodiments can be used for understanding each execution module in the embodiments of the present invention.
[0164] The beneficial effect of the present invention is that, in the process of image region annotation of the target image, based on the sparse representation method and utilizing the symbiotic relationship between the semantics of each test region, the unlabeled test region is sparsely reconstructed according to the semantics of the labeled test region, the label of the unlabeled test region is predicted, and the annotation result of the unlabeled test region is obtained, so that the annotation result of the unlabeled test region preferentially satisfies the constraint of the symbiotic relationship between the semantics and the semantics of each labeled test region, thereby improving the accuracy of image region annotation.
[0165] Figure 5 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 5 As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530 and a communication bus 540, wherein the processor 510, the communication interface 520 and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call the logic instructions in the memory 530 to execute the image region labeling method based on the semantic relationship, the method comprising: segmenting the target image to obtain a plurality of non-overlapping test regions, and obtaining a feature vector of each test region; for each first test region, sparsely reconstructing the first test region according to the feature vector of the first test region, the first semantics of each second test region and the symbiotic relationship between the semantics of the regions obtained in advance, and obtaining the labeling result of the first test region; wherein the first test region is an unlabeled test region; and the second test region is a labeled test region.
[0166] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0167] The processor 510 in the electronic device provided in the embodiment of the present invention can call the logic instructions in the memory 530. Its implementation method is consistent with the implementation method of the image area labeling method provided by the present invention and can achieve the same beneficial effects, which will not be repeated here.
[0168] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the image area labeling method based on semantic relationships provided by the above methods, and the method includes: segmenting the target image to obtain multiple non-overlapping test areas, and obtaining a feature vector for each test area; for each first test area, sparsely reconstructing the first test area according to the feature vector of the first test area, the first semantics of each second test area, and the symbiotic relationship between the semantics of the pre-obtained area, to obtain the labeling result of the first test area; wherein the first test area is an unlabeled test area; and the second test area is a labeled test area.
[0169] When the computer program product provided by the embodiment of the present invention is executed, the above-mentioned image region labeling method is implemented. Its specific implementation method is consistent with the implementation method described in the embodiment of the aforementioned method, and can achieve the same beneficial effects, which will not be repeated here.
[0170] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the image region labeling method based on semantic relationships provided by the above-mentioned methods, the method comprising: segmenting the target image to obtain multiple non-overlapping test regions, and obtaining a feature vector for each test region; for each first test region, sparsely reconstructing the first test region according to the feature vector of the first test region, the first semantics of each second test region, and the symbiotic relationship between the semantics of the pre-obtained region, to obtain the labeling result of the first test region; wherein the first test region is an unlabeled test region; and the second test region is a labeled test region.
[0171] When the computer program stored on the non-transitory computer-readable storage medium provided by the embodiment of the present invention is executed, the above-mentioned image area labeling method is implemented. Its specific implementation method is consistent with the implementation method described in the embodiment of the aforementioned method, and can achieve the same beneficial effects, which will not be repeated here.
[0172] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0173] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0174] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1A function specified in one or more boxes.
[0175] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0176] The foregoing description of specific exemplary embodiments of the present invention is for the purpose of illustration and demonstration. These descriptions are not intended to limit the present invention to the precise form disclosed, and it is clear that many changes and variations can be made based on the above teachings. The purpose of selecting and describing the exemplary embodiments is to explain the specific principles of the present invention and its practical application, so that those skilled in the art can realize and utilize various different exemplary embodiments of the present invention and various different selections and changes. The scope of the present invention is intended to be limited by the claims and their equivalents.
Claims
1. A method for labeling image regions based on semantic relations, characterized in that: include: Segmenting the target image to obtain a plurality of non-overlapping test areas, and obtaining a feature vector of each of the test areas; For each first test area, sparsely reconstruct the first test area according to the feature vector of the first test area, the first semantics of each second test area, and the symbiotic relationship between the semantics of the areas obtained in advance, to obtain a labeling result of the first test area; The first test area is an unmarked test area; and the second test area is a marked test area.
2. The image region labeling method based on semantic relationship according to claim 1, characterized in that: The step of performing sparse reconstruction on each first test area according to the feature vector of the first test area, the first semantics of each second test area, and the symbiotic relationship between the semantics of the areas obtained in advance to obtain the labeling result of the first test area includes: Based on the symbiotic relationship between the first semantics and the semantics of the region, an original dictionary is optimized to obtain a dictionary corresponding to the first test region; Performing sparse reconstruction on the first test area according to the feature vector of the first test area and a dictionary corresponding to the first test area to obtain a labeling result of the first test area; The original dictionary is obtained according to the semantics of all sample images and each sample area in the sample images.
3. The image region labeling method based on semantic relationship according to claim 2, characterized in that: The optimizing the original dictionary based on the symbiotic relationship between the first semantics and the semantics of the region to obtain a dictionary corresponding to the first test region includes: Determining a second semantics based on the symbiotic relationship between the first semantics and the semantics of the region; the second semantics being a semantics having a symbiotic relationship with the first semantics; The portion of the original dictionary that is irrelevant to the second semantics is removed to obtain a dictionary corresponding to the first test area.
4. The image region labeling method based on semantic relationship according to claim 2, characterized in that: The step of performing sparse reconstruction on the first test area according to the feature vector of the first test area and a dictionary corresponding to the first test area to obtain a labeling result of the first test area includes: Performing sparse reconstruction on the first test area according to the feature vector of the first test area and a dictionary corresponding to the first test area to obtain a sparse code of the first test area; Based on the sparse coding of the first test area, semantics of the first test area is obtained as a labeling result of the first test area.
5. The method for labeling image regions based on semantic relations according to any one of claims 1 to 4, characterized in that: After performing sparse reconstruction on each first test area according to the feature vector of the first test area, the first semantics of each second test area, and the symbiotic relationship between the semantics of the pre-obtained area, and obtaining the labeling result of the first test area, the method further includes: Acquire the actual position relationship between the first test area and each of the second test areas; When the labeling result of the first test area does not conform to the actual position relationship with the first semantics, each first test area is re-labeled.
6. The image region labeling method based on semantic relationship according to claim 5, characterized in that: The obtaining of the actual positional relationship between the first test area and each of the second test areas includes: Obtaining the centroid of the first test area and the centroid of each of the second test areas; For each of the second test areas, based on the positional relationship between the centroid of the first test area and the centroid of the second test area, an actual positional relationship between the first test area and the second test area is obtained.
7. An image region labeling device based on semantic relationship, characterized in that: include: An acquisition module, used for segmenting the target image to obtain a plurality of non-overlapping test areas, and obtaining a feature vector of each of the test areas; a labeling module, configured to perform sparse reconstruction on each first test area according to the feature vector of the first test area, the first semantics of each second test area, and the symbiotic relationship between the semantics of the areas obtained in advance, and obtain a labeling result of the first test area; The first test area is an unmarked test area; and the second test area is a marked test area.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the image region labeling method based on semantic relationship as claimed in any one of claims 1 to 6 are implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the image region labeling method based on semantic relationship as claimed in any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the image region labeling method based on semantic relationship as claimed in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Image area labeling method based on visual semantic relation graph
CN107967494A
Spatial position relationship graph matching-based image region tag correction method
CN108009279A
Method and device for image semantic annotation
CN108319985A