Superpixel segmentation method, device and server based on differential regional content
By extracting and utilizing the multi-dimensional differential feature matrix, combining clustering models and segmentation models based on multi-view tensors, the problem of insufficient utilization of initial image information by existing superpixel segmentation methods is solved, and the segmentation efficiency and accuracy are significantly improved.
Patent Information
- Application Number
- CN202411896849.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-23
AI Technical Summary
The existing superpixel segmentation method does not fully utilize the information of the initial image and fails to effectively fuse the differential features, resulting in insufficient adaptability and poor target boundary segmentation.
By obtaining the initial image for coarse segmentation, multi-dimensional input data is determined, and a multi-dimensional differential feature matrix is extracted, including target color features, texture features, and edge features. Then, the feature matrix is clustered using a clustering model based on multi-view tensors, the coefficient matrix is determined, and the coefficient matrix is graphically cut by a preset segmentation model to determine the target superpixel image.
It significantly improves the superpixel segmentation efficiency and segmentation accuracy, can better process multi-dimensional data, effectively utilize the complementary characteristics of multi-dimensional differentiated feature data, and improves the ability to judge irregular junctions of content, so as to better distinguish the content area and make it close to the original outline.
Smart Images

Figure CN119359745B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a superpixel segmentation method, device and server based on differential region content. Background Art
[0002] In traditional computer vision image segmentation, a single pixel is usually regarded as a processing unit. This segmentation method easily increases computing time and resource consumption. At present, relevant technologies propose that superpixel segmentation can be used to merge multiple pixels in a small area into a superpixel block based on the similarity of pixels and their neighborhoods, and use the superpixel block as a basic unit for subsequent segmentation processing based on gradient or graph theory methods. However, this image segmentation method does not fully utilize the information of the initial image and fails to effectively integrate differential features. Since its cutting criterion is more dependent on a specific affinity graph model, it leads to insufficient adaptability, and when processing the initial image, the segmentation performance of the target boundary is poor. In addition, factors such as large pixel differences between multiple targets, blurred differences between targets, and deviations between the internal and external worlds may have a significant impact on the segmentation effect of the above segmentation scheme. Summary of the invention
[0003] In view of this, an object of the present invention is to provide a superpixel segmentation method, device and server based on differential region content, which can significantly improve superpixel segmentation efficiency and segmentation accuracy.
[0004] In a first aspect, an embodiment of the present invention provides a superpixel segmentation method based on differential area content, the method comprising: acquiring an initial image, and performing a coarse segmentation process on the initial image to determine multidimensional input data, wherein the multidimensional input data is used to represent the differential area content of the image; performing feature extraction processing on the multidimensional input data to extract a multidimensional differential feature matrix of the multidimensional input data, wherein the multidimensional differential feature matrix includes target color features, target texture features, and target edge features of the multidimensional input data; performing data clustering processing on the multidimensional differential feature matrix through a preset clustering model based on a multi-view tensor to determine a coefficient matrix, and performing graph cutting processing on the coefficient matrix using a preset segmentation model to determine a target superpixel image.
[0005] In one embodiment, the steps of roughly segmenting the initial image and determining multidimensional input data include: roughly segmenting the initial image through a watershed transform to determine a marker image and a mask image corresponding to the initial image, and performing morphological dilation and erosion on the marker image and the mask image to determine a set of sheet regions; completing the remaining images for each sheet region to complete the blank areas in each sheet region with the same value, and stacking the completed images as multidimensional input data.
[0006] In one embodiment, the step of performing feature extraction processing on multidimensional input data to extract a multidimensional difference feature matrix of the multidimensional input data includes: obtaining the average signal, variance, peak value and distribution angle of frequency-oriented local edges of non-overlapping windows of different sizes in the multidimensional input data relative to their directions; extracting the color features of the multidimensional input data under different color models respectively, and determining the target color features by fusing the extracted color features; determining the roughness based on the average signal, determining the contrast based on the variance and the peak value, and determining the directionality based on the distribution angle through a preset texture feature extraction model, and determining the roughness, contrast and directionality as target edge features; performing edge feature extraction on the multidimensional input data based on the second-order derivative through a preset edge feature extraction model to determine the target edge features; determining the target color features, target texture features and target edge features as a multidimensional difference feature matrix.
[0007] In one embodiment, the color features of multidimensional input data under different color models are extracted respectively, and the steps of determining target color features by fusing the extracted color features include: converting the multidimensional input data from the original RGB color model to the HSV color model and the YUV color model respectively, wherein the HSV color model is a conical color model and the YUV color model is a linear color model; quantizing the HSV color model and dividing it into various color regions according to preset angle intervals to convert the multidimensional input data into a binary color index set, wherein each color region includes an index of a corresponding color component in a color space; determining the brightness component of the YUV color model as a grayscale feature of the multidimensional input data, and combining the grayscale feature with the color index set to determine it as a target color feature.
[0008] In one embodiment, a multi-dimensional difference feature matrix is subjected to data clustering processing through a preset clustering model based on a multi-view tensor, and the step of determining a coefficient matrix includes: obtaining local characteristic items of the multi-dimensional input data, wherein the local characteristic items are used to represent the spatial neighborhood attributes contained in the image itself; substituting the multi-dimensional difference feature matrix and the spatial neighborhood attributes into the preset clustering model based on a multi-view tensor, and applying a diagonalization item and a tensor nuclear norm to the clustering model after the data is substituted, so as to perform data clustering processing and determine a target clustering model; solving the target clustering model and determining the coefficient matrix.
[0009] In one embodiment, a preset segmentation model is used to perform graph cutting processing on a coefficient matrix to determine a target superpixel image, including: solving the coefficient matrix dimension by dimension using a preset weight calculation model and a mean calculation model to obtain a similarity matrix; performing similarity analysis processing on the similarity matrix using the preset segmentation model to determine a cutting judgment value, and performing graph cutting processing on the similarity matrix based on the cutting judgment value to determine the target superpixel image.
[0010] In one embodiment, a similarity analysis is performed on a similarity matrix using a preset segmentation model to determine a cutting judgment value, including: taking each multidimensional input data as a node set, performing similarity analysis on the similarity weights between any two nodes in the node set, determining a target similarity weight, and determining the target similarity weight as a cutting judgment value between nodes.
[0011] In a second aspect, an embodiment of the present invention further provides a superpixel segmentation device based on differential area content, the device comprising: a coarse segmentation module, which acquires an initial image and performs coarse segmentation processing on the initial image to determine multidimensional input data, wherein the multidimensional input data is used to represent the differential area content of the image; a feature extraction module, which performs feature extraction processing on the multidimensional input data to extract a multidimensional differential feature matrix of the multidimensional input data, wherein the multidimensional differential feature matrix includes color features, texture features and edge features of the multidimensional input data; a superpixel segmentation module, which performs data clustering processing on the multidimensional differential feature matrix through a preset clustering model based on a multi-view tensor, determines a coefficient matrix, and performs graph cutting processing on the coefficient matrix using a preset segmentation model to determine a target superpixel image.
[0012] In a third aspect, an embodiment of the present invention further provides a server, comprising a processor and a memory, wherein the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement any one of the methods provided in the first aspect.
[0013] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement any one of the methods provided in the first aspect.
[0014] The embodiments of the present invention bring the following beneficial effects:
[0015] The embodiment of the present invention provides a superpixel segmentation method, device and server based on differential regional content. After acquiring the initial image, the method performs a coarse segmentation process on the initial image, determines the multi-dimensional input data, and performs feature extraction process on the multi-dimensional input data to extract the multi-dimensional differential feature matrix of the multi-dimensional input data. Finally, a data clustering process is performed on the multi-dimensional differential feature matrix through a preset clustering model based on a multi-view tensor to determine the coefficient matrix, and a graph cutting process is performed on the coefficient matrix using a preset segmentation model to determine the target superpixel image. The embodiment of the present invention designs an MDTAL model (i.e., a clustering model). By introducing the tensor nuclear norm into the model, the multi-dimensional differential feature matrix can be effectively processed. At the same time, block diagonal constraints are adopted for the multi-view data so that each dimensional view is conducive to approaching the accurate block angle structure, and spatial local characteristics are introduced into each view to improve the aggregation of the content of different block regions. In addition, the watershed transform is used to perform a coarse segmentation of the initial image, and multi-features are used to extract data information after residual image completion, thereby improving the superpixel segmentation effect.
[0016] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description, claims and drawings.
[0017] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art are briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0019] Figure 1 A schematic flow chart of a superpixel segmentation method based on differential region content provided by an embodiment of the present invention;
[0020] Figure 2 A schematic flow chart of another superpixel segmentation method based on differential region content provided by an embodiment of the present invention;
[0021] Figure 3 A schematic diagram of the structure of a superpixel segmentation device based on differential region content provided by an embodiment of the present invention;
[0022] Figure 4A schematic diagram of the structure of a server provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution of the present invention will be clearly and completely described in combination with the embodiments below. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0024] At present, in traditional computer vision image segmentation, a single pixel is usually regarded as a processing unit. This segmentation method easily increases computing time and resource consumption. Related technologies propose that superpixel segmentation can be used to merge multiple pixels in a small area into a superpixel block based on the similarity of pixels and their neighborhoods, and then use the superpixel block as the basic unit for subsequent segmentation processing based on gradient or graph theory methods: the gradient-based method starts from the perspective of clustering, so that the pixels in the local area are clustered in the direction of the largest gradient change, and iteratively updates the cluster center and the pixel attributes until the stopping condition is met. Its representative methods include simple linear iterative clustering (SLIC), mean shift, level set and watershed; the graph theory-based method is to construct a heterogeneous undirected graph, regard image pixels as nodes, and the edges between nodes represent difference values. Then, according to different cutting criteria, the graph is segmented to achieve superpixels. Its representative methods include normalized cut sets, boundary evolution and superpixel network lattice method.
[0025] Superpixel segmentation technology improves the efficiency of image processing, reduces the computational burden, and retains important visual information. However, the above-mentioned image segmentation method still has the problem of insufficient utilization of the information of the initial image and failure to effectively integrate the differential features. Since its cutting criterion is more dependent on a specific affinity graph model, it leads to insufficient adaptability, and when processing the initial image, the segmentation performance of the target boundary is poor. In addition, factors such as large pixel differences between multiple targets, blurred differences between targets, and internal and external deviations may have a significant impact on the segmentation effect of the above-mentioned segmentation scheme. Based on this, the superpixel segmentation method based on differential regional content provided by the implementation of the present invention can improve the universality of the superpixel segmentation method so that it can process multi-dimensional data, and effectively utilize the complementary characteristics of multi-dimensional differential feature data to improve the judgment of irregular content boundaries, thereby better distinguishing content areas and making them close to the original contours, thereby improving segmentation efficiency and accuracy.
[0026] See also Figure 1 The flowchart of a superpixel segmentation method based on difference region content is shown in FIG. 1 , and the method mainly includes the following steps S102 to S106:
[0027] Step S102, obtaining an initial image, and performing a rough segmentation process on the initial image to determine multi-dimensional input data, wherein the multi-dimensional input data is used to represent the differential regional content of the image. In one embodiment, a watershed transform can be used to perform a rough segmentation on the initial image, and after morphological dilation and erosion operations, a plurality of irregular catchment areas and watersheds are formed, that is, regional content and boundaries, and then the residual image of a single area is completed to form multi-dimensional input data, wherein the marked image is the initial image composed of pixels, and the mask image is the image after the initial image is expanded. The erosion operation can eliminate the boundary points of the object and reduce the target, and can be used to eliminate noise points smaller than the structural elements. The dilation operation merges all background points in contact with the object into the object to increase the target, and can be used to fill the holes in the target.
[0028] Step S104, performing feature extraction processing on the multidimensional input data to extract a multidimensional difference feature matrix of the multidimensional input data, wherein the multidimensional difference feature matrix includes target color features, target texture features, and target edge features of the multidimensional input data. In one embodiment, multi-feature information extraction can be performed on each dimension of input data one by one, and target color features, target texture features, and target edge features can be extracted respectively, and the target color features, target texture features, and target edge features are determined as a multidimensional difference feature matrix.
[0029] Step S106, using a preset clustering model based on multi-view tensors, performs data clustering processing on the multi-dimensional difference feature matrix to determine the coefficient matrix, and uses a preset segmentation model to perform graph cutting processing on the coefficient matrix to determine the target super-pixel image, wherein the target super-pixel image contains multiple pixel blocks, and in an ideal segmentation state, each pixel block contains the same content. For example, the initial image contains houses and trees, and each pixel block of the target super-pixel image only contains houses or only contains trees, and there is no house and tree co-existing in the same pixel block.
[0030] The above-mentioned superpixel segmentation method based on differential area content provided by the embodiment of the present invention improves the universality of the superpixel segmentation method so that it can process multi-dimensional data, and effectively utilizes the complementary characteristics of multi-dimensional differential feature data to improve the judgment of irregular boundaries of content, thereby better distinguishing content areas and making them close to the original contours, thereby improving segmentation efficiency and accuracy.
[0031] See also Figure 2 The flowchart of another superpixel segmentation method based on differential region content is shown in FIG. 1 . The embodiment of the present invention also provides an implementation method of superpixel segmentation, and the details are as follows (1) to (4):
[0032] (1) The initial image to be processed is subjected to a watershed transformation, and after the rough segmentation, the single-region residual image is completed to meet the same size to form multi-dimensional input data. That is to say, the initial image can be roughly segmented by a watershed transformation to determine the marker image and mask image corresponding to the initial image, and the marker image and mask image are morphologically expanded and eroded to determine the set of sheet regions. After that, the remaining image is completed for each sheet region (that is, only the block area is contained in a single image, for example, only the tree area in the initial image or only the house area in the initial image) to complete the blank areas in each sheet region to the same value, and the completed images are accumulated as multi-dimensional input data.
[0033] Specifically, in the watershed transformation process, the initial image is changed into two input images, namely the marked image and the mask image. After morphological dilation and erosion operations, different sheet-like areas are formed, and the remaining images are completed for the contents of different blocks. That is, a single image only contains the block area, and the rest are the same value. Then the images generated by all blocks are stacked to form multi-dimensional input data.
[0034] In one embodiment, the expression of the reconstructed morphology after the morphological dilation and erosion operations is:
[0035]
[0036] Among them, q and h are two input images, o represents the basic morphological dilation operation, c represents the basic morphological erosion operation, R represents morphological reconstruction, and j represents point-by-point pixels.
[0037] (2) The extraction of target color features, target texture features and target edge features specifically includes the following (A) to (C):
[0038] (A) The color features of the multi-dimensional input data under different color models are extracted respectively, and the target color features are determined by fusing the extracted color features. Specifically, the multi-dimensional input data can be converted from the original RGB color model to the HSV color model and the YUV color model respectively, wherein the HSV color model is a conical color model and the YUV color model is a linear color model.
[0039] Furthermore, the HSV color model is quantized and divided into various color regions according to preset angle intervals to convert the multidimensional input data into a binary color index set. That is to say, the original RGB color model is converted into an HSV color model that conforms to human vision and quantized into multiple bins, and then divided into several regions. Each color model is indexed by a color component in the color space, and the image is expressed as a binary color index set, wherein each color region includes the index of the corresponding color component in the color space, and bins are the domains of the HSV color model divided by angle. Finally, the brightness component of the YUV color model is determined as the grayscale feature of the multidimensional input data, and the grayscale feature is combined with the color index set to be determined as the target color feature.
[0040] (B) Extracting three index parameters of texture features, mainly considering the average signal of non-overlapping windows of different sizes to calculate the roughness, using variance and peak value to define the contrast, and using the distribution angle of the frequency-oriented local edge relative to its direction to describe the directionality. That is to say, the average signal, variance, peak value and distribution angle of the frequency-oriented local edge relative to its direction of non-overlapping windows of different sizes in the multi-dimensional input data are obtained, and the roughness is determined based on the average signal, the contrast is determined based on the variance and peak value, and the directionality is determined based on the distribution angle through the preset texture feature extraction model, and the roughness, contrast and directionality are determined as the target edge features. Among them, a window (i.e., a 2^K pixel block) is used to sequentially slide (shift row by row) on the image to be processed to calculate the average grayscale value in the 2^K neighborhood of a single pixel point. A non-overlapping window is a window that has no intersecting or overlapping areas in the horizontal and vertical directions after sequential sliding.
[0041] (C) By using a preset edge feature extraction model, edge feature extraction is performed on multi-dimensional input data based on second-order derivatives to determine target edge features. In one embodiment, when edge feature extraction is performed on multi-dimensional input data using second-order derivatives, the Laplacian operator can be used to add the second-order partial derivatives of the horizontal and vertical coordinate values, and the Canny edge detection operator can be used to calculate the gradient strength and direction after Gaussian smoothing and denoising, and double threshold screening can be performed after non-maximum suppression to finally determine the target edge features.
[0042] (3) Design a clustering method based on multi-view tensor to perform clustering analysis on differential feature data, and input the obtained multi-dimensional differential feature matrix as the data matrix into the designed clustering model based on multi-view tensor to solve the representation coefficient matrix, so as to achieve the difference and identity of multi-dimensional data and solve a better multi-view representation coefficient matrix. Specifically, the local characteristic items of multi-dimensional input data can be obtained, and the multi-dimensional differential feature matrix and spatial neighborhood attributes can be substituted into the preset clustering model based on multi-view tensor, and the clustering after the data is substituted is analyzed. The model applies diagonalization terms and tensor nuclear norm to perform data clustering processing and determine the target clustering model. Finally, the target clustering model is solved to determine the coefficient matrix, where the local characteristic term is used to represent the spatial neighborhood attributes contained in the image itself. The nuclear norm of the tensor refers to the sum of the first k singular values of the matrix obtained by converting the tensor into a matrix and performing singular value decomposition on it, where k is a given positive integer. The nuclear norm is an important indicator to measure the low-rank nature of the tensor and is also the optimization target of many tensor decomposition algorithms. It can be applied to tasks such as data dimensionality reduction and feature extraction.
[0043] In one embodiment, a novel multi-view tensor-based clustering (MDTAL) model is designed to cluster multi-dimensional difference feature data, and the multi-dimensional difference feature matrix in the above step is used as input data. At the same time, considering the spatial neighborhood attributes contained in the image itself, local characteristics are incorporated into each view to characterize spatial correlation, and block diagonalization is applied to the representation coefficient matrix under each view ( ), the tensor kernel norm is applied to the stacked data ( ), the specific expression of the model is as follows:
[0044]
[0045] in, is the input data under the i-th view, is the representation coefficient matrix under the i-th view, is the noise term under the i-th view, is the local characteristic item under the i-th view, is the i-th view weight matrix, evaluating spatial correlation, Z is a tensor formed by stacking multiple views, is the balance parameter between different items.
[0046] (4) Using a spectral clustering algorithm to perform Ncut segmentation on the constructed similarity graph matrix to generate superpixel blocks, the similarity graph matrix is calculated according to a construction method such as weight calculation or mean calculation, and the matrix is used as input. Graph cutting is performed using Ncut, and different difference area contents are used as node sets. The similarity weight value between two points is used as the cutting judgment value, so as to obtain a superpixel image. Specifically, the coefficient matrix can be solved dimension by dimension using a preset weight calculation model and a mean calculation model to obtain a similarity matrix. Then, the similarity matrix is subjected to similarity analysis processing using a preset segmentation model to determine the cutting judgment value, and the similarity matrix is subjected to graph cutting processing based on the cutting judgment value to determine the target superpixel image. In one embodiment, each multidimensional input data is used as a node set, and similarity analysis processing is performed on the similarity weights between any two nodes in the node set to determine the target similarity weight, and the target similarity weight is determined as the cutting judgment value between the nodes.
[0047] In summary, the present invention can effectively reduce under-segmented areas. After the initial image is roughly segmented through watershed change, morphological corrosion and expansion operations are performed on the contents of different block areas, and multi-dimensional input data is formed through residual image completion, so as to more accurately segment the contents of different block areas. In addition, multi-dimensional difference feature data is adopted, three types of features are extracted, and color, edge, and texture feature information are fused to obtain multi-feature data, thereby improving the accuracy of super-pixel segmentation of the initial image, making the super-pixel blocks more uniform and the boundaries fit the regional edges in the original image as much as possible. Finally, the present invention also designs a new multi-view tensor clustering model, introduces the tensor nuclear norm to process multi-view data, applies block diagonal constraints to each view, and incorporates spatial local characteristics. At the same time, the deep complementary information between the data is utilized, so that single pixel points in the contents of different block areas can be more accurately assigned to the corresponding super-pixel blocks, so as to improve the super-pixel segmentation effect.
[0048] For the superpixel segmentation method based on difference region content provided in the above embodiment, the embodiment of the present invention provides a superpixel segmentation device based on difference region content, see Figure 3 A schematic diagram of the structure of a superpixel segmentation device based on difference region content is shown, and the device includes the following parts:
[0049] A coarse segmentation module 302 acquires an initial image and performs coarse segmentation processing on the initial image to determine multi-dimensional input data, wherein the multi-dimensional input data is used to represent the difference area content of the image;
[0050] The feature extraction module 304 performs feature extraction processing on the multi-dimensional input data to extract a multi-dimensional difference feature matrix of the multi-dimensional input data, wherein the multi-dimensional difference feature matrix includes color features, texture features and edge features of the multi-dimensional input data;
[0051] The superpixel segmentation module 306 performs data clustering processing on the multi-dimensional difference feature matrix through a preset clustering model based on multi-view tensors, determines the coefficient matrix, and performs graph cutting processing on the coefficient matrix using a preset segmentation model to determine the target superpixel image.
[0052] The above-mentioned superpixel segmentation device based on differential regional content provided in the embodiment of the present application can significantly improve the superpixel segmentation efficiency and segmentation accuracy.
[0053] In one embodiment, when performing a step of performing a coarse segmentation process on the initial image and determining multi-dimensional input data, the coarse segmentation module 302 is also used to: perform a coarse segmentation process on the initial image through a watershed transform, determine a marked image and a mask image corresponding to the initial image, and perform morphological dilation and erosion processes on the marked image and the mask image to determine a set of sheet regions; perform a remaining image completion process on each sheet region to complete the blank areas in each sheet region with the same value, and accumulate the completed images as multi-dimensional input data.
[0054] In one embodiment, when performing feature extraction processing on multidimensional input data to extract a multidimensional difference feature matrix of the multidimensional input data, the feature extraction module 304 is also used to: obtain the average signal, variance, peak value and distribution angle of the frequency-oriented local edge of non-overlapping windows of different sizes in the multidimensional input data relative to its direction; extract the color features of the multidimensional input data under different color models respectively, and determine the target color features by fusing the extracted color features; determine the roughness based on the average signal, determine the contrast based on the variance and peak value, determine the directionality based on the distribution angle through a preset texture feature extraction model, and determine the roughness, contrast and directionality as target edge features; perform edge feature extraction on the multidimensional input data based on the second-order derivative through a preset edge feature extraction model to determine the target edge features; determine the target color features, target texture features and target edge features as a multidimensional difference feature matrix.
[0055] In one embodiment, when extracting color features of multidimensional input data under different color models respectively and determining target color features by fusing the extracted color features, the feature extraction module 304 is also used to: convert the multidimensional input data from the original RGB color model to the HSV color model and the YUV color model respectively, wherein the HSV color model is a conical color model and the YUV color model is a linear color model; quantize the HSV color model and divide it into various color regions according to preset angle intervals to convert the multidimensional input data into a binary color index set, wherein each color region includes an index of a corresponding color component in the color space; determine the brightness component of the YUV color model as a grayscale feature of the multidimensional input data, and combine the grayscale feature with the color index set to determine it as the target color feature.
[0056] In one embodiment, when performing data clustering processing on the multi-dimensional difference feature matrix through a preset clustering model based on a multi-view tensor and determining the step of the coefficient matrix, the above-mentioned superpixel segmentation module 306 is also used to: obtain local characteristic items of the multi-dimensional input data, wherein the local characteristic items are used to represent the spatial neighborhood attributes contained in the image itself; substitute the multi-dimensional difference feature matrix and the spatial neighborhood attributes into the preset clustering model based on a multi-view tensor, and apply diagonalization terms and tensor nuclear norms to the clustering model after the data is substituted to perform data clustering processing and determine the target clustering model; solve the target clustering model and determine the coefficient matrix.
[0057] In one embodiment, when performing graph cutting processing on the coefficient matrix using a preset segmentation model to determine the target superpixel image, the superpixel segmentation module 306 is also used to: solve the coefficient matrix dimension by dimension using a preset weight calculation model and a mean calculation model to obtain a similarity matrix; perform similarity analysis processing on the similarity matrix using the preset segmentation model to determine a cutting judgment value, and perform graph cutting processing on the similarity matrix based on the cutting judgment value to determine the target superpixel image.
[0058] In one embodiment, when performing a similarity analysis on a similarity matrix using a preset segmentation model to determine a cutting judgment value, the superpixel segmentation module 306 is also used to: take each multidimensional input data as a node set, perform a similarity analysis on the similarity weights between any two nodes in the node set, determine a target similarity weight, and determine the target similarity weight as a cutting judgment value between nodes.
[0059] The device provided in the embodiment of the present invention has the same implementation principle and technical effects as those of the aforementioned method embodiment. For the sake of brief description, for matters not mentioned in the device embodiment, reference may be made to the corresponding contents in the aforementioned method embodiment.
[0060] An embodiment of the present invention provides a server. Specifically, the server includes a processor and a storage device. The storage device stores a computer program, and when the computer program is executed by the processor, it executes the method described in any one of the above-mentioned embodiments.
[0061] Figure 4 A structural diagram of a server provided in an embodiment of the present invention, the server 100 includes: a processor 40, a memory 41, a bus 42 and a communication interface 43, wherein the processor 40, the communication interface 43 and the memory 41 are connected via the bus 42; the processor 40 is used to execute an executable module stored in the memory 41, such as a computer program.
[0062] The memory 41 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 43 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used.
[0063] The bus 42 may be an ISA bus, a PCI bus, or an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0064] Among them, the memory 41 is used to store programs, and the processor 40 executes the program after receiving the execution instruction. The method executed by the device for flow process definition disclosed in any embodiment of the above-mentioned embodiment of the present invention can be applied to the processor 40 or implemented by the processor 40.
[0065] The processor 40 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 40. The above processor 40 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present invention can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in the embodiment of the present invention can be directly embodied as a hardware decoding processor to be executed, or the hardware and software modules in the decoding processor can be executed. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 41, and the processor 40 reads the information in the memory 41 and completes the steps of the above method in combination with its hardware.
[0066] The computer program product of the readable storage medium provided in the embodiment of the present invention includes a computer-readable storage medium storing program code, and the instructions included in the program code can be used to execute the methods described in the previous method embodiments. The specific implementation can be referred to the previous method embodiments, which will not be repeated here.
[0067] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program codes.
[0068] Finally, it should be noted that the above-described embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above-described embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-described embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A superpixel segmentation method based on differential region content, characterized in that: The method comprises: Acquire an initial image, and perform a rough segmentation process on the initial image to determine multi-dimensional input data, wherein the multi-dimensional input data is used to represent the difference area content of the image; Performing feature extraction processing on the multidimensional input data to extract a multidimensional difference feature matrix of the multidimensional input data, wherein the multidimensional difference feature matrix includes target color features, target texture features, and target edge features of the multidimensional input data; Performing data clustering processing on the multi-dimensional difference feature matrix through a preset clustering model based on multi-view tensors to determine a coefficient matrix, and performing graph cutting processing on the coefficient matrix using a preset segmentation model to determine a target superpixel image; The step of performing data clustering processing on the multi-dimensional difference feature matrix and determining the coefficient matrix by using a preset clustering model based on multi-view tensors includes: obtaining local characteristic items of the multi-dimensional input data, wherein the local characteristic items are used to represent the spatial neighborhood attributes contained in the image itself; substituting the multi-dimensional difference feature matrix and the spatial neighborhood attributes into the preset clustering model based on multi-view tensors, and applying diagonalization items and tensor nuclear norms to the clustering model after the data is substituted to perform data clustering processing and determine the target clustering model; solving the target clustering model to determine the coefficient matrix; The specific expression of the clustering model based on multi-view tensor is as follows: in, is the input data under the i-th view, is the representation coefficient matrix under the i-th view, is the noise term under the i-th view, is the local characteristic item under the i-th view, is the i-th view weight matrix, used to evaluate spatial correlation, Z is a tensor formed by stacking multiple views, is the balance parameter between different items.
2. The superpixel segmentation method based on difference region content according to claim 1, characterized in that: The step of performing a rough segmentation process on the initial image to determine multi-dimensional input data comprises: Performing a rough segmentation process on the initial image by watershed transformation to determine a marked image and a mask image corresponding to the initial image, and performing morphological expansion and corrosion processes on the marked image and the mask image to determine a set of sheet regions; The remaining image completion processing is performed for each sheet-like region to complete the blank areas in each sheet-like region to the same value, and the completed images are accumulated as the multi-dimensional input data.
3. The superpixel segmentation method based on difference region content according to claim 1, characterized in that: The step of performing feature extraction processing on the multi-dimensional input data to extract a multi-dimensional difference feature matrix of the multi-dimensional input data includes: Obtaining the average signal, variance, peak value and distribution angle of frequency-oriented local edges relative to their directions for non-overlapping windows of different sizes in the multi-dimensional input data; Extracting color features of the multi-dimensional input data under different color models respectively, and determining the target color features by fusing the extracted color features; By using a preset texture feature extraction model, the roughness is determined based on the average signal, the contrast is determined based on the variance and the peak value, the directionality is determined based on the distribution angle, and the roughness, the contrast and the directionality are determined as the target edge feature; By using a preset edge feature extraction model, edge feature extraction is performed on the multi-dimensional input data based on second-order derivatives to determine the target edge feature; The target color feature, the target texture feature and the target edge feature are determined as the multi-dimensional difference feature matrix.
4. The superpixel segmentation method based on difference region content according to claim 3, characterized in that: The step of respectively extracting the color features of the multi-dimensional input data under different color models and determining the target color features by fusing the extracted color features comprises: Convert the multi-dimensional input data from the original RGB color model to the HSV color model and the YUV color model, respectively, wherein the HSV color model is a conical color model and the YUV color model is a linear color model; Quantizing the HSV color model, dividing it into various color regions according to preset angle intervals, so as to convert the multi-dimensional input data into a binary color index set, wherein each color region includes an index of a corresponding color component in a color space; The brightness component of the YUV color model is determined as the grayscale feature of the multi-dimensional input data, and the grayscale feature is combined with the color index set to determine the target color feature.
5. The superpixel segmentation method based on difference region content according to claim 1, characterized in that: The step of performing graph cutting processing on the coefficient matrix using a preset segmentation model to determine a target superpixel image comprises: The coefficient matrix is solved dimension by dimension by a preset weight calculation model and a mean calculation model to obtain a similarity matrix; A preset segmentation model is used to perform similarity analysis processing on the similarity matrix to determine a cutting judgment value, and a graph cutting processing is performed on the similarity matrix based on the cutting judgment value to determine the target superpixel image.
6. The superpixel segmentation method based on difference region content according to claim 5, characterized in that: The step of performing similarity analysis on the similarity matrix using a preset segmentation model to determine a cutting judgment value comprises: Each item of the multidimensional input data is taken as a node set, and similarity analysis is performed on the similarity weights between any two nodes in the node set to determine a target similarity weight, and the target similarity weight is determined as the cutting judgment value between the nodes.
7. A superpixel segmentation device based on differential region content, characterized in that: The device comprises: A coarse segmentation module, which obtains an initial image and performs coarse segmentation processing on the initial image to determine multi-dimensional input data, wherein the multi-dimensional input data is used to represent the difference area content of the image; A feature extraction module performs feature extraction processing on the multidimensional input data to extract a multidimensional difference feature matrix of the multidimensional input data, wherein the multidimensional difference feature matrix includes color features, texture features and edge features of the multidimensional input data; The superpixel segmentation module performs data clustering processing on the multi-dimensional difference feature matrix through a preset clustering model based on multi-view tensors to determine a coefficient matrix, and performs graph cutting processing on the coefficient matrix using a preset segmentation model to determine a target superpixel image; The step of performing data clustering processing on the multi-dimensional difference feature matrix and determining the coefficient matrix by using a preset clustering model based on multi-view tensors includes: obtaining local characteristic items of the multi-dimensional input data, wherein the local characteristic items are used to represent the spatial neighborhood attributes contained in the image itself; substituting the multi-dimensional difference feature matrix and the spatial neighborhood attributes into the preset clustering model based on multi-view tensors, and applying diagonalization items and tensor nuclear norms to the clustering model after the data is substituted to perform data clustering processing and determine the target clustering model; solving the target clustering model to determine the coefficient matrix; The specific expression of the clustering model based on multi-view tensor is as follows: in, is the input data under the i-th view, is the representation coefficient matrix under the i-th view, is the noise term under the i-th view, is the local characteristic item under the i-th view, is the i-th view weight matrix, used to evaluate spatial correlation, Z is a tensor formed by stacking multiple views, is the balance parameter between different items.
8. A server, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Natural image segmentation method and system based on tensor subspace clustering
CN114782688A
Image segmentation method and device, vehicle and storage medium
CN115512145A