Image data enhancement method and system

By segmenting the environmental background image in the data enhancement method and stitching the images according to the foreground object categories, the problem of poor generalization ability of environmental changes in the prior art is solved. The generated data enhancement image is closer to the actual environment, and the generalization ability of the deep learning network is improved.

CN115170839BActive Publication Date: 2025-06-06BEIJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210849830.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-19
Publication Date
2025-06-06
Estimated Expiration
2042-07-19

AI Technical Summary

Technical Problem

Existing data enhancement methods have poor generalization capabilities when dealing with environmental changes and fail to fully consider environmental characteristics.

Method used

By obtaining the labeled image and environmental background image, segmenting the environmental background image to obtain the sky and ground area, cropping the labeled image to obtain the target area image, and separating the foreground and background by clustering, splicing it into the corresponding environmental area according to the category of the foreground object, and generating data augmented images.

Benefits of technology

The generalization ability of the data augmentation method for different environments is improved, and the generated data augmentation images are more in line with the actual environment, which enhances the generalization ability of the deep learning network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115170839B_ABST
    Figure CN115170839B_ABST
Patent Text Reader

Abstract

The present invention discloses an image data enhancement method and system, which relates to the field of image processing, and the method comprises: obtaining an annotated image and an environmental background image; segmenting the environmental background image to obtain a first sky region and a first ground region in the environmental background image; cropping the region annotated as a foreground object in the annotated image to obtain a target region image; clustering each pixel in the target region image to separate the foreground and the background, and obtaining a foreground image including only the foreground object; determining whether to splice the foreground image in the first ground region or the first sky region in the environmental background image according to the category of the object in the foreground image, and finally obtaining a data enhanced image. The present invention can obtain a data enhanced image that takes environmental characteristics into consideration, improves the generalization ability of the data enhancement method for different environments, and solves the problem that the generalization ability of the existing data enhancement method is poor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and in particular to an image data enhancement method and system. Background Art

[0002] In the supervised learning classification task of machine learning, some labeled data needs to be input for training to meet the classification requirements. In the case of insufficient data, data enhancement is required. Data enhancement is a method to make limited data produce the value equivalent to more data without increasing the amount of data. Data enhancement methods include supervised and unsupervised methods. Supervised learning uses preset data transformation rules to perform data augmentation on existing data. Supervised learning methods include geometric transformation and color transformation methods for single-sample data enhancement; linear interpolation methods for multi-sample data enhancement obtain new samples and increase the proportion of small sample data in the total data. In the unsupervised data learning method, GAN based on model learning is used to directly generate similar images, or to find the best image transformation strategy based on the data itself. However, the current data enhancement method mainly makes simple changes to the data without considering environmental changes and environmental adaptability, and has poor generalization ability for environmental changes. Summary of the invention

[0003] The purpose of the present invention is to provide an image data enhancement method and system, which can obtain a data enhanced image that takes environmental characteristics into consideration, thereby improving the generalization ability of the data enhancement method for different environments.

[0004] To achieve the above object, the present invention provides the following solutions:

[0005] A method for image data enhancement, the method comprising the following steps:

[0006] Acquire a labeled image and an environmental background image, wherein the labeled image is an image obtained by labeling the foreground object in the original image;

[0007] Segmenting the environment background image to obtain a first sky area and a first ground area in the environment background image;

[0008] Cropping the area marked as the foreground object in the marked image to obtain a target area image;

[0009] Clustering each pixel in the target area image to separate the foreground and the background, and obtaining a foreground image including only the foreground object;

[0010] Determining whether the object in the foreground image is a ground object;

[0011] If yes, the foreground image is spliced ​​into the first ground region in the environment background image to obtain a data enhanced image;

[0012] If not, the foreground image is spliced ​​into the first sky area in the environment background image to obtain a data enhanced image.

[0013] The present invention also provides an image data enhancement system, the system comprising:

[0014] An image acquisition unit, used to acquire a marked image and an environment background image, wherein the marked image is an image obtained by marking a foreground object in the original image;

[0015] An image segmentation unit, used to segment the environmental background image to obtain a sky area and a ground area in the environmental background image;

[0016] A target area image acquisition unit, used for cropping the area marked as a foreground object in the marked image to obtain a target area image;

[0017] a foreground image acquisition unit, configured to cluster pixels in the target area image to separate the foreground from the background, and acquire a foreground image including only the foreground object;

[0018] A judging unit, used for judging whether the object in the foreground image is a ground object;

[0019] A first splicing unit is used for splicing the foreground image into a ground area in the environment background image to obtain a data enhanced image when the judgment result is yes;

[0020] The second splicing unit is used to splice the foreground image into the sky area in the environment background image to obtain a data enhanced image when the judgment result is no.

[0021] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0022] The present invention discloses an image data enhancement method and system, the method comprising the following steps: obtaining an annotated image and an environmental background image, the annotated image being an image obtained after annotating a foreground object in an original image; segmenting the environmental background image to obtain a first sky region and a first ground region in the environmental background image; cropping the region annotated as a foreground object in the annotated image to obtain a target region image; clustering each pixel point in the target region image to separate the foreground and the background, and obtaining a foreground image including only the foreground object; determining whether the object in the foreground image is a ground object; if so, splicing the foreground image into the first ground region in the environmental background image to obtain a data enhanced image; if not, splicing the foreground image into the first sky region in the environmental background image to obtain a data enhanced image. The present invention obtains a sky region and a ground region in the environmental background image by segmenting the environmental background image, obtains a foreground image in the region of the foreground object by clustering, and then splices the foreground image into the ground region or the sky region according to the category of the object in the foreground image, and finally obtains a data enhanced image. Compared with the existing technology, it does not simply change the data, but takes the environmental characteristics into consideration, thus improving the generalization ability of the data enhancement method for different environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0024] Figure 1 A flowchart of an image data enhancement method provided in Embodiment 1 of the present invention;

[0025] Figure 2 Schematic diagram of the division of the optimal boundary line and the suboptimal boundary line;

[0026] Figure 3 This is a structural block diagram of an image data enhancement system provided in Example 2 of the present invention. DETAILED DESCRIPTION

[0027] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0028] The purpose of the present invention is to provide an image data enhancement method and system, which can obtain a data enhanced image that takes environmental characteristics into consideration, thereby improving the generalization ability of the data enhancement method for different environments.

[0029] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0030] Embodiment 1:

[0031] See also Figure 1 The present invention provides an image data enhancement method, the method comprising the following steps:

[0032] S1: Acquire a labeled image and an environmental background image, wherein the labeled image is an image obtained by labeling a foreground object in the original image;

[0033] S2: Segmenting the environment background image to obtain a first sky area and a first ground area in the environment background image;

[0034] S3: Crop the area marked as the foreground object in the marked image to obtain a target area image;

[0035] S4: clustering each pixel in the target area image to separate the foreground and the background, and obtaining a foreground image including only the foreground object;

[0036] S5: Determine whether the object in the foreground image is a ground object;

[0037] S6: If yes, then stitching the foreground image into the first ground region in the environment background image to obtain a data enhanced image;

[0038] S7: If not, splicing the foreground image into the first sky area in the environment background image to obtain a data enhanced image.

[0039] In the process of data enhancement, the method of fusing the foreground image in the target area image with the environmental background image is adopted. In order to ensure that the results after data enhancement are consistent with life logic and no unnecessary data affects the training effect, it is necessary to classify the sky and ground areas of the environmental background image and limit the data enhancement area. For example, vehicles, trees and other ground bosses, pits and other obstacles will not appear in the sky.

[0040] Considering that it is difficult to train the sky and ground segmentation using a deep learning network, it is necessary to collect enough samples for training in different environments, which violates the original intention of data enhancement to be simple and easy to use. The present invention uses a method based on numerical calculation to segment the sky and ground in the environmental background image. First, the pixel changes in the sky area are not obvious, which is much lower than the pixel changes in other areas of the ground. The change threshold can be calculated by mathematical methods to determine whether each column of pixels in the sky scanning environmental background image appears, and the gradient threshold of the pixel change is adjusted to obtain the boundary between the sky and the ground, and the optimal boundary is continuously optimized.

[0041] Specifically, in step S2, the environment background image is segmented to obtain a first sky area and a first ground area in the environment background image, which specifically includes:

[0042] The environment background image is subjected to matrix transformation to obtain a two-dimensional matrix of the image. A ;

[0043] The two-dimensional matrix of the image A conduct sobel The convolution calculation of the operator obtains the image gradient map, which is composed of the horizontal grayscale gradient map G x And the vertical gray gradient map G y constitute.

[0044] Among them, the gray gradient map G x The calculation formula is:

[0045]

[0046] Vertical grayscale gradient map G y The calculation formula is:

[0047]

[0048] Setting preset thresholds K , for the entire grayscale gradient image G x and G y and scanning from top to bottom to obtain the pixel value and position of each pixel point in the image gradient map;

[0049] Calculate the difference between the pixel values ​​of each pixel and its adjacent pixels respectively; the calculation formula of the difference is as follows:

[0050] between =| C now - Cnear |

[0051] in, between is the difference, C now is the pixel value of the current pixel, C near is the pixel value of the pixel point adjacent to the current pixel point. In this embodiment, the adjacent pixel points are pixel points with a straight line interval of 1 pixel, and there are 8 adjacent pixel points in total.

[0052] If the difference is not within the preset threshold K The pixels within the range are regarded as uneven pixels, and the positions of all uneven pixels are recorded ( x i , y i );

[0053] Connect all neighboring uneven pixel points to obtain several boundary lines, where the neighboring uneven pixel points are two uneven pixel points within 3 pixels of each other in Euclidean distance; specifically, connect the pixel points that meet the conditions from left to right and from top to bottom according to the pixel point positions to form a boundary line, and store the boundary line in the change boundary array edge As shown in the following formula:

[0054] ifbetween > K

[0055] Then ( x i , y i ) ofbetween Deposit edge

[0056] According to the uneven pixels, considering that the blue channel value of the sky pixel is relatively large, an energy function is constructed, and the energy function is a function related to the single pixel value of the blue channel of the upper and lower pixels of the boundary line and the upper and lower RGB channel colors of the boundary line; the calculation formula of the energy function is as follows:

[0057]

[0058] in, P edge is the energy function value, γ For variable parameters, n 1 is the number of pixels of the boundary line length, S b is the single pixel value of the blue channel of the pixel above the boundary line, G bThe single pixel value of the blue channel of the pixel below the boundary line is calculated based on the number of pixels above the boundary line. n 2 Calculate the average, S sky is the RGB channel color of the sky color, S It is the RGB channel color above the boundary line.

[0059] According to the energy function, calculating the energy value corresponding to each boundary line;

[0060] Selecting the boundary line with the largest energy value as the temporary optimal boundary line;

[0061] As shown in the following formula, for the energy values ​​of different boundary lines, the maximum energy value is found. The boundary line corresponding to the maximum energy value indicates that the average pixel value of the blue channel of the RGB pixels above the boundary line is higher. At the same time, it is not much different from the channel of the known sky color. At this time, the maximum energy value indicates that the pixel similarity in the sky set and the ground set is the highest. The boundary line at this time is called the temporary optimal boundary line true edge ;

[0062] true edge = edge [arg max( P edge )]

[0063] The environment background image is segmented according to the temporary optimal boundary line to obtain a first sky area and a first ground area in the environment background image.

[0064] However, since the boundary is segmented by a gradient threshold each time a scan is performed, there may be large threshold jumps in the presence of clouds, power lines, etc. For example, the optimal boundary line appears at the junction of the sky and the power lines, but there is also a sky area below the power lines. If only the optimal boundary line is calculated, a large number of sky boundaries will be discarded. In this case, the optimal boundary line and the suboptimal boundary line are selected according to the energy value of the boundary line. The suboptimal boundary line is the boundary line corresponding to the energy value that is only less than the maximum energy value. Figure 2 As shown, the image is divided into the first sky area M 1. Second Sky Area M 1+ M 2. The first ground area G + M 2 and second ground area G .

[0065] In order to distinguish M 1. M 1+ M2 Which is the real sky area? Calculate the first sky area respectively M 1 and the first ground area G + M 2 Mahalanobis distance and the second sky area M 1+ M 2 and the second ground area G The Mahalanobis distance is calculated as follows:

[0066]

[0067] in, D M ( m , n ) is the value of Mahalanobis distance, m is the pixel average matrix near the boundary sky, n is the pixel average matrix near the boundary line ground, is the covariance matrix, T is transposed.

[0068] The area with a large Mahalanobis distance is considered to be the real sky area ( M sure ), the rest is the ground area G If there is an area above the sky, then the area above is considered to be the sky.

[0069] Through the above method, the boundary between the sky area and the ground area is found in the environmental background image.

[0070] In the process of background separation, it is necessary to accurately identify the gradient boundary between the foreground and background with a small change threshold, and solve the problem of foreground pixel loss caused by the threshold being too close. In view of this difficulty, the present invention obtains the background color in the area of ​​the foreground object as a reference value, and combines the Kmeans clustering method to set reasonable clustering categories to extract the foreground image and the background image.

[0071] The annotated images come from the same environment and are used Kmeans Clustering method processing. The data that needs to be classified includes the foreground and background of the image. Kmeans During the clustering process, the foreground and background clusters are clustered to obtain the foreground result. The contour of the foreground is then extracted to fill in the missing points inside the foreground to enhance the extraction effect.

[0072] Specifically, the region marked as the foreground object in the marked image is cropped to obtain the target region image, wherein the mark is a rectangular mark frame surrounding the foreground object, and the rectangular mark frame includes the foreground image of the foreground object and the background image unrelated to the foreground object. In order to obtain the foreground image of the foreground object, the following method is used to separate the foreground and the background in the target region image:

[0073] Expand the area of ​​the foreground object in one dimension to obtain image expansion data Image , for a given number of clusters for foreground and background segmentation Type is 2.

[0074] use Kmeans The clustering algorithm clusters the image expansion data to obtain a clustering result, which includes: clustering result positions and cluster center arrays; the specific clustering process is as follows:

[0075] First, set the cluster labels. Clusters need to be divided into foreground and background clusters. Expand the data for the image Image Perform foreground and background clustering and expand the data by selecting an image Image Corner points as background initial center and image expansion data Image The center is used as the foreground initial center to obtain the initial center of the image data μ , using the RGB channel value of the initial center pixel as the initial value of the clustering sample vector, and comparing the RGB channel values ​​of other pixels x (i) The RGB channel values ​​of the initial center of the foreground and the initial center of the background μ rgb-j (RGB channel values ​​of the initial center of the foreground μ rgb-1 and the RGB channel values ​​of the background initial center μ rgb-2 ), calculate the difference between the two norms and store the minimum result in the foreground and background classification. c j In the array (the foreground is c 1 Array, background is c 2 Array), the calculation formula is as follows:

[0076] .

[0077] Set the number of iterations to 10 in a single Kmeans clustering process 6 Second, calculate the new foreground cluster center during the clustering process ( x c1mid , y c1mid) pixel position, the calculation formula is as follows:

[0078]

[0079] in, m 1 is the number of pixels in the foreground, ( X c1i , Y c1i ) is the use prospect c 1 The pixel location information stored in the array.

[0080] In the clustering process, new background cluster centers are calculated ( x c2mid , y c2mid ) pixel position, the calculation formula is as follows:

[0081]

[0082] in, m 2 is the number of background pixels, ( X c2i , Y c2i ) is used as background c 2 The pixel location information stored in the array.

[0083] If the number of iterations is reached or Kmeans The foreground cluster center of the clustering result ( x c1mid , y c1mid ) pixel position and image center ( μ X , μ Y ) The sum of squared errors of pixel position distances e ( i ) is less than 35 pixels, clustering is stopped; Kmeans The foreground cluster center of the clustering result ( x c1mid , y c1mid ) pixel position and image center ( μ X , μ Y ) The sum of squared errors of pixel position distances e ( i ) is calculated as follows:

[0084]

[0085] The clustering was repeated 10 times, and the energy density function was set to obtain the energy density J , select according to different environments, compare the foreground and background cluster centers with the image center ( μ X , μ Y ) The binary norm of the pixel position is selected, and the clustering information with the minimum energy density is selected as the optimal solution for output. The optimal clustering result is output, including two values: the clustering result position and the cluster center array. The calculation formula of the energy density function is as follows:

[0086]

[0087] Record the position of the optimal clustering result in the image and feed the clustered result position back to the image expansion data Image In the image, expand the data Image Make color masks with different RGB channel values, use black and white colors to distinguish, the foreground is white, the background is black, and the binary foreground image and background image are obtained. Expand the image data Image The calculation formula for color masking with different RGB channel values ​​is as follows:

[0088]

[0089] use sobel The operator performs image convolution extraction on the binary polarized foreground image to obtain foreground boundary elements;

[0090] dilating the foreground boundary element to obtain an expanded foreground boundary element, wherein the dilation is used to filter out discrete boundary interference points;

[0091] The pixels in the area enclosed by the expanded foreground boundary elements are extracted by using a mask technique to obtain a foreground image in the area of ​​the foreground object, wherein the foreground image is an image that only includes the foreground object in the area of ​​the foreground object.

[0092] As a possible implementation, in step S5, it is determined whether the object in the foreground image is a ground object, specifically including: constructing a database, the database including all objects belonging to the ground area and all objects belonging to the sky area; after extracting the foreground image in the target area image, the object in the foreground image is compared with the objects in the database, and then it is determined whether the object in the foreground image is a ground object or an object in the sky. For example: if the object in the foreground image is a vehicle, an oil drum, a person, a water horse, etc., after comparing these categories with the objects in the database, it can be known that these objects are ground objects; if the object in the foreground image is smoke, cloud, etc., after comparing these categories with the objects in the database, it can be known that these objects are objects in the sky area.

[0093] After obtaining the foreground image from a labeled image, an affine transformation is performed on the foreground image to make the data enhancement result closer to the actual situation.

[0094] Perform an affine transformation on the foreground image to obtain matrix pixel information; the calculation formula of the affine transformation is as follows:

[0095]

[0096] in, is the matrix pixel information, is the rotation matrix, ( x , y ) is the pixel position of the foreground image, is the translation matrix.

[0097] Determining the pixel position of the matrix pixel information in the first ground area specifically includes: randomly selecting a pixel area that conforms to life logic for enhancement according to the sky and ground separation results of the environment background image, and selecting the pixel points ( x label , y label ), and the new position of the upper left corner pixel of the foreground image in the first ground area or the first sky area is obtained from the foreground image coordinates after affine transformation ( x new , y new ), the calculation formula is as follows:

[0098] ( x new , y new )=( x label , y label ).

[0099] According to the new position and the determination result of the foreground object category, the matrix pixel information is respectively accumulated into the first sky area or the first ground area to obtain a data enhanced image.

[0100] By simultaneously inputting the original annotated images and the data augmented images into the deep learning network for training, the generalization ability of the deep learning network for different environments can be improved.

[0101] The present invention can obtain a data-enhanced image that takes environmental characteristics into consideration, thereby improving the generalization ability of the data enhancement method for different environments. At the same time, the data enhancement method is simple and easy to use, and a large amount of data enhancement can be completed by directly inputting the image.

[0102] In practical application, the image data enhancement method of the present invention is compared with the image data enhancement method of clipping, splicing and rotation provided by Yolo v5 based on the Yolo v5 network. In order to prove the effectiveness of the present invention in image training, the recall rate and precision rate change curves are obtained under 100 epochs of training. According to the change curve, the false positive rate and true positive rate under different accuracy thresholds are calculated, and the receiver operating characteristic curve (ROC curve / / Receiver Operating Characteristic Curve) with change characteristics is connected to represent the influence of the change of recall rate on the precision rate. In Table 1, the final regression rate, precision rate and AUR (Area Under Curve) area under the ROC curve formed around the coordinate axis of the two methods are statistically analyzed, and finally the effects of the two data enhancement methods are compared.

[0103] Table 1 Deep learning training results

[0104] Experimental data results parameter Yolo v5 original data enhancement method accuracy 89.976% Yolo v5 original data enhancement regression rate 92.456% Yolo v5 original data enhancement method AUR 0.827 Improved data enhancement method accuracy 93.234% Improved data enhancement method regression rate 98.721% Improved original data enhancement method AUR 0.983

[0105] By comparing the final regression rate, precision rate and AUR area formed by the ROC curve around the coordinate axis of two different data enhancement methods, the final effect of training the same deep learning network with the same training data under the condition of using different data enhancement methods is measured. It is found that the precision error of training using the data enhancement method proposed in the present invention is reduced by 36.36%; the regression rate error is reduced by 25.57%; and the AUR area of ​​the ROC curve is comprehensively improved by 18.86%. The use of the data enhancement method of the present invention enhances the generalization ability of the deep learning network for different environments.

[0106] Embodiment 2:

[0107] See also Figure 3 The present invention provides an image data enhancement system, the system comprising:

[0108] An image acquisition unit 1 is used to acquire a labeled image and an environment background image, wherein the labeled image is an image obtained by labeling a foreground object in an original image;

[0109] An image segmentation unit 2 is used to segment the environment background image to obtain a sky area and a ground area in the environment background image;

[0110] The target area image acquisition unit 3 is used to crop the area marked as the foreground object in the marked image to obtain the target area image;

[0111] A foreground image acquisition unit 4 is used to cluster each pixel in the target area image to separate the foreground and the background, and to acquire a foreground image including only the foreground object;

[0112] A judging unit 5, used for judging whether the object in the foreground image is a ground object;

[0113] A first splicing unit 6 is used for splicing the foreground image into the ground area in the environment background image to obtain a data enhanced image when the judgment result is yes;

[0114] The second splicing unit 7 is used for splicing the foreground image into the sky area in the environment background image to obtain a data enhanced image when the judgment result is no.

[0115] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0116] The principles and implementation methods of the present invention are described in this article using specific examples. The description of the above embodiments is only used to help understand the method and core idea of ​​the present invention. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A method for image data enhancement, It is characterized in that The following steps are involved: Acquire a labeled image and an environmental background image, wherein the labeled image is an image obtained by labeling the foreground object in the original image; Segmenting the environment background image to obtain a first sky area and a first ground area in the environment background image; Cropping the area marked as the foreground object in the marked image to obtain a target area image; Clustering each pixel in the target area image to separate the foreground and the background, and obtaining a foreground image including only the foreground object; Determining whether the object in the foreground image is a ground object; If yes, the foreground image is spliced ​​into the first ground region in the environment background image to obtain a data enhanced image; If not, the foreground image is spliced ​​into the first sky area in the environment background image to obtain a data enhanced image.

2. The image data enhancement method according to claim 1, It is characterized in that The segmenting of the environment background image to obtain a first sky area and a first ground area in the environment background image specifically includes: Performing matrix transformation on the environmental background image to obtain a two-dimensional image matrix; The image two-dimensional matrix is sobel The convolution calculation of the operator obtains the image gradient map; Scan the image gradient map to obtain the pixel value and position of each pixel point in the image gradient map; Calculate the difference between the pixel values ​​of each pixel and its adjacent pixels respectively; Pixels whose difference values ​​are not within a preset threshold range are regarded as uneven pixels, and the positions of all uneven pixels are recorded; Connect all neighboring uneven pixel points to obtain a plurality of boundary lines, wherein the neighboring uneven pixel points are two uneven pixel points whose Euclidean distance is within 3 pixels; Constructing an energy function, wherein the energy function is a function related to the single pixel value of the blue channel of the upper and lower pixel points of the boundary line and the upper and lower RGB channel colors of the boundary line; According to the energy function, calculating the energy value corresponding to each boundary line; Selecting the boundary line with the largest energy value as the temporary optimal boundary line; The environment background image is segmented according to the temporary optimal boundary line to obtain a first sky area and a first ground area in the environment background image.

3. The image data enhancement method according to claim 2, It is characterized in that After the step of "segmenting the environment background image according to the temporary optimal boundary line to obtain a first sky area and a first ground area in the environment background image", the method further includes: Determine a suboptimal boundary line according to the energy value corresponding to each boundary line, wherein the suboptimal boundary line is a boundary line corresponding to a boundary line when the energy value is only less than the maximum energy value; Segmenting the environment background image according to the suboptimal boundary line to obtain a second sky area and a second ground area; Calculating the Mahalanobis distance between the first sky area and the first ground area to obtain a first distance; Calculating the Mahalanobis distance between the second sky area and the second ground area to obtain a second distance; Determining whether the first distance is greater than the second distance; If yes, the temporary optimal boundary line is used as the optimal boundary line; If not, the suboptimal boundary line is used as the optimal boundary line; The environmental background image is segmented according to the optimal boundary line to obtain an optimal sky area and an optimal ground area.

4. The image data enhancement method according to claim 1, It is characterized in that The clustering of pixels in the target area image to separate the foreground and the background, and obtaining a foreground image including only the foreground object, specifically includes: Expanding the area of ​​the foreground object in one dimension to obtain image expansion data; use Kmeans The clustering algorithm clusters the image expansion data to obtain a clustering result, wherein the clustering result includes: a clustering result position and a cluster center array; According to the clustering result, the image expansion data is binary polarized to obtain a foreground image and a background image after binary polarization; use sobel The operator performs image convolution extraction on the binary polarized foreground image to obtain foreground boundary elements; dilating the foreground boundary element to obtain an expanded foreground boundary element, wherein the dilation is used to filter out discrete boundary interference points; The pixels in the area enclosed by the expanded foreground boundary elements are extracted using a masking technique to obtain a foreground image in the area of ​​the foreground object.

5. The image data enhancement method according to claim 1, It is characterized in that The step of splicing the foreground image into the first ground region in the environment background image to obtain a data enhanced image specifically includes: Performing affine transformation on the foreground image to obtain matrix pixel information; Determine pixel points of the matrix pixel information in the first ground area; According to the positions, the matrix pixel information is respectively accumulated into the first ground area to obtain a data enhanced image.

6. The image data enhancement method according to claim 1, It is characterized in that The step of splicing the foreground image into a first sky region in the environment background image to obtain a data enhanced image specifically includes: Performing affine transformation on the foreground image to obtain matrix pixel information; Determine pixel positions of the matrix pixel information in the first sky area; According to the positions, the matrix pixel information is respectively accumulated into the first sky area to obtain a data enhanced image.

7. The image data enhancement method according to claim 2, It is characterized in that The calculation formula of the energy function is as follows: in, P edge is the energy function value, γ For variable parameters, n 1 is the number of pixels of the boundary line length, S b is the single pixel value of the blue channel of the pixel above the boundary line, G b is the single pixel value of the blue channel of the pixel below the boundary line, S sky is the RGB channel color of the sky color, S It is the RGB channel color above the boundary line.

8. The image data enhancement method according to claim 3, It is characterized in that The calculation formula of the Mahalanobis distance is as follows: in, D M ( m , n ) is the value of Mahalanobis distance, m is the pixel average matrix near the boundary sky, n is the pixel average matrix near the boundary line ground, is the covariance matrix, T is transposed.

9. The image data enhancement method according to claim 5 or 6, It is characterized in that The calculation formula of the affine transformation is as follows: in, is the matrix pixel information, is the rotation matrix, ( x , y ) is the pixel position of the foreground image, is the translation matrix.

10. An image data enhancement system, It is characterized in that include: An image acquisition unit, used to acquire a marked image and an environment background image, wherein the marked image is an image obtained by marking a foreground object in the original image; An image segmentation unit, used to segment the environmental background image to obtain a sky area and a ground area in the environmental background image; A target area image acquisition unit, used for cropping the area marked as a foreground object in the marked image to obtain a target area image; A foreground image acquisition unit, used for clustering each pixel in the target area image to separate the foreground and the background, and acquiring a foreground image including only the foreground object; A judging unit, used for judging whether the object in the foreground image is a ground object; A first splicing unit is used for splicing the foreground image into a ground area in the environment background image to obtain a data enhanced image when the judgment result is yes; The second splicing unit is used to splice the foreground image into the sky area in the environment background image to obtain a data enhanced image when the judgment result is no.

Citation Information

Patent Citations

  • Water level monitoring method based on cluster partition and scale recognition

    US20210374466A1

  • Image background processing method, apparatus, electronic device, and storage medium

    WO2022142219A1