Weak semantic contour detection method, device, equipment and storage medium for live image
By using a pre-trained weak semantic contour detection model, the contour features in the live network image are extracted and the complete objects are identified, and the problems of low contour detection efficiency and high error detection rate in the prior art are solved, and more efficient and accurate object contour detection is achieved.
Patent Information
- Application Number
- CN202111057853.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-09
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-09-09
AI Technical Summary
The prior art has low contour detection efficiency, high error detection rate, and small detection range in network live broadcast, so it is impossible to effectively identify the contours of objects that need to be displayed.
The pre-trained weak semantic contour detection model is adopted to extract the contour features in the live image through the encoder. The classification module determines whether the complete object is included. The decoder extracts the object contour and only focuses on the complete and significant target object, reducing the calculation amount and improving detection efficiency.
It effectively reduces the error detection rate, increases the detection range, improves the detection efficiency, and ensures that the detected object profile is the object that needs to be displayed.
Smart Images

Figure CN113808151B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of network live broadcast, and in particular to a method, device, equipment and storage medium for detecting weak semantic contours of live broadcast images. Background Art
[0002] With the advancement of network communication technology, live streaming has become a new form of online interaction. It is also loved by more and more audiences because of its real-time and interactive features.
[0003] During live broadcasts, online hosts often need to interact with the audience. In some live broadcast scenarios, when the host shows an object to the audience, it is necessary to perform contour detection on the object. After the object is detected through contour detection, the object can be enlarged, special effects can be added, and it can be displayed separately.
[0004] During the research process, the inventors found that the current mainstream contour detection and recognition method is used to detect the contour of a certain type of specific objects with a small detection range, or contour detection is performed on all objects in the live image. There are many detection tasks, resulting in low detection efficiency, and the detected object contour is not necessarily the contour of the object that needs to be displayed, resulting in a high false detection rate. Summary of the invention
[0005] Based on this, the purpose of this application is to provide a method, device, equipment and storage medium for weak semantic contour detection of live images, which has the advantages of reducing false detection rate, increasing detection range and improving detection efficiency.
[0006] According to a first aspect of an embodiment of the present application, a method for detecting weak semantic contours of a live image is provided, and the method for detecting weak semantic contours of a live image comprises:
[0007] Acquire the live image to be detected;
[0008] Acquire contour features in the live image through an encoder in a pre-trained weak semantic contour detection model; wherein the weak semantic contour detection model includes an encoder, a classification module and a decoder;
[0009] Determining, by the classification module, whether the live broadcast image contains at least one complete object according to the contour feature;
[0010] If it is detected that the live image contains at least one complete object, the contour of the object in the live image is extracted by the decoder to obtain a contour map of the live image.
[0011] According to a second aspect of an embodiment of the present application, a device for detecting weak semantic contours of a live image is provided, and the device for detecting weak semantic contours of a live image comprises:
[0012] An acquisition module, used for acquiring a live image to be detected;
[0013] A contour feature acquisition module, used to acquire contour features in the live image through an encoder in a pre-trained weak semantic contour detection model; wherein the weak semantic contour detection model includes an encoder, a classification module and a decoder;
[0014] A complete object confirmation module, configured to determine whether the live broadcast image contains at least one complete object according to the contour features through the classification module;
[0015] The contour map acquisition module is used to extract the object contour in the live image through the decoder to acquire the contour map of the live image if it is detected that the live image contains at least one complete object.
[0016] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing any one of the methods for detecting weak semantic contours of live broadcast images.
[0017] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method for detecting weak semantic contours of a live image is implemented.
[0018] The present application obtains a live image to be detected, uses a classification module in a trained weak semantic contour detection model to determine whether the live image contains at least one complete object, and extracts the object contour in the live image through a decoder when detecting that the live image contains at least one complete object. The weak semantic contour detection model only focuses on complete and significant target objects, which can effectively reduce the amount of calculation in contour detection and improve detection efficiency. For live images without complete objects, contour detection is terminated in time, thereby reducing the false detection rate.
[0019] For better understanding and implementation, the present application is described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 A schematic diagram of an application scenario of a method for detecting weak semantic contours of a live broadcast image provided by an embodiment of the present application;
[0021] Figure 2 A flowchart of a method for detecting weak semantic contours of a live broadcast image provided by an embodiment of the present application;
[0022] Figure 3An example diagram of a weak semantic contour detection model for a live broadcast image provided by one embodiment of the present application;
[0023] Figure 4 A flowchart of a method for detecting weak semantic contours of a live broadcast image provided by another embodiment of the present application;
[0024] Figure 5 A schematic diagram of the structure of a weak semantic contour detection device for a live broadcast image provided by an embodiment of the present application;
[0025] Figure 6 A schematic block diagram of the structure of an electronic device provided for one embodiment of the present application. DETAILED DESCRIPTION
[0026] In order to make the objectives, technical solutions and advantages of the present application more clear, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0027] It should be clear that the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.
[0028] When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. Instead, they are only examples of devices and methods consistent with some aspects of the present application as detailed in the attached claims.
[0029] In the description of the present application, it should be understood that the terms "first", "second", "third", etc. are only used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to the specific circumstances. The singular forms of "a", "said", and "the" used in the present application and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. The words "if" / "if" used herein can be interpreted as "at the time of" or "when" or "in response to determination". In addition, in the description of the present application, unless otherwise specified, "multiple" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist, for example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the objects associated before and after are in an "or" relationship.
[0030] See also Figure 1 , which is a schematic diagram of an application scenario of the weak semantic contour detection method for a live broadcast image provided in the present application. The application scenario includes a live broadcast client 10 and a server 20, and the live broadcast client 10 interacts with the server 20.
[0031] The hardware pointed to by the live broadcast client 10 is essentially a computer device, specifically, it can be a computer device of the type of a smart phone, a smart interactive tablet, a personal computer, etc. The live broadcast client 10 can access the Internet through a well-known network access method and establish a data communication link with the server 20.
[0032] The server 20 is a service server, which can be responsible for further connecting to related audio data servers, video streaming servers and other servers providing related support services, so as to form a logically related service cluster to provide services for related terminal devices, such as Figure 1 The live broadcast client 10 shown in provides services.
[0033] The weak semantic contour detection method for live broadcast images can be run on the live broadcast client 10 and / or server 20. When the weak semantic contour detection method for live broadcast images is run on the live broadcast client 10, the live broadcast client 10 executes the weak semantic contour detection method for live broadcast images on the live broadcast images acquired locally to obtain the object contour detection result of the live broadcast images. When the weak semantic contour detection method for live broadcast images is run on the server, the server 20 acquires the live broadcast images from the live broadcast client, executes the weak semantic contour detection method for live broadcast images, obtains the contour map of the live broadcast images, and can return the detection result to the live broadcast client 10.
[0034] Embodiment 1:
[0035] The embodiment of the present application discloses a method for detecting weak semantic contours of a live broadcast image.
[0036] The following will be combined with the attached Figure 2 , a weak semantic contour detection method for a live broadcast image provided in an embodiment of the present application is introduced in detail.
[0037] The weak semantic contour detection method of the live broadcast image provided by the embodiment of the present application includes:
[0038] S101: Acquire a live image to be detected.
[0039] The live broadcast image may be a live broadcast image acquired by a live broadcast client, or a part of the live broadcast image, such as a partial screenshot of the live broadcast image.
[0040] S102: Acquire contour features in the live image through an encoder in a pre-trained weak semantic contour detection model; wherein the weak semantic contour detection model includes an encoder, a classification module and a decoder.
[0041] The weak semantic contour detection model is used to detect contours of significant complete objects in the live image. The weak semantic contour detection model does not care about the type of objects in the live image. As long as there are significant complete objects in the live image, the contours of the complete objects can be detected. The design of the weak semantic contour detection model is based on the encoder-decoder framework, which is a model architecture that uses different algorithms to solve different tasks, wherein encoding refers to an encoder converting an input sequence into a dense vector of a fixed dimension, and decoding refers to converting the dense vector obtained by encoding into target data.
[0042] In one embodiment, the encoder comprises an input layer and a plurality of encoding layers connected in sequence;
[0043] The step of obtaining the contour features in the live image through the encoder in the pre-trained weak semantic contour detection model comprises:
[0044] Convolving the live image through the input layer, downsampling to a first preset resolution and then outputting it to the plurality of encoding layers;
[0045] Separately convolve the live image through the plurality of coding layers to obtain a contour feature map in the live image;
[0046] The input layer downsamples the live image to a first preset resolution to reduce the amount of feature calculation output to the encoding layer and improve the efficiency of extracting object contours by the weak semantic contour detection model. The first preset resolution can be set according to the image size of the input live image.
[0047] The coding layer is used to perform convolution operation on the live image to obtain a contour feature map in the live image. Preferably, the convolution method of the coding layer is depthwise separable convolution. Compared with conventional convolution, depthwise separable convolution can reduce parameters and can improve the operation speed of the network to a certain extent.
[0048] S103: Determine, by the classification module, whether the live broadcast image contains at least one complete object according to the contour feature.
[0049] The classification module is a binary classification module, which divides each pixel of the input image into contour pixels and non-contour pixels, and determines the connectivity of each contour pixel to determine whether the live image contains at least one complete object. In the embodiment of the present application, whether the live image has significance is determined by judging whether the live image contains at least one or more complete objects: if the live image does not contain at least one or more complete objects, the live image has significance; if the live image does not contain a complete object, the live image has no significance. If the live image is judged to be significant, contour detection continues; if the live image is not significant, contour detection is terminated. The conclusion that the live image does not contain a complete object is drawn in advance, and contour detection is not required, which reduces contour detection of invalid images and reduces the false detection rate.
[0050] In one embodiment, the classification module includes an average pooling layer, a vector conversion layer, and a plurality of fully connected layers connected in sequence;
[0051] Downsampling the contour feature map output by the encoder to a second preset resolution through the average pooling layer;
[0052] The second preset resolution contour feature map is converted into a one-dimensional vector of a preset length through the vector conversion layer, and a binary classification value indicating whether the live broadcast image contains at least one complete object is obtained through the multiple fully connected layers.
[0053] The vector conversion layer converts the contour feature map of the second preset resolution into a one-dimensional vector with a length of 512. The classification module includes three fully connected layers with node numbers of 64, 16, and 1. The three fully connected layers connect the one-dimensional vectors to obtain a binary classification value indicating whether the live image contains at least one complete object.
[0054] S104: If it is detected that the live image contains at least one complete object, extract the contour of the object in the live image through the decoder to obtain a contour map of the live image.
[0055] The decoder is used to decode the contour feature map obtained by encoding by the encoder to obtain the contour of the object in the live image. In one embodiment, the decoder includes a plurality of decoding layers and output layers connected in sequence; wherein each decoding layer corresponds to an encoding layer; the step of extracting the contour of the object in the live image by the decoder includes:
[0056] Performing bilinear interpolation on the outputs of the corresponding encoding layers through the decoding layers respectively, and upsampling to the first preset resolution;
[0057] The object contours in the live image are extracted through the output layer.
[0058] Bilinear interpolation is an upsampling method that uses the existing four pixels in the original image to interpolate new pixels in two directions to improve the resolution of the image. Compared with other upsampling methods, bilinear interpolation is calculated based on the pixels in the original image to avoid jagged edges and obtain smoother, high-resolution images.
[0059] like Figure 3 As shown, it is a schematic diagram of the process of extracting the object contour using the weak semantic contour detection method of the live image described in the embodiment of the present application. Among them, the input image includes 256×192×3 feature points, and the weak semantic contour detection model includes an encoder, a classification module (cls_out), a decoder, and 5 connection layers (skip-layer5, skip-layer4, skip-layer3, skip-layer2, skip-layer1);
[0060] Among them, the encoder includes an input layer (InConv) and 5 encoding layers (Encoder1, Encoder2, Encoder3, Encoder4, Encoder5) connected in sequence, and the input layer and the 5 encoding layers are used to convolve and downsample the live image to obtain contour features in the live image.
[0061] The decoder includes 5 decoding layers (Decoder1, Decoder2, Decoder 3, Decoder4, Decoder 5) corresponding to the 5 coding layers and an output layer (OutConv). Each coding layer is connected to the corresponding decoding layer through a fully connected layer. Each decoding layer is used to upsample the output of its coding layer and the output of the previous coding layer to extract the contour of the object in the live image. Using a fully connected layer to connect the coding layer and the decoding layer can avoid information loss during the encoding process, so that the decoding layer can combine the information that has not been lost in the corresponding coding layer before encoding to encode when decoding, so that the extracted object contour is more accurate.
[0062] In an embodiment of the present application, a live image to be detected is obtained, and a classification module in a trained weak semantic contour detection model is used to determine whether the live image contains at least one complete object. When the live image is detected to contain at least one complete object, the object contour in the live image is extracted through a decoder. The weak semantic contour detection model only focuses on complete and significant target objects, which can effectively reduce the amount of calculation in contour detection and improve detection efficiency. For live images without complete objects, contour detection is terminated in time, thereby reducing the false detection rate.
[0063] In one embodiment, the method for detecting weak semantic contours of live broadcast images further includes the following steps:
[0064] Extracting straight line segments in the contour map of the live broadcast image based on a straight line segment detection algorithm;
[0065] Convert the straight line segments into straight lines, obtain the position information of the intersection points between every two straight lines, and merge the intersection points that meet the merging conditions according to a preset merging condition;
[0066] Obtain the area of a rectangle formed by every four intersections, and obtain the position information of the four intersections with the largest rectangular areas;
[0067] An affine transformation matrix is obtained based on preset target image position information and position information of the four intersection points, and the contour map of the live image is corrected using the affine transformation matrix to obtain a corrected contour map of the live image.
[0068] A Line Segment Detector (LSD) algorithm detects the gradient value and gradient direction of each pixel of the image, thereby performing straight line segmentation on the input grayscale image based on the gradient value and gradient direction to obtain a number of straight line segments. Specifically, based on the straight line segment detection algorithm, the steps of extracting straight line segments in the contour map of the live image include:
[0069] Based on the Gaussian downsampling method, downsampling the contour map of the live broadcast image to a preset image scale;
[0070] Obtaining the gradient value and gradient direction value of each pixel point of the contour map of the live image;
[0071] Eliminate pixels whose gradient values are less than the preset gradient threshold, and select the pixel with the maximum gradient value as the seed point;
[0072] Determine a direction value range based on the gradient direction value of the seed point and a preset range threshold, obtain pixel points whose gradient direction values are within the direction value range, and obtain a plurality of homogeneous points;
[0073] Based on the position information of the plurality of isotropic points, generating a rectangle including the plurality of isotropic points;
[0074] Obtain the length and width of the rectangle, and calculate the density of homosexual points of the rectangle according to the number of homosexual points in the rectangle;
[0075] If the density of the isotropic points is greater than or equal to a set density threshold, based on a fitted rectangle accuracy calculation function, an error value of the rectangle in the contour map is obtained;
[0076] If the error value is less than or equal to a preset threshold, the straight line segment of the rectangle is used as the straight line segment in the contour map of the live image; if the error value is greater than the preset threshold, the side length of the rectangle is adjusted until the error value of the rectangle obtained based on the fitting rectangle accuracy calculation function is less than or equal to the preset threshold.
[0077] The contour map of the live image is downsampled based on the Gaussian downsampling method, and the jagged phenomenon of the image can be effectively solved by reducing the image, thereby improving the accuracy of straight line segment detection. In the embodiment of the present application, the preset image scale can be 0.8.
[0078] Specifically, the gradient value can be calculated according to the grayscale values i(x+1, y), i(x, y+1) and i(x+1, y) of the pixel point (x, y) and its neighboring pixel points (x+1, y), (x, y+1) and (x+1, y) according to the gradient value calculation formula.
[0079]
[0080]
[0081]
[0082] Among them, G(x,y) is the gradient value, g x (x, y) is the first gray value, g y (x, y) is the second grayscale value.
[0083] The gradient direction value can be calculated according to the grayscale values i(x+1, y), i(x, y+1) and i(x+1, y) of the pixel point (x, y) and its neighboring pixel points (x+1, y), (x, y+1) and (x+1, y) according to the gradient direction value calculation formula.
[0084] The calculation formula of the gradient direction value is:
[0085]
[0086] Among them, θ is the gradient direction value.
[0087] The gradient threshold can be set according to actual needs.
[0088] The direction value range is determined based on the gradient direction value of the seed point and a preset range threshold. Specifically, the direction value range can be [at, a+t], where a is the gradient direction value of the seed point, and t is the preset range threshold. The range threshold can be set according to the actual needs of the user.
[0089] After determining the direction value range, the seed point is used as the starting point to search for each pixel point of the contour map of the live image to obtain pixel points (i.e., homogeneous points) whose gradient direction values are within the direction value range, and generate a rectangle containing all homogeneous points.
[0090] The density of homogeneous points is used to determine the number of homogeneous points in the rectangle, which can be obtained by dividing the number of homogeneous points by the area of the rectangle. The density threshold can be set according to the size of the input image and the actual needs of the user. When the homogeneous point density of a rectangle is less than the set density threshold, the rectangle can be truncated to convert it into multiple rectangles, and the homogeneous point density of the truncated rectangle is recalculated until the homogeneous point density of the obtained rectangle is greater than or equal to the set density threshold.
[0091] The fitted rectangle accuracy calculation function (Number of False Alarms, NFA) is used to evaluate the accuracy of the fitted rectangle. In the embodiment of the present application, when the error value is less than or equal to the preset threshold, it is judged that the current fitted rectangle meets the set requirements, and the straight line segments of the rectangle are used as the straight line segments in the contour map of the live image. If the error value is greater than the preset threshold, the side length of the rectangle is adjusted to cut it into multiple rectangular frames, and the error value is obtained based on the fitted rectangle accuracy calculation function until the error value is less than or equal to the preset threshold.
[0092] In one embodiment, the step of converting the straight line segment into a straight line specifically includes:
[0093] According to the position information of the endpoints of the straight line segment, the slope and intercept of the straight line corresponding to the straight line segment are obtained;
[0094] According to the slope and the intercept, a straight line corresponding to the straight line segment is obtained.
[0095] The position information of the endpoints of the straight line segment includes the position information of the two endpoints of the straight line segment, and the slope and intercept of the straight line corresponding to the two straight line segments are obtained according to the position information of the endpoints of the two straight line segments, and the straight line corresponding to the straight line segment is determined according to the slope and intercept. In a preferred embodiment, the amount of data calculation can be reduced and the correction efficiency can be improved by merging close straight lines. Therefore, after the step of obtaining the straight line corresponding to the straight line segment, the following steps are also included:
[0096] According to the slope and the intercept, merging straight lines that meet a preset merging condition;
[0097] The preset merging condition includes: the slope difference between at least two straight lines is within a preset slope difference range, and the intercept difference between the at least two straight lines is within an intercept difference range.
[0098] The intersection point is the intersection point between two straight lines. For adjacent intersection points, the amount of data calculation can be reduced by merging the intersection points to improve the correction efficiency. Specifically, the steps of merging the intersection points that meet the merging conditions according to the preset merging conditions include:
[0099] When there are at least two intersections whose distance is less than a set threshold, obtaining the average of the position information of the at least two intersections;
[0100] The at least two intersection points are merged, and a merged intersection point is generated according to an average of the position information of the at least two intersection points.
[0101] By merging adjacent intersections and generating a new intersection at the midpoint of the adjacent intersections according to the average of the position information of the adjacent intersections, the number of intersections is reduced and the correction efficiency is improved.
[0102] Every four intersections can determine a rectangle, and the rectangular area to be corrected in the contour map of the live image can be determined by the position information of the four intersections with the largest rectangular area. In one embodiment, after the step of obtaining the rectangular area formed by every four intersections, it also includes:
[0103] Obtain a rectangle formed by every four intersection points, and obtain a rectangle that satisfies a preset rectangle screening condition;
[0104] Among them, the preset rectangle screening conditions include: the angle between adjacent sides of the rectangle is greater than a set angle threshold, the length ratio of the opposite sides of the rectangle is greater than a set ratio threshold, the rectangle has at least one set of parallel opposite sides, and the aspect ratio of the rectangle is greater than a set aspect ratio threshold.
[0105] The rectangle is made to have at least one set of parallel opposite sides, so that the final corrected matrix area is a trapezoid or a parallelogram, which is more convenient for affine transformation. The angle threshold, the ratio threshold, and the aspect ratio threshold can be set according to the size of the input image and the size of the contour contained in the image. For example, the rectangle screening conditions can be set as follows: the angle between adjacent sides of the rectangle is greater than 4 degrees, the length ratio of the opposite sides of the rectangle is greater than 0.5, the rectangle has at least one set of parallel opposite sides, and the ratio of the shortest side to the longest side of the rectangle is greater than 0.15.
[0106] Affine transformation refers to a linear transformation between two-dimensional coordinates. After affine transformation, the lines in the image can maintain their original straightness and relative position relationship. Affine transformation includes transformations such as translation, scaling, flipping, rotation and shearing. In the embodiment of the present application, an affine transformation matrix can be constructed based on the target image position information and the position information of the four intersections. The affine transformation matrix is used to correct the contour map of the live image, so that the live contour map extracted by contour detection is clearer and easier to be recognized, thereby improving the detection efficiency of the contour image.
[0107] According to experiments, when the weak semantic contour detection method of live broadcast images described in this application is applied to most mid-to-high-end mobile phones (2000+ models) for contour detection, the number of floating point operations per second (flops) is 174M, the model parameter is 386.56k, and the calculation speed is about 30ms. It can be seen that the weak semantic contour detection method of live broadcast images described in this application can achieve ultra-real-time detection of contours.
[0108] Embodiment two:
[0109] The present embodiment is different from the first embodiment in that it also includes: a step of training a weak semantic contour detection model.
[0110] Optional, such as Figure 4 As shown, before the step of acquiring the live image to be detected, the method for detecting weak semantic contours of the live image further includes:
[0111] S201: Acquire a preset weak semantic training sample set; the weak semantic training sample set includes a plurality of images with complete objects and their corresponding contour images and a plurality of images with incomplete objects and their corresponding contour images;
[0112] S202: Based on the decoding-encoding framework, a weak semantic contour detection model with contour extraction function is constructed;
[0113] S203: Pre-training the weak semantic contour detection model using the weak semantic training sample set until the loss value of the weak semantic contour detection model meets the target loss, thereby obtaining the pre-trained weak semantic contour detection model.
[0114] The plurality of images with complete objects and the plurality of images with incomplete objects may be a collection of manually photographed images of indoor scenes, street scenes or landscapes, and the corresponding contour images may be images extracted by a contour detection algorithm. In an embodiment of the present application, the contour images corresponding to the images with incomplete objects are black images without contours.
[0115] In one embodiment, the weak semantic training sample set is an image collected from a live broadcast scene; or, the weak semantic training sample set is an image collected from a live broadcast room, and the plurality of images with complete objects include a combination of multiple live broadcast scenes and multiple complete objects. The live broadcast scene may include but is not limited to desktop scenes (such as solid color desktops, wooden desktops, floral desktops, desktops with sundries), handheld scenes (such as hand-held, hand-held), wall backgrounds (such as solid color walls, floral walls) and lighting change scenes (such as dark light, strong light, backlight, reflection), etc. The multiple objects may include but are not limited to cards (such as work cards, membership cards, ID cards), bills (such as receipts, tax bills, registration forms), electronic (mobile phones, computers, tablets) and other objects (such as notices, books, signs, paper boxes), etc.
[0116] The plurality of images with complete objects and their corresponding contour maps are used as positive samples of the weak semantic contour detection model, and the plurality of images with incomplete objects and their corresponding contour maps are used as negative samples of the weak semantic contour detection model; the weak semantic contour detection model is pre-trained by using the positive samples and negative samples until the loss value of the weak semantic contour detection model meets the target loss, thereby obtaining the pre-trained weak semantic contour detection model.
[0117] Specifically, the step of pre-training the weak semantic contour detection model using the weak semantic training sample set includes: adjusting model parameters of the weak semantic contour detection model based on the Adam optimization algorithm.
[0118] Model parameters may include parameters such as learning rate and training period.
[0119] The Adam optimization algorithm is a method for calculating an adaptive learning rate for each parameter, and can iteratively update the neural network weights based on training data; in an embodiment of the present application, the initial learning rate of the weak semantic contour detection model is set to 0.001, and the training is performed for 200 epochs; wherein, all weak semantic training sample sets are trained once in one training epoch, and then the learning rate is decayed to 0.0005, and the learning rate is decayed to 0.0001 after the 300th epoch, and fine-tuned after the 400th epoch, freezing all batch normalization layers (Batch Normalization, BN) in the network, and the learning rate is decayed to 0.00005 again, wherein freezing all batch normalization layers in the network means that the batch normalization layers are not allowed to participate in network training, that is, the parameters of the batch normalization layers are not updated.
[0120] The weak semantic contour detection model described in the embodiment of the present application performs contour detection by dividing each pixel of the input image into contour pixels and non-contour pixels. However, since the number of contour pixels is usually small, it is easy to cause imbalance in category prediction. Therefore, in a preferred embodiment, the loss function of the weak semantic contour detection model is:
[0121]
[0122] Where L represents the loss value, β represents the first coefficient, β=|Y - | / |Y + |,|Y _ | represents the number of non-contour pixels in the live image, |Y + | represents the number of contour pixels in the live image, y j represents the predicted value of the weak semantic contour detection model at pixel j. j =1|X)=σ(a j )∈[0,1],P(y j =0|X)=σ(a j )∈[0,1], σ(*) represents the sigmoid function, wherein the sigmoid function is an activation function of a neural network, which maps variables to between 0 and 1. By introducing a category-balanced cross entropy loss function as the loss function of the weak semantic contour detection model, the imbalance of category prediction can be avoided and the accuracy of contour pixel detection can be improved.
[0123] Embodiment three:
[0124] This embodiment provides a weak semantic contour detection device for live images, which can be used to execute the weak semantic contour detection method for live images of Embodiment 1 and Embodiment 2 of this application. For details not disclosed in this embodiment, please refer to Embodiment 1 and Embodiment 2 of this application.
[0125] See also Figure 5 , Figure 5 1 is a schematic diagram of the structure of a weak semantic contour detection device for a live broadcast image disclosed in an embodiment of the present application. The weak semantic contour detection device for a live broadcast image can be run in a server or a live broadcast client. The weak semantic contour detection device for a live broadcast image includes:
[0126] An acquisition module 301 is used to acquire a live image to be detected;
[0127] The contour feature acquisition module 302 is used to acquire the contour features in the live image through an encoder in a pre-trained weak semantic contour detection model; wherein the weak semantic contour detection model includes an encoder, a classification module and a decoder.
[0128] The complete object determination module 303 is used to determine whether the live broadcast image contains at least one complete object according to the contour features through the classification module.
[0129] The contour extraction module 304 is used to extract the contour of the object in the live image through the decoder to obtain a contour map of the live image if it is detected that the live image contains at least one complete object.
[0130] The embodiment of the present application obtains a live image to be detected, uses a classification module in a trained weak semantic contour detection model to determine whether the live image contains at least one complete object, and extracts the object contour in the live image through a decoder when detecting that the live image contains at least one complete object. The weak semantic contour detection model only focuses on complete and significant target objects, which can effectively reduce the amount of calculation in contour detection and improve detection efficiency. For live images without complete objects, contour detection is terminated in time, thereby reducing the false detection rate.
[0131] Embodiment four:
[0132] This embodiment provides an electronic device that can be used to perform all or part of the steps of the weak semantic contour detection method for live broadcast images of Embodiment 1 and Embodiment 2 of this application. For details not disclosed in this embodiment, please refer to Embodiment 1 and Embodiment 2 of this application.
[0133] See also Figure 6 , Figure 6 The electronic device 900 may be, but is not limited to, a combination of one or more of various servers, personal computers, laptops, smart phones, tablet computers, and the like.
[0134] In a preferred embodiment of the present application, the electronic device 900 includes a memory 901 , at least one processor 902 , at least one communication bus 903 and a transceiver 904 .
[0135] Those skilled in the art should understand that Figure 6 The structure of the electronic device shown does not constitute a limitation of the embodiments of the present application, and may be either a bus structure or a star structure. The electronic device 900 may also include more or less other hardware or software than shown in the figure, or a different component arrangement.
[0136] In some embodiments, the electronic device 900 is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application-specific integrated circuits, programmable gate arrays, digital processors, and embedded devices. The electronic device 900 may also include client devices, which include but are not limited to any electronic product that can interact with a client through a keyboard, mouse, remote control, touchpad, or voice-controlled device, such as a personal computer, tablet computer, smart phone, digital camera, etc.
[0137] It should be noted that the electronic device 900 is only an example, and other existing or future electronic products that are suitable for the present application should also be included in the protection scope of the present application and included here by reference.
[0138] In some embodiments, a computer program is stored in the memory 901, and when the computer program is executed by the at least one processor 902, all or part of the steps in the weak semantic contour detection method of the live image as described in the first embodiment and the second embodiment are implemented. The memory 901 includes a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable rewritable read-only memory (EEPROM), a compact disc (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.
[0139] In some embodiments, the at least one processor 902 is the control core (Control Unit) of the electronic device 900, and uses various interfaces and lines to connect the various components of the entire electronic device 900, and executes or executes the programs or modules stored in the memory 901, and calls the data stored in the memory 901 to execute various functions of the electronic device 900 and process data. For example, when the at least one processor 902 executes the computer program stored in the memory, it implements all or part of the steps of the weak semantic contour detection method of the live image described in the embodiment of the present application; or implements all or part of the functions of the weak semantic contour detection device of the live image. The at least one processor 902 can be composed of an integrated circuit, for example, it can be composed of a single packaged integrated circuit, or it can be composed of multiple integrated circuits with the same function or different functions, including one or more central processing units (CPU), microprocessors, digital processing chips, graphics processors, and various control chips. Combination, etc.
[0140] In some embodiments, the at least one communication bus 903 is configured to implement connection and communication between the memory 901 and the at least one processor 902, etc.
[0141] The electronic device 900 may also include a variety of sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.
[0142] Embodiment five:
[0143] This embodiment provides a computer-readable storage medium on which a computer program is stored. The instructions are suitable for being loaded by a processor and executing the weak semantic contour detection method for live broadcast images of Embodiment 1 and Embodiment 2 of the present application. The specific execution process can be found in the specific description of Embodiment 1 and Embodiment 2, which will not be repeated here.
[0144] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the components described as separate parts may or may not be physically separated, and the parts displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present application scheme. Ordinary technicians in this field can understand and implement it without paying creative work.
[0145] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0146] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0147] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. A method for detecting weak semantic contours in live broadcast images. It is characterized in that The following steps are involved: Acquire a preset weak semantic training sample set; the weak semantic training sample set includes a plurality of images with complete objects and their corresponding contour images and a plurality of images with incomplete objects and their corresponding contour images; the plurality of images with complete objects include a combination of multiple live broadcast scenes and multiple complete objects; Based on the decoding-encoding framework, a weak semantic contour detection model with contour extraction function is constructed; Pre-training a weak semantic contour detection model using the weak semantic training sample set until the loss value of the weak semantic contour detection model meets the target loss, thereby obtaining a pre-trained weak semantic contour detection model; the pre-trained weak semantic contour detection model is used to detect the contour of a complete object; Acquire the live image to be detected; Acquire contour features in the live image through an encoder in a pre-trained weak semantic contour detection model; wherein the weak semantic contour detection model includes an encoder, a classification module and a decoder; Determining, by the classification module, whether the live broadcast image contains at least one complete object according to the contour feature; If it is detected that the live image contains at least one complete object, extracting the contour of the object in the live image through the decoder to obtain a contour map of the live image; Extracting straight line segments in the contour map of the live broadcast image based on a straight line segment detection algorithm; Convert the straight line segments into straight lines, obtain the position information of the intersection points between every two straight lines, and merge the intersection points that meet the merging conditions according to a preset merging condition; Obtain the area of a rectangle formed by every four intersections, and obtain the position information of the four intersections with the largest rectangular areas; Acquire an affine transformation matrix based on preset target image position information and position information of the four intersection points, and use the affine transformation matrix to correct the contour map of the live image to obtain a corrected contour map of the live image; Based on the straight line segment detection algorithm, the step of extracting the straight line segments in the contour map of the live image includes: Based on the Gaussian downsampling method, downsampling the contour map of the live broadcast image to a preset image scale; Obtaining the gradient value and gradient direction value of each pixel point of the contour map of the live image; Eliminate pixels whose gradient values are less than the preset gradient threshold, and select the pixel with the maximum gradient value as the seed point; Determine a direction value range based on the gradient direction value of the seed point and a preset range threshold, obtain pixel points whose gradient direction values are within the direction value range, and obtain a plurality of homogeneous points; Based on the position information of the plurality of isotropic points, generating a rectangle including the plurality of isotropic points; Obtain the length and width of the rectangle, and calculate the density of homosexual points of the rectangle according to the number of homosexual points in the rectangle; If the density of the isotropic points is greater than or equal to a set density threshold, based on a fitted rectangle accuracy calculation function, an error value of the rectangle in the contour map is obtained; If the error value is less than or equal to a preset threshold, the straight line segment of the rectangle is used as the straight line segment in the contour map of the live image; if the error value is greater than the preset threshold, the side length of the rectangle is adjusted until the error value of the rectangle obtained based on the fitting rectangle accuracy calculation function is less than or equal to the preset threshold.
2. According to the method for detecting weak semantic contours of live broadcast images according to claim 1, Features: The encoder comprises an input layer and a plurality of encoding layers connected in sequence; The step of obtaining the contour features in the live image through the encoder in the pre-trained weak semantic contour detection model comprises: Convolving the live image through the input layer, downsampling to a first preset resolution and then outputting it to the plurality of encoding layers; Separate convolution is performed on the live image through the plurality of coding layers respectively to obtain a contour feature map in the live image.
3. The weak semantic contour detection method of live broadcast image according to claim 2, Features: The decoder includes a plurality of decoding layers and output layers connected in sequence; wherein each decoding layer corresponds to a coding layer; and the step of extracting the contour of the object in the live image by the decoder includes: Performing bilinear interpolation on the outputs of the corresponding encoding layers through the decoding layers respectively, and upsampling to the first preset resolution; The object contours in the live image are extracted through the output layer.
4. The weak semantic contour detection method of live broadcast image according to claim 3, Features: The classification module includes an average pooling layer, a vector conversion layer and several fully connected layers connected in sequence; Downsampling the contour feature map output by the encoder to a second preset resolution through the average pooling layer; The second preset resolution contour feature map is converted into a one-dimensional vector of a preset length through the vector conversion layer, and a binary classification value indicating whether the live broadcast image contains at least one complete object is obtained through the multiple fully connected layers.
5. The method for detecting weak semantic contours of live broadcast images according to claim 3, It is characterized in that The weak semantic contour detection model also includes a plurality of fully connected layers arranged between each encoding layer and each decoding layer.
6. The method for detecting weak semantic contours of live broadcast images according to claim 5, It is characterized in that The step of pre-training the weak semantic contour detection model using the weak semantic training sample set includes: Based on the Adam optimization algorithm, the model parameters of the weak semantic contour detection model are adjusted.
7. The method for detecting weak semantic contours of live broadcast images according to any one of claims 1 to 6, It is characterized in that The loss function of the weak semantic contour detection model is: Where L represents the loss value, β represents the first coefficient, β=|Y - | / |Y + |,|Y - | represents the number of non-contour pixels in the live image, |Y + | represents the number of contour pixels in the live image, y j represents the predicted value of the weak semantic contour detection model at pixel j, P(y j =1|X) means that when input X, y j =1, P(y j =0|X) means that when input X, y j =0 probability.
8. A device for detecting weak semantic contours of live images. It is characterized in that The device comprises: An acquisition module, used for acquiring a live image to be detected; A contour feature acquisition module, used to acquire contour features in the live image through an encoder in a pre-trained weak semantic contour detection model; wherein the weak semantic contour detection model includes an encoder, a classification module and a decoder; A complete object confirmation module, configured to determine whether the live broadcast image contains at least one complete object according to the contour features through the classification module; A contour map acquisition module, configured to extract the contour of the object in the live image through the decoder to acquire a contour map of the live image if at least one complete object is detected in the live image; The device is also used to obtain a preset weak semantic training sample set; the weak semantic training sample set includes a plurality of images with complete objects and their corresponding contour maps and a plurality of images with incomplete objects and their corresponding contour maps; the plurality of images with complete objects include a combination of multiple live broadcast scenes and multiple complete objects; based on a decoding-encoding framework, a weak semantic contour detection model with a contour extraction function is constructed; the weak semantic contour detection model is pre-trained using the weak semantic training sample set until the loss value of the weak semantic contour detection model meets the target loss, thereby obtaining the pre-trained weak semantic contour detection model; the pre-trained weak semantic contour detection model is used to detect the contour of a complete object; The device is also used to extract straight line segments in the contour map of the live image based on a straight line segment detection algorithm; convert the straight line segments into straight lines, obtain the position information of the intersections between every two straight lines, and merge the intersections that meet the merging conditions according to a preset merging condition; obtain the rectangular area formed by every four intersections, and obtain the position information of the four intersections with the largest rectangular area; obtain an affine transformation matrix based on the preset target image position information and the position information of the four intersections, and use the affine transformation matrix to correct the contour map of the live image to obtain the corrected contour map of the live image; Extracting straight line segments in the contour map of the live image based on a straight line segment detection algorithm includes: Based on the Gaussian downsampling method, downsampling the contour map of the live broadcast image to a preset image scale; Obtaining the gradient value and gradient direction value of each pixel point of the contour map of the live image; Eliminate pixels whose gradient values are less than the preset gradient threshold, and select the pixel with the maximum gradient value as the seed point; Determine a direction value range based on the gradient direction value of the seed point and a preset range threshold, obtain pixel points whose gradient direction values are within the direction value range, and obtain a plurality of homogeneous points; Based on the position information of the plurality of isotropic points, generating a rectangle including the plurality of isotropic points; Obtain the length and width of the rectangle, and calculate the density of homosexual points of the rectangle according to the number of homosexual points in the rectangle; If the density of the isotropic points is greater than or equal to a set density threshold, based on a fitted rectangle accuracy calculation function, an error value of the rectangle in the contour map is obtained; If the error value is less than or equal to a preset threshold, the straight line segment of the rectangle is used as the straight line segment in the contour map of the live image; if the error value is greater than the preset threshold, the side length of the rectangle is adjusted until the error value of the rectangle obtained based on the fitting rectangle accuracy calculation function is less than or equal to the preset threshold.
9. An electronic device, It is characterized in that include: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the weak semantic contour detection method for live images as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the method for detecting weak semantic contours of a live image as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
System for evaluating artistic quality
CN109118091A
Image palm area extraction method and device
CN110287771A