Method, apparatus and mechanical equipment for extracting difficult samples

By constructing a semantic segmentation model and filtering candidate keyframes, difficult samples are extracted, which solves the problem of high number of labels and difficult sample value in the existing technology, and improves the recognition efficiency of difficult samples and sample library quality.

CN114565803BActive Publication Date: 2025-06-20ZHONGKE YUNGU TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210065428.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-19
Publication Date
2025-06-20
Estimated Expiration
2042-01-19

AI Technical Summary

Technical Problem

The prior art cannot reduce the number of labels and cannot intuitively evaluate the sample value, resulting in low extraction efficiency of difficult samples.

Method used

By obtaining candidate keyframes, a semantic segmentation model is constructed, predicted difficult samples and edge mark samples are determined, and edge mark samples are filtered to obtain labeled difficult samples. Finally, candidate difficult samples are determined based on predicted difficult samples and labeled difficult samples.

Benefits of technology

The scale of manually confirmed pictures is reduced, the efficiency of identifying difficult samples is improved, and the quality of sample libraries for difficult samples is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114565803B_ABST
    Figure CN114565803B_ABST
Patent Text Reader

Abstract

The present application discloses a method, an apparatus, and a mechanical device for extracting difficult samples. The method includes: obtaining candidate key frames; constructing a semantic segmentation model based on the candidate key frames; determining predicted difficult samples and edge-labeled samples through the semantic segmentation model; screening the edge-labeled samples to obtain labeled difficult samples; and determining candidate difficult samples based on the predicted difficult samples and the labeled difficult samples. The present application determines predicted difficult samples and edge-labeled samples through a semantic segmentation model, and reduces the scale of pictures to be manually confirmed through various screening methods, improves the recognition efficiency of difficult samples, and improves the quality of the sample library of difficult samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent monitoring technology, and specifically, to a method, device, and mechanical equipment for extracting difficult samples. Background Art

[0002] The production requirements of semantic segmentation labels require accurate annotation of the edge point sets of each region of interest in the sample images. However, due to the high annotation cost, constructing a qualified sample library of a certain scale requires a large cost. Currently, the evaluation of high-value samples (difficult samples) is mainly based on labels, with the final loss function or its variants as the final quantitative evaluation standard. In the prior art, the extraction of difficult samples is to define relevant rules during model training to mine difficult samples, which cannot reduce the annotation amount; or manually observe data for selective annotation, which has a certain blindness for sample value recognition, cannot guarantee the quality of the sample library, and consumes a large amount of manpower. Therefore, the prior art cannot reduce the annotation quantity and cannot intuitively evaluate the sample value, resulting in low extraction efficiency of difficult samples. Summary of the Invention

[0003] The purpose of this application is to provide a method, device, and mechanical equipment for extracting difficult samples, so as to solve the problem that the prior art cannot reduce the annotation quantity and cannot intuitively evaluate the sample value, resulting in low extraction efficiency of difficult samples.

[0004] To achieve the above purpose, the first aspect of this application provides a method for extracting difficult samples, including:

[0005] Obtain candidate key frames;

[0006] Construct a semantic segmentation model according to the candidate key frames;

[0007] Determine predicted difficult samples and edge-labeled samples through the semantic segmentation model;

[0008] Screen the edge-labeled samples to obtain labeled difficult samples;

[0009] Determine candidate difficult samples according to the predicted difficult samples and the labeled difficult samples.

[0010] In the embodiments of this application, constructing a semantic segmentation model according to the candidate key frames includes:

[0011] Divide the candidate key frames into multiple groups of candidate key frames;

[0012] Select a preset group of candidate key frames for annotation to obtain an initial sample library;

[0013] Train a semantic segmentation model according to the initial sample library, and the semantic segmentation model is used to predict the remaining candidate key frames;

[0014] After predicting each group of remaining candidate key frames, update the initial sample library and retrain the semantic segmentation model to update the semantic segmentation model.

[0015] In the embodiment of the present application, after predicting each group of remaining candidate key frames, updating the initial sample library and retraining the semantic segmentation model to update the semantic segmentation model includes:

[0016] Add the predicted difficult samples corresponding to the current group to the updated initial sample library of the previous group to obtain the current sample library;

[0017] Retrain the semantic segmentation model according to the current sample library to obtain the current semantic segmentation model; the current semantic segmentation model is used to predict the next group of remaining candidate key frames.

[0018] In the embodiment of the present application, determining the predicted difficult samples and the edge marked samples through the semantic segmentation model includes:

[0019] For each candidate key frame, determine the difference map between the maximum probability layer and the second maximum probability layer of the current candidate key frame;

[0020] Count the number of target pixels in the difference map where the difference is less than the first threshold;

[0021] Determine the ratio of the number of target pixels to the total number of pixels of the current candidate key frame;

[0022] Judge whether the ratio is greater than the second threshold;

[0023] When the ratio is greater than the second threshold, determine that the current candidate key frame is a predicted difficult sample;

[0024] When the ratio is not greater than the second threshold, determine that the current candidate key frame is an edge marked sample.

[0025] In the embodiment of the present application, screening the edge marked samples to obtain the marked difficult samples includes:

[0026] Obtain the first edge marked sample and the second edge marked sample that are adjacent in time in sequence;

[0027] Determine the second edge marked sample as the target edge marked sample;

[0028] Determine the similarity between the target edge marked sample and the first edge marked sample;

[0029] Judge whether the target edge marked sample is a sample to be manually marked according to the similarity;

[0030] Obtain the marked difficult samples in the samples to be manually marked.

[0031] In an embodiment of the present application, determining whether a target edge marked sample is a sample to be manually marked according to the similarity includes:

[0032] Determining whether the similarity is less than a third threshold;

[0033] In the case where the similarity is less than the third threshold, determining that the target edge marked sample is a sample to be manually marked.

[0034] In an embodiment of the present application, obtaining candidate key frames includes:

[0035] Obtaining candidate key frames containing motion by using a three-frame difference method.

[0036] A second aspect of the present application provides a device for extracting difficult samples, including:

[0037] A memory configured to store instructions; and

[0038] A processor configured to call instructions from the memory and capable of implementing the above-mentioned method for extracting difficult samples when executing the instructions.

[0039] A third aspect of the present application provides a mechanical device, including:

[0040] A video acquisition device for acquiring a video of a moving scene with a fixed viewing angle;

[0041] The above-mentioned device for extracting difficult samples.

[0042] A fourth aspect of the present application provides a machine-readable storage medium, on which instructions are stored, and the instructions are used to cause a machine to execute the above-mentioned method for extracting difficult samples.

[0043] Through the above technical solutions, a semantic segmentation model is constructed according to the obtained candidate key frames; predicted difficult samples and edge marked samples are determined through the semantic segmentation model; the edge marked samples are screened to obtain marked difficult samples; and then candidate difficult samples are determined according to the predicted difficult samples and the marked difficult samples. Compared with directly using the model prediction results and marking them on the original image and then performing manual confirmation, the present application determines the predicted difficult samples and the edge marked samples through the semantic segmentation model, and through various screening methods, reduces the scale of the pictures to be manually confirmed, improves the recognition efficiency of difficult samples, and improves the quality of the sample library of difficult samples.

[0044] Other features and advantages of the present application will be described in detail in the subsequent specific implementation part. Description of the Drawings

[0045] The accompanying drawings are used to provide a further understanding of the present application and form a part of the specification. Together with the following specific embodiments, they are used to explain the present application, but do not constitute a limitation to the present application. In the accompanying drawings:

[0046] Figure 1 Schematically shows a flowchart of a method for extracting difficult samples according to an embodiment of the present application;

[0047] Figure 2 Schematically shows a structural block diagram of a device for extracting difficult samples according to an embodiment of the present application;

[0048] Figure 3 Schematically shows a structural diagram of a mechanical device according to an embodiment of the present application. Specific Embodiments

[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. It should be understood that the specific embodiments described herein are only used to illustrate and explain the embodiments of the present application, and are not used to limit the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts fall within the scope of protection of the present application.

[0050] It should be noted that if there are directional indications (such as up, down, left, right, front, back...) involved in the embodiments of the present application, such directional indications are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications will also change accordingly.

[0051] In addition, if there are descriptions such as "first", "second", etc. involved in the embodiments of the present application, such descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present application.

[0052] Figure 1 Schematically shows a flowchart of a method for extracting difficult samples according to an embodiment of the present application. As Figure 1As shown in the figure, an embodiment of the present application provides a method for extracting difficult samples, and the method may include the following steps:

[0053] Step 102: Obtain candidate key frames;

[0054] Step 104: Construct a semantic segmentation model according to the candidate key frames;

[0055] Step 106: Determine predicted difficult samples and edge-labeled samples through the semantic segmentation model;

[0056] Step 108: Screen the edge-labeled samples to obtain labeled difficult samples;

[0057] Step 110: Determine candidate difficult samples according to the predicted difficult samples and the labeled difficult samples.

[0058] The method for extracting difficult samples in the embodiment of the present invention can be applied to concrete mechanical equipment, and may include, but is not limited to, selecting key frames for the alignment of the feeding and discharging ports of a mixer truck in a mixing plant. In the embodiment of the present application, videos are all composed of static pictures, and these static pictures are called frames. Aiming at the problem of how to mine semantic segmentation difficult samples in a large number of videos in the case of fixed-view motion, the embodiment of the present application proposes an identification method for extracting semantic segmentation difficult samples. Aiming at the problem of how to quickly screen a large number of difficult samples (i.e., edge-labeled samples) under a fixed view, a fast screening method based on image structure similarity is proposed to improve the screening efficiency of labeled difficult samples.

[0059] In the embodiment of the present application, deep learning technology is an effective technical means to solve many tasks in the field of natural scenes of images. The generalization ability of a deep learning model comes from the structure of the model itself and model training techniques on the one hand, but ultimately determines the upper limit of the model's generalization ability by the sample library it uses. Therefore, constructing a high-value sample library is the key to completing a certain image task using deep learning technology. In today's big data era, the amount of data is large and the types of data are diverse. In most cases, model training requires labeled data, but easy samples are of no practical help in improving the training accuracy of the model, and their labeling ratio is high and the cost is large. Therefore, adding a difficult sample mining strategy in the model training process can improve the performance of the model, but cannot reduce the amount of labeling. Therefore, researching how to identify difficult samples (i.e., difficult samples) from a large amount of data is of great significance for improving the quality of the sample library, improving the performance of the model, and reducing the labeling cost. The method for extracting difficult samples in the present application includes three stages: initially selecting candidate key frames, determining predicted difficult samples and edge-labeled samples through a semantic segmentation model, and screening the edge-labeled samples to obtain labeled difficult samples, so as to determine candidate difficult samples according to the predicted difficult samples and the labeled difficult samples.

[0060] In the embodiments of the present application, the primary selection of candidate key frames is to preliminarily screen out motion frames from a large amount of videos, and preliminarily distinguish background frames and foreground frames. In one example, candidate key frames containing motion can be obtained by the three-frame difference method. Specifically, the method for obtaining candidate key frames containing motion by the three-frame difference method may include: obtaining a first video frame, a second video frame, and a third video frame that are adjacent in time in sequence, and determining the third video frame as the target video frame. Then, perform difference processing on the first video frame and the second video frame to obtain a first adjacent difference map, and at the same time perform difference processing on the second video frame and the target video frame to obtain a second adjacent difference map. Further, determine the similarity between the target video frame and the second video frame, and judge whether the target video frame contains motion according to the first adjacent difference map and the second adjacent difference map. In the case where the target video frame contains motion and the similarity is less than a first set value, determine the target video frame as a foreground key frame; in the case where the target video frame does not contain motion and the similarity is less than a second set value, determine the target video frame as a background key frame. Among them, the background key frame is a key frame that does not include motion, and the foreground key frame is a key frame that includes motion. What the embodiments of the present application need to obtain are foreground key frames containing motion. The embodiments of the present application use the three-frame difference method to determine motion frames mainly considering that it can achieve a balance between computational efficiency and motion perception. Extract three video frames that are adjacent in time in sequence, and perform image sharpening on three consecutive frames of images to reduce the impact of uneven illumination on motion detection. Further judge whether the target video frame contains motion and the similarity between the target video frame and the adjacent video frames, so as to be able to distinguish foreground key frames and background key frames while extracting frames, without the need to spend a lot of manpower for subsequent distinction, improving the efficiency and quality of selecting key frames.

[0061] In the embodiments of the present application, semantic segmentation is a basic task in image segmentation, which means that for an image, each pixel is labeled with a corresponding category without distinguishing individuals. Simply put, it is to divide the input data of a visual image into different semantically interpretable categories. After the processor obtains the candidate key frames, it can construct a semantic segmentation model according to the candidate key frames, and determine prediction difficult samples and edge marking samples through the semantic segmentation model. Among them, the prediction difficult sample is a difficult sample that can be directly determined by the semantic segmentation model; the edge marking sample is a sample to be manually marked.

[0062] In the embodiment of the present application, the construction of the semantic segmentation model may include the following steps: dividing the candidate key frames into multiple groups (such as m groups), taking the candidate key frames of several of them (such as the preset group) for annotation to obtain an initial sample library, and training the semantic segmentation model using the labeled data in the initial sample library. Identifying the remaining candidate key frames using the semantic segmentation model trained according to the initial sample library, where the remaining candidate key frames are the other candidate key frames except those included in the preset group. For example, using the semantic segmentation model to identify one group in the remaining groups; adding the remaining candidate key frames identified as difficult-to-predict samples among them to the initial sample library and retraining the semantic segmentation model, and so on in a loop until all the remaining candidate key frames are identified by the continuously updated semantic segmentation model. Taking 10,000 candidate key frames as an example, m is 100, dividing the 10,000 candidate key frames into 100 groups, with each group including 100 candidate key frames. Selecting 1 group for annotation to obtain an initial semantic segmentation model. Assuming there are 50 difficult samples in this group, then there are 50 difficult samples in the initial sample library. Using the constructed initial semantic segmentation model to predict one group among the remaining 99 groups to obtain the predicted difficult samples of this group. If 30 predicted difficult samples are obtained in this group, then adding these 30 predicted difficult samples to the initial sample library, and the number of difficult samples in the initial sample library increases to 80. Then retraining the semantic segmentation model according to the updated initial sample library. Further, using the semantic segmentation model trained twice to continue predicting one group of candidate key frames among the remaining 98 groups. Updating the initial sample library and retraining the semantic segmentation model in the above manner in sequence until all the candidate key frames are predicted.

[0063] In the embodiment of the present application, for each prediction of the semantic segmentation model, the evaluation criteria for difficult semantic segmentation samples can be improved based on the prediction confidence of a single pixel, and the corresponding index operator can be proposed. The quality of the semantic segmentation result can be evaluated from two aspects: the confidence distribution of the model predicting a single pixel and the deviation between the model prediction and the actual result. In one example, the processor can construct a prediction probability difference map, by determining the difference map between the maximum probability layer and the second maximum probability layer of the current candidate key frame, counting the number of target pixels in the difference map that are less than the first threshold thresh1, and determining the ratio of the number of target pixels to the total number of pixels of the current candidate key frame. By judging whether this ratio is greater than the second threshold thresh2, it is judged whether the current candidate key frame is a difficult-to-predict sample or an edge marking sample. Specifically, when the ratio is greater than the second threshold thresh2, it is determined that the current candidate key frame is a difficult-to-predict sample; when the ratio is not greater than the second threshold thresh2, it is determined that the current candidate key frame is an edge marking sample and needs to be further screened.

[0064] In the embodiments of the present application, the edge-marked samples are the samples to be manually confirmed. However, if the number of video frames of the edge-marked samples determined according to the semantic segmentation model is large, it is necessary to preliminarily screen the video frames by calculating the Structural Similarity (SSIM) algorithm of adjacent video frames. If there are still many edge-marked samples after screening, the original images of the remaining images can be further screened by SSIM, and the corresponding marked images can be screened out at the same time. After this screening process, the number of images to be manually confirmed will be greatly reduced, so as to improve the efficiency of manual screening. In one example, the processor can obtain a first edge-marked sample and a second edge-marked sample that are adjacent in time, and determine the second edge-marked sample as the target edge-marked sample. The first edge-marked sample is the previous picture adjacent to the target edge-marked sample in time. The first edge-marked sample at the earliest time is defaulted as the sample to be manually confirmed. The similarity between the target edge-marked sample and the previous adjacent edge-marked sample is determined, and then it is judged whether the target edge-marked sample is a sample to be manually marked according to the similarity. If the similarity is large, the current target edge-marked sample can be excluded and there is no need to repeat the manual marking. If the similarity is small, the current target edge-marked sample is quite different from the previous one and needs to be manually marked. For example, it is judged whether the similarity is less than the third threshold thresh3. In the case where the similarity is less than the third threshold thresh3, it is determined that the target edge-marked sample is a sample to be manually marked. In this way, the number of manual markings can be reduced, the scale of the images to be manually confirmed can be reduced, the efficiency of manual screening can be improved, and the difficult-to-mark samples in the samples to be manually marked can be obtained with higher efficiency. Further, the predicted difficult samples determined according to the semantic segmentation model and the difficult-to-mark samples determined by manual marking are determined as candidate difficult samples.

[0065] Through the above technical solution, a semantic segmentation model is constructed according to the obtained candidate key frames; predicted difficult samples and edge-marked samples are determined through the semantic segmentation model; the edge-marked samples are screened to obtain difficult-to-mark samples; and then candidate difficult samples are determined according to the predicted difficult samples and the difficult-to-mark samples. Compared with directly using the model prediction results and marking them on the original image and then performing manual confirmation, the present application determines the predicted difficult samples and edge-marked samples through the semantic segmentation model, and through various screening methods, reduces the scale of the images to be manually confirmed, improves the recognition efficiency of difficult samples, and improves the quality of the sample library of difficult samples.

[0066] In the embodiments of the present application, step 104, constructing a semantic segmentation model according to the candidate key frames may include:

[0067] Dividing the candidate key frames into multiple groups of candidate key frames;

[0068] Select candidate key frames of a preset group for annotation to obtain an initial sample library;

[0069] Train a semantic segmentation model according to the initial sample library, where the semantic segmentation model is used to predict the remaining candidate key frames;

[0070] After predicting each group of remaining candidate key frames, update the initial sample library and retrain the semantic segmentation model to update the semantic segmentation model.

[0071] Specifically, semantic segmentation is a basic task in image segmentation, which means that for an image, each pixel is labeled with the corresponding category without distinguishing individuals. Simply put, it is to divide the input data of a visual image into different semantically interpretable categories. After the processor obtains the candidate key frames, it can construct a semantic segmentation model according to the candidate key frames, and determine the difficult-to-predict samples and edge-labeled samples through the semantic segmentation model. Among them, the difficult-to-predict samples are the difficult samples that can be directly determined by the semantic segmentation model; the edge-labeled samples are the samples to be manually labeled.

[0072] In the embodiment of the present application, the construction of the semantic segmentation model may include the following steps: divide the candidate key frames into multiple groups (such as m groups), select the candidate key frames of several groups (such as a preset group) for annotation to obtain an initial sample library, and train the semantic segmentation model using the labeled data of the initial sample library. Use the semantic segmentation model trained according to the initial sample library to identify the remaining candidate key frames, where the remaining candidate key frames are the other candidate key frames except those included in the preset group. For example, use the semantic segmentation model to identify one group in the remaining groups; add the remaining candidate key frames identified as difficult-to-predict samples to the initial sample library and retrain the semantic segmentation model, and so on, until all the remaining candidate key frames are identified by the continuously updated semantic segmentation model. By updating the initial sample library and the semantic segmentation model after each group of predictions, the accuracy of the semantic segmentation model can be continuously improved and the model performance can be improved.

[0073] In the embodiment of the present application, after predicting each group of remaining candidate key frames, updating the initial sample library and retraining the semantic segmentation model to update the semantic segmentation model may include:

[0074] Add the difficult-to-predict samples corresponding to the current group to the updated initial sample library of the previous group to obtain the current sample library;

[0075] Retrain the semantic segmentation model according to the current sample library to obtain the current semantic segmentation model; the current semantic segmentation model is used to predict the next group of remaining candidate key frames.

[0076] Specifically, after each group of remaining candidate key frames is predicted, the remaining candidate key frames identified as difficult-to-predict samples among them are added to the initial sample library, and the semantic segmentation model is retrained. This process is repeated until all the remaining candidate key frames are identified by the continuously updated semantic segmentation model. Taking 10,000 candidate key frames as an example, with m being 100, the 10,000 candidate key frames are divided into 100 groups, with each group including 100 candidate key frames. One group is selected for annotation to obtain the initial semantic segmentation model. Assuming there are 50 difficult samples in this group, then there are 50 difficult samples in the initial sample library. The constructed initial semantic segmentation model is used to predict one of the remaining 99 groups to obtain the predicted difficult samples of this group. Suppose 30 predicted difficult samples are obtained in this group, then these 30 predicted difficult samples are added to the initial sample library, and the number of difficult samples in the initial sample library increases to 80. Then, the semantic segmentation model is retrained according to the updated initial sample library. Further, the secondarily trained semantic segmentation model is used to continue predicting one group of candidate key frames among the remaining 98 groups. The initial sample library is updated and the semantic segmentation model is retrained in the above manner until all the candidate key frames are predicted. In this way, the accuracy of the semantic segmentation model can be continuously improved and the model performance can be improved.

[0077] In the embodiment of the present application, step 106, determining the difficult-to-predict samples and the edge marking samples through the semantic segmentation model includes:

[0078] For each candidate key frame, determining the difference map between the maximum probability layer and the second maximum probability layer of the current candidate key frame;

[0079] Counting the number of target pixels in the difference map whose differences are less than the first threshold;

[0080] Determining the ratio of the number of target pixels to the total number of pixels of the current candidate key frame;

[0081] Judging whether the ratio is greater than the second threshold;

[0082] In the case where the ratio is greater than the second threshold, determining that the current candidate key frame is a difficult-to-predict sample;

[0083] In the case where the ratio is not greater than the second threshold, determining that the current candidate key frame is an edge marking sample.

[0084] Specifically, for each prediction of the semantic segmentation model, the evaluation criteria for difficult samples of semantic segmentation can be improved based on the prediction confidence of a single pixel, and the corresponding index operator can be proposed. The quality of the semantic segmentation result can be evaluated from two aspects: the confidence distribution of the model predicting a single pixel and the deviation between the model prediction and the actual result. In one example, the processor can construct a prediction probability difference map. By determining the difference map between the maximum probability layer and the second maximum probability layer of the current candidate key frame, counting the number of target pixels in the difference map that are less than the first threshold thresh1, and determining the ratio of the number of target pixels to the total number of pixels of the current candidate key frame. By determining whether this ratio is greater than the second threshold thresh2, it is determined whether the current candidate key frame is a difficult prediction sample or a marginal labeled sample. Specifically, when the ratio is greater than the second threshold thresh2, it is determined that the current candidate key frame is a difficult prediction sample; when the ratio is not greater than the second threshold thresh2, it is determined that the current candidate key frame is a marginal labeled sample and needs to be further screened. By determining difficult prediction samples from two dimensions, the quality of difficult prediction samples can be improved, and the value of difficult prediction samples can be intuitively evaluated.

[0085] In the embodiment of the present application, step 108 of screening the marginal labeled samples to obtain difficult labeled samples may include:

[0086] Obtain a first marginal labeled sample and a second marginal labeled sample that are adjacent in time in sequence;

[0087] Determine the second marginal labeled sample as the target marginal labeled sample;

[0088] Determine the similarity between the target marginal labeled sample and the first marginal labeled sample;

[0089] Determine whether the target marginal labeled sample is a sample to be manually labeled according to the similarity;

[0090] Obtain the difficult labeled samples in the samples to be manually labeled.

[0091] Specifically, the marginal labeled sample is a sample to be manually confirmed. If the number of video frames of the marginal labeled sample determined according to the semantic segmentation model is large, it is necessary to initially screen the video frames by calculating the Structural Similarity (SSIM) algorithm of the adjacent video frames. If there are still many screened marginal labeled samples, the original images of the remaining images can be further screened by SSIM, and the corresponding labeled maps can be screened at the same time. After this screening process, the number of images to be manually confirmed will be greatly reduced to improve the efficiency of manual screening.

[0092] In an embodiment of the present application, the processor may obtain a first edge marker sample and a second edge marker sample that are adjacent in time in sequence, and determine the second edge marker sample as the target edge marker sample. The first edge marker sample is the previous picture adjacent to the target edge marker sample in time. The first edge marker sample that is the earliest in time is defaulted as the sample to be manually confirmed. The similarity between the target edge marker sample and the previous adjacent edge marker sample is determined, and then it is judged whether the target edge marker sample is a sample to be manually marked according to the similarity. In this way, the number of manual markings can be reduced, the scale of the images to be manually confirmed can be reduced, the efficiency of manual screening can be improved, and the difficult-to-mark samples in the samples to be manually marked can be obtained with higher efficiency.

[0093] In an embodiment of the present application, judging whether the target edge marker sample is a sample to be manually marked according to the similarity may include:

[0094] Judging whether the similarity is less than a third threshold;

[0095] In the case where the similarity is less than the third threshold, it is determined that the target edge marker sample is a sample to be manually marked.

[0096] Specifically, it may be judged whether the target edge marker sample is a sample to be manually marked by SSIM. If the similarity is large, the current target edge marker sample can be excluded and does not need to be manually marked repeatedly. If the similarity is small, the current target edge marker sample is quite different from the previous one and needs to be manually marked. For example, it is judged whether the similarity is less than the third threshold thresh3. In the case where the similarity is less than the third threshold thresh3, it is determined that the target edge marker sample is a sample to be manually marked. By judging the similarity, the number of samples to be manually marked can be reduced.

[0097] In an embodiment of the present invention, the similarity may satisfy the following formula:

[0098]

[0099] c1 = (k1L) 2 ;

[0100] c2 = (k2L) 2 ;

[0101] where SSIM(x, y) is the similarity between the target edge marker sample and the first edge marker sample; x and y are the target edge marker sample and the first edge marker sample respectively; μ x and μ y are respectively the averages of the image gray matrices of the target edge marker sample and the first edge marker sample; σ x 2 and σ y 2They are the variance values of the image grayscale matrices of the target edge marker sample and the first edge marker sample respectively; σ xy is the covariance of the image grayscale matrices of the target edge marker sample and the first edge marker sample; c1 and c2 are constants used to maintain stability; L is the dynamic range of pixel values; k1 = 0.01; k2 = 0.03.

[0102] In the embodiment of the present application, step 102, obtaining candidate key frames may include:

[0103] Obtaining candidate key frames containing motion by the three-frame difference method.

[0104] The primary selection of candidate key frames is to initially screen out motion frames from a large amount of videos, and initially distinguish background frames and foreground frames. In the embodiment of the present application, candidate key frames containing motion can be obtained by the three-frame difference method. Specifically, the method of obtaining candidate key frames containing motion by the three-frame difference method may include: obtaining a first video frame, a second video frame, and a third video frame that are adjacent in time in sequence, and determining the third video frame as the target video frame. Then, performing difference processing on the first video frame and the second video frame to obtain a first adjacent difference map, and at the same time performing difference processing on the second video frame and the target video frame to obtain a second adjacent difference map. Further, determining the similarity between the target video frame and the second video frame, and judging whether the target video frame contains motion according to the first adjacent difference map and the second adjacent difference map. In the case where the target video frame contains motion and the similarity is less than a first set value, determining the target video frame as a foreground key frame; in the case where the target video frame does not contain motion and the similarity is less than a second set value, determining the target video frame as a background key frame. Among them, the background key frame is a key frame that does not include motion, and the foreground key frame is a key frame that includes motion. What the embodiment of the present application needs to obtain is the foreground key frame containing motion. The embodiment of the present application uses the three-frame difference method to determine motion frames mainly considering that it can achieve a balance between computational efficiency and motion perception, extracting three video frames that are adjacent in time in sequence, and performing image sharpening on three consecutive frames of images to reduce the impact of uneven illumination on motion detection. Further, judging whether the target video frame contains motion and the similarity between the target video frame and the adjacent video frames, so as to be able to distinguish foreground key frames and background key frames while extracting frames, without the need to spend a lot of manpower for subsequent distinction, improving the efficiency and quality of selecting key frames.

[0105] Figure 2 Schematically shows a structural block diagram of a device for extracting difficult samples according to an embodiment of the present application. As Figure 2 shown, the embodiment of the present application provides a device for extracting difficult samples, which may include:

[0106] A memory 210, configured to store instructions; and

[0107] A processor 220, configured to call instructions from a memory 210 and capable of implementing the above method for extracting difficult samples when executing the instructions.

[0108] Specifically, in an embodiment of the present application, the processor 220 may be configured to:

[0109] Obtain candidate key frames;

[0110] Construct a semantic segmentation model based on the candidate key frames;

[0111] Determine predicted difficult samples and edge-labeled samples through the semantic segmentation model;

[0112] Screen the edge-labeled samples to obtain labeled difficult samples;

[0113] Determine candidate difficult samples based on the predicted difficult samples and the labeled difficult samples.

[0114] Furthermore, the processor 220 may also be configured to:

[0115] Constructing a semantic segmentation model based on the candidate key frames includes:

[0116] Divide the candidate key frames into multiple groups of candidate key frames;

[0117] Select a preset group of candidate key frames for annotation to obtain an initial sample library;

[0118] Train a semantic segmentation model based on the initial sample library, and the semantic segmentation model is used to predict the remaining candidate key frames;

[0119] After predicting each group of remaining candidate key frames, update the initial sample library and retrain the semantic segmentation model to update the semantic segmentation model.

[0120] Furthermore, the processor 220 may also be configured to:

[0121] After predicting each group of remaining candidate key frames, updating the initial sample library and retraining the semantic segmentation model to update the semantic segmentation model includes:

[0122] Add the predicted difficult samples corresponding to the current group to the updated initial sample library of the previous group to obtain the current sample library;

[0123] Retrain the semantic segmentation model based on the current sample library to obtain the current semantic segmentation model; the current semantic segmentation model is used to predict the next group of remaining candidate key frames.

[0124] Furthermore, the processor 220 may also be configured to:

[0125] Determining predicted difficult samples and edge-labeled samples through the semantic segmentation model includes:

[0126] For each candidate key frame, determine the difference map between the maximum probability layer and the second maximum probability layer of the current candidate key frame;

[0127] Count the number of target pixels in the difference map whose differences are less than the first threshold;

[0128] Determine the ratio of the number of target pixels to the total number of pixels of the current candidate key frame;

[0129] Judge whether the ratio is greater than the second threshold;

[0130] If the ratio is greater than the second threshold, determine that the current candidate key frame is a difficult-to-predict sample;

[0131] If the ratio is not greater than the second threshold, determine that the current candidate key frame is an edge-labeled sample.

[0132] Further, the processor 220 can also be configured to:

[0133] Screen the edge-labeled samples to obtain labeled difficult samples, including:

[0134] Obtain a first edge-labeled sample and a second edge-labeled sample that are adjacent in time in sequence;

[0135] Determine the second edge-labeled sample as the target edge-labeled sample;

[0136] Determine the similarity between the target edge-labeled sample and the first edge-labeled sample;

[0137] Judge whether the target edge-labeled sample is a sample to be manually labeled according to the similarity;

[0138] Obtain the labeled difficult samples in the samples to be manually labeled.

[0139] Further, the processor 220 can also be configured to:

[0140] Judge whether the target edge-labeled sample is a sample to be manually labeled according to the similarity, including:

[0141] Judge whether the similarity is less than the third threshold;

[0142] If the similarity is less than the third threshold, determine that the target edge-labeled sample is a sample to be manually labeled.

[0143] Further, the processor 220 can also be configured to:

[0144] Obtain candidate key frames, including:

[0145] Obtain candidate key frames containing motion by using the three-frame difference method.

[0146] Through the above technical solution, a semantic segmentation model is constructed based on the obtained candidate key frames; the prediction difficult samples and the edge marked samples are determined through the semantic segmentation model; the edge marked samples are screened to obtain the marked difficult samples; and then the candidate difficult samples are determined according to the prediction difficult samples and the marked difficult samples. Compared with directly using the model prediction results and marking them on the original image and then performing manual confirmation, in this application, the prediction difficult samples and the edge marked samples are determined through the semantic segmentation model, and through various screening methods, the scale of the pictures to be manually confirmed is reduced, the recognition efficiency of the difficult samples is improved, and the quality of the sample library of the difficult samples is improved.

[0147] Figure 3 Schematically shows a structural schematic diagram of a mechanical device according to an embodiment of the present application. As Figure 3 shown, an embodiment of the present application further provides a mechanical device, which may include:

[0148] A video acquisition device 310 for acquiring a video of a moving scene with a fixed perspective;

[0149] The above device 320 for extracting difficult samples.

[0150] In an embodiment of the present invention, the video acquisition module 310 is electrically connected to the device 320 for extracting difficult samples. The video acquisition module 310 acquires a video of a moving scene with a fixed perspective, transmits the video to the device 320 for extracting difficult samples, and this device obtains candidate key frames; constructs a semantic segmentation model according to the candidate key frames; determines prediction difficult samples and edge marked samples through the semantic segmentation model; screens the edge marked samples to obtain marked difficult samples; and determines candidate difficult samples according to the prediction difficult samples and the marked difficult samples. In this way, for the problem of selecting difficult samples from a large number of videos of a moving scene with a fixed perspective, the prediction difficult samples and the edge marked samples are determined through the semantic segmentation model, and through various screening methods, the scale of the pictures to be manually confirmed is reduced, the recognition efficiency of the difficult samples is improved, and the quality of the sample library of the difficult samples is improved.

[0151] An embodiment of the present application further provides a machine-readable storage medium, on which instructions are stored, and these instructions are used to cause a machine to execute the above method for extracting difficult samples.

[0152] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0153] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce means for implementing the specified functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the specified functions in one block or multiple blocks.

[0154] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the specified functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the specified functions in one block or multiple blocks.

[0155] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the specified functions in one block or multiple blocks.

[0156] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0157] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0158] Computer readable media include permanent and non-permanent, removable and non-removable media, and can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0159] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0160] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.

Claims

1. A method for extracting difficult samples, characterized in that, Comprising: Obtaining candidate key frames; Constructing a semantic segmentation model according to the candidate key frames; Determining predicted difficult samples and edge marked samples through the semantic segmentation model; Screening the edge marked samples to obtain marked difficult samples; Determining candidate difficult samples according to the predicted difficult samples and the marked difficult samples; The determining the predicted difficult samples and the edge marked samples through the semantic segmentation model includes: For each candidate key frame, determining a difference map between the maximum probability layer and the second maximum probability layer of the current candidate key frame; Counting the number of target pixels in the difference map whose differences are less than a first threshold; Determining a ratio of the number of target pixels to the total number of pixels of the current candidate key frame; Judging whether the ratio is greater than a second threshold; When the ratio is greater than the second threshold, determining that the current candidate key frame is a predicted difficult sample; When the ratio is not greater than the second threshold, determining that the current candidate key frame is an edge marked sample; The screening the edge marked samples to obtain marked difficult samples includes: Obtaining a first edge marked sample and a second edge marked sample that are adjacent in time in sequence; Determining the second edge marked sample as the target edge marked sample; Determining a similarity between the target edge marked sample and the first edge marked sample; Judging whether the target edge marked sample is a sample to be manually marked according to the similarity; Obtaining the marked difficult samples in the samples to be manually marked.

2. The method according to claim 1, characterized in that, The constructing the semantic segmentation model according to the candidate key frames includes: Dividing the candidate key frames into multiple groups of candidate key frames; Selecting a preset group of candidate key frames for annotation to obtain an initial sample library; Training the semantic segmentation model according to the initial sample library, where the semantic segmentation model is used to predict the remaining candidate key frames; After predicting each group of remaining candidate key frames, updating the initial sample library and retraining the semantic segmentation model to update the semantic segmentation model.

3. The method according to claim 2, characterized in that, The after predicting each group of remaining candidate key frames, updating the initial sample library and retraining the semantic segmentation model to update the semantic segmentation model includes: Adding the predicted difficult samples corresponding to the current group to the updated initial sample library of the previous group to obtain the current sample library; Retraining the semantic segmentation model according to the current sample library to obtain the current semantic segmentation model; the current semantic segmentation model is used to predict the next group of remaining candidate key frames.

4. The method according to claim 1, characterized in that, The judging whether the target edge marked sample is a sample to be manually marked according to the similarity includes: Judging whether the similarity is less than a third threshold; When the similarity is less than the third threshold, determining that the target edge marked sample is a sample to be manually marked.

5. The method according to claim 1, characterized in that, The obtaining candidate key frames includes: Obtaining candidate key frames containing motion through a three-frame difference method.

6. A device for extracting difficult samples, characterized in that, Comprising: A memory configured to store instructions;And A processor configured to call the instructions from the memory and capable of implementing the method for extracting difficult samples according to any one of claims 1 to 5 when executing the instructions.

7. A mechanical device, characterized in that, Comprising: A video acquisition device for acquiring a video of a moving scene with a fixed perspective; The apparatus for extracting difficult samples according to claim 6.

8. A machine-readable storage medium, characterized in that, Instructions are stored on the machine-readable storage medium, and the instructions are used to cause the machine to execute the method for extracting difficult samples according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Chinese microblog emotion analysis method based on collaborative learning under loose condition

    CN108228569A

  • Medical image sample screening method and device, computer equipment and storage medium

    CN111666993A

  • Method and device for extracting video key frame and controller

    CN113794815A