Training Method, Device, Terminal Device and Medium of Feature Extraction Network

By determining the background interval and foreground pixel points of the video frame in self-supervised learning, acquiring the foreground images and constructing a sample pair, self-supervising training of the feature extraction network, the problem of incomplete features in the prior art is solved, and the matching degree and consistency of the feature extraction network are improved.

CN114419489BActive Publication Date: 2025-06-03SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111614706.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-25
Publication Date
2025-06-03
Estimated Expiration
2041-12-25

AI Technical Summary

Technical Problem

The existing self-supervised learning method cannot effectively obtain the complete information of the object during the image feature extraction process, resulting in incomplete features and low matching, affecting downstream tasks of object detection and segmentation.

Method used

By determining the background interval of each pixel point in the video frame sequence, identifying the foreground pixel points in the target video frame, acquiring the foreground image, and constructing positive and negative sample pairs, the feature extraction network is self-supervised.

Benefits of technology

The integrity and consistency of features learned by the feature extraction network is improved, and the matching between the feature extraction network and the object detection and segmentation tasks is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114419489B_ABST
    Figure CN114419489B_ABST
Patent Text Reader

Abstract

The embodiments of the present application are applicable to the field of image processing technology, and provide a training method, device, terminal device and medium for a feature extraction network. The method includes: determining a background interval of each pixel point in a video frame sequence; determining foreground pixel points in a target video frame according to the background interval, where the target video frame is any video frame in the video frame sequence; obtaining foreground images in each of the target video frames according to the foreground pixel points; determining positive samples and negative samples of the foreground image of the target video frame to construct positive sample pairs and negative sample pairs; and performing self-supervised training on a preset feature extraction network by using the positive sample pairs and the negative sample pairs. The feature extraction network trained by using the above method has a high matching degree in downstream tasks of target detection and segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of image processing, and particularly relates to a training method, device, terminal device and medium for a feature extraction network. Background Art

[0002] Self-supervised learning mainly uses auxiliary tasks to mine its own supervision information from large-scale unsupervised data, and trains the network through this constructed supervision information, so as to learn valuable representations for downstream tasks.

[0003] Self-supervised learning methods have been widely used in the field of object detection. However, the regions intercepted during sampling on the original image by existing self-supervised learning methods cannot reflect the complete information of the object or carry a large amount of irrelevant information. Therefore, the features learned currently are incomplete or redundant, resulting in a low matching degree of the downstream tasks of object detection and segmentation. Summary of the Invention

[0004] In view of this, the embodiments of this application provide a training method, device, terminal device and medium for a feature extraction network, so as to solve the problem that the features in the image extracted during the self-supervised training process in the prior art are incomplete and have a low matching degree for the downstream tasks of object detection and segmentation.

[0005] The first aspect of the embodiments of this application provides a training method for a feature extraction network, including:

[0006] Determine the background interval of each pixel point in the video frame sequence;

[0007] According to the background interval, determine the foreground pixel points in the target video frame, where the target video frame is any video frame in the video frame sequence;

[0008] According to the foreground pixel points, obtain the foreground images in each of the target video frames;

[0009] Determine the positive samples and negative samples of the foreground image of the target video frame to construct positive sample pairs and negative sample pairs;

[0010] Perform self-supervised training on a preset feature extraction network using the positive sample pairs and the negative sample pairs.

[0011] The second aspect of the embodiments of this application provides a training device for a feature extraction network, which is characterized by including:

[0012] A background interval determination module, configured to determine the background interval of each pixel point in the video frame sequence;

[0013] Foreground pixel determination module, configured to determine foreground pixels in a target video frame according to the background interval, where the target video frame is any video frame in the video frame sequence;

[0014] Foreground image acquisition module, configured to acquire foreground images in each of the target video frames according to the foreground pixels;

[0015] Sample pair determination module, configured to determine positive samples and negative samples of the foreground image of the target video frame to construct positive sample pairs and negative sample pairs;

[0016] Training module, configured to perform self-supervised training on a preset feature extraction network by using the positive sample pairs and the negative sample pairs. A third aspect of the embodiments of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the training method of the feature extraction network described in the first aspect above is implemented.

[0017] A fourth aspect of the embodiments of the present application provides a computer-readable storage medium storing a computer program, where when the computer program is executed by a processor, the training method of the feature extraction network described in the first aspect above is implemented.

[0018] A fifth aspect of the embodiments of the present application provides a computer program product, which when running on a terminal device, causes the terminal device to execute the training method of the feature extraction network described in the first aspect above.

[0019] Compared with the prior art, the embodiments of the present application include the following advantages:

[0020] In the embodiments of the present application, during the training process of the feature extraction network using the self-supervised method, first, the background interval of the pixels in the video frame sequence is determined, and then the foreground pixels of the video frame are determined according to the background interval; based on the foreground pixels, the foreground image in the video frame can be determined, so that the video frame can be sampled purposefully, and the sampling result includes the target of interest, making the features learned by the trained feature extraction network more complete, improving the matching degree between the feature extraction network and the downstream tasks of target detection and segmentation of the response, and at the same time, when the feature extraction network trained by this method is used for feature extraction, the extracted features have good consistency. Description of the Drawings

[0021] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0022] Figure 1 It is a schematic flowchart of the steps of a training method for a feature extraction network according to an embodiment of the present application;

[0023] Figure 2 It is a schematic flowchart of the steps of another training method for a feature extraction network according to an embodiment of the present application;

[0024] Figure 3 It is a schematic diagram of a method for determining a foreground object frame according to an embodiment of the present application;

[0025] Figure 4 It is a schematic diagram of image feature extraction according to an embodiment of the present application;

[0026] Figure 5 It is a schematic diagram of an apparatus for a feature extraction network according to an embodiment of the present application;

[0027] Figure 6 It is a schematic diagram of a terminal device according to an embodiment of the present application. Detailed implementation manners

[0028] In the following description, specific details such as specific system structures and technologies are proposed for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0029] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0030] It should also be understood that the term "and / or" used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0031] As used in the specification of this application and the appended claims, the term "if" may be construed contextually as "when", or "once", or "in response to determining", or "in response to detecting". Similarly, the phrase "if determined" or "if [the described condition or event] is detected" may be construed contextually to mean "once determined", or "in response to determining", or "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]".

[0032] In addition, in the description of the specification of this application and the appended claims, the terms "first", "second", "third", etc. are only used for differentiating descriptions and cannot be construed as indicating or implying relative importance.

[0033] The reference to "one embodiment" or "some embodiments" or the like described in the specification of this application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized.

[0034] Generally, in the process of target segmentation and detection, a feature extraction model is required to extract features; the feature extraction model can be obtained through self-supervised training. In the process of self-supervised training, the feature extraction network can be trained through an auxiliary task so that the features extracted by the feature extraction network from the image meet the expectations.

[0035] Generally, for the same target in different images, the features obtained by using the feature extraction network should have a certain degree of similarity; for different targets, the extracted features should have a certain degree of distinctiveness. If a feature network can meet the above requirements, it can be considered that the feature extraction network can perform deep learning on the information in the image and achieve the purpose of training.

[0036] Based on the above, in the embodiments of this application, the auxiliary task can be set as: the features of different images of the same target have the expected similarity; the features of images of different targets have a certain degree of distinctiveness.

[0037] In the embodiments of this application, when sampling, random sampling is not performed. Instead, the foreground and background of the image are distinguished. The foreground in the image is generally the region of interest when performing tasks such as target detection; the background in the image is generally the fixed scene of the picture shooting area.

[0038] In this embodiment, the foreground image is separated from the picture for sampling, so that the sample can contain the complete information of the foreground, avoiding information loss and redundancy, which is beneficial to achieving the feature consistency of the extracted features.

[0039] The technical solution of the present application will be described below through specific embodiments.

[0040] Referring to Figure 1 , a schematic flowchart of the steps of a training method for a feature extraction network according to an embodiment of the present application is shown, which may specifically include the following steps:

[0041] S101, determining the background interval of each pixel point in the video frame sequence.

[0042] The execution subject of this embodiment is a terminal device, such as a monitoring device, etc. The method in this embodiment can obtain an image feature extraction algorithm and be applied in the process of target segmentation and detection.

[0043] Specifically, the above video frame sequence can be obtained by processing monitoring data. Monitoring data generally includes a plurality of consecutive video frames arranged in chronological order. However, during the process of shooting video frames, there may be some special situations that cause the image information in the video frames to be missing. For example, in a dark environment, the pictures taken by the monitoring device are not clear, or when the lens of the monitoring device is blocked, the pictures taken do not contain the target content. Therefore, when continuing to train with monitoring data, it is necessary to eliminate the unavailable pictures from the monitoring data, collect the available picture information, and arrange it in chronological order to form the above video frame sequence.

[0044] The above background interval refers to the range of change of the pixel value of each pixel point in the background picture of the video frame picture. The corresponding video frames in the video frame sequence have the same size, the same number of pixel points in each video frame, and the same arrangement; for the pixel points at the same position in each video frame in the video frame sequence, the range of change can be calculated to determine the background interval of the pixel point. In this embodiment, the background interval is used to describe the parts that are generally not concerned when processing images. For example, when judging the target passing through a certain period of time in a monitoring video, generally, the background area of the monitoring video, such as walls, lawns, etc., is not concerned. These parts are generally unchanged. Therefore, when no one passes by, the images in the video frames taken are the same, but there will be some errors due to changes in temperature, weather, etc. Therefore, the values of the pixel points at each corresponding position of their images are within a certain fluctuation range. In this embodiment, the background interval is used to determine whether there is a foreground passing through the shooting area of the monitoring video.

[0045] S102. Determine the foreground pixel points in the target video frame according to the background interval, where the target video frame is any video frame in the video frame sequence.

[0046] Specifically, the target video frame refers to a video frame in the video frame sequence. For the target video frame, each pixel point therein can be compared with the background interval of the corresponding pixel point. If the pixel value of the pixel point is within the background interval, then the pixel point is regarded as a background pixel point; if the pixel value of the pixel point is not within the background interval, then the pixel point is regarded as a foreground pixel point. For example, the background interval of the pixel point arranged in the tenth row and tenth column is 345 - 567, and the pixel value of the pixel point arranged in the tenth row and tenth column in the target video is 789, which indicates that the pixel point arranged in the tenth row and tenth column in the target video frame is a foreground pixel point; the background interval of the pixel point arranged in the third row and fifth column is 123 - 367, and the pixel value of the pixel point arranged in the third row and fifth column in the target video is 234, which indicates that the pixel point arranged in the third row and fifth column in the target video frame is not a foreground pixel point but a background pixel point.

[0047] S103. Obtain the foreground images in each of the target video frames according to the foreground pixel points.

[0048] Specifically, the cameras of surveillance videos generally aim at the same shooting area, and the scene of this shooting area is fixed. It can be considered that the fixed scene in this shooting area is the background; when a target passes through the shooting area, the target captured in the video frame is the foreground in the video frame. In the previous step, the foreground pixel points in the target video frame were determined according to the background interval. The foreground pixel points can be a pixel point in the foreground. Therefore, the foreground image in the video frame can be determined according to the foreground pixel points.

[0049] Specifically, for each video frame, the foreground image in the target video frame can be determined according to the image area composed of the foreground pixel points.

[0050] Specifically, image binarization processing can be performed on the target video frame. For example, if the pixel point is a foreground pixel point, then the pixel value of the pixel point is set to 1; if the pixel point is not a foreground pixel point, then the pixel value of the pixel point can be set to 0. In this way, the pixel points with a pixel value of 1 can form a connected image in the target video frame, and the connected image is the foreground image.

[0051] Specifically, a video frame can include multiple foreground images, and each foreground image can correspond to a connected area. For example, if the pixel points with a pixel value of 1 in the video frame are combined into 3 connected areas, then it can be considered that these 3 connected areas respectively correspond to a foreground image.

[0052] S104. Determine the positive and negative samples of the foreground image of the target video frame to construct positive sample pairs and negative sample pairs.

[0053] Specifically, the above positive sample refers to an image region that has many identical features with the foreground image; the negative sample refers to an image region that is completely different from the foreground image.

[0054] Specifically, the foreground image corresponds to a target. Other images of this target in other video frames can be used as the positive samples of the foreground image in the target video frame; images of different targets in other video frames can be used as the negative samples of the foreground image in the target video frame. Exemplarily, adjacent video frames of this video frame can be selected. The same foreground in the target video frame and the adjacent video frames is used as a positive sample pair, and different foregrounds in the target video frame and the adjacent video frames are used as a negative sample pair.

[0055] Specifically, a video frame separated from the target video frame by a certain number of frames can also be selected, and then the same foreground is selected as the positive sample and different foregrounds are selected as the negative samples. The same foreground corresponds to the same target, and different foregrounds correspond to different targets.

[0056] Exemplarily, there is a foreground image Z1 in the target video frame, and its corresponding target is Z. In a sample video frame 5 video frames before the target video frame, there are two foreground images Z2 and Y1, and the corresponding targets are Z and Y respectively. Then Z2 can be used as the positive sample of Z1, and Y1 can be used as the negative sample of Z1; that is, Z1 and Z2 form a positive sample pair, and Z1 and Y1 form a negative sample pair.

[0057] S105. Use the positive sample pairs and the negative sample pairs to perform self-supervised training on a preset feature extraction network.

[0058] Specifically, in self-supervised learning training, it is generally necessary to construct an auxiliary task for training. Using positive sample pairs and negative sample pairs to continue training the feature extraction network based on the auxiliary task can make the trained feature extraction network have feature consistency.

[0059] In a possible implementation, positive sample pairs and negative sample pairs extracted from video frames are respectively used as training samples and input into a feature extraction network for training. After training the feature extraction network with a sample image pair, the trained feature extraction network can be used to extract the features of a foreground image, the positive sample corresponding to the foreground image, and the negative sample corresponding to the foreground image in a target video frame, obtaining corresponding foreground image features, positive sample features, and negative sample features. Then, calculate the positive sample feature difference between the foreground image feature and the positive sample feature; calculate the negative sample feature difference between the foreground image feature and the negative sample feature; when the negative sample feature difference is much larger than the positive sample feature difference, it can be shown that the feature extraction network can extract the same features for the same graphics and different features for different images. At this time, the feature extraction network is trained. When the gap between the negative sample feature difference and the positive sample feature difference does not reach the preset value, it indicates that the feature extraction network cannot recognize the same and different images, and the feature extraction network needs to be continuously trained with the images in the positive sample pairs and negative sample pairs. In this application, it is possible to detect whether the features in the positive sample pairs are the same and whether the features in the negative sample pairs are different, so as to determine whether the model is trained.

[0060] In addition, the auxiliary tasks can also include multiple types. For example, shuffling the order of the video frame sequence. If the video frame sequence can be reordered based on an algorithm, it can be shown that the content in the image can already be recognized and the training is completed.

[0061] In this embodiment, when sampling, instead of randomly selecting a region in the image, a region with a foreground is selected, so that the trained network can extract similar features for different images of the same target and different features for images of different targets. Furthermore, in subsequent object detection and object segmentation tasks, the network can better adapt to these downstream tasks.

[0062] Refer to Figure 2 , which shows the schematic flow chart of the steps of another method for training a feature extraction network according to an embodiment of the present application. Specifically, it can include the following steps:

[0063] S201, perform Gaussian fitting on the pixel points at each corresponding position of each video frame in the video frame sequence to obtain the fluctuation range of the pixel values of the pixel points.

[0064] The execution subject of this embodiment is a terminal device. By using the method in this embodiment, a feature extraction algorithm can be obtained; the feature extraction algorithm can be applied to fields such as image analysis and video analysis for object segmentation and detection.

[0065] Specifically, by performing Gaussian fitting on the pixel values of the corresponding positions of each pixel in multiple video frames in the video frame sequence, the fluctuation range of the pixel value of the pixel can be obtained.

[0066] In another possible implementation, a video frame image that only contains the background image in the surveillance video can be directly selected, and the maximum pixel value and the minimum pixel value corresponding to each pixel are determined; the fluctuation range of the pixel value of the pixel is determined according to the maximum pixel value and the minimum pixel value. Equivalently, the background interval is the change interval of the pixel values of the pixels in the background image.

[0067] S202, Use the fluctuation range of the pixel value of each pixel as the background interval of each pixel. Specifically, the calculated fluctuation range can be used as the background interval corresponding to the pixel.

[0068] S203, Determine whether the pixel value of the pixel in the target video frame is within the background interval corresponding to the pixel.

[0069] Specifically, the background interval refers to the possible range of the pixel values of each pixel in the background image captured by the surveillance video under normal conditions considering changes in light, weather, etc. When the pixel value of the pixel is not within the background interval, it indicates that the pixel corresponds to a foreground pixel.

[0070] Specifically, for the currently processed target video frame, it can be determined whether the pixel value of the pixel is within the background interval, so as to determine whether a certain pixel in the target video frame is a foreground pixel.

[0071] S204, If the pixel value of the pixel is not within the background interval corresponding to the pixel, determine that the pixel is a foreground pixel.

[0072] Specifically, if the pixel value of the pixel is not within the background interval, it means that the pixel does not belong to the background part but belongs to the foreground part and is a foreground pixel.

[0073] S205, Obtain the foreground image in each target video frame according to the foreground pixels.

[0074] Specifically, the area image obtained by connecting the foreground pixels can be framed as the border of the foreground object, and the position of the border is recorded, so as to select the complete foreground image according to the border.

[0075] Figure 3 It is a schematic diagram of the method for determining the foreground object frame in an embodiment of the present application, as Figure 3As shown in the figure, when obtaining the foreground object frame, Gaussian fitting is performed on each pixel point in the video frame sequence to obtain the background interval threshold T for each pixel point. If the pixel point corresponding to the processed video frame is within the background interval threshold T, then the pixel point is determined as a foreground pixel point, and the pixel points not within the background interval threshold T are determined as background pixel points. The connected foreground pixel points are boxed to obtain the foreground object frame. Record the coordinate position of the foreground object frame, so that the foreground image in the video frame can be obtained using the coordinate position.

[0076] S206, Extract a preset number of sample video frames from the video frame sequence at a preset interval.

[0077] Specifically, the preset interval can be 1. For example, the previous video frame of the target video frame is used as the sample video frame. Multiple foreground images can be extracted from the sample video frame. Some of these foreground images are the same as the foreground images in the target video frame image, and some are different.

[0078] Of course, a sample video frame may only include one foreground image. At this time, multiple sample video frames can be extracted at different intervals so that multiple video frames can include multiple foreground images, thus facilitating the determination of positive and negative samples.

[0079] S207, Determine whether the foreground images of each of the sample video frames are the same as the foreground image of the target video frame.

[0080] Specifically, the foreground images of the above sample video frames being the same as the foreground image of the target video frame can mean that they correspond to the same target. For example, they can be images of the same person captured in different video frames; the sizes and positions of these two images may be different, but since they correspond to the same person, they can be considered the same.

[0081] Specifically, each of the sample video frames can include the corresponding foreground image. For a foreground image in the current target video frame, a foreground image corresponding to the same target as this foreground image can be determined from the foreground images in the sample video frames; a foreground image corresponding to a different target from this foreground image can be determined from the foreground images in the sample video frames.

[0082] In a possible implementation, the foreground image in the sample video frame and the foreground image in the target video frame can be transformed so that their sizes become the same; then compare their shapes. If the shapes are the same, it can be considered that they correspond to the same target.

[0083] In another possible implementation, the first image frame of the foreground image in the sample video frame in the sample video frame can be determined, and the second image frame of the foreground image in the target video frame in the sample video frame can be determined; the time difference between the sample video frame and the target video frame can be determined. If the time difference is less than a preset time length, and the overlapping area of the first image frame and the second image frame is greater than a preset area, it can be determined that the foreground image of the sample video frame is the same as the foreground image of the target video frame. For example, the sample video frame can be the next video frame of the target video frame, and the time difference between the two video frames is very small, so the moving range of the foreground image is not large. If it can be determined that there is a large overlap between the first image frame and the second image frame, it means that the foreground images of the sample video frame and the target video frame correspond to the same target.

[0084] In the embodiments of the present application, by determining whether the foreground image of the sample video frame is the same as the foreground image of the target video frame, positive and negative samples for training the feature extraction network can be obtained. Specifically, if the foreground images of the sample video frame and the target video frame correspond to the same target, S208 can be executed to use the foreground image in the sample video frame as a positive sample. If the foreground images of the sample video frame and the target video frame correspond to different targets, S209 can be executed to use the foreground image of the different target in the sample video frame as a negative sample.

[0085] S208, use the foreground image of the sample video frame as a positive sample.

[0086] S209, use the foreground image of the sample video frame as a negative sample.

[0087] S210, form a positive sample pair according to the foreground image of the target video frame and the positive sample.

[0088] Specifically, the foreground image and its corresponding positive sample form a positive sample pair. The targets corresponding to the two images in the positive sample pair are the same, so they should have the same features or similar features.

[0089] In another possible implementation, the background regions in the target video frame and the sample video frame that do not contain the foreground image can also be selected as positive sample pairs. Specifically, a background region can be selected from the target video frame, and then the position coordinates of the background region are determined. Based on the position coordinates, a corresponding region is selected from the sample video frame, and then the two are used as a positive sample pair. Since both are the same background part, they should also have the same features. For example, in a surveillance video, the camera always captures flowers, roads, and mountains. Then the flowers, roads, and mountains are all background images that always appear in the video frames. Two image regions corresponding to the flowers in two different video frames can be selected as a positive sample pair. Since the two image regions correspond to the same flowers, they should correspond to the same features.

[0090] Specifically, for a video frame, when processing, multiple positive sample pairs can be selected.

[0091] S211, form negative sample pairs according to the foreground image of the target video frame and the negative samples.

[0092] Specifically, the foreground image and its corresponding negative sample form a negative sample pair. The targets corresponding to the two images in the negative sample pair are different, so they should have completely different features, or the feature gap between the two should be relatively large.

[0093] S212, use the feature extraction network to extract the features in the positive sample pairs and the negative sample pairs respectively. The features include foreground image features, positive sample features, and negative sample features.

[0094] Specifically, use the feature extraction network to be trained to extract image features from the positive sample pairs and the negative sample pairs respectively. The extracted image features include foreground image features, positive sample features, and negative sample features.

[0095] In addition, the extracted features can be transformed to be in the same dimension, so as to facilitate analysis and comparison.

[0096] S213, calculate the positive sample feature gap of the positive sample pairs and the negative sample feature gap of the negative sample pairs respectively. The positive sample feature gap is the gap between the foreground image feature and the positive sample feature, and the negative sample feature gap is the gap between the foreground image feature and the negative sample feature.

[0097] Specifically, if the feature extraction network can accurately identify image features, then the foreground image feature and the positive sample feature should be relatively similar; the difference between the foreground image feature and the negative sample feature is relatively large.

[0098] When determining whether the feature extraction network is trained, an auxiliary task can be constructed. In this embodiment, the auxiliary task can be whether the image features of different images of the same target are the same; whether the image features of different targets are different.

[0099] In another possible implementation, other auxiliary tasks can also be used to determine whether the network is trained.

[0100] S214. When the distance between the positive sample feature gap and the negative sample feature gap reaches a preset distance value, it is determined that the feature extraction network is trained; otherwise, adjust the parameters of the feature extraction network and continue to train the feature extraction network.

[0101] Specifically, when the negative sample feature gap is much larger than the positive sample feature gap, it can be considered that the features of the positive sample pairs are similar and the features of the negative sample pairs are different. At this time, it can be considered that the network is trained.

[0102] When the distance between the negative sample feature gap and the positive sample feature gap does not reach the preset distance value, it indicates that the feature network is not trained. The parameters of the feature network can be adjusted. For example, if the feature extraction network is a convolutional network, a preset method can be used to adjust the weights of each convolutional layer in the feature extraction network and adjust the hyperparameters of the neurons.

[0103] Exemplarily, for two adjacent frames, the corresponding same foreground can be extracted as positive sample pairs, and different foregrounds can be used as negative sample pairs. The features extracted in this way can improve the robustness to the angle and pose of the object. In addition, the same background regions at the same positions in the adjacent two video frames that do not contain foreground regions can be selected as positive sample pairs respectively. For two frames of pictures with a certain time interval, the same method is also used to extract background blocks as positive sample pairs. In this way, due to the change of lighting conditions within a day, the features extracted can improve the robustness to the lighting conditions and colors of the object. After the feature pairs are processed by ROIAlign (Region of Interest Alignment) and converted to the same dimension, based on their feature consistency, the features of the positive sample pairs are made as consistent as possible, and the feature differences of the negative sample pairs are increased for network training. Such self-supervised training of the auxiliary task can enable the network to extract the same features for the same object regions and different features for different object regions, so that the network can better adapt to these downstream tasks in subsequent object detection and object segmentation work.

[0104] Figure 4 is a schematic diagram of image feature extraction according to an embodiment of the present application; refer to Figure 4 , two adjacent video frames are determined from the video frame sequence, and then background subtraction is performed to obtain the foreground image in each video frame. Then through the backbone network f θObtain the corresponding feature blocks; cut the features at the corresponding positions of the foreground and background in the feature blocks. Use ROIAlign to transform features of different sizes to the same scale, reducing the gap for positive sample pairs and increasing the gap for negative sample pairs.

[0105] As Figure 4 shown, the three people outlined in these two adjacent video frames are the foreground in the video frames; these three people exist in both video frames, and the feature blocks obtained through extraction by the backbone network are respectively the extracted Vn and Vn+1. As shown in the figure, the three boxes A1, A2, and A3 in Vn correspond to the three foregrounds; the three boxes B1, B2, and B3 in Vn+1 correspond to the three foregrounds; A1 and B1 correspond to the same target, A2 and B2 correspond to the same target, and A3 and B3 correspond to the same target; therefore, positive sample pairs can be obtained: A1 and B1, A2 and B2, and A3 and B3; the two feature blocks of different targets can be regarded as negative sample pairs, such as: A1 and B2, A2 and B3, and A3 and B1, etc.

[0106] In addition, the borders A4 and A5 in Vn are the background parts that do not contain the foreground image; the borders B4 and B5 in Vn+1 are in the same positions as A4 and A5 respectively and have the same background. Therefore, positive sample pairs can be obtained: A4 and B4, A5 and B5.

[0107] For the intercepted feature blocks, region of interest alignment can be performed on them respectively, and then self-supervised training can be carried out to make the features of positive sample pairs the same and the features of negative sample pairs different, thereby training the network.

[0108] Specifically, the network trained by self-supervision can be applied to downstream tasks of object detection and segmentation to verify its performance on the COCO and Pascal VOC datasets.

[0109] In this embodiment, the method fully utilizes the inter-frame information of the surveillance video, collects the same foreground and background for feature consistency training, and achieves certain results on the detection and segmentation datasets. The effects of the method in this embodiment on downstream tasks on each dataset are as follows:

[0110] The result of object detection on COCO is:

[0111] AP AP50 AP75 37.443 56.068 40.088

[0112] The result of object segmentation on COCO is:

[0113]

[0114]

[0115] The results of object detection on Pascal VOC are as follows:

[0116] AP AP50 AP75 54.8224 81.9173 59.2730

[0117] After verifying the effectiveness of the method in this embodiment on the dataset, the method in this embodiment can be used to complete downstream tasks such as object segmentation and detection.

[0118] In this embodiment, the foreground and background in the video frame are distinguished, so as to select the complete foreground image as the sample image. The image is complete and does not contain redundancy, which is convenient for training. The algorithm obtained by training has a significant improvement effect in object segmentation and detection on each dataset.

[0119] It should be noted that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0120] Referring to Figure 5 , a schematic diagram of a training device for a feature extraction network according to an embodiment of the present application is shown. Specifically, it may include a background interval determination module 51, a foreground pixel point determination module 52, a foreground image acquisition module 53, a sample pair determination module 54, and a network training module 55, where:

[0121] The background interval determination module 51 is configured to determine the background interval of each pixel point in the video frame sequence;

[0122] The foreground pixel point determination module 52 is configured to determine the foreground pixel points in the target video frame according to the background interval, where the target video frame is any video frame in the video frame sequence;

[0123] The foreground image acquisition module 53 is configured to acquire the foreground images in each of the target video frames according to the foreground pixel points;

[0124] The sample pair determination module 54 is configured to determine the positive samples and negative samples of the foreground image of the target video frame to construct positive sample pairs and negative sample pairs;

[0125] The network training module 55 is configured to perform self-supervised training on a preset feature extraction network by using the positive sample pairs and the negative sample pairs.

[0126] The above background interval determination module 51 includes:

[0127] The fluctuation range determination sub-module is configured to perform Gaussian fitting on the pixel points at each corresponding position of each video frame in the video frame sequence to obtain the fluctuation range of the pixel values of the pixel points;

[0128] A background interval determination sub-module, which is used to take the fluctuation range of the pixel value of each pixel point as the background interval of each pixel point.

[0129] The above foreground pixel point determination module 52 includes:

[0130] A first judgment sub-module, which is used to judge whether the pixel value of the pixel point in the target video frame is within the background interval corresponding to the pixel point;

[0131] A foreground pixel point determination sub-module, which is used to determine the pixel point as a foreground pixel point if the pixel value of the pixel point is not within the background interval corresponding to the pixel point.

[0132] The above foreground image acquisition module 53 includes:

[0133] A setting sub-module, which is used to set the pixel value of the foreground pixel point in the target video frame to a preset value;

[0134] A coordinate determination sub-module, which is used to determine the coordinates of the border of the closed area connected by multiple pixel points with the preset value as the pixel value;

[0135] A foreground image acquisition sub-module, which is used to acquire the foreground image from the target video frame according to the coordinates.

[0136] The above sample pair determination module 54 includes:

[0137] A sample video extraction sub-module, which is used to extract a preset number of sample video frames from the video frame sequence at a preset interval;

[0138] A second judgment sub-module, which is used to judge whether the foreground image of each sample video frame is the same as the foreground image of the target video frame;

[0139] A sample determination sub-module, which is used to take the foreground image of the sample video frame as a positive sample if the foreground image of the sample video frame is the same as the foreground image of the target video frame, otherwise, take the foreground image of the sample video frame as a negative sample;

[0140] A positive sample pair construction sub-module, which is used to construct a positive sample pair according to the foreground image of the target video frame and the positive sample;

[0141] A negative sample pair construction sub-module, which is used to construct a negative sample pair according to the foreground image of the target video frame and the negative sample.

[0142] The above device further includes:

[0143] A target area determination module, which is used to select a target area that does not contain the foreground image from the target video frame;

[0144] A positive sample determination module, configured to use an image region in the sample video frame that is in the same position as the target region as a positive sample of the target region, and the target region and the positive sample of the target region form the positive sample pair.

[0145] The above-mentioned network training module 55 includes:

[0146] A feature extraction sub-module, configured to use the feature extraction network to extract features in the positive sample pair and the negative sample pair respectively, where the features include foreground image features, positive sample features, and negative sample features;

[0147] A calculation sub-module, configured to calculate the positive sample feature gap of the positive sample pair and the negative sample feature gap of the negative sample pair respectively, where the positive sample feature gap is the gap between the foreground image feature and the positive sample feature, and the negative sample feature gap is the gap between the foreground image feature and the negative sample feature;

[0148] A third judgment sub-module, configured to determine that the feature extraction network training is completed when the distance between the positive sample feature gap and the negative sample feature gap reaches a preset distance value; otherwise, adjust the parameters of the feature extraction network and continue to train the feature extraction network.

[0149] For the apparatus embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, refer to the description in the method embodiment section.

[0150] Figure 6 The structural schematic diagram of the terminal device provided by an embodiment of the present application is as follows. As Figure 6 shown, the terminal device 6 in this embodiment includes: at least one processor 60 ( Figure 6 only one is shown in the figure), a processor, a memory 61, and a computer program 62 stored in the memory 61 and executable on the at least one processor 60. When the processor 60 executes the computer program 62, the steps in any of the above method embodiments are implemented.

[0151] The terminal device 6 may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device may include, but is not limited to, a processor 60 and a memory 61. Those skilled in the art can understand that Figure 6 merely examples of the terminal device 6, and do not constitute a limitation on the terminal device 6. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0152] The so-called processor 60 may be a Central Processing Unit (CPU), and this processor 60 may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.

[0153] In some embodiments, the memory 61 may be an internal storage unit of the terminal device 6, such as the hard disk or memory of the terminal device 6. In other embodiments, the memory 61 may also be an external storage device of the terminal device 6, such as a plug-in hard disk equipped on the terminal device 6, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 61 may also include both the internal storage unit and the external storage device of the terminal device 6. The memory 61 is used to store an operating system, application programs, a BootLoader, data, and other programs, such as the program code of the computer program, etc. The memory 61 may also be used to temporarily store data that has been output or will be output.

[0154] An embodiment of the present application also discloses a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the property real-time communication method as described in the foregoing various embodiments.

[0155] An embodiment of the present application also discloses a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, it implements the property real-time communication method as described in the foregoing various embodiments.

[0156] An embodiment of the present application also discloses a computer program product. When the computer program product runs on a terminal device, it causes the terminal device to execute the property real-time communication method as described in the foregoing various embodiments.

[0157] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0158] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0159] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0160] In the embodiments provided in this application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.

[0161] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0162] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A training method for a feature extraction network, characterized in that, it includes: Determine the background interval of each pixel point in the video frame sequence; According to the background interval, determine the foreground pixel points in the target video frame, where the target video frame is any video frame in the video frame sequence; According to the foreground pixel points, obtain the foreground images in each of the target video frames; Determine the positive samples and negative samples of the foreground image of the target video frame to construct positive sample pairs and negative sample pairs; Use the positive sample pairs and the negative sample pairs to perform self-supervised training on a preset feature extraction network.

2. The method according to claim 1, characterized in that, the determining the background interval of each pixel point in the video frame sequence includes: Perform Gaussian fitting on the pixel points at each corresponding position of each video frame in the video frame sequence to obtain the fluctuation range of the pixel values of the pixel points; Take the fluctuation range of the pixel values of each pixel point as the background interval of each pixel point.

3. The method according to claim 1, characterized in that, the determining the foreground pixel points in the target video frame according to the background interval includes: Judge whether the pixel value of the pixel point in the target video frame is within the background interval corresponding to the pixel point; If the pixel value of the pixel point is not within the background interval corresponding to the pixel point, then determine that the pixel point is a foreground pixel point.

4. The method according to any one of claims 1-3, characterized in that, the obtaining the foreground images in each of the target video frames according to the foreground pixel points includes: Set the pixel values of the foreground pixel points in the target video frame to a preset value; Determine the coordinates of the border of the closed area connected by multiple pixel points with the preset value as the pixel value; Obtain the foreground image from the target video frame according to the coordinates.

5. The method according to any one of claims 1-3, characterized in that, the determining the positive samples and negative samples of the foreground image of the target video frame to construct positive sample pairs and negative sample pairs includes: Extract a preset number of sample video frames from the video frame sequence at a preset interval; Judge whether the foreground images of the sample video frames are the same as the foreground image of the target video frame; If the foreground image of the sample video frame is the same as the foreground image of the target video frame, then take the foreground image of the sample video frame as a positive sample, otherwise, take the foreground image of the sample video frame as a negative sample; Construct a positive sample pair according to the foreground image of the target video frame and the positive sample; Construct a negative sample pair according to the foreground image of the target video frame and the negative sample.

6. The method according to claim 5, characterized in that, after determining the positive samples and negative samples of the foreground image of the target video frame, it further includes: Select a target area in the target video frame that does not contain the foreground image; Take the image area in the sample video frame that is in the same position as the target area as the positive sample of the target area, and the target area and the positive sample of the target area form the positive sample pair.

7. The method according to claim 5, characterized in that, The self-supervised training of a preset feature extraction network using the positive sample pairs and the negative sample pairs includes: Using the feature extraction network to extract the features in the positive sample pairs and the negative sample pairs respectively, where the features include foreground image features, positive sample features, and negative sample features; Calculating the positive sample feature gap of the positive sample pairs and the negative sample feature gap of the negative sample pairs respectively. The positive sample feature gap is the gap between the foreground image feature and the positive sample feature, and the negative sample feature gap is the gap between the foreground image feature and the negative sample feature; When the distance between the positive sample feature gap and the negative sample feature gap reaches a preset distance value, it is determined that the training of the feature extraction network is completed; otherwise, the parameters of the feature extraction network are adjusted, and the training of the feature extraction network is continued.

8. A training device for a feature extraction network, characterized in that, it includes: A background interval determination module for determining the background interval of each pixel point in the video frame sequence; A foreground pixel point determination module for determining the foreground pixel points in the target video frame according to the background interval, where the target video frame is any video frame in the video frame sequence; A foreground image acquisition module for acquiring the foreground images in each of the target video frames according to the foreground pixel points; A sample pair determination module for determining the positive samples and negative samples of the foreground image of the target video frame to construct positive sample pairs and negative sample pairs; A network training module for performing self-supervised training on a preset feature extraction network using the positive sample pairs and the negative sample pairs.

9. A terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, when the processor executes the computer program, it implements the training method of the feature extraction network according to any one of claims 1-7.

10. A computer-readable storage medium storing a computer program, characterized in that, when the computer program is executed by a processor, it implements the training method of the feature extraction network according to any one of claims 1-7.

Citation Information

Patent Citations

  • Video foreground detection method based on full convolutional network and conditional adversarial network

    CN110580472A

  • Motion foreground detection method and device

    CN110879951A