A method for extracting sensitive video features
By introducing the FH-SFnet network into video feature extraction, combining spatial and timing features, positioning and extracting sensitive features in videos, the problems of weak anti-interference and high leakage judgment rate in the prior art are solved, and more efficient matching of sensitive video features is achieved.
Patent Information
- Application Number
- CN202210014500.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-07
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-01-07
AI Technical Summary
The prior art has poor anti-interference in video similarity traceability and is prone to missed judgments, especially when sensitive clips in violation videos account for a relatively small amount of time.
A sensitive video feature extraction method is proposed, using the feature extraction network FH-SFnet, to collect and preprocess videos, locate sensitive fragments, and perform feature extraction to output multiple sensitive video features. This method combines spatial and timing features, and uses technical means such as multi-layer convolutional networks and ternary loss functions to accurately extract sensitive features in videos.
This method effectively improves the matching rate of sensitive video features, reduces invalid features, and reduces the missed judgment rate, solving the problem that missed judgment is prone to missed judgment during video comparison in the prior art.
Smart Images

Figure CN114663797B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video features, and specifically to a method for extracting sensitive video features. Background Art
[0002] In recent years, with the explosive growth in the number of smartphones, pictures and videos have also become the main ways to record life. While enriching life, videos also carry a large amount of illegal information, such as sensitive segments involving porn, gambling, blood and gore, etc.
[0003] The large-scale spread of illegal videos also violates relevant laws. Currently, in the process of collecting evidence from mobile phones of lawbreakers, it is necessary to trace the similarity of videos in the mobile phone. Common technical means such as MD5 code comparison, by obtaining the MD5 code of the video and comparing it with the video codes in the database. In addition, a relatively novel method is to extract the global features of the video through a neural network and perform similarity matching with the features in the database.
[0004] In video similarity tracing, the conventional MD5 code comparison technology has weak anti-interference ability. When the video undergoes appropriate changes, such as cropping, reducing the resolution, splicing other videos, etc., the MD5 code of the video will change. When extracting video features through a neural network, the features often represent the global information of the video. If the time proportion of the illegal picture is relatively small, it will lead to more interference information in the features, and it is easy to miss judgments when comparing videos. Therefore, we make improvements on this and propose a method for extracting sensitive video features. Summary of the Invention
[0005] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0006] A method for extracting sensitive video features of the present invention includes a feature extraction network FH-SFnet, and is characterized in that the feature extraction network FH-SFnet extracts sensitive video features including the following steps:
[0007] S1. Collect various types of sensitive videos in a preset quantity, and preprocess the videos to obtain temporal features and spatial features;
[0008] S2. Sensitive segment positioning, using a sensitive video positioning network to perform sensitive video positioning on the temporal features and spatial features respectively, and perform position fusion to obtain the final position of the sensitive video segment;
[0009] S3. Sensitive feature extraction, using the feature extraction network to extract features from the sensitive segment, and finally output multiple sensitive video features.
[0010] As a preferred technical solution of the present invention, the spatial feature in S1 is to perform frame sampling on the video with a step size of 4, use the RGB image as the original input, and the dimension of the input data is [img_size, img_size, 3, n], where img_size is the size of the video frame and n is the number of sampled frames. Then, multiple pre-trained spatial convolutional networks are used to extract features, and finally a feature with a dimension of n×1024 is output.
[0011] As a preferred technical solution of the present invention, the temporal feature in S1 is to perform frame sampling on the video with a step size of 4, calculate the one-dimensional optical flow with 5 frames as a group, use the optical flow field image as the original input, and the dimension of the input data is [img_size, img_size, 1, n], where img_size is the size of the video frame and n is the number of sampled frames. Then, multiple pre-trained temporal convolutional networks are used, and finally a feature with a dimension of n×1024 is output.
[0012] As a preferred technical solution of the present invention, the specific process of sensitive segment localization in S2 is as follows:
[0013] S2-1. Construct a basic feature network, which includes 4 convolutional-batch normalization-activation function ReLu basic network groups, takes the 1024-dimensional feature of video preprocessing as the input, and outputs a feature with a dimension of [32, 512];
[0014] S2-2. Construct a backbone network on the basic feature network. The backbone network has three layers from top to bottom, and the basic feature network is the middle layer. Each layer has four outputs, and the output dimensions are: [16, 1024], [8, 1024], [4, 1024], [2, 1024]. One 1024-dimensional feature generates 5 candidate localizations, and the final number of generated candidate localizations is 150;
[0015] S2-3. Backbone network fusion. The 1024-dimensional feature in the backbone network passes through a convolution and finally obtains 2 two-dimensional vectors, including a two-dimensional class confidence and a two-dimensional action localization component;
[0016] S2-4. Two-stream fusion. The localization results are fused and fine-tuned by taking the mean value at the same position.
[0017] As a preferred technical solution of the present invention, when fusing the outputs of the backbone network in S2-3, the outputs of the three-layer structure are respectively defined as: the first layer is cls_clsBranch, loc_clsBranch, the second layer is cls_main, loc_main, and the third layer is others_propBranch, loc_propBranch, and the fused output is obtained through the following formula:
[0018] Cls_part = [(cls_main + cls_clsBranch) / 2, (loc_clsBranch + loc_main) / 2]
[0019] Loc_part = [others_propBranch, (loc_propBranch + loc_main) / 2]
[0020] Through the above fusion, and then decoding through the preselected positions (dboxes_w, dboxes_x) to obtain the output,
[0021] anchors_conf = Cls_part[2]
[0022] anchors_rx = Loc_part[2] * dboxes_w * 0.1 + dboxes_x
[0023] anchors_rw = e 0.1*Loc_part[4] * dboxes_w.
[0024] As a preferred technical solution of the present invention, the sensitive feature extraction in S3 is trained based on a pre-trained 3D convolutional feature network. By constructing a triplet training set, using the triplet loss TripletLoss and the cross-entropy loss CrossEntropyLoss as the loss functions of the network:
[0025] L total = L triplet + L crossentropy
[0026] Where:
[0027]
[0028] In the above formula, f θ () represents the output feature of the network, v i represents a video, represents a positive sample video similar to the video v i and represents a negative sample video different from the video v i Dissimilar negative sample videos, and D() represents the Euclidean distance between two output features.
[0029] The beneficial effects of the present invention are as follows:
[0030] This sensitive video feature extraction method locates, fuses, and extracts features from sensitive segments in the video. At the same time, by using the sensitive segment location technology, it realizes the accurate extraction of sensitive features in the video and reduces invalid features. By using the sensitive segment location fusion and similar video feature network strategy, it effectively improves the matching rate of sensitive video features and can effectively solve the problem of easy omission in video comparison in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings:
[0032] Figure 1 is a flowchart of the sensitive feature extraction of the video by FH-Sfnet of the present invention;
[0033] Figure 2 is a schematic diagram of the overall FH-SFnet of the present invention;
[0034] Figure 3 is a structural diagram of the sensitive segment location of a sensitive video feature extraction method of the present invention;
[0035] Figure 4 is a schematic diagram of the basic feature network structure of a sensitive video feature extraction method of the present invention;
[0036] Figure 5 is a schematic diagram of the backbone network structure of a sensitive video feature extraction method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] The following is a description of the preferred embodiments of the present invention with reference to the drawings. It should be understood that the preferred embodiments described herein are only for the purpose of illustrating and explaining the present invention and are not used to limit the present invention.
[0038] Embodiment: As Figures 1-3 shown, a sensitive video feature extraction method of the present invention includes a feature extraction network FH-SFnet, characterized in that the feature extraction network FH-SFnet extracts sensitive video features including the following steps:
[0039] S1. Collect various types of sensitive videos in a preset quantity and preprocess the videos to obtain temporal features and spatial features;
[0040] S2. Sensitive segment localization: Use a sensitive video localization network to perform sensitive video localization on temporal features and spatial features respectively, and perform position fusion to obtain the final position of the sensitive video segment.
[0041] S3. Sensitive feature extraction: Use a feature extraction network to extract features from the sensitive segment and finally output multiple sensitive video features.
[0042] Among them, the spatial features in S1 perform frame sampling on the video with a stride of 4, use the RGB image as the original input, and the dimension of the input data is [img_size, img_size, 3, n], where img_size is the size of the video frame and n is the number of sampled frames. Then use multiple pre-trained spatial convolutional networks to extract features, and finally output a feature with a dimension of n×1024.
[0043] Among them, the temporal features in S1 perform frame sampling on the video with a stride of 4, calculate the one-dimensional optical flow with 5 frames as a group, and use the optical flow field image as the original input. Then the dimension of the input data is [img_size, img_size, 1, n], where img_size is the size of the video frame and n is the number of sampled frames. Then use multiple pre-trained temporal convolutional networks, and finally output a feature with a dimension of n×1024.
[0044] Among them, the specific process of sensitive segment localization in S2 is as follows:
[0045] S2-1. As Figure 4 shown, construct a basic feature network. The basic feature network contains 4 convolutional-batch normalization-activation function ReLu basic network groups, uses the 1024-dimensional feature of video preprocessing as the input, and outputs a feature with a dimension of [32, 512].
[0046] S2-2. As Figure 5 shown, construct a backbone network on the basic feature network. The backbone network has three layers from top to bottom, and the basic feature network is the middle layer. Each layer has four outputs, and the output dimensions are: [16, 1024], [8, 1024], [4, 1024], [2, 1024]. One 1024-dimensional feature generates 5 candidate localizations, and the final number of generated candidate localizations is 150.
[0047] S2-3. Backbone network fusion: The 1024-dimensional feature in the backbone network passes through a convolution and finally obtains 2 two-dimensional vectors, including a 2-dimensional class confidence (sensitive and normal) and a 2-dimensional action localization component (the midpoint of the action, the time width of the action).
[0048] S2-4. Dual-stream fusion: After the backbone network fusion is completed, dual-stream fusion needs to be performed on the two-layer network for sensitive feature localization. We use the method of taking the average at the same position to fine-tune the localization results through fusion.
[0049] Among them, when fusing the outputs of the backbone network in S2-3, the outputs of the three-layer structure are respectively defined as follows: the first layer is cls_clsBranch, loc_clsBranch, the second layer is cls_main, loc_main, and the third layer is others_propBranch, loc_propBranch, and the fusion output is carried out through the following formula:
[0050] Cls_part = [(cls_main + cls_clsBranch) / 2, (loc_clsBranch + loc_main) / 2]
[0051] Loc_part = [others_propBranch, (loc_propBranch + loc_main) / 2]
[0052] Through the above fusion, and then decoding through the preselected positions (dboxes_w, dboxes_x) to obtain the output,
[0053] anchors_conf = Cls_part[2]
[0054] anchors_rx = Loc_part[2] * dboxes_w * 0.1 + dboxes_x
[0055] anchors_rw = e 0.1*Loc_part[4] * dboxes_w.
[0056] Among them, in S3, the extraction of sensitive features is trained based on a pre-trained 3D convolutional feature network. By constructing a triplet training set, the triplet loss TripletLoss and the cross-entropy loss CrossEntropyLoss are used as the loss functions of the network:
[0057] L total = L triplet + L crossentropy
[0058] Among them:
[0059]
[0060] In the above formula, f θ () represents the output features of the network, and v i represents a video, Representation and video v i Similar positive sample videos, Representation and video v i Dissimilar negative sample videos, D() represents the Euclidean distance between the two output features. The ternary loss function can shorten the distance between positive sample pairs and push the distance between negative sample pairs, thereby continuously reducing the feature distance of similar videos during network training.
[0061] Use the trained 3D convolutional network to extract features from sensitive video segments, extract each video segment in a uniform 24-frame manner, and finally output multiple sensitive video features. The above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions described in the aforementioned embodiments, or to make equivalent replacements for some of the technical features therein. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A sensitive video feature extraction method, comprising a feature extraction network FH-SFnet, It is characterized in that The feature extraction network FH-SFnet extracts video sensitive features including the following steps: S1. Collect a preset number of various sensitive videos and pre-process the videos to obtain temporal and spatial features; S2, sensitive segment positioning, using the sensitive video positioning network to position the sensitive video based on temporal features and spatial features respectively, and perform position fusion to obtain the final sensitive video segment position; S3, sensitive feature extraction, using the feature extraction network to extract features from sensitive clips, and finally output multiple sensitive video features; The specific process of locating the sensitive fragment in S2 is as follows: S2-1. Construct a basic feature network, which includes 4 convolution-batch normalization-activation function ReLu basic network groups, takes the 1024-dimensional features of video preprocessing as input, and outputs [32,512]-dimensional features; S2-2. A backbone network is constructed on the basic feature network. The backbone network has three layers from top to bottom, and the basic feature network is the middle layer. Each layer has four outputs, and the output dimensions are: [16, 1024], [8, 1024], [4, 1024], [2, 1024]. One 1024-dimensional feature generates 5 candidate locations, and the number of candidate locations generated is 150. S2-3, backbone network fusion, the 1024-dimensional features in the backbone network undergo a convolution, and finally obtain two 2-dimensional vectors, including a 2-dimensional category confidence and a 2-dimensional action localization component; S2-4, dual-stream fusion, the positioning results are fused and fine-tuned by taking the average at the same position; When the output of the backbone network is fused in S2-3, the outputs of the three-layer structure are defined as follows: the first layer is cls_clsBranch, loc_clsBranch, the second layer is cls_main, loc_main, and the third layer is others_propBranch, loc_propBranch, and the fusion output is performed by the following formula: Cls_part=[cls_main+cls_clsBranch) / 2,(loc_clsBranch+loc_main) / 2]Loc_part=[others_propBranch,(loc_propBranch+loc_main) / 2] Through the above fusion, the output is obtained by decoding the pre-selected position (dboxes_w, dboxes_x). anchors_conf = Cls_part[2] anchors_rx=Loc_part[2]*dboxes_w*0.1+dboxes_x anchors_rw = e 0.1*Loc_part[4] e 0.1*Loc_part[4] * dboxes_w。 2. A sensitive video feature extraction method according to claim 1, It is characterized in that The spatial feature in S1 is to sample the video frames with a step size of 4, take the RGB image as the original input, and the dimension of the input data is [img_size, img_size, 3, n], where img_size is the size of the video frame and n is the number of sampled frames. Then, multiple pre-trained spatial convolutional networks are used to extract features, and finally a feature of dimension n×1024 is output.
3. A sensitive video feature extraction method according to claim 1, characterized in that the temporal feature in S1 is to sample the video frames with a step size of 4, calculate the one-dimensional optical flow with 5 frames as a group, and take the optical flow field image as the original input. Then, the dimension of the input data is [img_size, img_size, 1, n], where img_size is the size of the video frame and n is the number of sampled frames. After that, multiple pre-trained temporal convolutional networks are used, and finally a feature of dimension n×1024 is output.
4. A sensitive video feature extraction method according to claim 1, characterized in that the sensitive feature extraction in S3 is trained based on a pre-trained 3D convolutional feature network. By constructing a triplet training set, the triplet loss TripletLoss and the cross-entropy loss CrossEntropyLoss are used as the loss functions of the network: L total = L triplet + L crossentropy where: In the above formula, f θ () represents the output feature of the network, v i represents a video, represents the positive sample video similar to the video v i , v; represents the negative sample video not similar to the video v i , and D() represents the Euclidean distance between two output features.
Citation Information
Patent Citations
Method for recognizing human body behaviors in video based on double-flow convolutional network
CN110909658A
Aerial video analysis method based on space-time 2D convolutional neural network
CN113269054A