Video Sequence Sample Enhancement Method Based on Non-Key Frame Perturbation
By extracting keyframes in video pedestrian recognition and applying random Gaussian noise to non-keyframes, the problem of neglecting frame quality in the prior art is solved, and the robustness and retrieval effect of video pedestrian recognition network is improved.
Patent Information
- Application Number
- CN202210808388.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-11
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-07-11
AI Technical Summary
The existing video pedestrian re-identification methods ignore frame quality when processing video sequences, resulting in weak learning capabilities of network models and unable to train a robust video pedestrian re-identification network.
A video sequence sample enhancement method based on non-keyframe perturbation is proposed. By calculating the gradient direction and influence degree of each frame in the video sequence, keyframes are extracted, and random Gaussian noise is applied to non-keyframes to construct a new video sequence sample to reduce the impact of non-keyframe data on the model.
Through confrontational learning, the robustness of the video pedestrian re-identification network is improved and the search effect in the video pedestrian re-identification scenario is improved.
Smart Images

Figure CN115205741B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and in particular to a method for enhancing video sequence samples based on non-key frame perturbation. Background Art
[0002] Video pedestrian re-identification is a hot topic in the field of computer vision, aiming to match pedestrians with continuous video sequences. Compared with the image-based pedestrian re-identification task, video pedestrian re-identification is closer to practical applications and can be used for video surveillance, finding lost people, etc. Existing video pedestrian re-identification methods focus on extracting features from time and space, ignoring the quality of each frame in the video sequence. Since there may be low-quality situations such as the target being occluded or lost in some frames of the continuous video sequence, if all video frames are regarded as training data of the same quality, it will weaken the learning ability of the network model and it is impossible to train a robust video pedestrian re-identification network. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to propose a method for enhancing video sequence samples based on non-key frame perturbation. First, the present invention calculates the gradient direction in the video sequence samples by using the sign function, and performs influence degree statistics on each frame in the video sequence, and extracts the first n_k frames with the highest influence degrees as key frames, and proposes a new method for extracting key frames of video sequences based on gradient direction. These key frames help the network learn discriminative information. In view of the low influence degree of non-key frames, the present invention applies random Gaussian noise to non-key frames to construct a new video sequence sample that highlights key frames, enabling the network to reduce the influence of non-key frame data on the model through adversarial learning, thereby improving the robustness of the video pedestrian re-identification network.
[0004] First, during the network training process, the input video sequence samples are fed into the video pedestrian re-identification network model, and the loss is calculated according to the network output result; subsequently, the sign function sign() is used to calculate the gradient direction of the video sequence samples; then, the sum() function is used for each video frame in the video sequence to calculate the sum of the absolute values of the gradient directions under this video frame; then, according to the sum value of each frame in the video sequence, the topk() function for finding the first n_k maximum values is used to obtain the indexes of the first n_k frames with the largest sum values in the video sequence, and these are regarded as the key frames in this video sequence; then, according to the indexes of the key frames, random Gaussian noise perturbation is performed on other non-key frames in the video sequence; finally, the perturbed non-key frames replace the frames corresponding to the indexes in the original video sequence to construct a new video sequence sample, which is fed into the video re-identification network again for subsequent training. The present invention can improve the retrieval effect in the video pedestrian re-identification scenario.
[0005] The present invention specifically adopts the following technical solutions:
[0006] A video sequence sample enhancement method based on non-critical frame perturbation, characterized by comprising the following steps:
[0007] Step S1: During network training, input the video sequence sample into the video pedestrian re-identification network model, and calculate the loss according to the network output result;
[0008] Step S2: Calculate the gradient direction of the video sequence sample;
[0009] Step S3: Calculate the sum of the absolute values of the gradient directions for each video frame in the video sequence;
[0010] Step S4: According to the sum value of each frame in the video sequence, calculate the indexes of the first n_k frames with the largest sum values in the video sequence, and regard them as the key frames in this video sequence;
[0011] Step S5: According to the indexes of the key frames, perform random Gaussian noise perturbation on other non-critical frames in the video sequence;
[0012] Step S6: Replace the frames corresponding to the indexes in the original video sequence with the perturbed non-critical frames, construct a new video sequence sample, and send it into the video re-identification network for subsequent training.
[0013] Further, step S1 is specifically as follows:
[0014] Step S11: During network training, input the video sequence sample n_x into the video pedestrian re-identification network model, and obtain the classification score n_α by the classifier in the network model, where the shape of n_x is a 5D tensor, namely batch, number of frames, number of channels, height, and width;
[0015] Step S12: Calculate the loss through the cross-entropy loss function according to the classification score n_α and the video sequence sample category label value n_y, and perform loss backpropagation. The formula is as follows:
[0016]
[0017] Where is the gradient of n_α, J() is the cross-entropy loss function, and model_θ represents the network parameters.
[0018] Further, step S2 is specifically as follows. Calculate the gradient direction n_v of the video sequence sample. The formula is as follows, where the shape of n_v is the same as the input video sequence sample n_x, and sign() represents performing a sign calculation on the gradient direction. For gradients greater than 0, the output is 1. For gradients less than 0, the output is -1. For gradients equal to 0, the output is 0:
[0019]
[0020] Further, step S3 specifically is to calculate the sum of the absolute values of the gradient direction n_v for each video frame in the video sequence. The formula is as follows, where abs() represents taking the absolute value of the value of the input gradient direction n_v, sum() represents summing the absolute values of the input gradient direction n_v, and dim represents the dimension selected by sum(). dim = [2, 3, 4] means selecting the number of channels, height, and width;
[0021] sum n_v = sum(abs(n_v)), dim = [2, 3, 4].
[0022] Further, step S4 specifically is to calculate, based on the sum value sum n_v for each frame in the video sequence, the index key of the first n_k frames with the largest sum values in the video sequence index , and regard the frames corresponding to the index as the key frames in this video sequence, and the rest as non-key frames. The formula is as follows, where topk() represents obtaining the first n_k maximum values in sum n_v , and dim represents the dimension selected by topk(). dim = [1] means sorting according to the summation results of each batch;
[0023] key index = topk(sum n_v ), dim = [1].
[0024] Further, it is characterized in that in step S5, according to the index key index of the key frames, random Gaussian noise perturbation is performed on other non-key frames in the video sequence. The formula is as follows, where the random Gaussian noise noise_δ follows a Gaussian distribution N with a mathematical expectation of μ and a standard deviation of σ 2 , and the shape and size are the same as those of the video sequence n_x. zero_like() represents generating all-0 data with the same shape as the input data:
[0025] noise_δ ~ N(μ, σ 2 )
[0026] noise_δ[key index = zero_like(noise_δ[key index ).
[0027] Further, step S6 specifically is to replace the frames corresponding to the index in the original video sequence n_x with the perturbed non-key frames to construct a new video sequence sample Among them, the all-zero part in noise_δ indicates that the corresponding frame is a key frame and no perturbation is performed. The formula is as follows. The new video sequence sample is fed into the video person re-identification network for subsequent training:
[0028]
[0029] In addition, a video sequence sample enhancement system based on non-key frame perturbation, characterized in that it includes a memory, a processor, and computer program instructions stored on the memory and executable by the processor. When the processor runs the computer program instructions, the above-mentioned method can be implemented.
[0030] A computer-readable storage medium stores computer program instructions executable by a processor. When the processor runs the computer program instructions, the above-mentioned method can be implemented.
[0031] Compared with the prior art, the present invention and its preferred solutions have the following beneficial effects:
[0032] 1. A method for extracting video key frames based on gradient direction is proposed. By using the sign function to calculate the gradient direction of the sample and summing and statistically analyzing the absolute values of the gradient directions in each frame, the key frames in the video sequence can be obtained;
[0033] 2. A method for enhancing video sequence samples based on random Gaussian noise is designed. Considering the low influence of non-key frames, random Gaussian noise is applied to non-key frames to construct a new video sequence sample that highlights key frames;
[0034] 3. A method for enhancing video sequence samples based on non-key frame perturbation is designed, enabling the network to reduce the influence of non-key frame data on the model through adversarial learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The following further elaborates on the present invention in detail in conjunction with the drawings and specific embodiments;
[0036] Figure 1 is a schematic diagram of the process and working principle of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] To make the features and advantages of this patent more obvious and understandable, specific embodiments are given below for detailed description as follows:
[0038] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0039] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0040] As Figure 1 shown, this embodiment provides a method for enhancing video sequence samples based on non-key frame perturbation, which specifically includes the following steps:
[0041] Step S1: During the network training process, input the video sequence sample into the video pedestrian re-identification network model, and calculate the loss according to the network output result;
[0042] Step S2: Calculate the gradient direction of the video sequence sample;
[0043] Step S3: Calculate the sum of the absolute values of the gradient directions for each video frame in the video sequence;
[0044] Step S4: According to the sum values of each frame in the video sequence, calculate the indices of the first n_k frames with the largest sum values in the video sequence, and regard them as the key frames of this video sequence;
[0045] Step S5: According to the indices of the key frames, perform random Gaussian noise perturbation on other non-key frames in the video sequence;
[0046] Step S6: Replace the frames corresponding to the indices in the original video sequence with the perturbed non-key frames, construct a new video sequence sample, and send it into the video re-identification network for subsequent training.
[0047] In this embodiment, step S1 is specifically as follows:
[0048] Step S11: During the network training process, input the video sequence sample n_x into the video pedestrian re-identification network model (this network model refers to the existing general video pedestrian re-identification model, that is, the method of this embodiment can be applied to all current video pedestrian re-identification methods), and obtain the classification score n_α from the classifier in the network model, where the shape of n_x is a 5-dimensional tensor, namely batch, number of frames, number of channels, height, and width;
[0049] Step S12: Calculate the loss through the cross-entropy loss function according to the classification score n_α and the video sequence sample category label value n_y, and perform loss backpropagation. The formula is as follows:
[0050]
[0051] where is the gradient of \(n_{\alpha}\), \(J()\) is the cross-entropy loss function, and \(model_{\theta}\) represents the network parameters.
[0052] In this embodiment, step S2 is specifically as follows: calculate the gradient direction \(n_v\) of the video sequence sample, and the formula is as follows. The shape of \(n_x\) is like the input video sequence sample. \(sign()\) represents performing a sign calculation on the gradient direction. For gradients greater than 0, the output is 1; for gradients less than 0, the output is -1; for gradients equal to 0, the output is 0.
[0053]
[0054] In this embodiment, step S3 is specifically as follows: calculate the sum of the absolute values of the gradient direction \(n_x\) for each video frame in the video sequence, and the formula is as follows. \(abs()\) represents taking the absolute value of the value of the input gradient direction \(n_v\), and \(sum()\) represents summing the absolute values of the input gradient direction \(n_v\). \(dim\) represents the dimension selected by \(sum()\). \(dim = [2, 3, 4]\) means selecting the number of channels, height, and width.
[0055] sum n_v = sum(abs(n_v)), dim = [2, 3, 4].
[0056] In this embodiment, step S4 is specifically as follows: calculate the index \(key\) of the first \(n_k\) frames with the largest sum values in the video sequence according to the sum value \(sum\) for each frame in the video sequence n_v , and regard the frames corresponding to the index as the key frames in this video sequence, and the rest as non-key frames. The formula is as follows. \(topk()\) represents obtaining the first \(n_k\) maximum values in \(sum\) index , and \(dim\) represents the dimension selected by \(topk()\). \(dim = [1]\) means sorting according to the summation results of each batch; n_v key
[0057] key index = topk(sum n_v ), dim = [1].
[0058] In this embodiment, step S5 is specifically as follows: perform random Gaussian noise perturbation on other non-key frames in the video sequence according to the index \(key\) of the key frames index , and the formula is as follows. Among them, the random Gaussian noise \(noise_{\delta}\) follows a Gaussian distribution \(N\) with a mathematical expectation of \(\mu\) and a standard deviation of \(\sigma\) 2 , and the shape and size are like the video sequence \(n_x\). \(zero_like()\) represents generating all 0 data with the same shape as the input data;
[0059] noise_δ ~ N(μ, σ 2 )
[0060] noise_δ[key index = zero_like(noise_δ[key index )。
[0061] In this embodiment, step S6 is specifically that the perturbed non-critical frames replace the frames with corresponding indexes in the original video sequence n_x to construct a new video sequence sample where the all-zero part in noise_δ indicates that the corresponding frame is a key frame and no perturbation is performed. The formula is as follows. The new video sequence sample is fed into the video person re-identification network for subsequent training;
[0062]
[0063] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a device, or a computer program product. Therefore, the present invention can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code
[0064] The present invention is described with reference to the flowcharts of methods, devices (apparatuses), and computer program products according to the embodiments of the present invention. It should be understood that each process in the flowchart can be implemented by computer program instructions, and the combination of processes in the flowchart can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one Figure 1 or more processes
[0065] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one Figure 1 or more processes in the flowchart
[0066] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide for implementing the steps of the functions specified in the process Figure 1 in one process or multiple processes.
[0067] As described above, it is only the preferred embodiment of the present invention, and it is not a limitation to the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still belong to the protection scope of the technical solution of the present invention.
[0068] This patent is not limited to the above best mode. Anyone inspired by this patent can derive various other forms of video sequence sample enhancement methods based on non-key frame perturbations. All equal changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope covered by this patent.
Claims
1. A video sequence sample enhancement method based on non-key frame perturbation, characterized in that It includes the following steps: Step S1: During the network training process, input video sequence samples are fed into the video person re-identification network model, and the loss is calculated based on the network output results; Step S2: Calculate the gradient direction of the video sequence samples; Step S3: Calculate the sum of the absolute values of the gradient directions for each video frame in the video sequence; Step S4: According to the sum values of each frame in the video sequence, calculate the indices of the first n_k frames with the largest sum values in the video sequence, and regard these frames as the key frames of this video sequence; Step S5: According to the indices of the key frames, perform random Gaussian noise perturbation on other non-key frames in the video sequence; Step S6: Replace the frames corresponding to the indices in the original video sequence with the perturbed non-key frames to construct new video sequence samples, and then feed them into the video re-identification network for subsequent training.
2. The method for enhancing a video sequence sample based on non-key frame perturbation according to claim 1, wherein Specifically, Step S1 is as follows: Step S11: During the network training process, input the video sequence sample n_x into the video person re-identification network model, and obtain the classification score n_α by the classifier in the network model. The shape of n_x is a 5D tensor, which are batch, number of frames, number of channels, height, and width respectively; Step S12: Calculate the loss according to the classification score n_α and the video sequence sample category label value n_y through the cross-entropy loss function, and perform loss backpropagation. The formula is as follows: where is the gradient of n_α, J() is the cross-entropy loss function, and model_θ represents the network parameters.
3. The method for enhancing a video sequence sample based on non-key frame perturbation according to claim 2, wherein Specifically, Step S2 is to calculate the gradient direction n_v of the video sequence samples. The formula is as follows. The shape of n_v is the same as the input video sequence sample n_x. sign() represents performing a sign calculation on the gradient direction. For gradients greater than 0, the output is 1. For gradients less than 0, the output is -1. For gradients equal to 0, the output is 0:
4. The method for enhancing a video sequence sample based on non-critical frame perturbation according to claim 3, wherein Specifically, Step S3 is to calculate the sum of the absolute values of the gradient direction n_v for each video frame in the video sequence. The formula is as follows. abs() represents taking the absolute value of the value of the input gradient direction n_v, and sum() represents summing the absolute values of the input gradient direction n_v. dim represents the dimension selected by sum(). dim = [2, 3, 4] means selecting the number of channels, height, and width; sum n_v = sum(abs(n_v)), dim = [2, 3, 4].
5. The method for enhancing a video sequence sample based on non-key frame perturbation according to claim 4, wherein Step S4 specifically is to calculate, according to the sum value sum of each frame in the video sequence n_v , to obtain the indexes key of the first n_k frames with the largest sum values in the video sequence index , and regard the frames corresponding to the indexes as the key frames in this video sequence, and the rest as non-key frames. The formula is as follows, where topk() represents obtaining the first n_k maximum values in sum n_v , dim represents the dimension selected by topk(), and dim = [1] means sorting according to the summation results of each batch; key index = topk(sum n_v ), dim = [1].
6. The method for enhancing a video sequence sample based on non-key frame perturbation according to claim 2, wherein In step S5, according to the index key of the key frame index , random Gaussian noise perturbation is performed on other non-key frames in the video sequence. The formula is as follows, where the random Gaussian noise noise_δ follows a Gaussian distribution N with a mathematical expectation of μ and a standard variance of σ 2 : the shape and size are like the input video sequence sample n_x, and zero_like() means generating all-zero data with the same shape as the input data noise_δ~N(μ,σ 2 ) noise_δ[key index = zero_like(noise_δ[key index )。 7. The method for enhancing a video sequence sample based on non-key frame perturbation according to claim 2, wherein Specifically, in step S6, the perturbed non-key frames replace the frames at the corresponding indices in the original input video sequence sample n_x to construct a new video sequence sample. Among them, the all-zero part in noise_δ indicates that the corresponding frame is a key frame and is not perturbed. The formula is as follows. The new video sequence sample is fed into the video person re-identification network for subsequent training:
8. A video sequence sample enhancement system based on non-key frame perturbation, characterized in that, It includes a memory, a processor, and computer program instructions stored on the memory and capable of being run by the processor. When the processor runs the computer program instructions, the method described in any one of claims 1-7 can be implemented.
9. A computer-readable storage medium, on which computer program instructions capable of being run by the processor are stored. When the processor runs the computer program instructions, the method described in any one of claims 1-7 can be implemented.
Citation Information
Patent Citations
Two-stage behavior identification method and system based on key frame sequence and behavior information
CN113239869A
Pedestrian re-identification model training and identification method and device based on time sequence diversity and correlation
CN113343810A