Feature construction method, device and equipment for video
By performing bidirectional feature splicing on short videos, the problem of insufficient accuracy of machine recognition method when labeling short videos is solved, and a more comprehensive feature extraction and labeling accuracy is achieved.
Patent Information
- Application Number
- CN202210102087.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-27
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-01-27
AI Technical Summary
In the prior art, the machine recognition method has insufficient accuracy when labeling short videos, and the sequence features of the video frames are easily forgotten, resulting in missing features and affecting the accuracy of labeling.
The bidirectional feature splicing method is used to slice the video into multiple segments, extract the positive and reverse sequence features respectively, and splice them to form a more comprehensive feature sequence, which is used to input bidirectional loop length and short time for training.
The accuracy of machine learning module labeling short videos is improved, the problem of feature forgetting in the prior art is avoided, and more comprehensive video features are obtained.
Smart Images

Figure CN114463679B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video features, and particularly relates to a method for constructing video features, a device for constructing video features, a device for constructing video features of a video, and a corresponding storage medium. Background Art
[0002] Generally, the methods for short video tagging include manual review method and machine recognition method. The disadvantages of the manual review method are very obvious. One is that the cost is very high. It requires manual viewing and understanding of the video to perform tag annotation, consuming a lot of manpower. The other is that it is not real-time. It cannot tag the user uploads in a timely manner. It requires manual browsing of the video before operation. When the number of simultaneously uploaded videos is large, the feedback time will be longer, resulting in a poor user experience. The machine recognition method refers to using artificial intelligence to let the machine automatically tag short videos. It can solve the problems of high cost and non-real-time of the manual method. Therefore, there is a trend to use the machine recognition method to solve the short video tagging problem. The advantage of the manual review method is relatively accurate. So now the problem becomes how to improve the accuracy of the machine recognition method.
[0003] Data and features determine the upper limit of machine learning, while models and algorithms only approach this upper limit. The cost of data acquisition is extremely high. Therefore, how to process features under the existing data is the key to further improving the upper limit of machine learning. The usual method for feature construction of this problem is to use a two-stream convolutional network. Its feature acquisition is unidirectional, and the video frames belong to sequential features. Sequential features have the property of forgetting. That is, after input into the network, the features of the frames in the front are easily forgotten, and more features are about the frames in the back. Therefore, features in the video are likely to be missed. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide a device, method and device for constructing video features to at least solve the above technical problems.
[0005] To achieve the above purpose, in a first aspect of the present invention, a method for constructing video features is provided. The method includes: slicing a video to obtain N video segments; respectively obtaining a first feature and a second feature of a single video segment; respectively performing forward superposition and reverse superposition on the N first features and the N second features to obtain a feature sequence of the video.
[0006] Preferably, the first feature includes the processing result of a visual processing algorithm on the single video segment; the second feature includes a subset of the content of the single video segment.
[0007] Preferably, the first feature includes: an optical flow map obtained from the single video clip based on the optical flow extraction algorithm in OpenCV; the second feature includes: a picture frame in the single video clip.
[0008] Preferably, the method further includes: determining whether the number of frames of the video is greater than the product of N and a preset frame number threshold; if the number of frames of the video is not greater than the product of N and the preset frame number threshold, then copying and splicing the video to make the video greater than the product of N and the preset frame number threshold.
[0009] Preferably, the method further includes: after slicing the video into N video clips, preprocessing the picture frames in the video clips into a preset specification.
[0010] In a second aspect of the present invention, there is also provided a device for constructing features of a video, the device includes: a video slicing module for slicing a video into N video clips; a clip feature module for respectively obtaining a first feature and a second feature of a single video clip; and a feature superposition module for respectively performing positive-order superposition and reverse-order superposition on the N first features and the N second features to obtain a feature sequence of the video.
[0011] Preferably, the first feature includes the processing result of the visual processing algorithm on the single video clip; the second feature includes a subset of the content of the single video clip.
[0012] Preferably, the first feature includes: an optical flow map obtained from the single video clip based on the optical flow extraction algorithm in OpenCV; the second feature includes: a picture frame in the single video clip.
[0013] Preferably, the device further includes a video length processing module; the video length processing module is used for: determining whether the number of frames of the video is greater than the product of N and a preset frame number threshold; if the number of frames of the video is not greater than the product of N and the preset frame number threshold, then copying and splicing the video to make the number of frames of the video greater than the product of N and the preset frame number threshold.
[0014] Preferably, the device further includes a preprocessing module; the preprocessing module is used for: after the video slicing module slices the video into N video clips, preprocessing the picture frames in the video clips into a preset specification.
[0015] In a third aspect of the present invention, there is provided a device for constructing features of a video, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the foregoing method for constructing features of a video is implemented.
[0016] In a fourth aspect of the present invention, there is provided a computer-readable storage medium storing instructions which, when run on a computer, cause the computer to execute the foregoing method for constructing features of a video.
[0017] In a fifth aspect of the present invention, there is provided a computer program product including a computer program which, when executed by a processor, implements the foregoing method for constructing features of a video.
[0018] The above technical solutions have the following beneficial effects:
[0019] The feature sequence constructed by the above embodiments can avoid the following defects of the existing two-stream convolutional network: the acquisition of its features is unidirectional, while the frames of a video belong to sequence features, and sequence features have the property of forgetting. That is, after being input into the network, the features of the frames at the front are prone to be forgotten, and more features are about the frames at the back. Therefore, the features in the video are likely to be missed. The feature construction method proposed in this embodiment, namely the bidirectional feature splicing method, on the basis of the original features, starts from the back of the video, extracts the features again in reverse, and then splices them with the features extracted forward before. Thus, more comprehensive video features are obtained, thereby improving the accuracy of tagging using the machine learning module.
[0020] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The drawings are used to provide a further understanding of the embodiments of the present invention, and constitute a part of the specification, and are used together with the following specific implementation to explain the embodiments of the present invention, but do not constitute a limitation to the embodiments of the present invention. In the drawings:
[0022] Figure 1 Schematically shows a step schematic diagram of the method for constructing features of a video according to an embodiment of the present application;
[0023] Figure 2 Schematically shows a structural schematic diagram of a bidirectional recurrent long short-term neural network according to an embodiment of the present application;
[0024] Figure 3 Schematically shows an overall architecture schematic diagram of the method for constructing features of a video according to an embodiment of the present application;
[0025] Figure 4 Schematically shows a structural schematic diagram of the device for constructing features of a video according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] The following will explain in detail the specific implementation manners of the embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific implementation manners described herein are only used to illustrate and explain the embodiments of the present invention, and are not used to limit the embodiments of the present invention.
[0027] In the technical solution of this application, the acquisition, storage, use, processing, etc. of data all comply with the relevant provisions of national laws and regulations.
[0028] Figure 1 Schematically shows a step schematic diagram of a method for constructing the features of a video according to an embodiment of this application. As Figure 1 shown, in an embodiment of this application, a method for constructing the features of a video, the method includes:
[0029] 101. Slice the video to obtain N video segments;
[0030] N takes a natural number, and its value determines the size of the feature sequence. The slicing method here can be average slicing or slicing according to a preset rule. After this step, a video is divided into N video segments.
[0031] 102. Obtain the first feature and the second feature of each individual video segment respectively; the features of the video segment include but are not limited to the attributes of the video segment itself, the snapshot of the content of the video segment, the slice of the content of the video segment, the pixel information, motion information, inter-frame relationship, etc. included in the video segment. This application extracts two features from the above multiple features in each individual video segment, and respectively records them as the first feature and the second feature.
[0032] 103. Stack the N first features and the N second features in positive order and reverse order respectively to obtain the feature sequence of the video.
[0033] The order of positive stacking among them can be the order of the video segments in the video, and the order of reverse stacking is just the opposite. The combination method can directly adopt sequential splicing, etc. The results after positive stacking and reverse stacking can be separated, or the positive stacking result and the reverse stacking result can be combined into a feature sequence, which depends on the input setting of the subsequent processing model.
[0034] The feature sequence constructed through the above implementation manner can avoid the following defects of the existing two-stream convolutional network: the acquisition of its features is unidirectional, while the frames of the video belong to sequence features, and sequence features have forgetfulness, that is, after being input into the network, the features of the frames in the front are prone to be forgotten, and more features are about the frames in the back, so the features in the video are likely to be missed.
[0035] The feature construction method proposed in this embodiment, namely the bidirectional feature splicing method, on the basis of the original features, starts from the back of the video, extracts the features again in reverse, and then splices them with the features extracted forward before. Thus, more comprehensive video features are obtained, thereby improving the accuracy of tagging using the machine learning module.
[0036] In some embodiments provided by the present invention, the first feature includes the processing result of the visual processing algorithm on the single video segment; the second feature includes a subset of the content of the single video segment. The visual processing algorithm here includes, but is not limited to, functions in machine vision image processing software such as OpenCV, Halcon, and visionpro. Using the visual processing algorithm in the prior art to process a single video segment has the advantage of being easy to use. The obtained processing result varies according to the output of the selected visual processing algorithm. A subset of the content of a single video segment is a part of the content of the single video segment. The selected part is selected according to a preset selection rule, and can be a multi-frame fusion value or a single-frame picture.
[0037] In some embodiments provided by the present invention, the first feature includes: an optical flow map obtained from the single video segment based on the optical flow extraction algorithm in OpenCV; the second feature includes: a picture frame in the single video segment. Optical flow is the instantaneous velocity of the pixels of a moving object in space on the observation imaging plane. The optical flow method is a method that uses the change of pixels in the time domain in the image sequence and the correlation between adjacent frames to find the corresponding relationship between the previous frame and the current frame, so as to calculate the motion information of the object between adjacent frames. This embodiment only provides a method for extracting an optical flow map, for example, by extracting based on the optical flow extraction algorithm in OpenCV. Specifically, using OpenCV (open source computer vision library) and python to extract video frames and extract TVL1 optical flow. Its main steps include: reading frames from the video, obtaining relevant attributes, and setting which frames to save; and extracting TVL1 optical flow through consecutive frames.
[0038] The picture frame can be selected from the video segment in a random selection manner, or a fixed position can be selected, such as uniformly selecting the nth picture frame in each video segment. The acquisition order of the optical flow map and the picture frame is not limited by the narrative order.
[0039] In some embodiments provided by the present invention, the method further includes: determining whether the number of frames of the video is greater than the product of N and a preset frame number threshold; if not, copying and splicing the video so that the number of frames of the video is greater than the product of N and the preset frame number threshold. For example, if the preset frame number threshold is 5, that is, each video segment has at least 5 frames, and N is 20, then the video needs to be at least 100 frames. If the video does not meet the length requirement of 100 frames, it needs to be copied and spliced until it is greater than 100 frames. Through this embodiment, it can be ensured that each video segment has enough frames to implement optical flow extraction.
[0040] In some embodiments provided by the present invention, the method further includes: after slicing the video into N video segments, preprocessing the picture frames in the video segments into a preset specification. For example, performing replaying processing on the image size of each frame in the video segment and scaling it into a picture of (224, 224, 3) to improve the calculation efficiency of subsequent tasks.
[0041] In some embodiments provided by the present invention, the feature sequence of the video is used as input to a Bi-directional LSTM RNN (Bidirectional Long Short-Term Memory Recurrent Neural Network) for training or judgment. LSTM (Long Short-Term Memory) is a special type of RNN, mainly to solve the problems of gradient vanishing and gradient explosion during the training of long sequences. Simply put, compared with ordinary RNNs, LSTM can perform better in longer sequences. An Attention mechanism can be introduced into LSTM. The basic idea of the Attention mechanism is to break the limitation that the traditional encoder-decoder structure depends on a fixed-length internal vector during encoding and decoding. The implementation of the Attention mechanism is to retain the intermediate output results of the LSTM encoder for the input sequence, and then train a model to selectively learn these inputs and associate the output sequence with them when the model outputs. Figure 2 Schematically shows a schematic structural diagram of a bidirectional long short-term memory recurrent neural network according to an embodiment of the present application, as Figure 2 shown. The bidirectional long short-term memory recurrent neural network includes an input layer, a forward layer, a backward layer, and an output layer, and w1-w6 in the figure respectively represent weights.
[0042] Figure 3 Schematically shows an overall architecture schematic diagram of a method for constructing features of a video according to an embodiment of the present application. As Figure 3 shown, taking a short video as an example, it includes the following steps:
[0043] 1. Obtain the short video to be processed, slice the short video into N segments to obtain video segments 1-N;
[0044] 2. Read the video segments sequentially for preprocessing, such as resizing; perform the processing of steps 3 and 4 on the preprocessed video segments respectively;
[0045] 3. Randomly select an original frame image from each video segment as the first feature;
[0046] 4. Calculate the optical flow map of each video segment as the second feature; this step and step 3 can be in a parallel relationship;
[0047] 5. Temporarily store the results obtained in steps 3 and 4 in the slice order until all N video segments are read and processed;
[0048] 6. Stack the optical flow map and the original frame image in the forward order and the reverse order respectively in the slice order to obtain the feature sequence of the short video. This feature sequence can be used as the input for subsequent algorithms such as Attention-LSTM.
[0049] Through the above embodiments, through the bidirectional splicing method, on the basis of the original features, extract the features again in reverse from the end of the video, and then splice them with the features extracted forward before. In this way, more comprehensive video features can be obtained, improving the accuracy of short video tagging.
[0050] Based on the same inventive concept, the embodiments of the present invention also provide a feature construction device for videos. Figure 4 Schematically shows a structural diagram of a feature construction device for videos according to an embodiment of the present application, as Figure 4 shown. A feature construction device for videos, the device includes: a video slicing module for slicing a video into N video segments; a segment feature module for respectively obtaining the first feature and the second feature of a single video segment; and a feature stacking module for respectively stacking the N first features and the N second features in the forward order and the reverse order to obtain the feature sequence of the video.
[0051] In some alternative embodiments, the first feature includes the processing result of the visual processing algorithm on the single video segment; the second feature includes a subset of the content of the single video segment.
[0052] In some alternative embodiments, the first feature includes: an optical flow map obtained from the single video segment based on the optical flow extraction algorithm in OpenCV; the second feature includes: a picture frame in the single video segment.
[0053] In some alternative embodiments, the device further includes a video length processing module; the video length processing module is configured to: determine whether the number of frames of the video is greater than the product of N and a preset frame number threshold; if the number of frames of the video is not greater than the product of N and the preset frame number threshold, then copy and splice the video so that the number of frames of the video is greater than the product of N and the preset frame number threshold.
[0054] In some alternative embodiments, the device further includes a preprocessing module; the preprocessing module is configured to: after the video slicing module slices the video into N video segments, preprocess the picture frames in the video segments into a preset specification.
[0055] For the specific limitations of each implementation step in the above video feature construction device, reference may be made to the limitations on the video feature construction method in the foregoing text, which will not be elaborated herein. Its beneficial effects can also be presumptively determined according to the foregoing video feature construction method.
[0056] The embodiments of the present application provide a device, which includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, the steps of the video feature construction method are implemented.
[0057] The present application also provides a computer program product, which is suitable for executing a program for initializing the steps including the above video feature construction method when executed on a data processing device.
[0058] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0059] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0060] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes Figure 1 or processes and / or boxes Figure 1 specified in one or more of the boxes.
[0061] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes Figure 1 or processes and / or boxes Figure 1 specified in one or more of the boxes.
[0062] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0063] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.
[0064] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0065] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.
[0066] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A method for constructing features of a video, characterized in that, The method includes: Slicing a video to obtain N video segments; Obtaining a first feature and a second feature of a single video segment, where the first feature includes the processing result of the visual processing algorithm on the single video segment; the second feature includes a subset of the content of the single video segment; the first feature includes: an optical flow map obtained from the single video segment based on the optical flow extraction algorithm in OpenCV; the second feature includes: a picture frame in the single video segment; Sequentially superimposing and reversely superimposing the N first features and the N second features respectively to obtain a feature sequence of the video; Depending on the input setting of the subsequent processing model, the results after sequential superposition and reverse superposition are separate or the results of sequential superposition and reverse superposition are combined into a single feature sequence.
2. The method according to claim 1, characterized in that, The method further includes: Judging whether the number of frames of the video is greater than the product of N and a preset frame number threshold; If the number of frames of the video is not greater than the product of N and the preset frame number threshold, then copy and splice the video to make the video greater than the product of N and the preset frame number threshold.
3. The method according to claim 1, wherein The method further includes: After slicing the video to obtain N video segments, preprocessing the picture frames in the video segments into a preset specification.
4. A feature construction device for a video, characterized in that, The apparatus includes: A video slicing module for slicing a video to obtain N video segments; A segment feature module for respectively obtaining a first feature and a second feature of a single video segment, where the first feature includes the processing result of the visual processing algorithm on the single video segment; the second feature includes a subset of the content of the single video segment; the first feature includes: an optical flow map obtained from the single video segment based on the optical flow extraction algorithm in OpenCV; the second feature includes: a picture frame in the single video segment; and A feature superposition module for sequentially superimposing and reversely superimposing the N first features and the N second features respectively to obtain a feature sequence of the video. Depending on the input setting of the subsequent processing model, the results after sequential superposition and reverse superposition are separate or the results of sequential superposition and reverse superposition are combined into a single feature sequence.
5. The device according to claim 4, characterized in that, The apparatus further includes a video length processing module; the video length processing module is used for: Judging whether the number of frames of the video is greater than the product of N and a preset frame number threshold; If the number of frames of the video is not greater than the product of N and the preset frame number threshold, then copy and splice the video to make the number of frames of the video greater than the product of N and the preset frame number threshold.
6. The device according to claim 5, characterized in that, The apparatus further includes a preprocessing module; the preprocessing module is used for: after the video slicing module slices the video to obtain N video segments, preprocessing the picture frames in the video segments into a preset specification.
7. A feature construction device for a video, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method for constructing the features of the video according to any one of claims 1 to 3.
8. A computer-readable storage medium, characterized in that, Instructions are stored in the storage medium, and when it runs on a computer, it causes the computer to execute the method for constructing the features of the video according to any one of claims 1 to 3.
9. A computer program product, comprising a computer program which, when executed by a processor, implements the method for constructing the characteristics of a video according to any one of claims 1 to 3.
Citation Information
Patent Citations
Video dense event description method based on generative adversarial network
CN111368142A