Method and device for extracting features of a target object in a video frame using a convolutional neural network
By calculating the proportion of the target object in the video frame and generating video frames of various sizes through different sampling, and then restoring and summing the features after extracting them in the receptive field, the problem of poor feature extraction effect and reduced speed of convolutional neural networks under different proportions of the target object is solved, and efficient feature extraction is achieved.
Patent Information
- Application Number
- CN202310618821.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-29
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2043-05-29
AI Technical Summary
In existing technologies, convolutional neural networks suffer from poor feature extraction and reduced operating speed when processing video frames with varying proportions of target objects.
By calculating the proportion of the target object in the video frame, various sizes of video frames are generated using different sampling methods. These video frames are then input into the corresponding receptive fields for feature extraction. Finally, the data is restored and summed to generate the target object features.
Without reducing the operating speed of the convolutional neural network, the effect of extracting image features with different proportions of target objects is improved.
Smart Images

Figure CN116778378B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of image feature extraction, in particular to a method and device for extracting features of a target object in a video frame by using a convolutional neural network. BACKGROUND
[0002] At present, due to the different distances of the shooting lens, the proportion of the target object in the image changes constantly, which is a great challenge for extracting the feature information of the target object in the video by using a convolutional neural network.
[0003] The existing technology mainly improves the feature extraction effect in the following two ways: the first method is to avoid the problem of different proportions of the target object in the image by using a fixed camera to shoot when generating the image. This method has poor universality and can only be used when the video to be recognized has not been shot, which does not solve the problem from the root. The second method is to deepen the structure of the convolutional neural network to improve the extraction effect of the convolutional neural network on the feature information. However, due to the deepening of the convolutional neural network, the running speed is reduced and the time delay is increased, which reduces the real-time performance of feature extraction.
[0004] Therefore, a method is needed that can improve the feature extraction effect of the convolutional neural network on images of different target object proportions without reducing the running speed of the convolutional neural network. SUMMARY
[0005] In view of the problems in the prior art, the present application provides a method for extracting features of a target object in a video frame by using a convolutional neural network. According to the method, different sampling is performed on the current video frame according to the proportion of the target object in the current video frame, and the video frames of multiple sizes after sampling are input into the corresponding receptive field for feature extraction. After obtaining multiple feature sets, the current video frame feature is obtained by restoration and summation.
[0006] According to a first aspect of the present application, a method for extracting features of a target object in a video frame by using a convolutional neural network is provided, comprising:
[0007] calculating the proportion of the target object in the current video frame;
[0008] determining the sampling method of the current video frame according to the proportion of the target object to generate video frames of multiple different sizes;
[0009] inputting the video frames of multiple different sizes into the corresponding receptive field respectively to generate multiple sampling feature sets;
[0010] generating a video frame feature set according to the multiple sampling feature sets.
[0011] According to some embodiments, the generating the video frame feature set according to the plurality of sampled feature sets comprises:
[0012] restoring the plurality of sampled feature sets to original sizes of the current video frame to generate a plurality of restored feature sets;
[0013] summing the plurality of restored feature sets to obtain the video frame feature set.
[0014] According to some embodiments, the restoring the plurality of sampled feature sets to original sizes of the current video frame comprises:
[0015] restoring according to the sampling manner.
[0016] According to some embodiments, the sampling manner comprises up-sampling, down-sampling and up-down-sampling, wherein:
[0017] the up-down-sampling comprises one or more times of enlarging and one or more times of reducing the current video frame;
[0018] the up-sampling comprises one or more times of enlarging the current video frame;
[0019] the down-sampling comprises one or more times of reducing the current video frame.
[0020] According to some embodiments, the determining the sampling manner of the video frame according to the proportion of the target object comprises:
[0021] in a case that the proportion of the target object is lower than a first proportion threshold, determining the sampling manner of the video frame as up-sampling;
[0022] in a case that the proportion of the target object is higher than a second proportion threshold, determining the sampling manner of the video frame as down-sampling;
[0023] in a case that the proportion of the target object is higher than the first proportion threshold and lower than the second proportion threshold, determining the sampling manner of the video frame as up-down-sampling.
[0024] According to some embodiments, the receptive field comprises a first receptive field, a second receptive field and a third receptive field, wherein:
[0025] the first receptive field is used for extracting target object features in a video frame with a size smaller than a first size threshold;
[0026] the second receptive field is used for extracting target object features in a video frame with a size larger than the first size threshold and smaller than a second size threshold;
[0027] The third receptive field is configured to extract the target feature in a video frame whose size is greater than the second size threshold.
[0028] According to a second aspect of the present application, a method for extracting a target feature in a target video by using a convolutional neural network is provided, comprising
[0029] receiving the target video;
[0030] decomposing the target video into a plurality of video frames;
[0031] extracting target features in the plurality of video frames according to the method of the first aspect of the present application to obtain a plurality of video frame feature sets, and taking the plurality of video frame feature sets as the target feature of the target video.
[0032] According to a third aspect of the present application, a device for extracting a target feature in a video frame by using a convolutional neural network is provided, comprising:
[0033] a proportion identification module configured to identify a proportion of a target object in a current video frame;
[0034] a sampling module configured to determine a sampling manner of the current video frame according to the proportion of the target object, and generate a plurality of video frames with different sizes;
[0035] a sampling feature generation module configured to input the plurality of video frames with different sizes into corresponding receptive fields respectively, and generate a plurality of sampling feature sets;
[0036] a video frame feature generation module configured to generate a video frame feature set according to the plurality of sampling feature sets.
[0037] According to a fourth aspect of the present application, an electronic device is provided, comprising:
[0038] a processor;
[0039] a memory storing a computer program, when the computer program is executed by the processor, the processor is caused to execute the method of the first aspect of the present application.
[0040] According to a fifth aspect of the present application, a non-transitory computer readable storage medium is provided, which stores computer readable instructions, when the instructions are executed by a processor, the processor is caused to execute the method of the first aspect of the present application.
[0041] The method and apparatus for extracting target object features from video frames using a convolutional neural network provided in this application perform different sampling based on the proportion of the target object in the current video frame. The sampled video frames of various sizes are input into the corresponding receptive fields for feature extraction. Multiple feature sets are obtained, and then the original data is restored and summed to obtain the target object features of the current video frame. According to the scheme provided in this application, the feature extraction effect of the convolutional neural network on frames with different proportions of the target object can be improved without reducing the running speed of the convolutional neural network. Attached Figure Description
[0042] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings, without exceeding the scope of protection claimed by this application.
[0043] Figure 1 A flowchart of the method for extracting target features from video frames using a convolutional neural network according to this application;
[0044] Figure 2 A flowchart of the method for extracting target object features from a target video using a convolutional neural network according to this application;
[0045] Figure 3 A schematic diagram of the apparatus for extracting target features from video frames using the convolutional neural network of this application;
[0046] Figure 4 This is a structural diagram of an electronic device provided in this application. Detailed Implementation
[0047] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0048] Figure 1 This is a flowchart illustrating the method for extracting target object features from video frames using the convolutional neural network described in this application. Figure 1 As shown, the method includes the following steps.
[0049] Step S101: Calculate the proportion of the target object in the current video frame.
[0050] In some embodiments, the target is a person. In some embodiments, a proportion of an image occupied by the target in the current video frame is identified. In some embodiments, the first proportion threshold is smaller than the second proportion threshold. In some embodiments, the first proportion threshold and the second proportion threshold are constants between 0 and 1.
[0051] In some embodiments, the proportion of the image occupied by the target is smaller than the first proportion threshold. In some embodiments, the proportion of the image occupied by the target is greater than the first proportion threshold and smaller than the second proportion threshold. In some embodiments, the proportion of the image occupied by the target is greater than the second proportion threshold.
[0052] In some embodiments, the proportion of the target in the current video frame is calculated by a software running a recognition algorithm. In some embodiments, the proportion of the target in the current video frame is calculated by manual recognition.
[0053] In step S102, a sampling manner of the current video frame is determined according to the proportion of the target, so as to generate video frames of multiple different sizes.
[0054] In some embodiments, the proportion of the image occupied by the target is smaller than the first proportion threshold, and the sampling manner is up-sampling. In some embodiments, the proportion of the image occupied by the target is greater than the second proportion threshold, and the sampling manner is down-sampling. In some embodiments, the proportion of the image occupied by the target is greater than the first proportion threshold and smaller than the second proportion threshold, and the sampling manner is up-down sampling.
[0055] In some embodiments, the up-down sampling includes one or more times of enlargement and one or more times of reduction of the current video frame. In some embodiments, the up-sampling includes one or more times of enlargement of the current video frame. In some embodiments, the down-sampling includes one or more times of reduction of the current video frame. In some embodiments, the original video frame is taken as one of the video frames of multiple different sizes.
[0056] In step S103, the multiple video frames of different sizes are respectively input into corresponding receptive fields, so as to generate multiple sets of sampling features.
[0057] In some embodiments, a first receptive field, a second receptive field and a third receptive field are preset. In some embodiments, the first receptive field is used to extract target features in a video frame with a size smaller than a first size threshold. In some embodiments, the second receptive field is used to extract target features in a video frame with a size greater than the first size threshold and smaller than a second size threshold. In some embodiments, the third receptive field is used to extract target features in a video frame with a size greater than the second size threshold.
[0058] In some embodiments, the size of the video frame is smaller than the first size threshold, the video frame is input into the first receptive field to obtain the sample feature set. In some embodiments, the size of the video frame is larger than the first size threshold and smaller than the second size threshold, the video frame is input into the second receptive field to obtain the sample feature set. In some embodiments, the size of the video frame is larger than the second size threshold, the video frame is input into the third receptive field to obtain the sample feature set. In some embodiments, the original video frame is taken as one of the video frames of different sizes to obtain the sample feature set.
[0059] Step S104, generating a video frame feature set according to the plurality of sample feature sets.
[0060] In some embodiments, the plurality of sample feature sets are restored to the original size of the current video frame to generate a plurality of restored feature sets. In some embodiments, the plurality of restored feature sets are summed to obtain the video frame feature set. In some embodiments, the sample feature set is restored according to the sampling manner to obtain the restored feature set. In some embodiments, the sample feature set is reduced or enlarged according to the ratio of enlargement or reduction of the sampling manner to restore the sample feature set to a size corresponding to the original size of the current video frame. In some embodiments, the sample feature set obtained by inputting the original video frame into the receptive field is taken as one of the restored feature sets of the current video frame.
[0061] The method for extracting the feature of the target object in the video frame provided by the application extracts the feature of the target object in the video frame according to the proportion of the target object in the current video frame, inputs the video frames of different sizes obtained by sampling into the corresponding receptive field to extract the feature, restores and sums the plurality of feature sets to obtain the feature of the target object in the current video frame. According to the scheme provided by the application, the feature extraction effect of the convolutional neural network on the picture with different proportions of the target object can be improved without reducing the running speed of the convolutional neural network.
[0062] Figure 2 The flowchart of the method for extracting the feature of the target object in the video frame by the convolutional neural network provided by the application is shown in FIG. 1. As shown in the figure, the method comprises the following steps. Figure 2
[0063] Step S201, receiving a target video.
[0064] Step S202, decomposing the target video into a plurality of video frames.
[0065] Step S203, inputting the plurality of video frames into the convolutional neural network to obtain a plurality of sample feature sets, wherein the convolutional neural network comprises a plurality of receptive fields, and the plurality of receptive fields are configured to extract the feature of the target object in the video frame. Figure 1 The method extracts target object features in the plurality of video frames in sequence to obtain a plurality of video frame feature sets, and takes the plurality of video frame feature sets as the target object features of the target video.
[0066] In some embodiments, a target video to be subjected to feature extraction is received. In some embodiments, the target video is decomposed into a plurality of video frames. In some embodiments, target object features in the plurality of video frames decomposed from the target video are extracted in sequence. In some embodiments, a plurality of video frame feature sets are taken as the target object features of the target video.
[0067] In some embodiments, the picture proportion of all video frames in the target video satisfies that it is greater than a first proportion threshold and lower than a second proportion threshold. In some embodiments, all video frames in the target video are subjected to up-sampling and down-sampling. In some embodiments, the video frames are subjected to up-sampling and down-sampling, specifically, one time of up-sampling and one time of down-sampling.
[0068] In some embodiments, the video frames are subjected to one time of up-sampling, input into a corresponding receptive field to obtain a first sampling feature set, and subjected to one time of down-sampling, input into a corresponding receptive field to obtain a second sampling feature set. In some embodiments, the original video frames obtained after the decomposition of the target video are input into a corresponding receptive field to obtain a third sampling feature set. In some embodiments, the first sampling feature set is up-sampled to obtain a first restored feature set. In some embodiments, the second sampling feature set is down-sampled to obtain a second restored feature set. In some embodiments, the third sampling feature set is taken as a third restored feature set. In some embodiments, a combination of the first restored feature set, the second sampling feature set and the third sampling feature set is taken as a video frame feature set of the current video frame.
[0069] Figure 3 A schematic diagram of an apparatus for extracting target object features in video frames by a convolutional neural network of the present application.
[0070] As shown in Figure 3 An apparatus for extracting target object features in video frames by a convolutional neural network includes a proportion identification module, a sampling module, a sampling feature generation module and a video frame feature generation module.
[0071] The proportion identification module is configured to identify the proportion of the target object in the current video frame.
[0072] In some embodiments, the target object is a person. In some embodiments, a proportion of an image frame occupied by the target object in the current video frame is identified. In some embodiments, the first proportion threshold is smaller than the second proportion threshold. In some embodiments, the first proportion threshold and the second proportion threshold are constants between 0 and 1.
[0073] In some embodiments, the proportion of the image frame occupied by the target object is smaller than the first proportion threshold. In some embodiments, the proportion of the image frame occupied by the target object is greater than the first proportion threshold and smaller than the second proportion threshold. In some embodiments, the proportion of the image frame occupied by the target object is greater than the second proportion threshold.
[0074] In some embodiments, the proportion of the target object in the current video frame is calculated by a software running a recognition algorithm. In some embodiments, the proportion of the target object in the current video frame is calculated by manual recognition.
[0075] The sampling module is configured to determine a sampling method of the current video frame according to the proportion of the target object, and generate a plurality of video frames with different sizes.
[0076] In some embodiments, the proportion of the image frame occupied by the target object is smaller than the first proportion threshold, and the sampling method is up-sampling. In some embodiments, the proportion of the image frame occupied by the target object is greater than the second proportion threshold, and the sampling method is down-sampling. In some embodiments, the proportion of the image frame occupied by the target object is greater than the first proportion threshold and smaller than the second proportion threshold, and the sampling method is up-down sampling.
[0077] In some embodiments, the up-down sampling includes one or more times of enlargement and one or more times of reduction of the current video frame. In some embodiments, the up-sampling includes one or more times of enlargement of the current video frame. In some embodiments, the down-sampling includes one or more times of reduction of the current video frame. In some embodiments, the original video frame is one of the plurality of video frames with different sizes.
[0078] The sampling feature generation module is configured to input the plurality of video frames with different sizes into corresponding receptive fields respectively, and generate a plurality of sampling feature sets.
[0079] In some embodiments, a first receptive field, a second receptive field and a third receptive field are preset. In some embodiments, the first receptive field is configured to extract a target object feature in a video frame with a size smaller than a first size threshold. In some embodiments, the second receptive field is configured to extract a target object feature in a video frame with a size greater than the first size threshold and smaller than a second size threshold. In some embodiments, the third receptive field is configured to extract a target object feature in a video frame with a size greater than the second size threshold.
[0080] In some embodiments, the video frame has a size smaller than a first size threshold, and the video frame is input into a first receptive field to obtain the set of sampled features. In some embodiments, the video frame has a size larger than the first size threshold and smaller than a second size threshold, and the video frame is input into a second receptive field to obtain the set of sampled features. In some embodiments, the video frame has a size larger than the second size threshold, and the video frame is input into a third receptive field to obtain the set of sampled features. In some embodiments, the original video frame is treated as one of a plurality of video frames of different sizes to obtain the set of sampled features.
[0081] The video frame feature generation module is configured to generate a set of video frame features based on the plurality of sets of sampled features.
[0082] In some embodiments, the plurality of sets of sampled features are restored to the original size of the current video frame to generate a plurality of sets of restored features. In some embodiments, the plurality of sets of restored features are summed to obtain the set of video frame features. In some embodiments, a set of sampled features is restored according to the sampling manner to obtain a set of restored features. In some embodiments, a set of sampled features is reduced or enlarged according to a ratio of enlargement or reduction of the sampling manner to restore the set of sampled features to a size corresponding to the original size of the current video frame. In some embodiments, a set of sampled features obtained by inputting the original video frame into a receptive field is taken as one of the sets of restored features of the current video frame.
[0083] Reference should be made to Figure 4 , Figure 4 An electronic device is provided, including a processor and a memory. The memory stores computer instructions which, when executed by the processor, cause the processor to perform the method and refinements as shown in Figure 1 or Figure 2 .
[0084] It should be understood that the above-mentioned apparatus embodiments are only illustrative, and the apparatus disclosed in the present application can also be implemented in other manners. For example, the division of the units / modules in the above-mentioned embodiments is only a logical function division, and actual implementation can have another division manner. For example, a plurality of units / modules or components can be combined, or can be integrated into another system, or some features can be ignored or not executed.
[0085] In addition, each functional unit / module in each embodiment of the present application can be integrated in one unit / module, or each unit / module can exist physically, or two or more units / modules can be integrated together. The above-mentioned integrated unit / module can be realized in the form of hardware or in the form of a software program module.
[0086] The integrated units / modules, if implemented in the form of hardware, can be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the processor or chip can be any appropriate hardware processor, such as a CPU, a GPU, an FPGA, a DSP, an ASIC, etc. Unless otherwise specified, the on-chip cache, off-chip memory, storage can be any appropriate magnetic storage medium or magneto-optical storage medium, such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high-bandwidth memory (HBM), hybrid memory cube (HMC), etc.
[0087] The integrated units / modules, if implemented in the form of software program modules and sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the present application, essentially or the part that contributes to the prior art, or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present disclosure. The aforementioned storage medium includes: a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.
[0088] The embodiments of the present application also provide a non-transitory computer storage medium storing a computer program, which, when executed by a plurality of processors, causes the processors to perform the method and detailed solutions shown in Figure 1 or Figure 2 .
[0089] The above has carried out the detailed introduction to the embodiment of the application, the principle and implementation mode of the application have been described by applying specific examples in this paper, the above embodiment description is only used for helping understanding the method of the application and its core idea. At the same time, the changes or deformations made by the person skilled in the art on the basis of the specific implementation mode and the application range of the application according to the idea of the application all belong to the protection scope of the application. In summary, the content of the specification should not be understood as the limitation of the application.
Claims
1. A method for extracting target object features from video frames using a convolutional neural network, comprising: Calculate the proportion of the target object in the current video frame; The sampling method of the current video frame is determined based on the proportion of the target object to generate video frames of various sizes. The various video frames of different sizes are input into their respective receptive fields to generate multiple sets of sampling features; A video frame feature set is generated based on the multiple sampling feature sets; The step of generating a video frame feature set based on the plurality of sampled feature sets includes: The multiple sampled feature sets are restored to the original size of the current video frame to generate multiple restored feature sets; The summation of the multiple restored feature sets yields the video frame feature set.
2. The method as described in claim 1, characterized in that, Restoring the multiple sampled feature sets to the original size of the current video frame includes: The restoration is performed according to the sampling method described above.
3. The method as described in claim 1, characterized in that, The sampling methods include upsampling, downsampling, and oversampling, wherein: The upsampling and downsampling includes scaling up the current video frame once or multiple times and scaling it down once or multiple times. The upsampling includes magnifying the current video frame once or multiple times; The downsampling includes downsampling the current video frame once or multiple times.
4. The method as described in claim 3, characterized in that, The method of determining the sampling of the video frame based on the proportion of the target object includes: If the proportion of the target object is lower than the first proportion threshold, the sampling method of the video frame is determined to be upsampling; If the proportion of the target object is higher than the second proportion threshold, the sampling method of the video frame is determined to be downsampling; If the proportion of the target object is higher than a first proportion threshold and lower than a second proportion threshold, the sampling method of the video frame is determined to be upsampling or downsampling.
5. The method as described in claim 1, characterized in that, The receptive field includes a first receptive field, a second receptive field, and a third receptive field, wherein: The first receptive field is used to extract target features in video frames whose frame size is smaller than a first size threshold; The second receptive field is used to extract target features in video frames whose size is greater than the first size threshold and less than the second size threshold; The third receptive field is used to extract target features from video frames whose video frame size is greater than the second size threshold.
6. A method for extracting target object features from a target video using a convolutional neural network, comprising receiving the target video; The target video is decomposed into multiple video frames; As claimed in claim 1 The method described in any one of the 5 methods sequentially extracts target object features from the plurality of video frames to obtain a plurality of video frame feature sets, and uses the plurality of video frame feature sets as the target object features of the target video.
7. An apparatus for extracting features of a target object in a video frame using a convolutional neural network, comprising: The proportion recognition module is used to identify the proportion of the target object in the current video frame; The sampling module is used to determine the sampling method of the current video frame based on the proportion occupied by the target object, and generate video frames of various sizes; The sampling feature generation module is used to input the various video frames of different sizes into their corresponding receptive fields to generate multiple sampling feature sets; and A video frame feature generation module is used to generate a video frame feature set based on the plurality of sampled feature sets; Specifically, the video frame feature generation module is used for: The multiple sampled feature sets are restored to the original size of the current video frame to generate multiple restored feature sets; The summation of the multiple restored feature sets yields the video frame feature set.
8. An electronic device, comprising: processor; A memory storing a computer program, which, when executed by the processor, causes the processor to perform as claimed in claim 1. The method described in any one of the 6.
9. A non-transitory computer-readable storage medium having stored thereon computer-readable instructions that, when executed by a processor, cause the processor to perform the action claimed in claim 1. The method described in any one of the 6.
Citation Information
Patent Citations
Video size switching system and video size switching method facing display terminal
CN102541494A
Real-time object tracking method capable of supporting target size change
CN105488815A