Video recognition method, device, electronic device and storage medium
By performing stable frame group detection and keyframe image recognition on video frame images, the problems of large amount of video recognition calculation and low recognition accuracy in the prior art are solved, and more efficient and accurate video recognition is achieved.
Patent Information
- Application Number
- CN202210065095.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-18
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-01-18
AI Technical Summary
In the prior art, video recognition calculations are large and the recognition accuracy is not high.
By acquiring multiple video frame images in the video to be identified, detecting them, obtaining at least one stable frame group, acquiring the keyframe images in the stable frame group, and targeting the keyframe images.
The calculation amount of video recognition is reduced, and the recognition efficiency and accuracy are improved.
Smart Images

Figure CN114494954B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a video recognition method, device, electronic device and storage medium. Background Art
[0002] In the related technology, objects in video scenes are detected and classified, and a deep learning model is used to recognize images. Since a video is composed of many image frames, a deep learning model is used to recognize each image frame in the video, which requires a large amount of calculation and has a low recognition accuracy. Summary of the invention
[0003] The present disclosure provides a video recognition method, device, electronic device and storage medium to at least solve the problem of large amount of calculation and low recognition accuracy in related technologies. The technical solution of the present disclosure is as follows:
[0004] According to a first aspect of an embodiment of the present disclosure, a video recognition method is provided, comprising: obtaining a video to be recognized; wherein the video to be recognized includes a plurality of video frame images; detecting the video frame images to obtain at least one stable frame group; wherein the stable frame group includes a plurality of continuous stable frame images; obtaining a key frame image in the stable frame group; and performing target recognition on the key frame images.
[0005] In some embodiments, detecting the video frame image to obtain at least one stable frame group includes:
[0006] Acquire matching feature points between a current video frame image and a previous video frame image and motion parallax of the matching feature points; determine the stabilized frame image based on the motion parallax and the matching feature points; and when there are multiple consecutive stabilized frame images and the number of consecutive images is greater than or equal to a first threshold, determine the multiple consecutive stabilized frame images as one stabilized frame group.
[0007] In some embodiments, determining the stable frame image based on the motion parallax and the matching feature points includes: when the first number of the matching feature points is greater than or equal to a second threshold, obtaining the second number of the matching feature points whose motion parallax is greater than or equal to a third threshold; and when the second number is greater than or equal to a fourth threshold, determining that the current video frame image is the stable frame image.
[0008] In some embodiments, the method further includes: in the absence of a previous video frame image, determining that the current video frame image is not the key frame image.
[0009] In some embodiments, the method further includes: when the first number is less than the second threshold, or the second number is less than the fourth threshold, determining that the current video frame image is not the key frame image.
[0010] In some embodiments, obtaining the key frame image in the stable frame group includes: randomly selecting a target stable frame image from a plurality of consecutive stable frame images in the stable frame group, or selecting a target stable frame image corresponding to when the number of the stable frame images reaches the first threshold from a plurality of consecutive stable frame images in the stable frame group, and determining that the video frame image corresponding to the target stable frame image is the key frame image.
[0011] In some embodiments, the method further includes: when selecting a target stable frame image corresponding to the time when the number of the stable frame images reaches the first threshold, and determining that the video frame image corresponding to the target stable frame image is the key frame image, determining that the video frame image corresponding to the stable frame image located after the first threshold among multiple consecutive stable frame images is not the key frame image.
[0012] In some embodiments, the target recognition of the key frame image includes: processing the key frame image to generate an image to be recognized; inputting the image to be recognized into at least one deep learning model to obtain a target recognition result.
[0013] In some embodiments, the method further includes: performing deduplication aggregation processing on the multiple target recognition results of the multiple key frame images to obtain the recognition result of the video to be recognized.
[0014] According to a second aspect of an embodiment of the present disclosure, a video recognition device is provided, comprising: a video acquisition unit, used to acquire a video to be recognized; wherein the video to be recognized includes multiple video frame images; a first detection unit, used to detect the video frame images and acquire at least one stable frame group; wherein the stable frame group includes multiple continuous stable frame images; a key frame acquisition unit, used to acquire key frame images in the stable frame group; and a key frame recognition unit, used to perform target recognition on the key frame images.
[0015] In some embodiments, the first detection unit includes: a data acquisition module, used to obtain matching feature points between a current video frame image and a previous video frame image and motion parallax of the matching feature points; a stable frame determination module, used to determine the stable frame image based on the motion parallax and the matching feature points; a stable frame group determination module, used to determine multiple consecutive stable frame images as one stable frame group when there are multiple consecutive stable frame images and the number of consecutive images is greater than or equal to a first threshold.
[0016] In some embodiments, the stable frame determination module includes: a stable frame counting submodule, used to obtain the second number of matching feature points whose motion parallax is greater than or equal to a third threshold when the first number of matching feature points is greater than or equal to a second threshold; and a stable frame determination submodule, used to determine that the current video frame image is the stable frame image when the second number is greater than or equal to a fourth threshold.
[0017] In some embodiments, the first detection unit further includes: a non-key frame determination module, which is further used to determine that the current video frame image is not the key frame image when there is no previous video frame image.
[0018] In some embodiments, the stable frame determination module further includes: an unstable frame determination submodule, which is used to determine that the current video frame image is not the key frame image when the first number is less than the second threshold or the second number is less than the fourth threshold.
[0019] In some embodiments, the key frame acquisition unit is specifically used to randomly select a target stable frame image from a plurality of consecutive stable frame images in the stable frame group, or to select a target stable frame image corresponding to when the number of the stable frame images reaches the first threshold from a plurality of consecutive stable frame images in the stable frame group, and determine that the video frame image corresponding to the target stable frame image is the key frame image.
[0020] In some embodiments, the key frame acquisition unit is also used to, when selecting a target stable frame image corresponding to the time when the number of the stable frame images reaches the first threshold, determine that the video frame image corresponding to the target stable frame image is the key frame image, and determine that the video frame image corresponding to the stable frame image located after the first threshold among multiple consecutive stable frame images is not the key frame image.
[0021] In some embodiments, the key frame recognition unit includes: a processing module for processing the key frame image to generate an image to be recognized; and a first recognition module for inputting the image to be recognized into at least one deep learning model to obtain a target recognition result.
[0022] In some embodiments, the first recognition module is further used to perform deduplication and aggregation processing on the multiple target recognition results of the multiple key frame images to obtain the recognition result of the video to be recognized.
[0023] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the video recognition method as described in the first aspect above.
[0024] According to a fourth aspect of an embodiment of the present disclosure, a storage medium is provided. When instructions in the storage medium are executed by a processor of an electronic device, the electronic device can execute the video recognition method as described in the first aspect above.
[0025] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, including a computer program, which, when executed by a processor, implements the video recognition method as described in the first aspect above.
[0026] The technical solution provided by the embodiments of the present disclosure brings at least the following beneficial effects:
[0027] The video recognition method provided in the embodiment of the present disclosure obtains a video to be recognized by implementing the video recognition method provided in the embodiment of the present disclosure; wherein the video to be recognized includes multiple video frame images; the video frame images are detected to obtain at least one stable frame group; wherein the stable frame group includes multiple continuous stable frame images; key frame images in the stable frame group are obtained; and target recognition is performed on the key frame images. Thus, the amount of computation for video recognition can be reduced, recognition efficiency can be improved, and recognition accuracy can be improved.
[0028] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.
[0030] Figure 1 is a flow chart of a video recognition method according to an exemplary embodiment;
[0031] Figure 2 is a flow chart of another video recognition method according to an exemplary embodiment;
[0032] Figure 3 is a flowchart of S20 in a video recognition method according to an exemplary embodiment;
[0033] Figure 4 is a structural diagram of a video recognition device according to an exemplary embodiment;
[0034] Figure 5 is a structural diagram of a first detection unit in a video recognition device according to an exemplary embodiment;
[0035] Figure 6 is a structural diagram of a stable frame determination module in another video recognition device according to an exemplary embodiment;
[0036] Figure 7 is a structural diagram of a key frame recognition unit in another video recognition device according to an exemplary embodiment;
[0037] Figure 8 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0038] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.
[0039] Unless the context requires otherwise, throughout the specification and claims, the term "including" is to be interpreted as an open, inclusive meaning, that is, "including, but not limited to". In the description of the specification, the term "some embodiments" and the like are intended to indicate that specific features, structures, materials or characteristics associated with the embodiment or example are included in at least one embodiment or example of the present disclosure. The schematic representation of the above terms does not necessarily refer to the same embodiment or example. In addition, the specific features, structures, materials or characteristics may be included in any one or more embodiments or examples in any appropriate manner.
[0040] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the attached claims.
[0041] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0042] It should be noted that the video recognition method of the embodiment of the present disclosure can be executed by the video recognition device of the embodiment of the present disclosure, and the video recognition device can be implemented by software and / or hardware, and the video recognition device can be configured in an electronic device, wherein the electronic device can install and run a video display program. The electronic device can include but is not limited to hardware devices with various operating systems such as smart phones and tablet computers.
[0043] Figure 1 The figure is a flowchart of a video recognition method according to an exemplary embodiment.
[0044] like Figure 1 As shown, the video recognition method provided by the embodiment of the present disclosure includes but is not limited to the following steps:
[0045] S1: Obtain a video to be identified; wherein the video to be identified includes multiple video frame images.
[0046] It is understandable that the video to be identified can be a video file uploaded by a user, a local video file, or a real-time video stream collected by an image acquisition device such as a camera, etc. The execution subject of the video recognition method in the embodiment of the present disclosure can obtain the video to be identified.
[0047] The video to be identified includes multiple video frame images. It can be understood that video generally refers to various technologies that capture, record, process, store, transmit and reproduce a series of static images (such as video frame images) in the form of electrical signals. When the continuous image changes more than 24 frames per second, according to the principle of visual persistence, the human eye cannot distinguish a single static picture; it looks like a smooth and continuous visual effect, and such a continuous picture is called a video.
[0048] S2: Detect the video frame images to obtain at least one stable frame group; wherein the stable frame group includes a plurality of continuous stable frame images.
[0049] In the embodiment of the present disclosure, the video frame image is detected, and multiple video frame images in the video to be identified can be detected in sequence. The detection method can be to compare the video frame image of the current frame with the video frame image of the previous frame to determine whether the video frame image of the current frame is a stable frame image.
[0050] It is understandable that when the continuous image changes exceed 24 frames per second, the human eye cannot distinguish a single static image according to the principle of visual persistence; it looks like a smooth and continuous visual effect, and such a continuous image is called a video. Therefore, there are multiple consecutive frames in the video with similar image content.
[0051] Based on this, in the embodiment of the present disclosure, the video frame image of the current frame is determined to be a stable frame image, which means that the content in the video frame image of the current frame changes little relative to the content of the video frame image of the previous frame, and the contents of the video frame images of the two frames are similar.
[0052] Furthermore, the plurality of video frame images are detected in sequence, and when it is detected that the plurality of consecutive video frame images are all stable frame images, the plurality of consecutive video frame images that are stable frame images are determined as a stable frame group.
[0053] Wherein, a stable frame group includes a plurality of continuous video frame images, and a stable frame group may include two or more continuous video frame images.
[0054] In order to reduce the amount of calculation, in the embodiment of the present disclosure, the number of continuous video frame images included in a stable frame group can be determined according to the total number of frames of the video to be identified. When the total number of frames is large, the number of continuous video frame images included in a stable frame group can be appropriately increased. When the total number of frames is small, the number of continuous video frame images included in a stable frame group can be appropriately reduced. It can be determined according to the specific situation, and the embodiment of the present disclosure does not impose specific restrictions on this.
[0055] In the disclosed embodiment, the video frame images are detected, and the multiple video frame images in the video to be identified can be sampled at intervals of preset frames. For example, the multiple video frame images in the video to be identified are sampled at intervals of 5 frames, some video frame images are screened out from the multiple video frame images, and then the screened out video frame images are detected in turn, thereby reducing the number of video frame images to be detected and improving detection efficiency.
[0056] It should be noted that the above examples are for illustration only and are not intended to be specific limitations on the embodiments of the present disclosure. In the embodiments of the present disclosure, multiple video frame images in the video to be identified may be sampled at intervals of 10 or 20 frames, etc., and may be arbitrarily set as needed, and the embodiments of the present disclosure do not impose specific limitations on this.
[0057] S3: Acquire key frame images in the stable frame group.
[0058] In the embodiment of the present disclosure, after a stable frame group including a plurality of continuous stable frame images is acquired, a key frame image may be selected from the plurality of continuous stable frame images in the stable frame group.
[0059] S4: Perform target recognition on key frame images.
[0060] In the disclosed embodiment, the key frame image may be identified by inputting the key frame image into a deep learning model, identifying the content in the key frame image, and obtaining a result label of the key frame image.
[0061] In some embodiments, target recognition is performed on a key frame image, including: processing the key frame image to generate an image to be recognized; and inputting the image to be recognized into at least one deep learning model to obtain a target recognition result.
[0062] In the disclosed embodiment, the key frame image is processed by scaling the key frame image and adjusting it to the image input size of the trained deep learning model. Furthermore, the scaled key frame image is converted from the YUV color space to the RGB color space to obtain the image to be identified.
[0063] It should be noted that the above-mentioned step of processing the key frame image to generate the image to be identified may not be necessary. When the size of the key frame image meets the image input size of the deep learning model and its format is the RGB color space format, the key frame image may not be processed and the key frame image may be determined as the image to be identified.
[0064] In one possible implementation, when there is only one deep learning model, the image to be identified is input into the deep learning model to obtain a target recognition result of the image to be identified. The target recognition result may be the distribution probability of the image to be identified belonging to each category. The classification results are threshold filtered, and the category with the highest distribution probability is selected as the target recognition result.
[0065] In some embodiments, when multiple deep learning models are included, the multiple deep learning models have a hierarchical relationship; according to the hierarchical relationship, the image to be recognized is input into the corresponding deep learning model in sequence to obtain a target recognition result.
[0066] In one possible implementation, when there are multiple deep learning models, the multiple deep learning models have a hierarchical relationship. After obtaining the classification result of the deep learning model of the first level, the applicable deep learning model of the second level can be determined according to the classification result of the deep learning model of the first level, and further, until the deep learning model of the last level is input to obtain the target recognition result.
[0067] It is understandable that by using multiple levels of deep learning models to identify the image layer by layer, more accurate recognition results can be obtained and the accuracy of the recognition results can be improved.
[0068] It should be noted that in the embodiments of the present disclosure, the key frame images can also be identified by using image recognition methods in related technologies, which are not limited to deep learning models and can be set as needed. The embodiments of the present disclosure do not impose specific restrictions on this.
[0069] In the disclosed embodiment, when there is only one key frame image, the target recognition result of the key frame image is determined as the recognition result of the video to be recognized.
[0070] In some embodiments, when target recognition results of multiple key frame images are obtained, deduplication and aggregation processing is performed on the multiple target recognition results to obtain the recognition result of the video to be recognized.
[0071] In the disclosed embodiment, when there are multiple key frame images, multiple target recognition results of the multiple key frame images are deduplicated and aggregated, so as to obtain the recognition result of the video to be recognized.
[0072] By implementing the video recognition method provided by the embodiment of the present disclosure, a video to be recognized is obtained; wherein the video to be recognized includes multiple video frame images; the video frame images are detected to obtain at least one stable frame group; wherein the stable frame group includes multiple continuous stable frame images; key frame images in the stable frame group are obtained; and target recognition is performed on the key frame images. In this way, the amount of computation for video recognition can be reduced, recognition efficiency can be improved, and recognition accuracy can be improved.
[0073] Figure 2 The figure is a flowchart of another video recognition method according to an exemplary embodiment.
[0074] like Figure 2 As shown, including but not limited to the following steps:
[0075] S10: Obtain a video to be identified; wherein the video to be identified includes a plurality of video frame images.
[0076] The description of S10 in the embodiment of the present disclosure can refer to the description of S1 in the above embodiment, which will not be repeated here.
[0077] S20: Obtain matching feature points and motion disparity of the matching feature points between the current video frame image and the previous video frame image.
[0078] It is understandable that in the process of determining whether a video frame image is a stable frame image, the current video frame image needs to be compared with its adjacent previous video frame image. Only when the content difference between the two is small can the current video frame image be determined to be a stable frame image.
[0079] In some embodiments, the video frame images are also preprocessed in the embodiments of the present disclosure.
[0080] It can be understood that the format of the video frame image can be RGB color space or YUV color space, and the size of the video frame image also has various specifications. In the embodiment of the present disclosure, in order to facilitate unified processing, the video frame image is pre-processed before being recognized to generate a video frame image that can be processed uniformly.
[0081] In some embodiments, preprocessing the video frame image includes: scaling the video frame image to a target size image; converting the target size image to a grayscale image to generate the video frame image.
[0082] In the disclosed embodiment, the video frame image is preprocessed to scale the video frame image to a target size image, and further, the target size image is converted from a YUV color space or an RGB color space to a grayscale image, thereby obtaining a video frame image.
[0083] It should be noted that the above-mentioned step of preprocessing the video frame image to generate the video frame image may not be necessary. When the size of the video frame image meets the size required for subsequent processing, the video frame image may not be preprocessed.
[0084] Based on this, Figure 3 As shown, in some embodiments, S20 includes:
[0085] S201: Acquire multiple first feature points of a previous video frame image.
[0086] In the disclosed embodiment, the video frame image is a grayscale image, and multiple first feature points of the video frame image are obtained. By performing feature point detection on the video frame image, the Shi-Tomasi corner point detection algorithm is selected to obtain multiple first feature points of the video frame image.
[0087] The number of first feature points may be selected according to the size of the video frame image which is a grayscale image, for example, 100.
[0088] It should be noted that in the embodiments of the present disclosure, feature point detection of the video frame image can also be performed through the ORB (Oriented FAST and Rotated BRIEF) feature point algorithm, or, the SIFT feature point algorithm or the SURF feature point algorithm can also be used to obtain multiple first feature points of the video frame image.
[0089] In some embodiments, when there is no previous video frame image, it is determined that the current video frame image is not a key frame image.
[0090] It can be understood that the video frame image of the first frame in the video to be identified does not have a previous video frame image. In this case, it is determined that the video frame image of the first frame is not a key frame image, that is, it is determined that the current video frame image that does not have a previous video frame image is not a key frame image.
[0091] In the disclosed embodiment, based on determining that the current video frame image without the previous video frame image is not a key frame image, multiple first feature points of the video frame image of the first frame are obtained and saved for processing the video frame images of subsequent frames.
[0092] S202: Acquire matching feature points that match the first feature points according to the previous video frame image, the multiple first feature points, and the current video frame image.
[0093] S203: Obtain the image position distance of the matching feature points between the current video frame image and the previous video frame image to determine the motion parallax.
[0094] In the disclosed embodiment, a previous video frame image, a plurality of first feature points, and a current video frame image are used as inputs for optical flow calculation, and a pyramid Lucas-Kanade optical flow algorithm is selected to obtain matching feature points that match the first feature points, as well as the image position distance of the matching feature points between the current video frame image and the previous video frame image, to determine motion parallax.
[0095] It should be noted that in the embodiments of the present disclosure, when multiple first feature points of a video frame image are obtained by using a SIFT feature point algorithm or a SURF feature point algorithm, a pyramid-free Lucas-Kanade optical flow algorithm can also be used to obtain matching feature points that match the first feature points, as well as the image position distance between the matching feature points in the current video frame image and the previous video frame image to determine the motion parallax, or other methods can be used, and the embodiments of the present disclosure do not impose specific restrictions on this.
[0096] Please continue to see Figure 2 , S30: Determine a stable frame image according to the motion parallax and matching feature points.
[0097] In the disclosed embodiment, when obtaining data on matching feature points and motion parallax of a current video frame image compared to a previous video frame image, it is possible to determine whether the current video frame image is a stable frame image based on the obtained data on matching feature points and motion parallax.
[0098] Specifically, how to determine whether the current video frame image is a stable frame image according to the matching feature points and motion parallax of the current video frame image compared with the previous video frame image.
[0099] In some embodiments, when the first number of matching feature points is greater than or equal to a second threshold, a second number of matching feature points whose motion parallax is greater than or equal to a third threshold is obtained; when the second number is greater than or equal to a fourth threshold, the current video frame image is determined to be a stable frame image.
[0100] In the disclosed embodiment, matching feature points that match the first feature points are obtained based on a previous video frame image, a plurality of first feature points, and a current video frame image, and a first number of matching feature points is counted. When the first number of matching feature points of the current video frame image is greater than or equal to a second threshold, it is considered that the content change of the current video frame image is smaller than that of the previous video frame image. Furthermore, a second number of matching feature points whose motion parallax is greater than or equal to a third threshold is obtained. When the second number is greater than or equal to a fourth threshold, it is considered that the content change of the current video frame image is small enough compared to that of the previous video frame image, and the current video frame image is determined to be a stable frame image.
[0101] The value of the second threshold can be determined based on the total number of first feature points. Exemplarily, it is obtained by multiplying the total number of first feature points by a certain ratio, for example, the total number of first feature points multiplied by 0.8.
[0102] The value of the third threshold may be determined according to the number of pixels of the size of the video frame image which is a grayscale image. Exemplarily, the third threshold may be 16 pixels.
[0103] The value of the fourth threshold can also be determined based on the total number of first feature points obtained, or can also be determined based on the value of the second threshold. Exemplarily, it is obtained by multiplying the value of the second threshold by a certain ratio, for example, multiplying the value of the second threshold by 0.8.
[0104] It should be noted that the above examples are for illustration only. In the embodiment of the present disclosure, the values of the second threshold and the third threshold may also be other parameters and may be set as needed. The embodiment of the present disclosure does not impose any specific limitation on this.
[0105] In some embodiments, when the first number is smaller than the second threshold, or the second number is smaller than the fourth threshold, it is determined that the current video frame image is not a key frame image.
[0106] In the disclosed embodiment, matching feature points that match the first feature points are obtained based on the previous video frame image, multiple first feature points, and the current video frame image, and the first number of matching feature points is counted. When the first number of matching feature points of the current video frame image is less than a second threshold, it is considered that the content of the current video frame image is significantly changed compared to the previous video frame image, and it is determined that the current video frame image is not a key frame image.
[0107] In addition, a second number of matching feature points whose motion parallax is greater than or equal to the third threshold is obtained. When the second number is less than the fourth threshold, it is considered that the content of the current video frame image has changed significantly compared to the previous video frame image, and it is determined that the video frame image corresponding to the current video frame image is not a key frame image.
[0108] Please continue to see Figure 2 , S40: when there are multiple continuous stable frame images and the number of continuous images is greater than or equal to a first threshold, the multiple continuous stable frame images are determined as a stable frame group.
[0109] In the embodiment of the present disclosure, when there are multiple continuous stable frame images and the number of continuous images is greater than or equal to a first threshold, the multiple continuous stable frame images are determined as a stable frame group.
[0110] It is understandable that in the embodiments of the present disclosure, for a video to be identified, if the content differences of multiple consecutive video frame images are small, multiple consecutive video frame images with similar content can be determined as a stable frame group, that is, multiple consecutive stable frame images are determined as a stable frame group. For the consecutive number of multiple consecutive stable frame images, if the consecutive number is less than a certain number, the content of the video frame image corresponding to the stable frame image may not be very useful for the identification result of the video to be identified. On the contrary, there may be cases of misidentification. Based on this, in the embodiments of the present disclosure, multiple consecutive stable frame images are determined as a stable frame group only when there are multiple consecutive stable frame images and the consecutive number is greater than or equal to the first threshold.
[0111] In some embodiments, when there are multiple consecutive stable frame images and the number of consecutive images is less than a first threshold, it is determined that the video frame image corresponding to the stable frame image is not a key frame image.
[0112] Based on the description in the above example, for the number of consecutive stable frame images, if the number of consecutive images is less than the first threshold, the content of the video frame image corresponding to the stable frame image may not be very useful for the recognition result of the video to be recognized. On the contrary, there may still be misrecognition. Based on this, in the embodiment of the present disclosure, when there are multiple consecutive stable frame images and the number of consecutive images is less than the first threshold, it is determined that the video frame image corresponding to the stable frame image is not a key frame image. This can reduce the misrecognition rate and improve the recognition accuracy.
[0113] S50: Acquire key frame images in the stable frame group.
[0114] In some embodiments, a target stable frame image is randomly selected from a plurality of consecutive stable frame images in a stable frame group, or a target stable frame image corresponding to a time when the number of stable frame images reaches a first threshold is selected from a plurality of consecutive stable frame images in a stable frame group, and a video frame image corresponding to the target stable frame image is determined as a key frame image.
[0115] It can be understood that the contents of multiple consecutive stable frame images of a stable frame group are similar. A target stable frame image can be randomly selected and its corresponding video frame image can be determined as a key frame image. The content of the key frame image can characterize the content of multiple consecutive stable frame images of the stable frame group, thereby reducing the number of detected video frame images, reducing the amount of calculation, and improving detection efficiency.
[0116] In some embodiments, when a target stable frame image corresponding to a target stable frame image is selected when the number of stable frame images reaches a first threshold, and when it is determined that the video frame image corresponding to the target stable frame image is a key frame image, it is determined that the video frame image corresponding to the stable frame image located after the first threshold among multiple consecutive stable frame images is not a key frame image.
[0117] In the disclosed embodiment, when a target stable frame image corresponding to the time when the number of stable frame images reaches a first threshold is selected from a plurality of continuous stable frame images in a stable frame group, and when it is determined that the video frame image corresponding to the target stable frame image is a key frame image, it is determined that the video frame image corresponding to the stable frame image located after the first threshold in the plurality of continuous stable frame images is not a key frame image. Therefore, after determining a key frame image in a stable frame group, the disclosed embodiment determines that the other stable frame images in the stable frame group are not key frame images, which can reduce the computational complexity of video recognition and improve recognition efficiency.
[0118] In order to better understand the method for obtaining key frame images in a video to be identified in the video identification method provided in the embodiment of the present disclosure, the following example is provided:
[0119] Step 1: Determine whether the current video frame image has the previous video frame image and its multiple first feature points.
[0120] If it does not exist, the feature point detection is performed on the current video frame image. The feature point detection method uses the Shi-Tomasi corner point detection algorithm. The number of feature points can be selected according to the grayscale image size, for example, 100 is selected here. The first feature point detected is cached together with the current video frame image for use in the next frame. The calculation of the frame is completed, and the video frame image corresponding to the current video frame image is determined not to be a key frame image.
[0121] Among them, if it exists, the previous video frame image, multiple first feature points, and the current video frame image are taken as input, and the pyramid Lucas-Kanade optical flow algorithm is selected to obtain whether there is a matching feature point of the first feature point of the previous video frame image in the current video frame image, and the image pixel position information in the current video frame image.
[0122] Step 2: Count the first number of matching feature points of the previous video frame image that match the current video frame image. If the first number is less than the second threshold, it is considered that the content of the current video frame image has changed significantly compared to the previous video frame image, and the video frame image corresponding to the current video frame image is determined not to be a key frame image. The feature point information of the current video frame image is cleared, and the count of the stable frame image is cleared to 0. The calculation of the current video frame image is completed, and the video frame image corresponding to the current video frame image is determined not to be a key frame image.
[0123] Step 3: Calculate the image pixel position information of the matching feature points in the current video frame image and the previous video frame image, obtain the image position distance between the matching feature points in the current video frame image and the previous video frame image, and determine the motion parallax.
[0124] Step 4: Obtain a second number of matching feature points whose motion parallax is greater than or equal to the third threshold. The value of the third threshold can be determined according to the number of pixels of the size of the video frame image which is a grayscale image. Exemplarily, the third threshold can be 16 pixels. When the second number is greater than or equal to the fourth threshold, the current video frame image is determined to be a stable frame image. The value of the fourth threshold can also be determined based on the total number of first feature points obtained, or can also be determined based on the value of the second threshold. Exemplarily, it is obtained by multiplying the value of the second threshold by a certain ratio, for example, the value of the second threshold is multiplied by 0.8.
[0125] When the second number is smaller than the fourth threshold, the count of the stable frame image is cleared to 0, the calculation of the current video frame image is terminated, and the current video frame image is determined not to be a key frame image.
[0126] When the second number is greater than or equal to the fourth threshold, the count of the stable frame image is increased by 1.
[0127] Step 5: The count of the stable frame images is judged against the first threshold. If there are multiple continuous stable frame images and the number of continuous images is greater than or equal to the first threshold, the multiple continuous stable frame images are determined as a stable frame group. If there are multiple continuous stable frame images and the number of continuous images is less than the first threshold, it is determined that the video frame images corresponding to the multiple continuous stable frame images are not key frame images.
[0128] Then, a target stable frame image is randomly selected from the multiple continuous stable frame images of the stable frame group, or a target stable frame image corresponding to the time when the number of stable frame images reaches a first threshold is selected from the multiple continuous stable frame images of the stable frame group, and the video frame image corresponding to the target stable frame image is determined as the key frame image.
[0129] Exemplarily, when a target stable frame image corresponding to a first threshold is selected from a plurality of continuous stable frame images of a stable frame group, and a video frame image corresponding to the target stable frame image is determined to be a key frame image, the count of the stable frame images and the first threshold are judged, and when there are a plurality of continuous stable frame images, and the number of continuous images is equal to the first threshold, the video frame image corresponding to a target stable frame image corresponding to the time when the number of stable frame images reaches the first threshold is determined to be a key frame image. And in the subsequent determination of the stable frame images, when the video frame image corresponding to a target stable frame image corresponding to the first threshold output before the determination is a key frame image, it is determined that the video frame image corresponding to the stable frame image located after the first threshold among the plurality of continuous stable frame images is not a key frame image.
[0130] S60: Perform target recognition on the key frame image.
[0131] The description of S60 in the embodiment of the present disclosure can refer to the relevant description of S4 in the above embodiment, which will not be repeated here.
[0132] For example, taking a 30-second video with a frame rate of 30fps as an example, the original video has a total of 900 frames of images. Under certain circumstances, if the video content has been stable and has not changed, only one key frame image needs to be identified. Under the normal condition of constantly changing video content, according to different key frame image judgment parameter settings, generally only a maximum of dozens of frames of images need to be identified, and the amount of calculation is greatly reduced.
[0133] In summary, in the disclosed embodiments, key frame images suitable for image recognition can be effectively obtained, and the number of key frame images is much smaller than the number of video frame images in the video to be recognized, thereby significantly reducing the amount of computation for video recognition and improving recognition efficiency. At the same time, the content of the key frame images is stable, and generally no blurred images caused by motion are selected, and the recognition accuracy is high, so that the overall recognition result accuracy of the video to be recognized is high.
[0134] Figure 4 The figure is a structural diagram of a video recognition device according to an exemplary embodiment.
[0135] like Figure 4 As shown, the video recognition device 1 includes: a video acquisition unit 11, a first detection unit 12, a key frame acquisition unit 13 and a key frame recognition unit 14.
[0136] The video acquisition unit 11 is used to acquire a video to be identified; wherein the video to be identified includes a plurality of video frame images.
[0137] The first detection unit 12 is used to detect the video frame image and obtain at least one stable frame group; wherein the stable frame group includes a plurality of continuous stable frame images.
[0138] The key frame acquisition unit 13 is used to acquire the key frame images in the stable frame group.
[0139] The key frame recognition unit 14 is used to perform target recognition on the key frame image.
[0140] like Figure 5 As shown, in some embodiments, the first detection unit 12 includes: a data acquisition module 121, a stable frame determination module 122 and a stable frame group determination module 123.
[0141] The data acquisition module 121 is used to acquire matching feature points and motion parallax of the matching feature points between the current video frame image and the previous video frame image.
[0142] The stable frame determination module 122 is used to determine a stable frame image according to the motion parallax and the matching feature points.
[0143] The stable frame group determination module 123 is configured to determine the multiple continuous stable frame images as a stable frame group when there are multiple continuous stable frame images and the number of continuous stable frame images is greater than or equal to a first threshold.
[0144] like Figure 6 As shown, in some embodiments, the stable frame determination module 122 includes: a stable frame counting submodule 1231 and a stable frame determination submodule 1222 .
[0145] The stable frame counting submodule 1221 is configured to obtain a second number of matching feature points having a motion parallax greater than or equal to a third threshold value when the first number of matching feature points is greater than or equal to a second threshold value.
[0146] The stable frame determination submodule 1222 is used to determine that the current video frame image is a stable frame image when the second number is greater than or equal to a fourth threshold.
[0147] Please continue to see Figure 5 In some embodiments, the first detection unit 12 further includes: a non-key frame determination module 124, which is used to determine that the current video frame image is not a key frame image when there is no previous video frame image.
[0148] Please continue to see Figure 6 In some embodiments, the stable frame determination module 122 further includes: an unstable frame determination submodule 1223 .
[0149] The unstable frame determination submodule 1223 is used to determine that the current video frame image is not a key frame image when the first number is smaller than the second threshold or the second number is smaller than the fourth threshold.
[0150] In some embodiments, the key frame acquisition unit 13 is specifically used to randomly select a target stable frame image from a plurality of consecutive stable frame images in a stable frame group, or to select a target stable frame image corresponding to when the number of stable frame images reaches a first threshold from a plurality of consecutive stable frame images in a stable frame group, and determine that the video frame image corresponding to the target stable frame image is a key frame image.
[0151] In some embodiments, the key frame acquisition unit 13 is also used to select a target stable frame image corresponding to when the number of stable frame images reaches a first threshold, and determine that the video frame image corresponding to the target stable frame image is a key frame image, and determine that the video frame image corresponding to the stable frame image located after the first threshold among multiple consecutive stable frame images is not a key frame image.
[0152] like Figure 7 As shown, in some embodiments, the key frame identification unit 14 includes: a processing module 141 and a first identification module 142 .
[0153] The processing module 141 is used to process the key frame image to generate an image to be recognized.
[0154] The first recognition module 142 is used to input the image to be recognized into at least one deep learning model to obtain a target recognition result.
[0155] In some embodiments, the first recognition module 142 is further used to perform deduplication and aggregation processing on multiple target recognition results of multiple key frame images to obtain recognition results of the video to be recognized.
[0156] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0157] The beneficial effects that can be achieved by the video recognition device provided in the embodiment of the present disclosure are the same as the beneficial effects achieved by the video recognition method provided in the above example, and will not be repeated here.
[0158] Figure 8 is a block diagram of an electronic device 100 for a video recognition method according to an exemplary embodiment.
[0159] Exemplarily, the electronic device 100 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, and the like.
[0160] like Figure 8 As shown, the electronic device 100 may include one or more of the following components: a processing component 101, a memory 102, a power component 103, a multimedia component 104, an audio component 105, an input / output (I / O) interface 106, a sensor component 107, and a communication component 108.
[0161] The processing component 101 generally controls the overall operation of the electronic device 100, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 101 may include one or more processors 1011 to execute instructions to complete all or part of the steps of the above-mentioned method. In addition, the processing component 101 may include one or more modules to facilitate the interaction between the processing component 101 and other components. For example, the processing component 101 may include a multimedia module to facilitate the interaction between the multimedia component 104 and the processing component 101.
[0162] The memory 102 is configured to store various types of data to support operations on the electronic device 100. Examples of such data include instructions for any application or method operating on the electronic device 100, contact data, phone book data, messages, pictures, videos, etc. The memory 102 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as SRAM (Static Random-Access Memory), EEPROM (Electrically Erasable Programmable read only memory), EPROM (Erasable Programmable Read-Only Memory), PROM (Programmable read-only memory), ROM (Read-Only Memory), magnetic storage, flash memory, magnetic disk or optical disk.
[0163] The power supply assembly 103 provides power to various components of the electronic device 100. The power supply assembly 103 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 100.
[0164] The multimedia component 104 includes a touch screen that provides an output interface between the electronic device 100 and the user. In some embodiments, the touch screen may include an LCD (Liquid Crystal Display) and a TP (Touch Panel). The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor can not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 104 includes a front camera and / or a rear camera. When the electronic device 100 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.
[0165] The audio component 105 is configured to output and / or input audio signals. For example, the audio component 105 includes a MIC (Microphone), and when the electronic device 100 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 102 or sent via the communication component 108. In some embodiments, the audio component 105 also includes a speaker for outputting audio signals.
[0166] I / O interface 2112 provides an interface between processing component 101 and peripheral interface modules, which may be keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: home button, volume button, start button, and lock button.
[0167] The sensor assembly 107 includes one or more sensors for providing various aspects of status assessment for the electronic device 100. For example, the sensor assembly 107 can detect the open / closed state of the electronic device 100, the relative positioning of the components, such as the display and keypad of the electronic device 100, and the sensor assembly 107 can also detect the position change of the electronic device 100 or a component of the electronic device 100, the presence or absence of contact between the user and the electronic device 100, the orientation or acceleration / deceleration of the electronic device 100, and the temperature change of the electronic device 100. The sensor assembly 107 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 107 may also include an optical sensor, such as a CMOS (Complementary Metal Oxide Semiconductor) or CCD (Charge-coupled Device) image sensor for use in imaging applications. In some embodiments, the sensor assembly 107 may also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0168] The communication component 108 is configured to facilitate wired or wireless communication between the electronic device 100 and other devices. The electronic device 100 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 108 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 108 also includes an NFC (Near Field Communication) module to facilitate short-range communication. For example, the NFC module can be implemented based on RFID (Radio Frequency Identification) technology, IrDA (Infrared Data Association) technology, UWB (Ultra Wide Band) technology, BT (Bluetooth) technology and other technologies.
[0169] In an exemplary embodiment, the electronic device 100 can be implemented by one or more ASICs (Application Specific Integrated Circuit), DSPs (Digital Signal Processor), digital signal processing devices (DSPDs), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), controllers, microcontrollers, microprocessors or other electronic components to execute the above-mentioned video recognition method.
[0170] It should be noted that the implementation process and technical principles of the electronic device of this embodiment refer to the aforementioned explanation of the video recognition method of the embodiment of the present disclosure, and will not be repeated here.
[0171] The electronic device provided by the embodiments of the present disclosure can execute the video recognition method as described in some of the above embodiments, and its beneficial effects are the same as the beneficial effects of the above-mentioned video recognition method, which will not be repeated here.
[0172] In order to implement the above embodiments, the present disclosure also proposes a storage medium.
[0173] When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device can perform the video recognition method as described above. For example, the storage medium can be ROM (Read Only Memory Image), RAM (Random Access Memory), CD-ROM (Compact Disc Read-Only Memory), magnetic tape, floppy disk and optical data storage device.
[0174] In order to implement the above embodiments, the present disclosure further provides a computer program product. When the computer program is executed by a processor of an electronic device, the electronic device can execute the video recognition method as described above.
[0175] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.
[0176] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A video recognition method, characterized in that: include: Acquire a video to be identified; wherein the video to be identified includes multiple video frame images; Detecting the video frame images to obtain at least one stable frame group; wherein the stable frame group includes a plurality of continuous stable frame images; Acquire a key frame image in the stable frame group; Performing target recognition on the key frame image; The detecting the video frame image to obtain at least one stable frame group includes: Obtaining matching feature points between a current video frame image and a previous video frame image and motion parallax of the matching feature points, including: obtaining a plurality of first feature points of the previous video image, obtaining matching feature points that match the first feature points according to the previous video frame image, the plurality of first feature points, and the current video frame image, obtaining an image position distance of the matching feature points between the current video image and the previous video image, and determining motion parallax; Determining the stable frame image according to the motion parallax and the matching feature points; When there are a plurality of continuous stable frame images, and the number of continuous images is greater than or equal to a first threshold, the plurality of continuous stable frame images are determined as one stable frame group.
2. The method according to claim 1, characterized in that The step of determining the stable frame image according to the motion parallax and the matching feature points comprises: When the first number of the matching feature points is greater than or equal to the second threshold, obtaining a second number of the matching feature points whose motion parallax is greater than or equal to a third threshold; When the second number is greater than or equal to the fourth threshold, the current video frame image is determined to be the stable frame image.
3. The method according to claim 2, characterized in that The method further comprises: In the case that there is no previous video frame image, it is determined that the current video frame image is not the key frame image.
4. The method according to claim 2, characterized in that: The method further comprises: When the first number is smaller than the second threshold, or the second number is smaller than the fourth threshold, it is determined that the current video frame image is not the key frame image.
5. The method according to claim 2, characterized in that: The step of acquiring the key frame image in the stable frame group includes: A target stable frame image is randomly selected from a plurality of continuous stable frame images in the stable frame group, or a target stable frame image corresponding to the time when the number of the stable frame images reaches the first threshold is selected from a plurality of continuous stable frame images in the stable frame group, and the video frame image corresponding to the target stable frame image is determined as the key frame image.
6. The method according to claim 5, characterized in that The method further comprises: When a target stable frame image corresponding to the time when the number of the stable frame images reaches the first threshold is selected and the video frame image corresponding to the target stable frame image is determined to be the key frame image, it is determined that the video frame image corresponding to the stable frame image located after the first threshold among multiple consecutive stable frame images is not the key frame image.
7. The method according to claim 1, characterized in that The performing target recognition on the key frame image comprises: Processing the key frame image to generate an image to be recognized; The image to be recognized is input into at least one deep learning model to obtain a target recognition result.
8. The method according to claim 7, characterized in that The method further comprises: Deduplication and aggregation processing is performed on the multiple target recognition results of the multiple key frame images to obtain the recognition result of the video to be recognized.
9. A video recognition device, characterized in that: include: A video acquisition unit, used to acquire a video to be identified; wherein the video to be identified includes a plurality of video frame images; A first detection unit, configured to detect the video frame image and obtain at least one stable frame group; wherein the stable frame group includes a plurality of continuous stable frame images; A key frame acquisition unit, used for acquiring key frame images in the stable frame group; A key frame recognition unit, used for performing target recognition on the key frame image; The first detection unit comprises: A data acquisition module, used to acquire matching feature points between a current video frame image and a previous video frame image and motion parallax of the matching feature points, including: acquiring a plurality of first feature points of the previous video image, acquiring matching feature points that match the first feature points according to the previous video frame image, the plurality of first feature points, and the current video frame image, acquiring an image position distance of the matching feature points between the current video image and the previous video image, and determining motion parallax; A stable frame determination module, used to determine the stable frame image according to the motion parallax and the matching feature points; The stable frame group determination module is used to determine the multiple continuous stable frame images as one stable frame group when there are multiple continuous stable frame images and the number of continuous images is greater than or equal to a first threshold.
10. The device according to claim 9, characterized in that The stable frame determination module comprises: a stable frame counting submodule, configured to obtain, when the first number of the matching feature points is greater than or equal to the second threshold, a second number of the matching feature points whose motion parallax is greater than or equal to a third threshold; The stable frame determination submodule is used to determine that the current video frame image is the stable frame image when the second number is greater than or equal to a fourth threshold.
11. The device according to claim 10, characterized in that The first detection unit further includes: The non-key frame determination module is further used to determine that the current video frame image is not the key frame image when there is no previous video frame image.
12. The device according to claim 10, characterized in that The stable frame determination module further includes: The unstable frame determination submodule is used to determine that the current video frame image is not the key frame image when the first number is smaller than the second threshold or the second number is smaller than the fourth threshold.
13. The device according to claim 9, characterized in that The key frame acquisition unit is specifically used to randomly select a target stable frame image from multiple consecutive stable frame images in the stable frame group, or to select a target stable frame image corresponding to when the number of the stable frame images reaches the first threshold from multiple consecutive stable frame images in the stable frame group, and determine that the video frame image corresponding to the target stable frame image is the key frame image.
14. The device according to claim 13, characterized in that The key frame acquisition unit is also used to select a target stable frame image corresponding to the time when the number of the stable frame images reaches the first threshold, and determine that the video frame image corresponding to the target stable frame image is the key frame image, and then determine that the video frame image corresponding to the stable frame image located after the first threshold among multiple consecutive stable frame images is not the key frame image.
15. The device according to claim 9, characterized in that The key frame identification unit comprises: A processing module, used for processing the key frame image to generate an image to be recognized; The first recognition module is used to input the image to be recognized into at least one deep learning model to obtain a target recognition result.
16. The device according to claim 15, characterized in that The first recognition module is further used to perform deduplication and aggregation processing on the multiple target recognition results of the multiple key frame images to obtain the recognition result of the video to be recognized.
17. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the method according to any one of claims 1 to 8.
18. A storage medium, characterized in that: When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method as claimed in any one of claims 1 to 8.
19. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
A video data processing method and a related device
CN109697416A