Video sampling method and device

By dynamically adjusting the size and position of the video sampling window and obtaining the action completion rate based on the AI ​​model, the problem of inaccurate action recognition in existing technologies is solved, achieving high-precision action recognition and positioning, and reducing computational redundancy.

CN113128256BActive Publication Date: 2026-04-03BEIJING SAMSUNG TELECOM R&D CENT +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-12-30
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In existing technologies, the fixed size of the video sampling window and the fixed movement step size lead to inaccurate behavior recognition, which cannot effectively cover the entire process of the action, resulting in low recognition accuracy and computational redundancy.

Method used

By obtaining the completion rate of the action based on an AI model, the size and position of the sampling window are dynamically adjusted to achieve high-precision sampling of the video and ensure that the window covers the entire action process.

Benefits of technology

It achieves high-precision recognition and localization of actions, improves the accuracy of video sampling, and reduces computational redundancy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113128256B_ABST
    Figure CN113128256B_ABST
Patent Text Reader

Abstract

A video sampling method and apparatus are provided. The video sampling method includes: sampling a video based on a sampling window to obtain a current sampled image sequence; obtaining motion parameters corresponding to the current sampled image sequence; adjusting the sampling window according to the motion parameters corresponding to the current sampled image sequence; and sampling the video based on the adjusted sampling window.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of video processing technology. More specifically, this disclosure relates to a video sampling method and apparatus. Background Technology

[0002] Action recognition in videos has a wide range of applications, primarily including human-computer intelligent interaction, intelligent video editing and editing, etc.

[0003] In the active perception application scenarios of intelligent robots, existing robots can only passively interact with humans through voice. Giving robots the ability to recognize human actions is crucial for their perception, imitation, and reasoning about human behavior, and will be an important foundation for the future development of intelligent robots. In the application scenarios of monitoring occupants in autonomous vehicles, the behavior analysis of the driver and passengers is an important component of autonomous driving. It enables the monitoring of abnormal driver behavior and the analysis of driver and passenger behavioral characteristics, thereby realizing an intelligent and personalized in-vehicle perception and interaction system.

[0004] With the development of mobile networks, video has gradually become a new medium for information sharing and dissemination. Identifying and locating people's behavior and actions in videos can enable functions such as automatic editing of exciting actions, slow-motion playback, and editing of special effects. Summary of the Invention

[0005] Exemplary embodiments of this disclosure provide a video sampling method and apparatus for optimizing existing video sampling methods.

[0006] According to an exemplary embodiment of the present disclosure, a video sampling method is provided, comprising: sampling a video based on a sampling window to obtain a current sampled image sequence; obtaining motion parameters corresponding to the current sampled image sequence; adjusting the sampling window according to the motion parameters corresponding to the current sampled image sequence; and sampling the video based on the adjusted sampling window.

[0007] Optionally, the action parameters may include: the probability that the current sampled image sequence contains an action and / or the degree of completion of the contained action.

[0008] Optionally, the step of obtaining the action parameters corresponding to the current sampled image sequence may include: extracting features from the current sampled image sequence to obtain the features of the current sampled image sequence; and performing feature recognition on the features of the current sampled image sequence to obtain the action parameters in the current sampled image sequence.

[0009] Optionally, the action parameters may include: the probability that the current sampled image sequence contains an action and the degree of completion of the contained action. The step of performing feature recognition on the features of the current sampled image sequence may include: performing action recognition on the features of the current sampled image sequence to obtain the probability that the current sampled image sequence contains an action; and performing action completion recognition on the features of the current sampled image sequence to obtain the degree of completion of the contained action in the current sampled image sequence.

[0010] Optionally, the step of adjusting the sampling window may include: calculating the size change value of the sampling window and / or the position of the sampling window movement based on the motion parameters in the current sampling image sequence; and adjusting the sampling window based on the size change value of the sampling window and / or the position of the sampling window movement.

[0011] Optionally, the step of calculating the size change value of the sampling window and / or the position moved by the sampling window may include: when it is determined from the action parameters that the action contained in the current sampling image sequence is not completed, calculating the increment value of the sampling window according to the action parameters; and determining the position moved by the sampling window according to the increment value of the sampling window.

[0012] Optionally, the video sampling method may further include: smoothing the incremental values ​​of the sampling window based on the historical adjustment information of the sampling window.

[0013] Optionally, the step of determining the position of the sampling window movement may include: determining the number of frames of the augmented image to be sampled in the augmented window corresponding to the incremental value of the sampling window based on the incremental value of the sampling window; and determining the position of the sampling window movement based on the determined number of frames.

[0014] Optionally, the step of sampling the video based on the adjusted sampling window may include: sampling the video in the augmentation window according to a determined number of frames to obtain an augmented image sequence; obtaining the sampled image sequence corresponding to the adjusted sampling window in the current sampled image sequence; and using the sampled image sequence corresponding to the adjusted sampling window and the augmented image sequence as the sampled image sequence obtained based on the adjusted sampling window.

[0015] Optionally, the video sampling method may further include: selecting the action with the highest probability of being included in the current sampled image sequence from various actions as the action included in the current sampled image sequence; determining that the action included in the current sampled image sequence has been completed when the probability of the action included in the current sampled image sequence is greater than a first threshold and the completion degree of the action included in the current sampled image sequence is greater than a second threshold; and / or determining that the current sampled image sequence does not contain any action when the probability of the action included in the current sampled image sequence is less than a third threshold and the completion degree of the action included in the current sampled image sequence is less than a fourth threshold; and / or determining that the action included in the current sampled image sequence is not completed when the probability of the action included in the current sampled image sequence is between the first threshold and the third threshold and the completion degree of the action included in the current sampled image sequence is between the second threshold and the fourth threshold.

[0016] Optionally, the steps of calculating the size change value of the sampling window and / or the position moved by the sampling window may include: when it cannot be determined whether the action in the current sampled image sequence is completed based on the action parameters, calculating the decrease value of the sampling window based on the action parameters; and determining the position moved by the sampling window based on the decrease value of the sampling window.

[0017] Optionally, the step of determining the position of the sampling window based on the reduction value of the sampling window may include: subtracting a window with a length equal to the reduction value of the sampling window from both ends of the sampling window to obtain a first sampling window and a second sampling window; using one of the first sampling window and the second sampling window as the updated sampling window; and determining the position of the current sampling window based on the updated sampling window.

[0018] Optionally, the step of using one of the first sampling window and the second sampling window as the updated sampling window may include: sampling a preset number of frames of sampled image sequences from the first sampling window and the second sampling window respectively; extracting features and recognizing features from the sampled image sequences sampled from the first sampling window and the sampled image sequences sampled from the second sampling window respectively; and selecting one of the first sampling window and the second sampling window as the updated sampling window based on the recognition results.

[0019] Optionally, the position to which the sampling window moves is the position to which the starting point of the sampling window moves.

[0020] Optionally, the step of performing action recognition on the features of the current sampled image sequence may include: performing three-dimensional linear gating processing on the features of the current sampled image sequence to obtain action recognition features; and obtaining the probability that the current sampled image sequence contains an action based on the obtained action recognition features.

[0021] Optionally, the step of identifying the degree of action completion of the features of the current sampled image sequence may include: performing three-dimensional linear gating processing on the features of the current sampled image sequence to obtain action completion recognition features; and obtaining the degree of action completion contained in the current sampled image sequence based on the obtained action completion recognition features.

[0022] Optionally, the step of performing three-dimensional linear gating on the features of the current sampled image sequence may include: generating temporal attention weights on the features of the current sampled image sequence in the temporal dimension; performing spatial convolution on the features of the current sampled image sequence in the spatial dimension; and performing dot product between the temporal attention weights and the spatially convolved features to obtain the features after three-dimensional linear gating.

[0023] Optionally, the video is sampled based on the sampling window to obtain a preset number of frames.

[0024] According to an exemplary embodiment of the present disclosure, a video sampling apparatus is provided, comprising: a first sampling unit configured to sample a video based on a sampling window to obtain a current sampled image sequence; a parameter acquisition unit configured to acquire action parameters corresponding to the current sampled image sequence; a window adjustment unit configured to adjust the sampling window according to the action parameters corresponding to the current sampled image sequence, so as to sample the video based on the adjusted sampling window; and a second sampling unit configured to sample the video based on the adjusted sampling window.

[0025] Optionally, the action parameters may include: the probability that the current sampled image sequence contains an action and / or the degree of completion of the contained action.

[0026] Optionally, the parameter acquisition unit may include: a feature extraction unit configured to extract features from the current sampled image sequence to obtain the features of the current sampled image sequence; and a feature recognition unit configured to recognize the features of the current sampled image sequence to obtain the action parameters in the current sampled image sequence.

[0027] Optionally, the action parameters may include: the probability that the current sampled image sequence contains an action and the degree of completion of the contained action. The feature recognition unit is configured to: perform action recognition on the features of the current sampled image sequence to obtain the probability that the current sampled image sequence contains an action; and perform action completion recognition on the features of the current sampled image sequence to obtain the degree of completion of the contained action in the current sampled image sequence.

[0028] Optionally, the window adjustment unit can be configured to: calculate the size change value of the sampling window and / or the position of the sampling window movement based on the motion parameters in the current sampling image sequence; and adjust the sampling window based on the size change value of the sampling window and / or the position of the sampling window movement.

[0029] Optionally, the window adjustment unit can also be configured to: when it is determined from the action parameters that the action contained in the current sampled image sequence is incomplete, calculate the incremental value of the sampling window based on the action parameters; and determine the position to be moved by the sampling window based on the incremental value of the sampling window.

[0030] Optionally, the video sampling device may further include a smoothing and adjustment unit configured to smooth the incremental values ​​of the sampling window based on historical adjustment information of the sampling window.

[0031] Optionally, the window adjustment unit can also be configured to: determine the number of frames of the augmented image sampled in the augmented window corresponding to the incremental value of the sampling window based on the incremental value of the sampling window; and determine the position to be moved by the sampling window based on the determined number of frames.

[0032] Optionally, the second sampling unit can be configured to: sample the video in the augmentation window according to a determined number of frames to obtain an augmented image sequence; obtain the sampled image sequence corresponding to the adjusted sampling window in the current sampled image sequence; and use the sampled image sequence corresponding to the adjusted sampling window and the augmented image sequence as the sampled image sequence obtained based on the adjusted sampling window.

[0033] Optionally, the video sampling device may further include a determining unit configured to: select the action with the highest probability of being included in the current sampled image sequence from various actions as the action included in the current sampled image sequence; determine that the action included in the current sampled image sequence has been completed when the probability of the action included in the current sampled image sequence is greater than a first threshold and the completion degree of the action included in the current sampled image sequence is greater than a second threshold; and / or determine that the current sampled image sequence does not contain any action when the probability of the action included in the current sampled image sequence is less than a third threshold and the completion degree of the action included in the current sampled image sequence is less than a fourth threshold; and / or determine that the action included in the current sampled image sequence has not been completed when the probability of the action included in the current sampled image sequence is between the first threshold and the third threshold and the completion degree of the action included in the current sampled image sequence is between the second threshold and the fourth threshold.

[0034] Optionally, the window adjustment unit can also be configured to: calculate the decrement value of the sampling window based on the action parameters when it cannot be determined whether the action in the current sampled image sequence is completed based on the action parameters; and determine the position to be moved by the sampling window based on the decrement value of the sampling window.

[0035] Optionally, the window adjustment unit can also be configured to: subtract a window with a length equal to the reduction value of the sampling window from both ends of the sampling window to obtain a first sampling window and a second sampling window; use one of the first sampling window and the second sampling window as the updated sampling window; and determine the position to which the current sampling window moves based on the updated sampling window.

[0036] Optionally, the window adjustment unit can also be configured to: sample a preset number of frames of sampled image sequences from the first sampling window and the second sampling window respectively; extract features and perform feature recognition on the sampled image sequences sampled from the first sampling window and the sampled image sequences sampled from the second sampling window respectively; and select one of the first sampling window and the second sampling window as the updated sampling window based on the recognition result.

[0037] Optionally, the position to which the sampling window moves is the position to which the starting point of the sampling window moves.

[0038] Optionally, the feature recognition unit can also be configured to: perform three-dimensional linear gating processing on the features of the current sampled image sequence to obtain action recognition features; and based on the obtained action recognition features, obtain the probability that the current sampled image sequence contains an action.

[0039] Optionally, the feature recognition unit can also be configured to: perform three-dimensional linear gating processing on the features of the current sampled image sequence to obtain action completion recognition features; and based on the obtained action completion recognition features, obtain the completion degree of the actions contained in the current sampled image sequence.

[0040] Optionally, the feature recognition unit can also be configured to: generate temporal attention weights for the features of the current sampled image sequence in the time dimension; perform spatial convolution on the features of the current sampled image sequence in the spatial dimension; and perform dot product between the temporal attention weights and the spatially convolved features to obtain the features after three-dimensional linear gating.

[0041] Optionally, the first sampling unit and the second sampling unit respectively sample a preset number of frames of images.

[0042] According to exemplary embodiments of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements a video sampling method according to exemplary embodiments of the present disclosure.

[0043] According to an exemplary embodiment of the present disclosure, an electronic device is provided, including: a processor; and a memory storing a computer program, which, when executed by the processor, implements a video sampling method according to an exemplary embodiment of the present disclosure.

[0044] The video sampling method and apparatus according to exemplary embodiments of the present disclosure obtain a current sampled image sequence by sampling a video based on a sampling window; obtain action parameters corresponding to the current sampled image sequence; adjust the sampling window according to the action parameters corresponding to the current sampled image sequence; and sample the video based on the adjusted sampling window, thereby achieving high-precision recognition and positioning of actions, and thus achieving accuracy in video sampling.

[0045] Further aspects and / or advantages of the general concept of this disclosure will be set forth in part in the description which follows, and in part will be clear from the description or may be learned by practice of the general concept of this disclosure. Attached Figure Description

[0046] The above and other objects and features of exemplary embodiments of this disclosure will become clearer from the following description taken in conjunction with the accompanying drawings, which exemplarily illustrate the embodiments, wherein:

[0047] Figure 1a This diagram illustrates the process of recognizing actions in a video using existing techniques.

[0048] Figure 1b A schematic diagram illustrating the movement of a window in small steps according to the prior art is shown;

[0049] Figure 1c This diagram illustrates the process of obtaining the final recognition result by integrating the results of multiple windows based on existing technologies.

[0050] Figure 1d This diagram illustrates motion recognition in a video using a sliding window based on existing technology.

[0051] Figure 1e The results of an experiment using a video containing four actions, based on existing technology, are shown.

[0052] Figure 1f This diagram illustrates the recognition of actions in a video according to an exemplary embodiment of the present disclosure.

[0053] Figure 2 A flowchart illustrating a video sampling method according to an exemplary embodiment of the present disclosure is shown;

[0054] Figure 3 A schematic diagram illustrating a video sampling process according to an exemplary embodiment of the present disclosure;

[0055] Figure 4 A schematic diagram showing the internal structures of M2 and M3 according to exemplary embodiments of the present disclosure;

[0056] Figure 5A schematic diagram showing the internal structure of a 3D GLU according to an exemplary embodiment of the present disclosure;

[0057] Figure 6 A schematic diagram showing the internal structure of the portion of M4 used for incremental window adjustment according to an exemplary embodiment of the present disclosure;

[0058] Figure 7 A schematic diagram showing the internal structure of M5 according to an exemplary embodiment of the present disclosure;

[0059] Figure 8 A schematic diagram illustrating incremental window smoothing according to an exemplary embodiment of the present disclosure is shown;

[0060] Figure 9 A schematic diagram illustrating window adaptation according to an exemplary embodiment of the present disclosure is shown;

[0061] Figure 10 A schematic diagram showing the internal structure of the portion of M4 used for adjusting the reduction window according to an exemplary embodiment of the present disclosure;

[0062] Figure 11 A schematic diagram showing the internal structure of M4 according to an exemplary embodiment of the present disclosure;

[0063] Figure 12 A schematic diagram illustrating the process of calculating the reduction window according to an exemplary embodiment of the present disclosure;

[0064] Figure 13 This diagram illustrates how a window is adjusted for a truncated long action according to an exemplary embodiment of the present disclosure;

[0065] Figure 14 This diagram illustrates adjusting a window for a short action according to an exemplary embodiment of the present disclosure;

[0066] Figure 15 A block diagram showing a video sampling apparatus according to exemplary embodiments of the present disclosure; and

[0067] Figure 16 A schematic diagram of an electronic device according to an exemplary embodiment of the present disclosure is shown. Detailed Implementation

[0068] Exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings, examples of which are illustrated in the drawings, wherein the same reference numerals always refer to the same components. The embodiments will now be described with reference to the accompanying drawings in order to explain the present disclosure.

[0069] like Figure 1a As shown, to recognize actions in a video, the video must first be sampled. Figure 1aAction recognition algorithms sample certain frames from a video sequence as input. Existing techniques use a fixed-time window (sliding window) for sampling, sampling a fixed number of images (e.g., N frames) within that window for action recognition. The window is then moved forward at a fixed time step, and a new number of images are sampled from the corresponding video segment within the new window region for a new round of action recognition. Figure 1a In this context, deep neural networks can be used for action recognition. For example, for each N frames of images obtained from sampling, each frame is input into the deep neural network to obtain action classification results. Based on the action classification results corresponding to each of the N frames of images obtained from sampling, the following can be obtained: Figure 1a The action recognition results shown indicate that the action with the highest probability of being recognized can be obtained from the N frames of images sampled each time. For example... Figure 1a In the first N frames of images obtained by the sliding window, the action recognition result is "no action", the action recognition result is "action 1" for the second N frames of images obtained by the sliding window, the action recognition result is "action 2" for the third N frames of images obtained by the sliding window, and the action recognition result is "action 1" for the fourth N frames of images obtained by the sliding window.

[0070] Fixed sampling window size and step size cannot meet the requirements for high-precision recognition of actions of different lengths. This often results in short actions having a large amount of noise in the sampling window, and long actions being truncated into multiple windows, leading to low action recognition accuracy. For example... Figure 1a In the process, the actual action corresponding to the N frames of images obtained by the first sampling based on the sliding window is "Action 1", the actual action corresponding to the N frames of images obtained by the second sampling is "Action 2", the actual action corresponding to the N frames of images obtained by the third sampling is "Action 2", and the actual action corresponding to the N frames of images obtained by the fourth sampling is "Action 2".

[0071] Moving the sampling window with a very small step size can improve the performance of video segments that contain many other actions or no actions at all. These video segments, for the target action that needs to be identified, are called background noise. Figure 1b As shown, by moving the window in small steps, window 3 contains less background noise than window 0, and its action recognition performance is better.

[0072] However, this leads to a lot of redundant calculations. Because there is a lot of overlap between adjacent sampling windows, the same information needs to be processed multiple times to get an accurate recognition result, resulting in a lot of redundant calculations and failing to meet the needs of rapid recognition of behavior and actions.

[0073] To address the issue of long animations being truncated, existing technologies can combine the results from multiple windows for judgment, such as... Figure 1cAs shown. However, the problem with this method is that a single window can only receive information from a fragment, resulting in low accuracy. For a long action, a fragment of the action may be identified as another action, leading to poor accuracy when combining multiple erroneous results. Moreover, different actions have different durations, making it difficult to determine exactly how many windows' results to combine.

[0074] The existing solutions have two main problems:

[0075] (1) Fixed window size leads to inaccurate behavior and action recognition.

[0076] In the real world, the duration of different human behaviors varies greatly, and it is difficult to select the appropriate window size when sampling with a fixed window size.

[0077] If the window is too small, long-duration actions will be truncated into several windows, such as... Figure 1d Action 2 in the model cannot fully capture the entire process of the action's progression within each window, resulting in low recognition accuracy. For example, for a walking action that lasts for a long time, it is difficult to distinguish it from other stationary standing actions when segmented into a short window.

[0078] If the window is too large, short-duration actions will mix a lot of background and other actions within a single window, for example... Figure 1d Both sampling windows for action 1 contain a large amount of background, which can lead to low recognition accuracy. For example, in a crucial shooting motion in a sports match, the duration is very short, and if it is segmented into different windows, it will be identified as background motions such as running and jumping, making it difficult to accurately identify the shooting motion.

[0079] (2) The fixed window movement step size leads to inaccurate behavior recognition.

[0080] Moving the window with a fixed step size makes it difficult to ensure that the starting point of the window coincides with the starting point of the action. This introduces significant noise into the sampled data. Furthermore, the frequency of actions in the real world is unpredictable, making it difficult to determine the appropriate step size. A large step size results in more background noise within the window, impacting the accuracy of action recognition. Conversely, a small step size leads to significant overlap between sampled windows, causing data duplication and redundant computation for subsequent action recognition, hindering fast and effective recognition.

[0081] For action recognition, it is crucial that the sampling window accurately covers the entire action process. Adjusting the sampling window to cover the entire action results in high accuracy in action recognition and classification; conversely, adjusting the sampling window to not cover the entire action process results in low accuracy. Figure 1eAs shown, an experiment was conducted using a video containing four actions, with an action classification performed within a 1-second sampling window. Figure 1e Points marked with asterisks represent correctly identified windows, while points marked with solid circles represent incorrectly identified windows. Statistically, the accuracy rate was only 28.57%. However, if the sampling window was adjusted to cover the entire action process, the accuracy rate reached 100%. This demonstrates the importance of the sampling window accurately covering the entire action process for action recognition.

[0082] In exemplary embodiments of this disclosure, such as Figure 1f As shown, an AI model can be used to obtain the completion level of an ongoing action, and the size and position of the sampling window can be adjusted based on the completion level. For example, for action 1 or action 2, additional windows of different sizes can be dynamically added based on the completion level of action 1 or action 2 in the sampling window (corresponding to...). Figure 1f The sampling window (amplification 1 and amplification 2) automatically deletes the previous window according to the size of the sampling window, so that the adjusted sampling window includes the complete action 1 or action 2, thereby achieving high-precision action recognition and positioning by containing the complete action process in the final window.

[0083] Figure 2 A flowchart illustrating a video sampling method according to an exemplary embodiment of this disclosure is provided. (Refer to...) Figure 2 In step S201, the video is sampled based on the sampling window to obtain the current sampled image sequence. Here, the size of the sampling window can be adjusted according to the length of the action.

[0084] In step S202, the action parameters corresponding to the current sampled image sequence are obtained. Here, the action parameters may include: the probability that the current sampled image sequence contains an action and / or the completion degree of the contained action.

[0085] In an exemplary embodiment of this disclosure, when obtaining the action parameters corresponding to the current sampled image sequence, feature extraction can be performed on the current sampled image sequence first to obtain the features of the current sampled image sequence, and then feature recognition can be performed on the features of the current sampled image sequence to obtain the action parameters in the current sampled image sequence.

[0086] For example, features can be extracted from the current sampled image sequence using features such as histogram of Oriented Gradient (HOG), scale-invariant feature transform (SIFT), speed-up robust features (SURF), difference of Gaussian (DOG), local binary pattern (LBP), haar-like features (HAAR), and features extracted by deep neural networks, such as a combination of multi-frame image features extracted by a two-dimensional deep neural network or a feature extraction algorithm that directly extracts image sequences from a three-dimensional neural network.

[0087] In an exemplary embodiment of this disclosure, feature recognition is performed on the features of the current sampled image sequence to obtain action parameters in the current sampled image sequence. Here, action parameters refer to action-related feature parameters identified from image features, which may include, for example, action probability and action completion degree. Action probability refers to the probability that the current sampled image sequence contains an action, and action completion degree refers to the degree of completion of the action contained in the current sampled image sequence.

[0088] In an exemplary embodiment of this disclosure, the action parameter may include: the probability that the current sampled image sequence contains an action and the degree of completion of the contained action. When performing feature recognition on the features of the current sampled image sequence, action recognition may first be performed on the features of the current sampled image sequence to obtain the probability that the current sampled image sequence contains an action, and then the degree of completion of the action may be recognized on the features of the current sampled image sequence to obtain the degree of completion of the action contained in the current sampled image sequence.

[0089] In an exemplary embodiment of this disclosure, the action with the highest probability of being included in the current sampled image sequence can be selected from various actions as the action included in the current sampled image sequence; when the probability of the action included in the current sampled image sequence is greater than a first threshold and the completion degree of the action included in the current sampled image sequence is greater than a second threshold, it is determined that the action included in the current sampled image sequence has been completed; and / or, when the probability of the action included in the current sampled image sequence is less than a third threshold and the completion degree of the action included in the current sampled image sequence is less than a fourth threshold, it is determined that the current sampled image sequence does not contain any action; and / or, when the probability of the action included in the current sampled image sequence is between the first threshold and the third threshold and the completion degree of the action included in the current sampled image sequence is between the second threshold and the fourth threshold, it is determined that the action included in the current sampled image sequence has not been completed.

[0090] In an exemplary embodiment of this disclosure, when performing action recognition on the features of the current sampled image sequence, the features of the current sampled image sequence can first be subjected to three-dimensional linear gating processing to obtain action recognition features, and then the probability that the current sampled image sequence contains an action can be obtained based on the obtained action recognition features.

[0091] In an exemplary embodiment of this disclosure, when identifying the degree of action completion of the features of the current sampled image sequence, three-dimensional linear gating processing can be performed on the features of the current sampled image sequence first to obtain action completion recognition features, and then the degree of action completion contained in the current sampled image sequence can be obtained based on the obtained action completion recognition features.

[0092] In an exemplary embodiment of this disclosure, when performing three-dimensional linear gating processing on the features of the current sampled image sequence, a temporal attention weight can first be generated for the features of the current sampled image sequence in the time dimension, and a spatial convolution can be performed on the features of the current sampled image sequence in the spatial dimension. Then, the temporal attention weight and the spatially convolved features are multiplied by a dot to obtain the features after three-dimensional linear gating processing.

[0093] In step S203, the sampling window is adjusted according to the action parameters corresponding to the current sampled image sequence.

[0094] In an exemplary embodiment of this disclosure, when adjusting the sampling window, the size change value and / or the position moved by the sampling window can first be calculated based on the motion parameters in the current sampled image sequence. Then, the sampling window is adjusted based on the size change value and / or the position moved by the sampling window. Here, the size change value of the sampling window refers to the change in the window size for the next sampling relative to the current sampling window size. In an exemplary embodiment of this disclosure, the position moved by the sampling window can be the position moved from the starting point of the sampling window.

[0095] In an exemplary embodiment of this disclosure, when calculating the size change value of the sampling window and / or the position of the sampling window movement, the incremental value of the sampling window can be calculated first based on the action parameters when it is determined that the action contained in the current sampling image sequence is incomplete. Then, the position of the sampling window movement is determined based on the incremental value of the sampling window. In an exemplary embodiment of this disclosure, the incremental value of the sampling window can be smoothed based on the historical adjustment information of the sampling window.

[0096] In an exemplary embodiment of this disclosure, when determining the position to which the sampling window moves, the number of frames of the augmented image sampled in the augmented window corresponding to the incremental value of the sampling window can be determined based on the incremental value of the sampling window, and the position to which the sampling window moves can be determined based on the determined number of frames.

[0097] In an exemplary embodiment of this disclosure, when calculating the size change value of the sampling window and / or the position of the sampling window movement, if it cannot be determined whether the action in the current sampling image sequence is completed based on the action parameters, the decrease value of the sampling window can be calculated based on the action parameters, and the position of the sampling window movement can be determined based on the decrease value of the sampling window.

[0098] In an exemplary embodiment of this disclosure, when determining the position of the sampling window based on the decrement value of the sampling window, a window with a length equal to the decrement value of the sampling window can be subtracted from both ends of the sampling window to obtain a first sampling window and a second sampling window. Then, one of the first sampling window and the second sampling window is used as the updated sampling window, and the position of the current sampling window is determined based on the updated sampling window.

[0099] In an exemplary embodiment of this disclosure, when using one of the first sampling window and the second sampling window as the updated sampling window, a preset number of sampled image sequences can be sampled from both the first and second sampling windows respectively. Feature extraction and feature recognition are then performed on both the sampled image sequences sampled from the first and second sampling windows, and finally, based on the recognition results, one of the first and second sampling windows is selected as the updated sampling window. In step S204, the video is sampled based on the adjusted sampling window.

[0100] In an exemplary embodiment of this disclosure, when sampling a video based on an adjusted sampling window, the video can first be sampled in an augmentation window according to a determined number of frames to obtain an augmented image sequence. Then, the sampled image sequence corresponding to the adjusted sampling window in the current sampled image sequence is obtained. Finally, the sampled image sequence corresponding to the adjusted sampling window and the augmented image sequence are used as the sampled image sequence obtained based on the adjusted sampling window.

[0101] In an exemplary embodiment of this disclosure, a preset number of frame images are obtained by sampling the video based on a sampling window, and a preset number of frame images are also obtained by sampling the video based on an adjusted sampling window.

[0102] Figure 3 A schematic diagram illustrating the video sampling process according to an exemplary embodiment of the present disclosure.

[0103] Reference Figure 3 The video sampling process can be as follows:

[0104] 1) Initial sampling is performed using a window of size W. N frames of images are uniformly sampled within the initial window and fed into the deep learning feature extraction network M1 to obtain features representing the information in the current window.

[0105] 2) The feature is fed into M2 for action recognition to obtain the categories and probabilities of various actions, and then fed into M3 for action completion recognition to obtain the completion degree of the current action. For example, the degree of completion of a specific action can be represented by a percentage.

[0106] 3) Based on the action recognition result and the completion rate of the current action, M4 can calculate the window size to be increased for the next sampling. If the current action has been completed or the current window does not contain the action of interest, there is no need to expand the current window; simply output the result of the current action recognition and restart step 1. If the action in the current window is not completed, calculate the size of the window to be expanded.

[0107] 4) To ensure robustness and stability, the estimated amplification window value is further smoothed and adjusted in M5 to ensure stable growth of the window size. Based on the adjusted amplification window size, the position of the window's starting point is calculated.

[0108] 5) Based on the starting point of the adjusted window, perform uniform sampling in the new window, execute step 1), and further iterate to perform action recognition and action completion estimation.

[0109] In exemplary embodiments of this disclosure, M1, M2, M3, M4 and M5 may be units or modules implemented by software and / or hardware.

[0110] M1 is a feature extraction network that takes multiple frames of images as input and outputs features representing the input information. This feature extraction network can be implemented using a deep neural network.

[0111] M2 is used for action recognition. Its input is features, and its output window contains the probability of various actions.

[0112] M3 is used for action completion recognition. Its input is features, and its output is the degree of progress of the action within the window.

[0113] M4 is used for window growth decisions. Its inputs are the action recognition result and the action completion recognition result. Its outputs are whether to directly output the result in the next step, and whether to expand the window length in the next step if the current action is not completed.

[0114] M5 is used for sampling window adjustment. Its input is the length of the amplification window determined by M4, and its output is the next round of input image sequence sampled within the window after further adjustment.

[0115] Figure 4 A schematic diagram showing the internal structure of M2 and M3 according to an exemplary embodiment of the present disclosure is provided.

[0116] In exemplary embodiments of this disclosure, action recognition and action completion recognition can be performed by different independent algorithm networks, or a multi-task network structure design can be adopted. If a multi-task network structure is used, refer to... Figure 4 M2 and M3 use a common base network M1 to extract features. The advantage of this design is that it can greatly reduce the amount of computation in the network. Furthermore, since the development process of the same type of action is often similar, the two branches sharing the underlying feature extraction network can extract consistent features that can represent both the category and the degree of completion of the current action. The sharing of the underlying network structure and parameters between the two branches has practical significance.

[0117] Specifically, the size of the input feature can be C×T×H×W. Here, C is the number of feature channels, T is the length of the feature time axis, H is the height in the feature space, and W is the width in the feature space.

[0118] In an exemplary embodiment of this disclosure, in M2, the input features first pass through a 3D-gated linear unit (3D GLU) to become discriminative features for action recognition, and then pass through a pooling layer, a fully connected layer, and a softmax layer to output their probabilities across all action categories. In M3, the input features pass through a 3D GLU to become discriminative features for action completion recognition, and then pass through a pooling layer, a fully connected layer, and a softmax layer to output the action completion value.

[0119] In M2 and M3, the input to the 3D GLU is the same underlying network's input features, and the output is discriminative features tailored to different tasks. The function of 3D GLU is to generate discriminative output features based on the specific task. The discriminativeness of these features is reflected in the following two aspects:

[0120] i. For action recognition and action completion recognition, different levels of temporal attention are given to make the output features of the two branches discriminative. For action recognition, the recognition result of the same action with different completion levels should be the same; therefore, the 3D GLU branch for action recognition will pay more attention to the starting position of the action. For action completion recognition, the recognition result of the same action with the same starting point but different ending points will be different; therefore, the 3D GLU branch for action recognition will pay more attention to the ending position of the action.

[0121] ii. By using different convolutional kernel parameters on the features, task-related parameters are increased, making the features themselves more discriminative. In 3D GLU, the parameters of the convolutional kernels in the two branches naturally differ during training due to different tasks. Therefore, the output features of the same feature will be different after convolution. This is equivalent to deepening the network in each of the two branches, giving the features discriminative expressive power.

[0122] Figure 5 A schematic diagram showing the internal structure of a 3D GLU according to an exemplary embodiment of the present disclosure is provided.

[0123] Reference Figure 5 The 3D GLU is internally divided into two branches: the temporal gating branch and the feature convolution branch.

[0124] i. Time-domain gating support

[0125] Generate attention weights for different times in the time dimension.

[0126] The temporal gating branch consists of two layers: a convolutional layer W with a temporal dimension. G Its convolution kernel has the shape [K t [,1,1], here, K t This is the size parameter of the convolution kernel on the time axis. The convolution kernel in the time dimension performs convolution on the time axis T.

[0127] The next layer is a Sigmoid non-linear layer. After passing through the gated non-linear layer, attention weights are generated on the time axis.

[0128] ii. Feature convolution branch

[0129] This branch directly performs convolution on the input features in the spatial dimension, with the convolution kernel having a dimension of [1, K]. s ,K s ], here, K s This represents the size of the convolution kernel in space. This increases the computational depth of the network, giving the features richer expressiveness.

[0130] Both branches output features of the same dimension. Finally, the dot product operation is used to combine the outputs of the two branches, thereby enabling different weights to be applied to the features in the time dimension.

[0131] Figure 6 A schematic diagram showing the internal structure of the portion of M4 used for incremental window adjustment according to an exemplary embodiment of the present disclosure.

[0132] In an exemplary embodiment of this disclosure, M4 determines the next sampling strategy based on the result of current action recognition and the result of action completion. Its output has two modes: if there is no action in the current window or the action has ended, a new round of sampling begins with the initial window size; if the action in the current window has not ended, the size of the window that needs to be increased is determined based on the current window size. When the action completion is high, a larger window is added. When the action completion is low, a smaller window is added.

[0133] Specifically, refer to Figure 6 The specific calculation process in M4 is as follows:

[0134] The maximum value of the action recognition results across different action categories is taken to obtain the predicted action category and its probability. Combined with the action completion result, three cases are identified:

[0135] If the probability of behavior recognition is p class and action completion rate p finished If all values ​​are greater than the threshold Thres1, it means that the action within the current window has been well identified and the action has been basically completed, and a new round of initialization of fixed window size sampling begins.

[0136] If the probability of behavior recognition is p class and action completion rate p finished If all values ​​are less than the threshold Thres2, it means that the probability of the current window containing all action categories is very low, and the action completion rate is also low. This means that the current window may not contain any action, and a new round of initialization of fixed window size sampling begins.

[0137] If the probability of behavior recognition is p class and action completion rate p finished If the thresholds are between Thres2 and Thres1, then the current window is considered to contain an ongoing action, and the parameter α is calculated using the following formula:

[0138] (ω is an empirical parameter).

[0139] The size of the incremental window is: IW = α * W, where W is the size of the window corresponding to this sampling.

[0140] Figure 7 A schematic diagram showing the internal structure of M5 according to an exemplary embodiment of the present disclosure is provided.

[0141] In an exemplary embodiment of this disclosure, M5 serves to further adjust the size of the incremental window and determine the starting position of the new sampling window. Figure 7In M5, the process is divided into two sub-parts: the incremental window smoothing part and the window adaptation part. The following will refer to... Figure 8 and Figure 9 Each will be explained separately.

[0142] Figure 8 A schematic diagram illustrating incremental window smoothing according to an exemplary embodiment of the present disclosure is shown.

[0143] The purpose of incremental window smoothing is to smooth the size of the currently calculated incremental window based on the records of previous adjustments, ensuring that the growth of the incremental window is robust and stable in relation to changes in action completion. The inputs to incremental window smoothing are the incremental window size IW obtained from M4 and the action completion p within the current window. finished The output is the final incremental window size IW after smoothing. new .

[0144] The specific calculation process for incremental window smoothing is as follows:

[0145] i. The system stores the size (w1, w2) and corresponding completion degree (f1, f2) of the first two sampling windows;

[0146] ii. Linearly fit the relationship between window size and completion rate, based on the current window's action completion rate p. finished The size w of the window corresponding to the current action completion level is obtained through reasoning. predict ;

[0147] iii. The window size obtained by M4 is w caculate =w² + IW, the final adjusted window size is w final =0.5*(w predict +w caculate The corresponding incremental window size is IW. new =w final -w2.

[0148] Figure 9 A schematic diagram illustrating window adaptation according to an exemplary embodiment of the present disclosure is shown.

[0149] For action recognition networks, the size of the input data is fixed. Therefore, for windows of different sizes, the number of sampled image frames is a fixed number N, which can be adjusted according to different systems and different target actions.

[0150] The purpose of window adaptation is to sample N in the newly added window based on the size of the incremental window. add A new image of the frame, and in the N frames of the previously sampled image, N are removed from the header. add The images of each frame are then merged into a new sequence of N sampled images. addThe calculation formula is: N add =P*IW / ((w2+IW) / N). Here, P is an empirical coefficient and P>1, ensuring that relatively more information is added in the new window, IW is the size of the incremental window, w2 is the size of the previous sampling window, and N is the number of sampling frames. These frames are added while the images of these frames are deleted from the previous sampling frames, ensuring that the final input data for the new round is also N frames.

[0151] The incremental window approach primarily addresses the issue of truncated long actions. In practical applications, the initial window size is often small, so increasing the window length improves the accuracy reduction caused by truncated long actions. For very short actions, whose length may be less than the initial window length, this invention designs a reducible window adjustment scheme to address the recognition problem of short actions, thus avoiding simply reducing the initial window size to handle short actions. This design is more efficient.

[0152] Figure 10 A schematic diagram showing the internal structure of the portion of M4 used for reducing window adjustment according to an exemplary embodiment of the present disclosure.

[0153] In addition to the incremental window adjustment, a discrimination branch is added. When the parameter α is between (a, b), it indicates that the algorithm has difficulty identifying the type of action within the window and whether the action has been completed. This situation is often caused by background noise mixed in within the window. Therefore, in this case, the window size SW that needs to be subtracted is calculated. The formula for calculating SW is: SW = α * W / N. Here, W is the length of the current sampling window, and N is the number of image frames sampled in the current window.

[0154] Figure 11 A schematic diagram of the internal structure of M4 for incremental window adjustment and decremental window adjustment according to an exemplary embodiment of the present disclosure is shown. Figure 12 A schematic diagram illustrating the process of calculating a reduction window according to an exemplary embodiment of the present disclosure.

[0155] Reference Figure 11 and Figure 12Based on the size of the reduction window calculated by M4, windows of length SW are subtracted from the head and tail of the original sampling window, respectively, resulting in SSW1 and SSW2. N frames of images are then sampled within each of these two newly obtained windows for action recognition and action completion assessment. Finally, the window with the higher action recognition probability between SSW1 and SSW2 is selected. This design can handle short actions located anywhere within the window. If the action is closer to the start point of the window, SSW1 contains more action information, resulting in more accurate recognition; if the action is closer to the end point of the initial window, SSW2 will contain more action information, leading to more accurate recognition. Therefore, for short actions, there is no need to significantly reduce the size of the initial window; noise within the window can be effectively reduced, and short actions located at any position within the window can be pinpointed, achieving high-precision recognition and localization of short actions. Figure 12 In this case, the action is a short action relative to the window size, so the noise part in the window is subtracted so that the adjusted window includes the entire action and includes as little noise as possible.

[0156] Figure 13 This diagram illustrates how a window is adjusted for a truncated long action according to an exemplary embodiment of the present disclosure.

[0157] Reference Figure 13 For truncated long actions, the starting position of the window can be iteratively adjusted using the window expansion method described above, so that the final window contains the complete action process, achieving high-precision action recognition and localization. Figure 13 In this context, action 2 is a long action relative to the window size, so the sampling window is amplified so that the amplified sampling window includes the entire action 2.

[0158] Figure 14 This diagram illustrates adjusting a window for a short action according to an exemplary embodiment of the present disclosure.

[0159] Reference Figure 14 For short actions shorter than the initial window, there's no need to significantly reduce the initial window size. This effectively reduces noise within the window and locates the short action at any position within it. The adjusted window thus encompasses the entire action's occurrence, enabling high-precision recognition and localization of short actions. Figure 14 In this process, action 1 is a short action relative to the window size. Therefore, the noise portion in the window is subtracted so that the adjusted window includes the entire action 1 and contains as little noise as possible.

[0160] The above has been combined with Figure 1 to... Figure 14 A video sampling method according to exemplary embodiments of the present disclosure has been described. Hereinafter, reference will be made to... Figure 15A video sampling apparatus and its units according to exemplary embodiments of the present disclosure will be described.

[0161] Figure 15 A block diagram of a video sampling apparatus according to an exemplary embodiment of the present disclosure is shown.

[0162] Reference Figure 15 The video sampling device includes a first sampling unit 151, a parameter acquisition unit 152, a window adjustment unit 153, and a second sampling unit 154.

[0163] The first sampling unit 151 is configured to sample the video based on the sampling window to obtain the current sampled image sequence.

[0164] The parameter acquisition unit 152 is configured to acquire the action parameters corresponding to the current sampled image sequence.

[0165] In an exemplary embodiment of this disclosure, the action parameters may include: the probability that the current sampled image sequence contains an action and / or the degree of completion of the contained action.

[0166] In an exemplary embodiment of this disclosure, the parameter acquisition unit 152 may include: a feature extraction unit configured to extract features from the current sampled image sequence to obtain features of the current sampled image sequence; and a feature recognition unit configured to recognize features of the current sampled image sequence to obtain action parameters in the current sampled image sequence.

[0167] In an exemplary embodiment of this disclosure, the action parameters may include: the probability that the current sampled image sequence contains an action and the degree of completion of the contained action. The feature recognition unit may be configured to: perform action recognition on the features of the current sampled image sequence to obtain the probability that the current sampled image sequence contains an action; and perform action completion recognition on the features of the current sampled image sequence to obtain the degree of completion of the action contained in the current sampled image sequence.

[0168] In an exemplary embodiment of this disclosure, the video sampling device may further include a determining unit (not shown), configured to: select the action with the highest probability of being included in the current sampled image sequence from various actions as the action included in the current sampled image sequence; determine that the action included in the current sampled image sequence has been completed when the probability of the action included in the current sampled image sequence is greater than a first threshold and the completion degree of the action included in the current sampled image sequence is greater than a second threshold; and / or determine that the current sampled image sequence does not contain any action when the probability of the action included in the current sampled image sequence is less than a third threshold and the completion degree of the action included in the current sampled image sequence is less than a fourth threshold; and / or determine that the action included in the current sampled image sequence has not been completed when the probability of the action included in the current sampled image sequence is between the first threshold and the third threshold and the completion degree of the action included in the current sampled image sequence is between the second threshold and the fourth threshold.

[0169] The window adjustment unit 153 is configured to adjust the sampling window according to the motion parameters corresponding to the current sampled image sequence, so as to sample the video based on the adjusted sampling window.

[0170] In an exemplary embodiment of this disclosure, the window adjustment unit 153 may be configured to: calculate the size change value of the sampling window and / or the position of the sampling window movement based on the motion parameters in the current sampling image sequence; and adjust the sampling window based on the size change value of the sampling window and / or the position of the sampling window movement.

[0171] In an exemplary embodiment of this disclosure, the window adjustment unit 153 may also be configured to: calculate the incremental value of the sampling window based on the action parameters when it is determined that the action contained in the current sampled image sequence is incomplete; and determine the position to be moved by the sampling window based on the incremental value of the sampling window.

[0172] In an exemplary embodiment of this disclosure, the video sampling apparatus may further include a smoothing and adjustment unit (not shown), configured to smooth the incremental values ​​of the sampling window based on historical adjustment information of the sampling window.

[0173] In an exemplary embodiment of this disclosure, the window adjustment unit 153 may also be configured to: determine the number of frames of the augmented image sampled in the augmented window corresponding to the incremental value of the sampling window based on the incremental value of the sampling window; and determine the position to be moved by the sampling window based on the determined number of frames.

[0174] In an exemplary embodiment of this disclosure, the window adjustment unit 153 may also be configured to: calculate a reduction value of the sampling window based on the action parameters when it cannot be determined whether the action in the current sampled image sequence is completed based on the action parameters; and determine the position to which the sampling window is moved based on the reduction value of the sampling window.

[0175] In an exemplary embodiment of this disclosure, the window adjustment unit 153 may also be configured to: subtract a window with a length equal to the reduction value of the sampling window from both ends of the sampling window to obtain a first sampling window and a second sampling window; use one of the first sampling window and the second sampling window as the updated sampling window; and determine the position to which the current sampling window moves based on the updated sampling window.

[0176] In an exemplary embodiment of this disclosure, the window adjustment unit 153 may also be configured to: sample a preset number of frames of sampled image sequences from the first sampling window and the second sampling window respectively; extract features and perform feature recognition on the sampled image sequences sampled from the first sampling window and the sampled image sequences sampled from the second sampling window respectively; and select one of the first sampling window and the second sampling window as the updated sampling window based on the recognition result.

[0177] In an exemplary embodiment of this disclosure, the position to which the sampling window moves is the position to which the starting point of the sampling window moves.

[0178] In an exemplary embodiment of this disclosure, the feature recognition unit may also be configured to: perform three-dimensional linear gating processing on the features of the current sampled image sequence to obtain action recognition features; and based on the obtained action recognition features, obtain the probability that the current sampled image sequence contains an action.

[0179] In an exemplary embodiment of this disclosure, the feature recognition unit may also be configured to: perform three-dimensional linear gating processing on the features of the current sampled image sequence to obtain action completion recognition features; and obtain the completion degree of the action contained in the current sampled image sequence based on the obtained action completion recognition features.

[0180] In an exemplary embodiment of this disclosure, the feature recognition unit may also be configured to: generate temporal attention weights for the features of the current sampled image sequence in the time dimension; perform spatial convolution on the features of the current sampled image sequence in the spatial dimension; and perform dot product between the temporal attention weights and the spatially convolved features to obtain the features after three-dimensional linear gating.

[0181] The second sampling unit 154 is configured to sample the video based on the adjusted sampling window.

[0182] In an exemplary embodiment of this disclosure, the second sampling unit 154 may be configured to: sample the video in an amplification window according to a determined number of frames to obtain an amplified image sequence; obtain the sampled image sequence corresponding to the adjusted sampling window in the current sampled image sequence; and use the sampled image sequence corresponding to the adjusted sampling window and the amplified image sequence as the sampled image sequence obtained based on the adjusted sampling window.

[0183] In an exemplary embodiment of this disclosure, the first sampling unit and the second sampling unit respectively sample a preset number of frame images.

[0184] Furthermore, according to exemplary embodiments of the present disclosure, a computer-readable storage medium is also provided, on which a computer program is stored, which, when executed, implements the video sampling method according to exemplary embodiments of the present disclosure.

[0185] In an exemplary embodiment of this disclosure, the computer-readable storage medium may carry one or more programs that, when executed, perform the following steps: sampling the video with the current sampling window to obtain a current sampled image sequence of a preset number of frames; extracting features from the current sampled image sequence to obtain features of the current sampled image sequence; performing feature recognition on the features of the current sampled image sequence to obtain action parameters in the current sampled image sequence; calculating the window size change value and the position of the window starting point movement for the next sampling based on the action parameters in the current sampled image sequence; and using the next sampling window adjusted based on the window size change value and the position of the window starting point movement to perform the next sampling of the video.

[0186] Computer-readable storage media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In embodiments of this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a computer program that can be used by or in conjunction with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable storage medium can be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof. A computer-readable storage medium can be included in any apparatus; it can also exist independently without being assembled into that apparatus.

[0187] The above has been combined Figure 15 A video sampling apparatus according to exemplary embodiments of the present disclosure has been described. Next, in conjunction with… Figure 16 An electronic device according to exemplary embodiments of the present disclosure will be described.

[0188] Figure 16A schematic diagram of an electronic device according to an exemplary embodiment of the present disclosure is shown.

[0189] Reference Figure 16 An electronic device 16 according to an exemplary embodiment of the present disclosure includes a memory 161 and a processor 162. The memory 161 stores a computer program that, when executed by the processor 162, implements a video sampling method according to an exemplary embodiment of the present disclosure.

[0190] In an exemplary embodiment of this disclosure, when the computer program is executed by the processor 162, the following steps can be implemented: sampling the video with the current sampling window to obtain a current sampling image sequence of a preset number of frames; extracting features from the current sampling image sequence to obtain features of the current sampling image sequence; performing feature recognition on the features of the current sampling image sequence to obtain motion parameters in the current sampling image sequence; calculating the window size change value and the position of the window starting point movement for the next sampling based on the motion parameters in the current sampling image sequence; and using the next sampling window adjusted based on the window size change value and the position of the window starting point movement to perform the next sampling of the video.

[0191] The electronic devices in this disclosure may include, but are not limited to, devices such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), and desktop computers. Figure 16 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0192] The above references are shown in Figures 1 to 12. Figure 16 A video sampling method and apparatus according to exemplary embodiments of the present disclosure have been described. However, it should be understood that: Figure 15 The video sampling device and its units shown can be configured as software, hardware, firmware, or any combination thereof to perform specific functions. Figure 16 The electronic device shown is not limited to the components shown above, but some components may be added or removed as needed, and the above components may also be combined.

[0193] The video sampling method and apparatus according to exemplary embodiments of the present disclosure obtain a current sampled image sequence by sampling a video based on a sampling window; obtain action parameters corresponding to the current sampled image sequence; adjust the sampling window according to the action parameters corresponding to the current sampled image sequence; and sample the video based on the adjusted sampling window, thereby achieving high-precision recognition and positioning of actions, and thus realizing the accuracy of video sampling.

[0194] Although this disclosure has been specifically shown and described with reference to exemplary embodiments thereof, those skilled in the art should understand that various changes in form and detail may be made therein without departing from the spirit and scope of this disclosure as defined by the claims.

Claims

1. A video sampling method, comprising: The video is sampled based on the sampling window to obtain the current sampled image sequence; Obtain the action parameters corresponding to the current sampled image sequence; Select the action with the highest probability of being included in the current sampled image sequence from among all actions; Based on a comparison between the probability of an action being contained in the current sampled image sequence and at least one threshold, and a comparison between the completion rate of the action contained in the current sampled image sequence and at least one threshold, it is determined whether the action contained in the current sampled image sequence has been completed, and whether the current sampled image sequence contains an action. Based on the determination that the action contained in the current sampled image sequence is incomplete, the sampling window is adjusted according to the action parameters corresponding to the current sampled image sequence; The video is sampled based on the adjusted sampling window.

2. The video sampling method according to claim 1, wherein, The action parameters include: the probability that the current sampled image sequence contains an action and / or the completion rate of the action contained in the current sampled image sequence.

3. The video sampling method according to claim 1, wherein, The steps to obtain the action parameters corresponding to the current sampled image sequence include: Feature extraction is performed on the current sampled image sequence to obtain the features of the current sampled image sequence; Feature recognition is performed on the features of the current sampled image sequence to obtain the action parameters in the current sampled image sequence.

4. The video sampling method according to claim 3, wherein, The steps for feature recognition of the current sampled image sequence include: Perform action recognition on the features of the current sampled image sequence to obtain the probability that the current sampled image sequence contains an action; The degree of action completion is identified by analyzing the features of the current sampled image sequence to obtain the degree of action completion contained in the current sampled image sequence.

5. The video sampling method according to claim 1, wherein, The steps for adjusting the sampling window include: Calculate the size change of the sampling window and / or the position of the sampling window based on the motion parameters in the current sampled image sequence; The sampling window is adjusted based on the change in the size of the sampling window and / or the position of the sampling window movement.

6. The video sampling method according to claim 5, wherein, The steps for calculating the change in the size of the sampling window and / or the position moved by the sampling window include: The incremental value of the sampling window is calculated based on the action parameters; The position to which the sampling window moves is determined based on the incremental value of the sampling window.

7. The video sampling method according to claim 6, further comprising: Based on the historical adjustment information of the sampling window, the incremental values ​​of the sampling window are smoothed.

8. The video sampling method according to claim 6, wherein, The steps to determine the position of the sampling window include: The number of frames of the amplified image sampled in the amplification window corresponding to the increment value of the sampling window is determined based on the increment value of the sampling window. The position to move the sampling window is determined based on the number of frames.

9. The video sampling method according to claim 8, wherein, The steps for sampling video based on the adjusted sampling window include: The video is sampled within the augmentation window according to a determined number of frames to obtain an augmented image sequence; Obtain the sampled image sequence corresponding to the adjusted sampling window in the current sampled image sequence; The sampled image sequence corresponding to the adjusted sampling window and the amplified image sequence are used as the sampled image sequence obtained based on the adjusted sampling window.

10. The video sampling method according to claim 1, the step of determining whether an action in the current sampled image sequence has been completed and whether the current sampled image sequence contains an action based on a comparison between the probability of an action being contained in the current sampled image sequence and at least one threshold, and a comparison between the completion degree of the action contained in the current sampled image sequence and at least one threshold, includes: Based on the probability that the current sampled image sequence contains an action greater than a first threshold, and the completion rate of the action contained in the current sampled image sequence greater than a second threshold, it is determined that the action contained in the current sampled image sequence has been completed; and / or Based on the probability that the current sampled image sequence contains an action being less than a third threshold, and the completion rate of the action contained in the current sampled image sequence being less than a fourth threshold, it is determined that the current sampled image sequence does not contain an action; and / or Based on the probability that the current sampled image sequence contains an action between a first threshold and a third threshold, and the completion rate of the action contained in the current sampled image sequence between a second threshold and a fourth threshold, it is determined that the action contained in the current sampled image sequence is incomplete.

11. The video sampling method according to claim 5, wherein, The steps for calculating the change in the size of the sampling window and / or the position moved by the sampling window include: When it is impossible to determine whether an action has been completed in the current sampled image sequence based on the action parameters, the decrement value of the sampling window is calculated based on the action parameters. The position to move the sampling window is determined based on the decrease value of the sampling window.

12. The video sampling method according to claim 11, wherein, The steps for determining the position to move the sampling window based on the decrease value of the sampling window include: Subtract a window of length equal to the reduction value of the sampling window from both ends of the sampling window to obtain the first sampling window and the second sampling window; Use one of the first and second sampling windows as the updated sampling window; The position to move the current sampling window is determined based on the updated sampling window.

13. The video sampling method according to claim 12, wherein, The steps of using one of the first and second sampling windows as the updated sampling window include: A sequence of sampled images of a preset number of frames is sampled from the first sampling window and the second sampling window respectively, wherein each of the first sampling window and the second sampling window includes a preset number of sampled images; Feature extraction and feature recognition are performed on the sampled image sequences sampled from the first sampling window and the sampled image sequences sampled from the second sampling window, respectively. Based on the recognition results, one of the first and second sampling windows is selected as the updated sampling window.

14. The video sampling method according to claim 5, wherein, The position to which the sampling window moves is the position to which the starting point of the sampling window moves.

15. The video sampling method according to claim 4, wherein, The steps for action recognition based on the features of the current sampled image sequence include: Three-dimensional linear gating is applied to the features of the current sampled image sequence to obtain action recognition features; Based on the obtained action recognition features, the probability that the current sampled image sequence contains an action is obtained.

16. The video sampling method according to claim 4, wherein, The steps for identifying the degree of action completion based on the features of the current sampled image sequence include: Three-dimensional linear gating processing is applied to the features of the current sampled image sequence to obtain action completion recognition features; Based on the obtained action completion recognition features, the completion level of the actions contained in the current sampled image sequence is obtained.

17. The video sampling method according to claim 15 or 16, wherein, The steps for performing three-dimensional linear gating on the features of the current sampled image sequence include: Generate temporal attention weights for the features of the current sampled image sequence in the temporal dimension; Spatial convolution is performed on the features of the current sampled image sequence in the spatial dimension; Multiply the temporal attention weights by the spatially convolved features to obtain the features after three-dimensional linear gating.

18. The video sampling method according to claim 1, wherein, The video is sampled using a sampling window to obtain a preset number of frames.

19. The video sampling method according to claim 1, wherein, The steps for adjusting the sampling window include: Calculate the increment value of the sampling window based on the motion parameters in the current sampled image sequence; Based on the historical adjustment information of the sampling window, the incremental values ​​of the sampling window are smoothed. The position of the sampling window is determined based on the smoothed incremental value.

20. The video sampling method according to claim 1, wherein, The step of adjusting the sampling window based on the motion parameters corresponding to the current sampled image sequence includes: Based on the determination that the action contained in the current sampled image sequence is incomplete, the incremental value of the sampling window is calculated according to the action parameters; The position to which the sampling window moves is determined based on the incremental value of the sampling window.

21. A video sampling device, comprising: The first sampling unit is configured to sample the video based on the sampling window to obtain the current sampled image sequence; The parameter acquisition unit is configured to acquire the action parameters corresponding to the current sampled image sequence; The determining unit is configured to select the action with the highest probability of being included in the current sampled image sequence from each action as the action included in the current sampled image sequence, and determine whether the action included in the current sampled image sequence has been completed and whether the current sampled image sequence contains an action based on a comparison between the probability of the action included in the current sampled image sequence and at least one threshold, and a comparison between the completion degree of the action included in the current sampled image sequence and at least one threshold. The window adjustment unit is configured to adjust the sampling window based on the action parameters corresponding to the current sampling image sequence, based on the determination that the action contained in the current sampling image sequence is incomplete, so as to sample the video based on the adjusted sampling window; and The second sampling unit is configured to sample the video based on the adjusted sampling window.

22. A computer-readable storage medium storing a computer program, wherein, When the computer program is executed by a processor, it implements the video sampling method according to any one of claims 1 to 20.

23. An electronic device, comprising: A memory, a processor, and a computer program stored on the memory, characterized in that the processor executes the computer program to implement the steps of the method according to any one of claims 1-20.

Citation Information

Patent Citations

  • Method used for identifying object motion directions and equipment

    CN105225248A

  • Video signal sampling device

    JP1991145391A