Sound production action time sequence positioning method and system based on inflection point flow

Through deep learning methods based on inflection point flow, combined with optical flow estimation algorithm and self-supervised learning, cross-kinematic features in video are extracted, and the problem of insufficient accuracy of vocal movement timing positioning in the existing technology is solved, and high-precision vocal movement frame positioning and automated dubbing support are realized.

CN120047866AActive Publication Date: 2025-05-27SOUTH CHINA UNIV OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510067040.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-27
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

The existing timing action positioning methods are difficult to capture specific moments within the action, especially the high-precision timing positioning of the vocalization action, resulting in insufficient automatic dubbing quality and accuracy.

Method used

The vocal action timing positioning method based on inflection point flow is adopted. Through deep learning technology, combined with optical flow estimation calculation method and self-supervised learning, image context features, bidirectional motion features and bidirectional velocity inflection point features are extracted, cross-kinematic aggregation features are calculated, and cross-video comparison learning and internal smooth constraints of video are carried out to realize frame-level vocal action frame positioning.

Benefits of technology

It realizes high-precision vocalization timing positioning, reduces dependence on manual operations, is suitable for video content of various action types, improves work efficiency, and ensures the consistency and authenticity of audio and visual elements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047866A_ABST
    Figure CN120047866A_ABST
Patent Text Reader

Abstract

The invention discloses a sound production action time sequence positioning method and system based on inflection point flow. The method comprises the following steps: firstly, reading a silent video containing a sound production action; carrying out inflection point flow analysis according to a kinematic law to obtain a bidirectional motion flow and a bidirectional inflection point flow; performing feature extraction to obtain image context features, bidirectional motion features and bidirectional speed inflection point features; calculating and splicing cross-kinematics features from the image to the motion and from the image to the inflection point to obtain a cross-kinematics aggregation feature; meanwhile, a discrimination graph is obtained from the cross-kinematics aggregation features, cross-video comparative learning is carried out on the motion area and the non-motion area on the discrimination graph, video internal smooth constraint is carried out, and activated motion area features are obtained; and finally, performing spatial-temporal feature fusion and frame-by-frame classification prediction, and identifying whether each image frame is a sounding action frame or not. According to the invention, frame-level collision sounding action frame positioning is realized, and the video frame corresponding to the collision sounding action can be accurately identified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision and video processing, and particularly relates to a method and system for temporal positioning of sound-producing actions based on inflection point flow. Background Art

[0002] With the booming development of social media platforms, video content has become an indispensable part of daily life, especially on popular video sharing platforms. For visual media such as movies and TV dramas, dubbing is one of the key steps in the production process. For dubbing sound-producing actions such as the collision of musical instruments or objects, this process relies on professionals to determine the precise moment for audio insertion, which is particularly complex and challenging.

[0003] Technologies specifically designed for video clip dubbing applications, especially for precisely locating the temporal positions of sound-producing actions in videos, are still in the initial exploration stage. The most recent research is on temporal action localization methods, which automatically identify specific actions occurring in videos through algorithms and accurately mark the temporal boundaries of these actions; such technologies aim to confirm the category of action instances and determine their start and end time points in the video stream. However, existing temporal action localization methods mainly focus on detecting the overall process of a single action, such as the entire cycle of a kicking action, with the focus mainly on the semantic aspects of the overall action rather than its specific kinematic characteristics. This emphasis has led to the neglect of internal changes in the action, which are crucial for identifying and differentiating the sound-producing moments of actions. In addition, these methods do not consider explicitly learning the spatial motion of objects and fail to capture the spatio-temporal dynamics of specific motion frames that are crucial for detecting audible actions. In contrast, the problem of temporal positioning of sound-producing actions is more refined, which requires the algorithm to be able to capture specific instants within the action cycle, such as the exact moment when the foot touches the ball in a kicking action. This high-precision temporal positioning ability ensures the quality and accuracy of automated dubbing, making the video content more vivid and realistic, and also providing assistance for other post-production tasks and applications that require high-precision temporal positioning.

[0004] In summary, the present invention aims to fill an important gap in existing research and technologies, that is, to provide an efficient and reliable solution to achieve high-precision temporal positioning of sound-producing actions and support the automatic dubbing of video clips. Summary of the Invention

[0005] The main objective of the present invention is to overcome the disadvantages and deficiencies of the prior art, and provide a method and system for temporal positioning of sound-producing actions based on inflection point flow. Through deep learning technology, frame-level localization of sound-producing action frames during collisions is achieved, significantly reducing the dependence on manual operations, being applicable to video content of various action types, and improving work efficiency.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] On the one hand, the present invention provides a method for timing positioning of vocalization actions based on inflection point flow, including the following steps:

[0008] Read a silent video containing vocalization actions;

[0009] Regard the silent video as an object displacement map changing with time, and perform inflection point flow analysis on the silent video according to the laws of kinematics to obtain a bidirectional motion flow and a bidirectional inflection point flow;

[0010] Extract features from the image frames, bidirectional motion flow, and bidirectional inflection point flow in the silent video respectively to obtain image context features, bidirectional motion features, and bidirectional velocity inflection point features;

[0011] Calculate cross-kinematic features from image to motion and image to inflection point based on the image context features, bidirectional motion features, and bidirectional velocity inflection point features, and splice them to obtain cross-kinematic aggregation features;

[0012] Obtain a discriminant map from the cross-kinematic aggregation features, perform cross-video contrast learning on the motion region and non-motion region on the discriminant map and perform intra-video smoothing constraints to obtain activated motion region features; the discriminant map contains a motion region and a non-motion region;

[0013] Encode the cross-kinematic aggregation features and the activated motion region features respectively, then perform spatio-temporal feature fusion and frame-by-frame classification prediction to identify whether each image frame is a vocalization action frame. As a preferred technical solution, the inflection point flow analysis of the silent video using the optical flow estimation algorithm according to the laws of kinematics is specifically:

[0014] Regard the silent video as an object displacement map changing with time, use the optical flow to represent the velocity, and use the optical flow difference to represent the acceleration;

[0015] Use the optical flow estimation algorithm to calculate the forward optical flow of an image frame and its next adjacent image frame as its corresponding forward motion flow; calculate the backward optical flow of the image frame and its previous adjacent image frame as its corresponding backward motion flow;

[0016] Perform a difference operation on the forward motion flow of the image frame and the forward motion flow of its next adjacent image frame to obtain its corresponding forward inflection point flow, and perform a difference operation on the backward motion flow of the image frame and the backward motion flow of its previous adjacent image frame to obtain its corresponding backward inflection point flow;

[0017] Splice the forward motion flow and the backward motion flow of the image frame to obtain its bidirectional motion flow, and splice the forward inflection point flow and the backward inflection point flow of the image frame to obtain its bidirectional inflection point flow.

[0018] As a preferred technical solution, the feature extraction uses three independent ResNet encoder networks to extract the context features of each image frame and its bidirectional motion flow and bidirectional inflection point flow respectively. After feature splicing, the image context features, bidirectional motion features, and bidirectional speed inflection point features of each image frame are obtained.

[0019] As a preferred technical solution, the splicing obtains the cross-kinematic aggregation features, specifically:

[0020] By projecting the image context features, bidirectional motion features, and bidirectional speed inflection point features of each image frame into a linear layer, the attention matrix from image to motion and the attention matrix from image to inflection point are calculated;

[0021] After passing through the Softmax activation layer respectively, the cross-kinematic features from image to motion and from image to inflection point are obtained;

[0022] The image context features, cross-kinematic features from image to motion, and cross-kinematic features from image to inflection point of each image frame are feature-connected to obtain the cross-kinematic aggregation features of each image frame.

[0023] As a preferred technical solution, the cross-kinematic aggregation features pass through a 3×3 convolutional layer, a batch normalization layer, and a Sigmoid operation to obtain the corresponding discriminant map. The high-probability regions in the discriminant map are motion regions, and the low-probability regions are non-motion regions.

[0024] As a preferred technical solution, cross-video contrast learning and video-internal smoothing constraints are performed on the motion regions and non-motion regions in the discriminant map, specifically:

[0025] The motion regions and non-motion regions in the discriminant map are used to mask the cross-kinematic aggregation features to obtain motion region features and non-motion region features;

[0026] Based on the motion region features and non-motion region features, positive contrast target pairs and negative contrast target pairs are constructed;

[0027] According to the positive contrast target pairs and negative contrast target pairs, a cross-video contrast learning loss function is constructed and the cross-kinematic features are optimized in a self-supervised learning manner to obtain activated motion region features;

[0028] The cross-video contrast learning loss function is expressed as:

[0029]

[0030] Among them, is the negative contrast target loss function, is the positive contrast target loss function where the positive contrast target pair is the motion region feature, is the positive contrast target loss function for the positive contrast target pair being the non-motion region feature;

[0031] The internal video smoothing constraint optimizes by constructing a video temporal regularization loss function based on the discriminant maps of the image frames in the same silent video, expressed as:

[0032]

[0033] where D p+2 is the discriminant map of the (p + 2)-th image frame, and D p+1 is the discriminant map of the (p + 1)-th image frame.

[0034] As a preferred technical solution, the negative contrast target loss function is constructed based on the motion region features and non-motion region features of the image frames sampled from multiple silent videos:

[0035]

[0036] where k is the number of image frames sampled from multiple silent videos, is the inner product, expressed as:

[0037]

[0038] is the motion region feature of the p-th image frame, is the non-motion region feature of the q-th image frame;

[0039] The positive contrast target loss function is obtained by constructing based on the motion region features or non-motion region features of the image frames sampled from multiple silent videos:

[0040]

[0041] where, is the positive contrast target loss function for the positive contrast target pair being the motion region feature or non-motion region feature, is the motion region feature or non-motion region feature of the p-th image frame, is the motion region feature or non-motion region feature of the q-th image frame; is 1 when p = q, otherwise 0; w p,q is the penalty factor, expressed as:

[0042]

[0043] α is the hyperparameter for smoothness control, represents the inner product in the ranking of the cosine similarity pairs of the motion region features or non-motion region features of all image frames.

[0044] As a preferred technical solution, 3D convolution is used to encode the cross-kinematic aggregation features and the activated motion region features respectively, and then the encoded output is input into a Transformer network of a self-attention module for spatio-temporal feature fusion to obtain spatio-temporal fusion features;

[0045] Then, through a fully connected layer, each image frame is classified and predicted frame by frame to identify whether each image frame is an action-occurring frame, and an identification result is obtained.

[0046] As a preferred technical solution, when performing classification and prediction frame by frame through the fully connected layer, cross-entropy loss and focal loss are used for supervision:

[0047]

[0048] where λ ce and λ focal are respectively the weight parameters of the cross-entropy loss and the focal loss ; the cross-entropy loss is calculated using a soft label technology based on Gaussian enhancement.

[0049] On the other hand, a vocalization action timing positioning system based on inflection point flow is provided, which is applied to the above-mentioned vocalization action timing positioning method based on inflection point flow, and includes a video reading module, an inflection point flow analysis module, a feature extraction module, a feature aggregation module, a self-supervised optimization module, and a classification prediction module;

[0050] The video reading module is used to read a silent video containing vocalization actions;

[0051] The inflection point flow analysis module is used to regard the silent video as an object displacement map changing with time, and perform inflection point flow analysis on the silent video according to the kinematic law to obtain a bidirectional motion flow and a bidirectional inflection point flow;

[0052] The feature extraction module is used to extract features from the image frames, the bidirectional motion flow, and the bidirectional inflection point flow in the silent video respectively to obtain image context features, bidirectional motion features, and bidirectional velocity inflection point features;

[0053] The feature aggregation module is used to calculate cross-kinematic features from image to motion and from image to inflection point based on the image context features, bidirectional motion features, and bidirectional velocity inflection point features, and splice them to obtain cross-kinematic aggregation features;

[0054] The self-supervised optimization module is used to obtain a discriminant map from the cross-kinematic features, perform cross-video contrast learning on the motion region and the non-motion region on the discriminant map, and perform video internal smoothing constraints to obtain activated motion region features; the discriminant map contains a motion region and a non-motion region;

[0055] The classification prediction module is used to encode the cross-kinematic features and the activated motion area features respectively, then perform spatio-temporal feature fusion and frame-by-frame classification prediction to identify whether each image frame is a vocalization action frame.

[0056] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0057] 1. According to the kinematic law, the present invention performs inflection point flow analysis by combining the second derivative of the position-time image with the optical flow estimation algorithm, aiming to enrich the kinematic data by capturing the details of the object state change, so as to simulate the collision vocalization action.

[0058] 2. The present invention proposes an auxiliary optimization method based on self-supervised spatial localization features. By performing cross-video contrast learning on the motion area and non-motion area in the discriminant map, and constructing an internal video temporal regularization loss function for internal video smoothing constraint, the obtained spatial information is used to enhance the representation ability of the network, and the spatial position where the collision action occurs is additionally identified as a secondary output.

[0059] 3. The present invention realizes frame-level collision vocalization action frame localization, can accurately identify the video frames corresponding to the collision vocalization action, ensures the consistency and authenticity of the audio-visual elements, and provides technical support for subsequent dubbing.

[0060] 4. Through deep learning technology, the present invention greatly reduces the dependence on manual operations, can be applied to video content of various action types, and improves work efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0062] Figure 1 It is the overall flowchart of the vocalization action temporal localization method based on inflection point flow in the embodiment of the present invention.

[0063] Figure 2 It is the flow framework diagram of the vocalization action temporal localization method based on inflection point flow in the embodiment of the present invention.

[0064] Figure 3 It is the structural schematic diagram of the vocalization action temporal localization system based on inflection point flow in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0065] To enable those skilled in the art to better understand the solution of this application, the following will clearly and completely describe the technical solution in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of this application.

[0066] In this application, the mention of "embodiment" means that the specific features, structures or characteristics described in combination with the embodiment may be included in at least one embodiment of this application. The appearance of this phrase at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described in this application can be combined with other embodiments.

[0067] The goal of the present invention is to determine the moment when the vocalization action occurs, and the goal is to accurately determine the moment when the vocalization action occurs at the frame level, that is, to judge whether each frame contains an action that can produce sound, which helps to achieve lip-sync. Vocalization actions, especially collisions, are usually caused by sudden changes in the force acting on an object, which are manifested as inflection points in velocity; in addition, these collisions are the main source of sound. Based on this assumption, the present invention deeply studies the velocity inflection points and motion characteristics from the perspectives of time and space. However, the previous action recognition has not fully explored and reflected the potential of velocity inflection points from motion. Therefore, the present invention proposes to introduce an inflection point flow to supplement the traditional kinematic analysis (i.e., only including the motion flow). The present invention proposes an inflection point flow to represent the sudden change in the velocity of an object, to enhance the traditional kinematic analysis by serving as prior motion information, and at the same time combines feature extraction and feature aggregation, aiming to extract context features guided by motion. And this application also introduces a self-supervised auxiliary optimization strategy, which constrains the prediction of vocalization actions through the combination of internal and external perspectives of the video. The specific steps and methods of the present invention are as Figure 1 、 Figure 2 shown, including the following steps:

[0068] S1. Read the silent video I = {I t ∈R H×W×3 |t = 1, 2,..., T}, where T is the number of image frames, and H and W are the height and width of the image frame.

[0069] S2. Inflection point flow analysis: Regard the silent video as a map of object displacement changing with time, and perform inflection point flow analysis on the silent video according to the kinematic law to obtain a bidirectional motion flow and a bidirectional inflection point flow.

[0070] The inflection point flow proposed in this application is used to represent the sudden change in the velocity of an object, and it enhances traditional kinematic analysis by serving as prior information on motion. The principle is as follows:

[0071] The silent video is regarded as a map of object displacement changing over time, that is, a two-dimensional displacement-time graph (i.e., an x-t graph, where x is a two-dimensional displacement vector). According to the laws of kinematics, by performing a first-order derivative operation on the time series of the object's displacement, the corresponding velocity-time graph (v-t graph) can be obtained:

[0072]

[0073] In addition, by performing a second-order derivative operation on the displacement-time graph, the acceleration-time graph (a-t graph) is obtained:

[0074]

[0075] According to Newton's first law, the sudden application of an external force will cause a sudden change in acceleration, thereby changing the motion state of the object. The acceleration function a(t) provides an important clue for identifying vocalization actions; therefore, the velocity inflection point represented by the acceleration a(t) is called the "inflection point flow". Then, the inflection point flow a t : = a(t) and the motion flow v t : = v(t) are combined as a new type of kinematic prior. Thus, in practical applications, according to the laws of kinematics, inflection point flow analysis is performed on the silent video, specifically as follows:

[0076] S201. Regard the silent video as a map of object displacement changing over time, use optical flow to represent velocity, and use optical flow difference to represent acceleration.

[0077] S202. Use an optical flow estimation algorithm to calculate the forward optical flow of an image frame and its adjacent subsequent image frame as its corresponding forward motion flow; calculate the backward optical flow of an image frame and its adjacent previous image frame as its corresponding backward motion flow.

[0078] S203. Perform a difference operation on the forward motion flow of an image frame and the forward motion flow of its adjacent subsequent image frame to obtain its corresponding forward inflection point flow, and perform a difference operation on the backward motion flow of an image frame and the backward motion flow of its adjacent previous image frame to obtain its corresponding backward inflection point flow;

[0079] S204. Concatenate the forward motion flow and the backward motion flow of an image frame to obtain its bidirectional motion flow, and concatenate the forward inflection point flow and the backward inflection point flow of an image frame to obtain its bidirectional inflection point flow.

[0080] For example, for three adjacent image frames I i-1 、I i 、I i+1 in the silent video, calculate the image frame Ii The forward optical flow with the subsequent adjacent image frame I i+1 is used as the forward motion flow of the image frame I i Calculate I i The backward optical flow with the previous adjacent image frame I i-1 is used as its backward motion flow Then, the forward motion flow of the image frame I i is differenced with the forward motion flow of the subsequent adjacent image frame I i+1 to obtain the forward inflection point flow of the image frame I i Similarly, the backward motion flow of the image frame I i is differenced with the backward motion flow of the previous adjacent image frame I i-1 to obtain the backward inflection point flow of the image frame I i

[0081] S3. Feature extraction: Feature extraction is respectively performed on the image frames, bidirectional motion flows, and bidirectional inflection point flows in the silent video to obtain image context features, bidirectional motion features, and bidirectional speed inflection point features.

[0082] Specifically, the present invention uses three independent ResNet encoder networks (E X , E M and E C ) to respectively extract the context features of each image frame and its bidirectional motion flow and bidirectional inflection point flow. After feature concatenation, the image context feature f i , bidirectional motion feature m i and bidirectional speed inflection point feature c i of each image frame are obtained.

[0083] S4. Cross-kinematics aggregation: Cross-kinematic features from image to motion and image to inflection point are calculated based on the image context features, bidirectional motion features, and bidirectional speed inflection point features, and the cross-kinematic aggregation features are obtained by concatenation.

[0084] To better utilize the guidance of kinematic priors, the present invention introduces a cross-kinematics aggregation strategy, aiming to extract motion and speed inflection point information from the image context features; specifically:

[0085] S401. First, by projecting the image context features, bidirectional motion features, and bidirectional speed inflection point features of each image frame into a linear layer, the attention matrices from image to motion and image to inflection point are calculated, and the calculation formulas are as follows:

[0086] ​​​​​​​

[0087] Among them, A mf(i) is the attention matrix from image to motion, Q m(i) is the bidirectional motion feature query vector, K f(i) is the image context feature key vector, T is the transpose operation, and A cf(i) is the attention matrix from image to inflection point, Q c(i) is the bidirectional speed inflection point feature query vector, and d is the dimension of the query vector and the key vector.

[0088] S402. After passing through the Softmax activation layer respectively, the cross-kinematic features from image to motion and from image to inflection point are obtained:

[0089] h m(i) = V f(i) softmax(A mf(i) ), h c(i) = V f(i) softmax(A cf(i) ),

[0090] where h m(i) and h c(i) respectively represent the cross-kinematic features from image to motion and from image to inflection point, and V f(i) is obtained by applying a linear layer to the image context feature f i .

[0091] S403. Finally, the image context feature, the cross-kinematic feature from image to motion, and the cross-kinematic feature from image to inflection point of each image frame are concatenated to obtain the cross-kinematic aggregation feature F i of each image frame.

[0092] S5. Self-supervised spatial auxiliary optimization: Obtain a discriminative map from the cross-kinematic features, perform cross-video contrast learning on the motion area and the non-motion area on the discriminative map, and perform intra-video smoothing constraints to obtain the activated motion area features; among them, the discriminative map contains the motion area and the non-motion area.

[0093] In previous temporal action localization methods, the spatial motion of the learning object was not explicitly considered. Therefore, the present invention introduces spatial motion localization as an auxiliary optimization task. This strategy can effectively improve the model's ability to identify the most discriminative areas of the object's motion, thereby better locating potential vocal action frames. First, a discriminative map needs to be obtained; the present invention obtains the corresponding discriminative map by passing the cross-kinematic aggregation feature through a 3×3 convolutional layer, a batch normalization layer, and a Sigmoid operation. The high-probability area in the discriminative map is the motion area, and the low-probability area is the non-motion area.

[0094] Next, cross-video contrast learning and intra-video smoothing constraints are performed, that is, cross-video contrast learning and intra-video smoothing constraints are carried out on the discriminant graph for the motion region (high-probability region) and the non-motion region (low-probability region). Specifically:

[0095] S501. Use the motion region and non-motion region in the discriminant graph to mask the cross-kinematic aggregation features, obtaining the motion region features and non-motion region features, expressed as:

[0096]

[0097] Among them, represents the Hadamard product, D p is the discriminant graph of the p-th image frame, F p is the cross-kinematic aggregation feature of the p-th image frame, is the motion region feature of the p-th image frame, is the non-motion region feature of the p-th image frame.

[0098] S502. Construct positive contrast target pairs and negative contrast target pairs based on the motion region features and non-motion region features. Among them, the negative contrast target pairs are constructed based on the motion region features and non-motion region features of the image frames sampled from multiple silent videos, and the positive contrast target pairs are constructed based on the motion region features or non-motion region features of the image frames sampled from multiple silent videos.

[0099] S503. According to the positive contrast target pairs and negative contrast target pairs, construct a cross-video contrast learning loss function and optimize the cross-kinematic features in a self-supervised learning manner to obtain the activated motion region features. Among them, the cross-video contrast learning loss function is expressed as:

[0100]

[0101] Among them, is the negative contrast target loss function, is the positive contrast target loss function when the positive contrast target pair is the motion region feature, is the positive contrast target loss function when the positive contrast target pair is the non-motion region feature.

[0102] Specifically, the negative contrast target loss function is constructed based on the negative contrast target pairs:

[0103]

[0104] Among them, k is the number of image frames sampled from the read silent videos, is the inner product, expressed as:

[0105]

[0106] is the motion region feature of the p-th image frame, is the non-motion region feature of the q-th image frame;

[0107] And the positive contrast target loss function is constructed based on the positive contrast target pair:

[0108]

[0109] Wherein, is the positive contrast target loss function when the positive contrast target pair is the motion region feature or the non-motion region feature, is the motion region feature or the non-motion region feature of the p-th image frame, is the motion region feature or the non-motion region feature of the q-th image frame; is 1 when p = q, otherwise 0; w p,q is the penalty factor, which is used to penalize the positive contrast target pairs with lower similarity in the ranking, and is expressed as:

[0110]

[0111] α is the hyperparameter for smoothness control, represents the inner product represents the ranking in the cosine similarity pairs of the motion region features or the non-motion region features of all image frames.

[0112] S504. Based on the fact that the actions in the video are usually continuous, an internal video smoothing constraint is performed to ensure that there is no significant change in the localization region between adjacent frames, thereby ensuring the localization accuracy. Specifically: According to the discriminant map of the image frames in the same silent video, an internal video temporal regularization loss function is constructed for optimization, and is expressed as:

[0113]

[0114] Wherein, D p+2 is the discriminant map of the (p + 2)-th image frame, D p+1 is the discriminant map of the (p + 1)-th image frame.

[0115] Through the above self-supervised auxiliary optimization strategy, the spatial localization result is fed back to the vocalization action prediction part and trained in a self-supervised learning manner, and the vocalization action frames can be assisted in detection by identifying the action region.

[0116] S6. Spatiotemporal feature fusion and prediction: Encode the cross-kinematic aggregation feature and the activated motion region feature respectively, then perform spatiotemporal feature fusion and classify and predict frame by frame to identify whether each image frame is a vocalization action frame.

[0117] After obtaining the cross-kinematic aggregation features and the activated motion region features, 3D convolution is used to encode the temporal and spatial motion information, encoding the cross-kinematic aggregation features and the activated motion region features respectively; in order to extract more detailed information about the vocalization actions, the output of the 3D convolution is input into a Transformer network of a self-attention module for spatio-temporal feature fusion, obtaining spatio-temporal fusion features, and effectively locating the potential instances of vocalization actions in the video.

[0118] Finally, a fully connected layer is used to classify and predict each image frame frame by frame, identifying whether each image frame is a frame of vocalization action, and obtaining the recognition result. When performing classification prediction, cross-entropy loss and focal loss are used for supervision:

[0119]

[0120] where λ ce and λ focal are the weight parameters of the cross-entropy loss and the focal loss respectively; at the same time, the cross-entropy loss is calculated using the soft label technology based on Gaussian enhancement. In this embodiment, λ ce = 1, λ focal = 0.1.

[0121] Therefore, the total loss function of this method is obtained by weighted summation of the above three loss functions:

[0122]

[0123] where λ action , λ cont and λ temp are the weight parameters of the loss terms respectively.

[0124] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously.

[0125] Based on the same idea as the vocalization action timing localization method based on inflection point flow in the above embodiment, the present invention also provides a vocalization action timing localization system based on inflection point flow, and this system can be used to execute the above vocalization action timing localization method based on inflection point flow. For the sake of convenience of description, in the structural schematic diagram of the vocalization action timing localization system embodiment based on inflection point flow, only the parts related to the embodiment of the present invention are shown, and those skilled in the art can understand that the illustrated structure does not constitute a limitation on the device, and it can include more or fewer components than those illustrated, or combine certain components, or different component arrangements.

[0126] As Figure 3 shown in the figure, another embodiment of the present invention provides a temporal positioning system for vocalization actions based on inflection point flow, including a video reading module, an inflection point flow analysis module, a feature extraction module, a feature aggregation module, a self-supervised optimization module, and a classification prediction module;

[0127] Among them, the video reading module is used to read a silent video containing vocalization actions;

[0128] The inflection point flow analysis module is used to regard the silent video as an object displacement map changing with time, perform inflection point flow analysis on the silent video according to the kinematic law, and obtain a bidirectional motion flow and a bidirectional inflection point flow;

[0129] The feature extraction module is used to extract features from the image frames, bidirectional motion flow, and bidirectional inflection point flow in the silent video respectively, and obtain image context features, bidirectional motion features, and bidirectional velocity inflection point features;

[0130] The feature aggregation module is used to calculate cross-kinematic features from image to motion and image to inflection point based on the image context features, bidirectional motion features, and bidirectional velocity inflection point features, and splice them to obtain cross-kinematic aggregated features;

[0131] The self-supervised optimization module is used to obtain a discriminant map from the cross-kinematic features, perform cross-video contrast learning on the motion area and non-motion area on the discriminant map and perform intra-video smoothing constraints to obtain activated motion area features; the discriminant map contains a motion area and a non-motion area;

[0132] The classification prediction module is used to encode the cross-kinematic features and the activated motion area features respectively, then perform spatio-temporal feature fusion and perform frame-by-frame classification prediction to identify whether each image frame is a vocalization action frame.

[0133] It should be noted that the temporal positioning system for vocalization actions based on inflection point flow of the present invention corresponds one-to-one with the temporal positioning method for vocalization actions based on inflection point flow of the present invention. The technical features and their beneficial effects described in the embodiments of the above temporal positioning method for vocalization actions based on inflection point flow are applicable to the embodiments of the temporal positioning system for vocalization actions based on inflection point flow. For specific content, reference can be made to the description in the method embodiments of the present invention, which will not be elaborated here. This is hereby declared.

[0134] In addition, in the implementation of the vocalization action timing positioning system based on inflection point flow in the above embodiments, the logical division of each program module is only an example. In actual applications, according to needs, for example, considering the configuration requirements of the corresponding hardware or the convenience of software implementation, the above functions can be assigned to different program modules to complete, that is, the internal structure of the vocalization action timing positioning system based on inflection point flow is divided into different program modules to complete all or part of the functions described above.

[0135] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0136] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A method for locating the timing sequence of vocalization actions based on inflection point flow, characterized in that: The steps include: Read silent videos containing sound-generating actions; The silent video is regarded as a displacement map of an object changing with time, and the inflection point flow analysis is performed on the silent video according to the kinematic law to obtain a bidirectional motion flow and a bidirectional inflection point flow. Feature extraction is performed on image frames, bidirectional motion flows and bidirectional inflection point flows in silent videos respectively to obtain image context features, bidirectional motion features and bidirectional speed inflection point features; Based on the image context features, bidirectional motion features and bidirectional velocity inflection point features, the image-to-motion and image-to-inflection point cross-kinematic features are calculated and spliced ​​to obtain the cross-kinematic aggregate features; Obtaining a discriminant map from the cross-kinematic aggregated features, performing cross-video comparative learning on the motion area and the non-motion area on the discriminant map and performing internal smoothing constraints on the video to obtain activated motion area features; the discriminant map contains the motion area and the non-motion area; The cross-kinematic aggregation features and activated motion area features are encoded separately, and then the spatiotemporal features are fused and classified and predicted frame by frame to identify whether each image frame is a sound action frame.

2. According to claim 1, the method for locating the timing of vocalization actions based on inflection point flow is characterized in that: The inflection point flow analysis of the silent video is performed using an optical flow estimation algorithm according to the kinematic law, specifically: The silent video is considered as a displacement map of an object changing over time, the optical flow is used to represent the velocity, and the optical flow difference is used to represent the acceleration; The optical flow estimation algorithm is used to calculate the forward optical flow between the image frame and its subsequent adjacent image frame as the corresponding forward motion flow; the reverse optical flow between the image frame and its previous adjacent image frame is calculated as the corresponding reverse motion flow; Performing a differential operation on the forward motion flow of the image frame and the forward motion flow of the next adjacent image frame to obtain the corresponding forward inflection point flow, and performing a differential operation on the reverse motion flow of the image frame and the reverse motion flow of the previous adjacent image frame to obtain the corresponding reverse inflection point flow; The forward motion flow and the reverse motion flow of the image frame are spliced ​​to obtain its bidirectional motion flow, and the forward inflection point flow and the reverse inflection point flow of the image frame are spliced ​​to obtain its bidirectional inflection point flow.

3. The method for locating the timing of vocalization actions based on inflection point flow according to claim 1, characterized in that: The feature extraction uses three independent ResNet encoder networks to respectively extract the context features of each image frame and its bidirectional motion flow and bidirectional inflection point flow. After feature splicing, image context features, bidirectional motion features and bidirectional speed inflection point features of each image frame are obtained.

4. According to claim 1, the method for locating the timing sequence of vocalization actions based on inflection point flow is characterized in that: The splicing obtains cross-kinematic aggregate features, specifically: The image-to-motion attention matrix and the image-to-inflection-point attention matrix are calculated by projecting the image context features, bidirectional motion features, and bidirectional velocity inflection-point features of each image frame to the linear layer; After passing through the Softmax activation layer, the cross-kinematic features from image to motion and image to inflection point are obtained respectively; The image context features, the image-to-motion cross-kinematic features, and the image-to-inflection point cross-kinematic features of each image frame are feature-connected to obtain the cross-kinematic aggregate features of each image frame.

5. According to claim 1, the method for locating the timing sequence of vocalization actions based on inflection point flow is characterized in that: The cross-kinematics aggregation feature is passed through a 3×3 convolution layer, a batch normalization layer and a Sigmoid operation to obtain a corresponding discriminant map, in which the high-probability area is the motion area, and the low-probability area is the non-motion area.

6. The method for locating the timing sequence of vocalization actions based on inflection point flow according to claim 1, characterized in that: The cross-video contrast learning of the motion area and the non-motion area on the discriminant map and the internal smoothness constraint of the video are specifically as follows: Use the motion area and non-motion area in the discriminant map to mask the cross-kinematic aggregation features to obtain motion area features and non-motion area features; Construct positive contrast target pairs and negative contrast target pairs based on motion region features and non-motion region features; According to the positive contrast target pairs and the negative contrast target pairs, a cross-video contrast learning loss function is constructed and the cross-kinematic features are optimized by self-supervised learning to obtain the activated motion area features. The cross-video contrastive learning loss function is expressed as: in, is the negative contrast objective loss function, is the positive contrast target loss function for the motion region feature, is the positive contrast target loss function for the positive contrast target to the non-motion area feature; The video internal smoothness constraint is optimized by constructing a video timing regularization loss function based on the discriminant graph of the image frame in the same silent video, which is expressed as: Among them, D p+2 is the discriminant map of the p+2th image frame, D p+1 is the discriminant map of the p+1th image frame.

7. The method for locating the timing sequence of vocalization actions based on inflection point flow according to claim 6, characterized in that: The negative contrast target loss function is constructed based on the motion region features and non-motion region features of image frames sampled from multiple silent videos: Where k is the number of image frames sampled from multiple silent videos, is the inner product, expressed as: is the motion region feature of the p-th image frame, is the non-motion area feature of the qth image frame; The positive contrast target loss function is constructed based on the motion region features or non-motion region features of image frames sampled from multiple silent videos: in, is the positive contrast target loss function for the positive contrast target pair, which is the motion region feature or the non-motion region feature. is the motion region feature or non-motion region feature of the p-th image frame, is the motion region feature or non-motion region feature of the qth image frame; When p=q, it is 1, otherwise it is 0; w p,q is the penalty factor, expressed as: α is a hyperparameter for smoothness control, Inner product Ranking among the cosine similarity pairs of motion region features or non-motion region features in all image frames.

8. The method for locating the timing sequence of vocalization actions based on inflection point flow according to claim 1, characterized in that: 3D convolution is used to encode cross-kinematic aggregation features and activated motion region features respectively, and then the encoded output is input into a Transformer network of a self-attention module for spatiotemporal feature fusion to obtain spatiotemporal fusion features; Then, a fully connected layer is used to classify and predict each image frame frame by frame to identify whether each image frame is an action frame and obtain the recognition result.

9. The method for locating the timing of an action based on an inflection point flow according to claim 8, characterized in that: When performing classification prediction frame by frame through the fully connected layer, cross entropy loss and focal loss are used for supervision: Among them, λ ce and λ focal They are cross entropy losses and focal loss The cross entropy loss is calculated using a soft label technique based on Gaussian enhancement.

10. The vocalization action timing positioning system based on inflection point flow is characterized by: The method for locating the timing of vocalization actions based on inflection point flow as described in any one of claims 1 to 9 comprises a video reading module, an inflection point flow analysis module, a feature extraction module, a feature aggregation module, a self-supervised optimization module and a classification prediction module; The video reading module is used to read the silent video containing the sound-generating action; The inflection point flow analysis module is used to regard the silent video as a displacement map of an object that changes over time, and perform inflection point flow analysis on the silent video according to the law of kinematics to obtain a bidirectional motion flow and a bidirectional inflection point flow; The feature extraction module is used to extract features from image frames, bidirectional motion streams and bidirectional inflection point streams in silent videos respectively, to obtain image context features, bidirectional motion features and bidirectional speed inflection point features; The feature aggregation module is used to calculate the cross-kinematic features from image to motion and from image to inflection point based on the image context features, the bidirectional motion features and the bidirectional velocity inflection point features, and to splice the cross-kinematic aggregation features; The self-supervised optimization module is used to obtain a discriminant map from cross-kinematic features, perform cross-video comparative learning on the motion area and the non-motion area on the discriminant map and perform internal smoothing constraints on the video to obtain activated motion area features; the discriminant map contains motion areas and non-motion areas; The classification prediction module is used to encode cross-kinematic features and activated motion area features respectively, and then perform spatiotemporal feature fusion and classification prediction frame by frame to identify whether each image frame is a sound action frame.

Citation Information

Patent Citations

  • Action recognition method based on double-flow convolution attention

    CN112926396A

  • Sound effect generation method and device based on visual semantics

    CN114399984A

  • AI assisted sound effect generation for silent video

    CN115428469A

  • Audio importing method and device and electronic equipment

    CN115665356A

  • Motion analysis apparatus, motion analysis method, and program

    JP2021174048A