Camera motion detection method, device and equipment and computer readable storage medium

By training the camera motion detection model for video frame feature extraction and loss optimization, the shortcomings of camera motion detection in video are solved, efficient and accurate camera motion recognition is achieved, and the effects of video generation and creation analysis are improved.

CN120510558AActive Publication Date: 2025-08-19SHANDONG HAILIANG INFORMATION TECH RES INST
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511005705.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-08-19
Estimated Expiration
2045-07-22

AI Technical Summary

Technical Problem

The lack of effective detection schemes for camera movement in videos in the prior art, resulting in insufficient video generation quality and difficulty for creators to systematically analyze and apply camera movement techniques.

Method used

By using video samples annotated with camera motion categories, the camera motion detection model is trained, and feature extraction of continuous video frames and cross-video frames is performed, combined with loss optimization, efficient and accurate identification of camera motion categories is achieved.

Benefits of technology

It realizes efficient and accurate identification of camera motion categories in video, which helps improve the quality of video generation and creators' analysis capabilities, and assists in the performance improvement of downstream tasks such as motion trajectory analysis of complex objects and video mirror analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510558A_ABST
    Figure CN120510558A_ABST
Patent Text Reader

Abstract

The invention discloses a camera motion detection method, a camera motion detection device, camera motion detection equipment and a computer readable storage medium, and relates to the technical field of video processing. The camera motion detection model is used for carrying out continuous video frame feature extraction on a video sample to obtain a first video feature, cross-video frame feature extraction is carried out on the video sample to obtain a second video feature, and local and global dynamic features of camera motion in a video can be effectively captured through the two feature extraction modes. Therefore, the dynamic process of camera motion is described more comprehensively, classification calculation is carried out according to extracted video features to obtain a first camera motion category detection result, loss optimization is carried out on a camera motion detection model, and then camera motion category detection is carried out on an input video by using the trained camera motion detection model. Therefore, efficient and accurate identification of the camera motion type in the video can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video processing technology, and in particular to a camera motion detection method, apparatus, device, and computer-readable storage medium. Background Art

[0002] Video is an important information carrier. The camera motion during the shooting process contains rich semantic information, which is crucial for understanding video content, improving video generation quality, and assisting video creation and analysis. Currently, there is a lack of camera motion detection solutions in video processing technology. Summary of the Invention

[0003] The present invention provides a camera motion detection method, apparatus, device and computer-readable storage medium to at least solve the problem in the related art of lacking a solution for detecting camera motion in a video.

[0004] The present invention provides a camera motion detection method, comprising: Obtain video samples with camera motion category annotations; Using a camera motion detection model, extracting features of continuous video frames of the video sample to obtain a first video feature, extracting features across video frames of the video sample to obtain a second video feature, and performing classification calculation based on the first video feature and the second video feature to obtain a first camera motion category detection result; Calculating a loss value based on the first camera motion category detection result and the corresponding camera motion category label, and using the loss value to update model parameters of the camera motion detection model until an iterative training end condition is met, thereby obtaining the trained camera motion detection model; The trained camera motion detection model is used to calculate a second camera motion category detection result based on the input video.

[0005] The present invention also provides a camera motion detection device, comprising: An acquisition module is used to obtain video samples with camera motion category annotations; A model training module is configured to use a camera motion detection model to perform feature extraction of continuous video frames on the video sample to obtain a first video feature, perform feature extraction across video frames on the video sample to obtain a second video feature, perform classification calculation based on the first video feature and the second video feature to obtain a first camera motion category detection result; calculate a loss value based on the first camera motion category detection result and the corresponding camera motion category label, and use the loss value to update model parameters of the camera motion detection model until an iterative training end condition is met, thereby obtaining the trained camera motion detection model; The monitoring module is configured to calculate a second camera motion category detection result based on the input video using the trained camera motion detection model.

[0006] The present invention also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned camera motion detection methods when executing the computer program.

[0007] The present invention also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned camera motion detection methods are implemented.

[0008] Through the present invention, a camera motion detection model is trained by using video samples with camera motion category annotations, and the camera motion detection model is used to extract features of continuous video frames of the video samples to obtain a first video feature, and feature extraction is performed on the video samples across video frames to obtain a second video feature. The two feature extraction methods can effectively capture the local and global dynamic characteristics of the camera motion in the video, thereby more comprehensively characterizing the dynamic process of the camera motion. Classification calculation is performed based on the extracted video features to obtain the first camera motion category detection result and the camera motion detection model is loss optimized. Then, the trained camera motion detection model is used to detect the camera motion category of the input video, thereby achieving efficient and accurate recognition of the camera motion category in the video, which is helpful to improve the performance of multiple downstream tasks such as auxiliary video generation, complex object motion trajectory analysis, and video camera movement analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0010] Figure 1 A flowchart of a camera motion detection method provided by an embodiment of the present invention; Figure 2 A schematic structural diagram of a camera motion detection system provided by an embodiment of the present invention; Figure 3 A schematic structural diagram of a camera motion detection model provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0011] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0012] It should be noted that, in the description of the present invention, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. The terms "first," "second," etc., in the present invention are used to distinguish similar objects, and are not used to describe a particular order or precedence.

[0013] In order to enable those skilled in the art to better understand the solutions of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0014] Here, some key terms used in the embodiments of the present invention are explained.

[0015] Accurately detecting and analyzing camera motion in videos is crucial for understanding video content and improving video processing techniques.

[0016] For example, one of the goals of video generation is to produce realistic and expressive dynamic videos. However, related technologies lack research on camera motion in videos, resulting in a lack of three-dimensional consistency in the generated results or being limited to static scenes, making it difficult to generate dynamic scenes with complex object motion. Furthermore, generating consistent multi-view videos of the same scene from different camera trajectories is also extremely challenging. Due to the lack of large-scale, wild-field multi-view video datasets, current multi-view video generation is primarily limited to nearly static scenes or synthetic objects.

[0017] One type of video generation solution in related technologies learns a 3D generative model from large-scale unlabeled video data and uses visual conditioning techniques to control camera orientation to generate multi-view images. However, these approaches only consider 3D model generation and viewpoint control, lacking information about camera motion, resulting in poor video generation performance. Therefore, effectively detecting and analyzing camera motion types can provide important guidance for video generation models, such as for training data augmentation or as a control signal for the generation model, thereby improving the realism and expressiveness of generated videos.

[0018] Furthermore, for video and film creators, understanding and analyzing camera motion is crucial for improving the quality of their work. Excellent filmmakers skillfully utilize camera motion to direct the viewer's attention, enhance narrative impact, and define visual style. However, there is currently a lack of effective tools to help creators analyze and quantify camera motion in videos, making it difficult for creators to systematically learn and apply camera motion techniques.

[0019] Therefore, it is necessary to provide a solution for detecting camera motion in videos.

[0020] To this end, embodiments of the present invention provide a camera motion detection method, apparatus, device, and computer-readable storage medium. A camera motion detection model is trained by using video samples with camera motion category annotations. The camera motion detection model is used to extract features of continuous video frames of the video samples to obtain a first video feature. Feature extraction is performed across video frames of the video samples to obtain a second video feature. The two feature extraction methods can effectively capture the local and global dynamic characteristics of camera motion in the video, thereby more comprehensively characterizing the dynamic process of camera motion. Classification and calculation are performed based on the extracted video features to obtain a first camera motion category detection result and the camera motion detection model is loss optimized. The trained camera motion detection model is then used to detect the camera motion category of the input video, thereby achieving efficient and accurate recognition of the camera motion category in the video, which helps to improve the performance of multiple downstream tasks such as auxiliary video generation, complex object motion trajectory analysis, and video camera movement analysis.

[0021] An embodiment of the present invention provides a camera motion detection method. The method is described in detail below in conjunction with the execution flow of the camera motion detection method.

[0022] Figure 1 A flowchart of a camera motion detection method provided by an embodiment of the present invention; Figure 2 A schematic structural diagram of a camera motion detection system provided by an embodiment of the present invention.

[0023] like Figure 1 As shown, the camera motion detection method provided by the embodiment of the present invention includes: S101: obtaining a video sample with a camera motion category label.

[0024] S102: Using a camera motion detection model, extract features of continuous video frames of the video sample to obtain a first video feature, extract features across video frames of the video sample to obtain a second video feature, and perform classification calculation based on the first video feature and the second video feature to obtain a first camera motion category detection result.

[0025] S103: Calculate a loss value based on the first camera motion category detection result and the corresponding camera motion category label, and use the loss value to update the model parameters of the camera motion detection model until the iterative training end condition is met to obtain a trained camera motion detection model.

[0026] S104: Calculate a second camera motion category detection result based on the input video using the trained camera motion detection model.

[0027] In order to realize the camera motion detection of the video, the camera motion detection method provided by the embodiment of the present invention mainly includes two steps: training the camera detection model and using the trained camera detection model to perform the camera motion detection task. In this regard, the camera motion detection method provided by the embodiment of the present invention can be applied to the following: Figure 2 The camera motion detection system shown includes a model training system and a camera motion detection device.

[0028] like Figure 2 As shown, the model training system can be provided with computing power and storage support by multiple artificial intelligence servers (artificial intelligence servers 1~n), and the camera motion detection model can be trained based on the model training system using video samples. In some optional implementations of the embodiments of the present invention, a computing resource pool can be constructed based on the model training system, and the computing resource pool includes multiple computing threads, so that the training task of the camera motion detection model can be split into multiple subtasks for parallel execution. The model parallel training method can be adopted, and the data parallel training method can also be adopted. The model training system can be deployed in a cloud computing center. The equipment structure of the cloud computing center can be similar to Figure 2 The camera motion detection device shown here includes basic infrastructure such as an AI processor, storage, input / output devices, a communication bus, and a communication interface. It also comprises a software environment including an operating system, database, and middleware, as well as application software. The model training system receives manually annotated video samples and can fine-tune a camera motion detection model using a stored pre-trained model or train a randomly initialized model framework.

[0029] The camera motion detection device can be a computing device or terminal device deployed in a cloud computing center, and its structure is as follows: Figure 2As shown in the camera motion detection device in . If a computing device in a cloud computing center is used, then after deploying the trained camera motion detection model, the terminal device uploads the input video to be detected to the cloud computing center, and uses the trained camera motion detection model to detect the camera motion category of the input video to be detected, and outputs a second camera motion category detection result. If the camera motion detection device uses a terminal device, the model training system trains the visual camera motion detection model and sends it to the terminal device via the communication unit and deploys it in the storage of the terminal device. The terminal device receives the input video to be detected, performs camera motion category detection using the local camera motion detection model, and outputs a second camera motion category detection result.

[0030] In an embodiment of the present invention, by pre-labeling the camera motion categories of video samples and using the video samples with the camera motion category labels to train the camera motion detection model, the camera motion detection model can learn the relationship between the video features in the video samples and the camera motion categories.

[0031] For S101, multiple camera motion categories can be classified according to the camera motion mode. In some optional implementations of the embodiments of the present invention, the camera motion categories of the video sample may include: camera view rotation around the target object, camera view rotation and sweeping, camera view movement and tracking toward the target object, camera translation, camera static or micro-motion, and camera rotation.

[0032] The target object is a reference object in the video. When classifying the camera motion categories, the motion mode of the target object may be considered or not.

[0033] The method for determining whether the camera view rotates around the target object can be: the camera moves relatively stably around the target object (facing the target), the target object does not move or moves slightly, and the camera angle changes by more than about 30 degrees (visually obvious rotation); the camera does not move or moves slightly, the target object rigidly rotates, and the rotation angle of the target object is greater than about 90 degrees (i.e., visually obvious rotation).

[0034] The camera rotation sweep may be determined by: the camera position remains unchanged and rotates, the movement of one or more captured target objects is very small (does not affect the main body of the captured scene), and the camera rotation angle is greater than approximately 45 degrees.

[0035] The method for determining the movement and tracking of the camera angle of view toward the target object may be: the camera moves toward the target object, and the angle of view rotates toward the target object.

[0036] The camera translation can be determined by: the camera moves in parallel, the viewing angle is fixed in one direction, and there may be no target object.

[0037] Methods for determining camera static or slight motion include: the camera is stationary or at a very small angle, the camera is nearly stationary; the camera is stationary or at a very small angle, the camera is nearly stationary, and the target object is moving; the camera is only focused, and the zoom ratio exceeds 30% (the pixel ratio at the start and end times is less than 70%, and there is very obvious zoom).

[0038] The camera rotation may be determined by rotating the camera while aiming at a certain position without moving or scanning.

[0039] The above six camera motion categories are camera motion modes in natural scenes. In addition, the camera motion categories of video samples may also include: non-natural scene videos; non-natural scene videos include at least one of two-dimensional animations, videos that have been superimposed, videos with markers, multi-screen videos, videos with the main screen blocked, videos obtained by screenshots or recordings of the screen, and videos produced through special effects.

[0040] In this embodiment of the present invention, a certain number of video samples may be labeled for each camera motion category. For example, the number of video samples for each camera motion category may be no less than 2000. To simplify calculations, each video sample is labeled with only one camera motion category. The labeler may select the most appropriate camera motion category to label the video sample.

[0041] For S102, the camera motion detection model used in embodiments of the present invention may employ a deep learning model, such as a convolutional neural network (CNN) or a Transformer model built using an attention network. An attention network is a network model trained using an attention mechanism. This model assigns different weights to each part of an input sequence, thereby extracting more important feature information from the input sequence, ultimately resulting in a more accurate output.

[0042] Embodiments of the present invention provide a method for feature extraction based on continuous video frames and feature extraction across video frames. The former can effectively capture local temporal relationships, while the latter can establish longer-range temporal dependencies. This multi-scale temporal feature extraction method can more comprehensively characterize the dynamic characteristics of camera motion in videos, enabling camera motion detection models to more effectively learn and represent camera motion patterns at different scales in videos.

[0043] In S103, a loss value for the current iterative training is calculated using the first camera motion category detection result output by the camera motion detection model and the camera motion category label of the input video sample. The loss value is then used to update model parameters of the camera motion detection model until an iterative training termination condition is met, thereby obtaining a trained camera motion detection model. The iterative training termination condition may be that the number of iterative training iterations of the camera motion detection model reaches a preset number of iterations, or the loss value of the camera motion detection model is less than a preset loss value.

[0044] In S104 , the input video to be detected is input into the trained camera motion detection model, and the model calculation is performed as in S102 , and a second camera motion category detection result is output.

[0045] An embodiment of the present invention provides a camera motion detection method, which trains a camera motion detection model using video samples with camera motion category annotations, uses the camera motion detection model to extract features of continuous video frames of the video samples to obtain a first video feature, and extracts features across video frames of the video samples to obtain a second video feature. The two feature extraction methods can effectively capture the local and global dynamic characteristics of camera motion in the video, thereby more comprehensively characterizing the dynamic process of camera motion. Classification calculation is performed based on the extracted video features to obtain a first camera motion category detection result and the camera motion detection model is optimized for loss. Then, the trained camera motion detection model is used to detect the camera motion category of the input video, thereby achieving efficient and accurate recognition of the camera motion category in the video, which helps to improve the performance of multiple downstream tasks such as auxiliary video generation, complex object motion trajectory analysis, and video camera movement analysis.

[0046] Based on the above embodiment, the embodiment of the present invention continues to describe the model structure of the camera detection model.

[0047] Figure 3 A schematic structural diagram of a camera motion detection model provided by an embodiment of the present invention.

[0048] like Figure 3 As shown, the camera motion detection model adopted in the embodiment of the present invention can integrate video preprocessing, feature extraction and classification into a unified framework, realize end-to-end training and reasoning, and thus realize end-to-end video camera motion detection based on deep learning.

[0049] In an embodiment of the present invention, a camera motion detection model is used to perform feature extraction on continuous video frames of a video sample to obtain a first video feature, and feature extraction across video frames of the video sample is performed to obtain a second video feature. This can include: uniformly cropping the video frames of the video sample to obtain multiple video frame slices; flattening the video frame slices corresponding to a video frame into a first sequence, performing attention coding calculation on the position coding of the first sequence and the video frame to obtain video frame features corresponding to the video frame; and performing feature extraction on the video frame features of multiple video frames of the video sample to obtain a first video feature and a second video feature.

[0050] Before extracting features from a video, the input video (video sample or input video to be detected) may be preprocessed.

[0051] Video preprocessing can include: Frame extraction: For a video, multiple frames (e.g., 36 frames) can be extracted at equal intervals. Cropping: All data is cropped to a 256*256 size around a central point. Regularization: The pixel values of all frames are normalized (using red, green, and blue (RGB)). Each video frame is cropped to a 16*16 patch size. This yields video frame slices corresponding to the extracted frames, resulting in 256 slices per frame.

[0052] Then, the video frame slices of each video frame are input into the frame-level attention calculation network (Frame Transformer) of the camera motion detection model. In an embodiment of the present invention, the frame-level attention calculation network can be composed of a plurality of (for example, 6) Transformer encoders (Transformer Encoders) stacked together, and each Transformer encoder can adopt a multi-head encoder, and the number of attention heads (Attention head) can be set to 8. The video frame slices corresponding to a video frame are flattened into a sequence, recorded as the first sequence, and the first sequence is added with the position encoding (position embedding) of the video frame and then input into the frame-level attention calculation network, and all the output video frame slices are averaged to obtain the video frame features of the video frame. , Indicates the frame number.

[0053] After the extracted video frames are processed by the frame-level attention network, the corresponding video frame features are output. Frame-level feature representation based on video frame slices, that is, by dividing the video frames into video frame slices and encoding them using Transformer, can capture spatial information within the frame in a more fine-grained manner, providing richer feature representations for subsequent time series modeling.

[0054] In an embodiment of the present invention, performing feature extraction of continuous video frames of a video sample to obtain a first video feature may include: performing multiple rounds of feature extraction of video frame features of multiple video frames of the video sample at different scales, and performing feature fusion on the extracted first local sequence features to obtain the first video feature; wherein different scales correspond to different numbers of continuous video frames in the video sample.

[0055] After obtaining the video frame features of each video frame, this embodiment of the present invention designs a sliding window attention calculation module (Slide Window Transformer) to further extract video features. Since the goal of this embodiment of the present invention is to capture changes in video frames to detect the camera's perspective, modeling frame representations within a specific window (this can be denoted as the first window) may be more meaningful than observing frame changes globally. Therefore, this embodiment of the present invention designs an attention window based on the Transformer encoder. This window allows a video frame to only consider the attention of adjacent frames when calculating the contextual representation.

[0056] In the embodiment of the present invention, the window size of the sliding window attention calculation module can be set to 5, that is, for the frame , just need to consider Then, by stacking the attention, the receptive field (i.e., the size of the first window) can be gradually expanded to perform multiple rounds of feature extraction at different scales.

[0057] In an embodiment of the present invention, a sliding window attention calculation module can be designed to stack 4 layers of Transformer encoders, and the number of attention heads can be set to 8.

[0058] In an embodiment of the present invention, performing feature extraction across video frames of a video sample to obtain a second video feature may include: selecting a number of video frames corresponding to the second preset window size based on a second preset window size and a preset step size, and the spacing between adjacent video frames in the video sample is the preset step size; performing feature extraction on the video frame features of the selected video frames to obtain the second video feature.

[0059] The embodiment of the present invention also designs a cross-window attention calculation module (Cross Window Transformer) to model larger-scale video frame features. The main purpose of this module is to establish the representation context between video frames at a large time scale. By setting the window size (which can be recorded as the second window size) and the compensation between the video frames, the video frame is allowed to take the video frame of the window size at a certain step length for attention calculation. In the embodiment of the present invention, the second window size can be set to 5, and the preset step size can be 3, that is, for the video frame , need to consider attention.

[0060] In an embodiment of the present invention, extracting the video frame features of a selected video frame to obtain the second video feature may include: performing multiple rounds of feature extraction at different scales on multiple video frames of the video sample, and fusing the extracted second local sequence features to obtain the second video feature; wherein the different scales correspond to different numbers of video frames spaced at a preset step size within the video sample. In practical applications, the cross-window attention calculation module stacks two layers of Transformer encoders and sets the number of attention heads to 8.

[0061] In the camera motion detection model, the sliding window attention calculation module and the cross-window attention calculation module are parallel modules.

[0062] By combining two different-scale temporal modeling modules, a sliding window attention calculation module and a cross-window attention calculation module, the embodiment of the present invention can effectively capture the local and global dynamic characteristics of camera motion in the video, and effectively capture the relationship between frames in different time ranges, thereby more comprehensively characterizing the dynamic process of camera motion.

[0063] Then, the output of the sliding window attention calculation module is averaged to obtain the first video feature, and the output of the cross-window attention calculation module is averaged to obtain the second video feature. The first video feature and the second video feature are spliced together to obtain a representation of the entire video. The representation is input into the classifier of the camera motion detection model to obtain the probabilities of multiple camera motion categories. The output camera motion category detection result is determined according to the size of the probability value.

[0064] In the embodiment of the present invention, the classifier may be composed of a multilayer perceptron (MLP) and an activation function (softmax function).

[0065] In an embodiment of the present invention, when training a camera motion detection model, an Adam Weight Decay Optimizer (AdamW) may be used to perform loss optimization calculation, and the learning rate and weight decay may be set to 0.00001 and 0.001, respectively.

[0066] The video feature extraction method of the camera motion detection model introduced in the above embodiment is to extract the video features of the video frame as a whole. In addition, the camera motion detection model can also be assisted in judging the camera motion category based on the features of the target object in the video frame and / or the background features of the video frame. The camera motion detection model is used to extract the features of the continuous video frames of the video sample to obtain the first video feature, and to extract the features across the video frames of the video sample to obtain the second video feature. It can include: extracting the foreground object features and the background object features of the video frames of the video sample to obtain the foreground object features and background features of the multiple video frames as the video frame features corresponding to the video frames; extracting the video frame features of the multiple video frames of the video sample to obtain the first video feature and the second video feature.

[0067] Based on this, when extracting the video frame features of multiple video frames of a video sample to obtain the first video feature and the second video feature, the first video feature and the second video feature can be extracted from the foreground object features and background features of the selected video frame respectively, and the first video features corresponding to the foreground object features and background features output by the sliding window attention calculation module are spliced, and the second video features corresponding to the foreground object features and background features output by the cross-window attention calculation module are spliced, and then the spliced first video features and the spliced second video features are spliced, and the obtained results are input into the classifier for classification calculation.

[0068] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0069] An embodiment of the present invention also provides a camera motion detection device, which may include: an acquisition module for acquiring video samples with camera motion category annotations; a model training module for using a camera motion detection model to perform feature extraction of continuous video frames on the video samples to obtain a first video feature, perform feature extraction across video frames on the video samples to obtain a second video feature, perform classification calculation based on the first video feature and the second video feature to obtain a first camera motion category detection result; calculate a loss value based on the first camera motion category detection result and the corresponding camera motion category annotation, and use the loss value to update the model parameters of the camera motion detection model until the iterative training end condition is met to obtain a trained camera motion detection model; a monitoring module for using the trained camera motion detection model to calculate a second camera motion category detection result based on the input video.

[0070] In an embodiment of the present invention, the camera motion categories of the video sample may include: camera view rotation around the target object, camera view rotation and sweeping, camera view movement and tracking toward the target object, camera translation, camera static or micro motion, and camera rotation.

[0071] In an embodiment of the present invention, the camera motion category of the video sample may also include: non-natural scene video; non-natural scene video includes at least one of two-dimensional animation, video after superposition processing, marker video, multi-screen video, video with the main screen blocked, video obtained by screenshot or recording the screen, and video produced by special effects.

[0072] In an embodiment of the present invention, the model training module uses a camera motion detection model to perform feature extraction on continuous video frames of a video sample to obtain a first video feature, and performs feature extraction across video frames of the video sample to obtain a second video feature, which may include: uniformly cropping the video frames of the video sample to obtain multiple video frame slices; flattening the video frame slices corresponding to a video frame into a first sequence, performing attention coding calculation on the position coding of the first sequence and the video frame to obtain video frame features corresponding to the video frame; and performing feature extraction on the video frame features of multiple video frames of the video sample to obtain a first video feature and a second video feature.

[0073] In an embodiment of the present invention, the model training module performs feature extraction of continuous video frames on a video sample to obtain a first video feature, which may include: performing multiple rounds of feature extraction of video frame features of multiple video frames of the video sample at different scales, and performing feature fusion on the extracted first local sequence features to obtain the first video feature; wherein different scales correspond to different numbers of continuous video frames in the video sample.

[0074] In an embodiment of the present invention, the model training module performs feature extraction across video frames on the video sample to obtain a second video feature, which may include: selecting a number of video frames corresponding to the second preset window size according to the second preset window size and the preset step size, and the spacing between adjacent video frames in the video sample is the preset step size; performing feature extraction on the video frame features of the selected video frames to obtain the second video feature.

[0075] In an embodiment of the present invention, the model training module performs feature extraction on the video frame features of the selected video frames to obtain the second video features, which may include: performing multiple rounds of feature extraction at different scales on multiple video frames of the video sample, and performing feature fusion on the extracted second local sequence features to obtain the second video features; wherein different scales correspond to different numbers of video frames spaced at preset step sizes in the video sample.

[0076] For the description of the features in the embodiment corresponding to the camera motion detection device, reference may be made to the relevant description of the embodiment corresponding to the camera motion detection method, which will not be repeated here.

[0077] An embodiment of the present invention further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps of any of the above-mentioned camera motion detection method embodiments.

[0078] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned camera motion detection method embodiments when running.

[0079] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0080] An embodiment of the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned camera motion detection method embodiments are implemented.

[0081] An embodiment of the present invention further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned camera motion detection method embodiments are implemented.

[0082] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0083] The above describes in detail the camera motion detection method, apparatus, device, and computer-readable storage medium provided by the present invention. This document uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is intended only to facilitate understanding of the method and core concepts of the present invention. It should be noted that those skilled in the art will be able to make various improvements and modifications to the present invention without departing from the principles of the present invention, and such improvements and modifications are also within the scope of protection of the present invention.

Claims

1. A camera motion detection method, characterized in that: include: Obtain video samples with camera motion category annotations; Using a camera motion detection model, extracting features of continuous video frames of the video sample to obtain a first video feature, extracting features across video frames of the video sample to obtain a second video feature, and performing classification calculation based on the first video feature and the second video feature to obtain a first camera motion category detection result; Calculating a loss value based on the first camera motion category detection result and the corresponding camera motion category label, and using the loss value to update model parameters of the camera motion detection model until an iterative training end condition is met, thereby obtaining the trained camera motion detection model; The trained camera motion detection model is used to calculate a second camera motion category detection result based on the input video.

2. The camera motion detection method according to claim 1, wherein: The camera motion categories of the video samples include: camera view rotating around the target object, camera view rotating and sweeping, camera view moving and tracking toward the target object, camera translation, camera static or micro-motion, and camera rotation.

3. The camera motion detection method according to claim 2, wherein: The camera motion categories of the video samples also include: unnatural scene videos; The non-natural scene video includes at least one of a two-dimensional animation, a video that has been superimposed, a marker video, a multi-screen video, a video with the main screen blocked, a video obtained by screenshot or recording the screen, and a video produced by special effects.

4. The camera motion detection method according to claim 1, wherein: Using a camera motion detection model to perform feature extraction of continuous video frames on the video sample to obtain a first video feature, and performing feature extraction across video frames on the video sample to obtain a second video feature, including: Evenly cropping the video frames of the video sample to obtain a plurality of video frame slices; Flattening a video frame slice corresponding to the video frame into a first sequence, performing attention coding calculation on the first sequence and the position coding of the video frame to obtain a video frame feature corresponding to the video frame; Feature extraction is performed on video frame features of the plurality of video frames of the video sample to obtain the first video feature and the second video feature.

5. The camera motion detection method according to claim 1, wherein: Extracting features of continuous video frames of the video sample to obtain a first video feature includes: performing multiple rounds of feature extraction at different scales on video frame features of multiple video frames of the video sample, and performing feature fusion on the extracted first local sequence features to obtain the first video features; Different scales correspond to different numbers of continuous video frames in the video sample.

6. The camera motion detection method according to claim 1, wherein: Performing feature extraction across video frames on the video sample to obtain a second video feature includes: According to a second preset window size and a preset step size, selecting a number of video frames corresponding to the second preset window size, wherein a distance between adjacent video frames in the video sample is the preset step size; Feature extraction is performed on the video frame features of the selected video frame to obtain the second video features.

7. The camera motion detection method according to claim 6, wherein: Extracting the video frame features of the selected video frame to obtain the second video features includes: performing multiple rounds of feature extraction at different scales on the plurality of video frames of the video sample, and performing feature fusion on the extracted second local sequence features to obtain the second video features; Different scales correspond to different numbers of video frames spaced by the preset step length in the video samples.

8. A camera motion detection device, characterized in that: include: An acquisition module is used to obtain video samples with camera motion category annotations; A model training module is configured to use a camera motion detection model to perform feature extraction of continuous video frames on the video sample to obtain a first video feature, perform feature extraction across video frames on the video sample to obtain a second video feature, perform classification calculation based on the first video feature and the second video feature to obtain a first camera motion category detection result; calculate a loss value based on the first camera motion category detection result and the corresponding camera motion category label, and use the loss value to update model parameters of the camera motion detection model until an iterative training end condition is met, thereby obtaining the trained camera motion detection model; The monitoring module is configured to calculate a second camera motion category detection result based on the input video using the trained camera motion detection model.

9. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the camera motion detection method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the camera motion detection method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Camera movement analyzing method and device in video

    CN102737383A

  • Video action recognition model training method and device, computing equipment and storage medium

    CN114266997A

  • Abnormal data intelligent identification method, device and equipment based on video stream data

    CN116310985A

  • Moving target detection method based on multi-frame space attention mechanism

    CN117541854A

  • Classroom complete meta-action recognition method based on dynamic position embedding

    CN118823636A