Camera motion detection method, apparatus, device and computer readable storage medium

By training the camera motion detection model to extract and classify video frame features, the shortcomings of camera motion detection are solved, efficient recognition of camera motion in videos is achieved, and the effects of video generation and creative analysis are improved.

CN120510558BActive Publication Date: 2025-10-17SHANDONG HAILIANG INFORMATION TECH RES INST
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511005705.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-10-17
Estimated Expiration
2045-07-22

AI Technical Summary

Technical Problem

The existing technology lacks effective camera motion detection solutions, which leads to insufficient video generation quality and makes it difficult for creators to systematically analyze and apply camera motion techniques.

Method used

By using video samples with camera motion category annotations to train the camera motion detection model, feature extraction of continuous video frames and across video frames is performed. Combined with deep learning models such as convolutional neural networks and Transformer models, the local and global dynamic characteristics of camera motion are captured, and classification calculation and loss optimization are performed.

Benefits of technology

It achieves efficient and accurate recognition of camera motion categories in videos, which helps improve the quality of video generation and the performance of downstream tasks such as complex object motion trajectory analysis and video camera movement analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510558B_ABST
    Figure CN120510558B_ABST
Patent Text Reader

Abstract

The application discloses a camera motion detection method and device, equipment and a computer readable storage medium, and relates to the technical field of video processing. The camera motion detection model is trained by using a video sample with camera motion category labels. The camera motion detection model is used to extract features of continuous video frames of the video sample to obtain first video features, and extract features across video frames of the video sample to obtain second video features. The two feature extraction methods can effectively capture the local and global dynamic characteristics of the camera motion in the video, thereby more comprehensively depicting the dynamic process of the camera motion. The first camera motion category detection result is obtained by classification calculation according to the extracted video features, and the camera motion detection model is loss optimized. Then, the trained camera motion detection model is used to detect the camera motion category of an input video, so that efficient and accurate identification of the camera motion category in the video can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video processing, and particularly relates to a camera motion detection method and device, equipment and a computer readable storage medium. BACKGROUND

[0002] As an important information carrier, the camera motion in the shooting process of a video contains rich semantic information, which is crucial for understanding the video content, improving the video generation quality and assisting video creation analysis. At present, there is a lack of detection scheme for the camera motion mode in the video processing technology. SUMMARY

[0003] The present application provides a camera motion detection method, device, equipment and computer readable storage medium to at least solve the problem of lacking a detection scheme for the camera motion mode in the related art.

[0004] The present application provides a camera motion detection method, comprising:

[0005] obtaining a video sample with camera motion category annotation;

[0006] extracting features of consecutive video frames of the video sample by using a camera motion detection model to obtain first video features, extracting features across video frames of the video sample to obtain second video features, and performing classification calculation according to the first video features and the second video features to obtain a first camera motion category detection result;

[0007] calculating a loss value according to the first camera motion category detection result and the corresponding camera motion category annotation, updating model parameters of the camera motion detection model by using the loss value until an iterative training end condition is reached to obtain a trained camera motion detection model;

[0008] calculating a second camera motion category detection result according to an input video by using the trained camera motion detection model.

[0009] The present application also provides a camera motion detection device, comprising:

[0010] an obtaining module configured to obtain a video sample with camera motion category annotation;

[0011] a model training module configured to perform feature extraction on consecutive video frames of the video sample by using the camera motion detection model to obtain first video features, perform feature extraction across video frames of the video sample to obtain second video features, perform classification calculation according to the first video features and the second video features to obtain a first camera motion category detection result, calculate a loss value according to the first camera motion category detection result and the corresponding camera motion category label, update model parameters of the camera motion detection model by using the loss value until an iterative training end condition is reached, and obtain the trained camera motion detection model;

[0012] a monitoring module configured to calculate a second camera motion category detection result according to an input video by using the trained camera motion detection model.

[0013] The present application also provides an electronic device, which comprises a memory configured to store a computer program and a processor configured to execute the computer program to implement the steps of any of the camera motion detection methods.

[0014] The present application also provides a computer readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement the steps of any of the camera motion detection methods.

[0015] According to the present application, the camera motion detection model is trained by using video samples with camera motion category labels, the camera motion detection model is used to perform feature extraction on consecutive video frames of the video sample to obtain first video features, feature extraction is performed across video frames of the video sample to obtain second video features, the local and global dynamic characteristics of camera motion in the video can be effectively captured by the two feature extraction methods, so that the dynamic process of camera motion can be more comprehensively described, the first camera motion category detection result is obtained by classification calculation according to the extracted video features, and the camera motion detection model is loss optimized, and then the trained camera motion detection model is used to perform camera motion category detection on an input video, so that efficient and accurate identification of the camera motion category in the video can be realized, which is helpful to performance improvement of video generation, complex object motion trajectory analysis, video panning analysis and other downstream tasks. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0017] Figure 1A flow chart of a camera motion detection method provided for an embodiment of the present application;

[0018] Figure 2 A structural schematic diagram of a camera motion detection system provided for an embodiment of the present application;

[0019] Figure 3 A structural schematic diagram of a camera motion detection model provided for an embodiment of the present application. DETAILED DESCRIPTION

[0020] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0021] It should be noted that, in the description of the present application, the terms “comprise”, “contain” or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. The terms “first”, “second” and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.

[0022] In order for those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the drawings and specific embodiments.

[0023] Some key terms used in the embodiments of the present application will be explained first.

[0024] Accurate detection and analysis of camera motion in a video are crucial for understanding video content and improving video processing technology.

[0025] For example, one of the goals of video generation tasks is to generate dynamic videos with realistic and expressive performances, but the related art lacks research on camera motion in videos, resulting in a lack of three-dimensional consistency in the generation results, or being limited to static scenes, making it difficult to generate dynamic scenes containing complex object motion. In addition, it is also extremely challenging to generate multi-view consistent videos with different camera trajectories in the same scene. Due to the lack of large-scale outdoor multi-view video datasets, current multi-view video generation is mainly limited to near-static scenes or synthetic objects.

[0026] A type of video generation scheme in the related art is to learn a three-dimensional generation model from large-scale unlabeled video data, and generate multi-view images by controlling the camera direction through visual condition technology, but it only considers the generation of a three-dimensional model and the control of a view angle, and lacks the acquisition of camera motion information, resulting in poor video generation performance. Therefore, effectively detecting and analyzing the camera motion type can provide important guidance information for the video generation model, such as for training data enhancement or as a control signal for the generation model, thereby improving the realism and expressiveness of the generated video.

[0027] In addition, for video and film creators, understanding and analyzing camera motion is crucial to improving the quality of works. Excellent film producers skillfully use camera motion to guide audience attention, enhance narrative effects, and define visual style. However, there is currently a lack of effective tools to help creators analyze and quantify camera motion in videos, making it difficult for creators to systematically learn and apply camera motion techniques.

[0028] Therefore, it is necessary to provide a camera motion detection scheme for the manner of camera motion in a video.

[0029] To this end, the embodiments of the present application provide a camera motion detection method, device, equipment and computer readable storage medium, by training a camera motion detection model using video samples with camera motion category labels, extracting features of consecutive video frames of the video samples using the camera motion detection model to obtain first video features, and extracting features across video frames of the video samples to obtain second video features, the two feature extraction methods can effectively capture the local and global dynamic characteristics of camera motion in the video, thereby more comprehensively depicting the dynamic process of camera motion, and the first camera motion category detection result is obtained by classification calculation according to the extracted video features and the camera motion detection model is loss optimized, and then the trained camera motion detection model is used to detect the camera motion category of the input video, so that efficient and accurate recognition of the camera motion category in the video can be realized, which is helpful to assist video generation, complex object motion trajectory analysis, video panning analysis and performance improvement of other downstream tasks.

[0030] The embodiments of the present application provide a camera motion detection method, and the method is described in detail below in combination with the execution process of the camera motion detection method.

[0031] Figure 1 A flowchart of a camera motion detection method provided by the embodiments of the present application; Figure 2 A structural schematic diagram of a camera motion detection system provided by the embodiments of the present application.

[0032] As Figure 1As shown, the camera motion detection method provided by the embodiment of the present invention includes: S101: obtaining a video sample with a camera motion category label.

[0033] S102: Using a camera motion detection model, extract features of continuous video frames of the video sample to obtain a first video feature, extract features across video frames of the video sample to obtain a second video feature, and perform classification calculation based on the first video feature and the second video feature to obtain a first camera motion category detection result.

[0034] S103: Calculate a loss value based on the first camera motion category detection result and the corresponding camera motion category label, and use the loss value to update the model parameters of the camera motion detection model until the iterative training end condition is met to obtain a trained camera motion detection model.

[0035] S104: Calculate a second camera motion category detection result based on the input video using the trained camera motion detection model.

[0036] In order to realize the camera motion detection of the video, the camera motion detection method provided by the embodiment of the present invention mainly includes two steps: training the camera detection model and using the trained camera detection model to perform the camera motion detection task. In this regard, the camera motion detection method provided by the embodiment of the present invention can be applied to the following: Figure 2 The camera motion detection system shown includes a model training system and a camera motion detection device.

[0037] like Figure 2 As shown, the model training system can be provided with computing power and storage support by multiple artificial intelligence servers (artificial intelligence servers 1~n), and the camera motion detection model can be trained based on the model training system using video samples. In some optional implementations of the embodiments of the present invention, a computing resource pool can be constructed based on the model training system, and the computing resource pool includes multiple computing threads, so that the training task of the camera motion detection model can be split into multiple subtasks for parallel execution. The model parallel training method can be adopted, and the data parallel training method can also be adopted. The model training system can be deployed in a cloud computing center. The equipment structure of the cloud computing center can be similar to Figure 2 The camera motion detection device shown here includes basic infrastructure such as an AI processor, storage, input / output devices, a communication bus, and a communication interface. It also comprises a software environment including an operating system, database, and middleware, as well as application software. The model training system receives manually annotated video samples and can fine-tune a camera motion detection model using a stored pre-trained model or train a randomly initialized model framework.

[0038] The camera motion detection device can be a computing device deployed in a cloud computing center or a terminal device, and the structure of the camera motion detection device is as shown in Figure 2 If the computing device of the cloud computing center is adopted, after the trained camera motion detection model is deployed, the input video to be detected is uploaded to the cloud computing center by the terminal device, the trained camera motion detection model is used to perform camera motion category detection on the input video to be detected, and a second camera motion category detection result is output. If the terminal device is adopted as the camera motion detection device, after the camera motion detection model is trained by the model training system and sent to the terminal device through the communication unit and deployed in the storage of the terminal device, the terminal device receives the input video to be detected, performs camera motion category detection on the input video to be detected through the local camera motion detection model, and outputs a second camera motion category detection result.

[0039] In the embodiment of the present application, by pre-annotating the camera motion category of the video sample, the camera motion detection model is trained using the video sample with camera motion category annotation, so that the camera motion detection model can learn the relationship between the video features in the video sample and the camera motion category.

[0040] For S101, a plurality of camera motion categories can be divided according to the manner of camera motion. In some optional embodiments of the embodiment of the present application, the camera motion category of the video sample can include: camera perspective rotation around the target object, camera perspective self-rotation, camera perspective moving and tracking towards the target object, camera translation, camera static or micro-motion, and camera rotation.

[0041] The target object is a reference object in the video. When dividing the camera motion category, the movement manner of the target object can be considered or not.

[0042] The determination method of the camera perspective rotation around the target object can be that the camera moves around the target object stably (facing the target), the target object is static or moves slightly, and the camera angle changes by more than about 30 degrees (obviously rotated visually); the camera is static or moves slightly, the target object rotates rigidly, and the rotation angle of the target object is greater than about 90 degrees (i.e., obviously rotated visually).

[0043] The determination method of the camera perspective self-rotation can be that the camera position is unchanged and rotates, and the movement amplitude of one or more target objects captured is very small (without affecting the main body of the shooting scene), and the camera rotation angle is greater than about 45 degrees.

[0044] The determination method of the camera perspective moving and tracking towards the target object can be that the camera moves towards the target object and the perspective is aligned with the target object.

[0045] The determination method of the camera translation can be that the camera moves in parallel, the angle of view is fixed in one direction, and there is no target object.

[0046] The determination method of the camera static or micro-motion can be that the camera is close to static, the camera is close to static, and the target object moves, or the camera only performs focusing, and the zooming ratio exceeds 30% (the pixel ratio at the start and end time is less than 70%, and there is very obvious zooming).

[0047] The determination method of the camera rotation can be that the camera is rotated to align with a certain position, and there is no movement or scanning.

[0048] The above six camera motion categories are camera motion modes in natural scenes, in addition, the camera motion category of the video sample can also include at least one of the following: a two-dimensional animation, a video after superposition processing, a marker video, a multi-screen video, a video in which the main body screen is blocked, a video obtained by screen capturing or screen recording, and a video obtained by special effect production.

[0049] In the embodiment of the application, a certain number of video samples can be labeled for each camera motion category, for example, the number of video samples of each camera motion category can be not less than 2000. To simplify the calculation, only one camera motion category is labeled for each video sample, and the labeler can select the most suitable camera motion category to label the video sample.

[0050] For S102, the camera motion detection model used in the embodiment of the application can use a deep learning model, such as a convolutional neural network (CNN), a Transformer model constructed by an attention network, etc. The attention network refers to a network model trained by using an attention mechanism, which extracts more important feature information in the input sequence by giving different weights to each part of the input sequence, so that the model finally obtains more accurate output.

[0051] The embodiment of the application provides a method based on feature extraction of continuous video frames and cross-video frame feature extraction. The former can effectively capture local temporal relationships, and the latter can establish longer temporal dependencies. This multi-scale time feature extraction method can more comprehensively depict the dynamic characteristics of camera motion in the video, so that the camera motion detection model can more effectively learn and represent different scale camera motion patterns in the video.

[0052] For S103, the loss value of the current iteration training is calculated by using the first camera motion category detection result output by the camera motion detection model and the camera motion category label of the video sample input by the camera motion detection model, and the model parameter of the camera motion detection model is updated by using the loss value until the iteration training end condition is reached, and the trained camera motion detection model is obtained. The iteration training end condition can be that the iteration training times of the camera motion detection model reaches the preset iteration times, or the loss value of the camera motion detection model is less than the preset loss value.

[0053] For S104, the input video to be detected is input into the trained camera motion detection model, and the model calculation in S102 is performed to output the second camera motion category detection result.

[0054] The camera motion detection method provided by the embodiment of the application trains the camera motion detection model by using the video sample with camera motion category label, extracts the features of the continuous video frames of the video sample by using the camera motion detection model to obtain the first video features, extracts the features across the video frames of the video sample to obtain the second video features, and the local and global dynamic characteristics of the camera motion in the video can be effectively captured by the two feature extraction methods, so that the dynamic process of the camera motion is more comprehensively described. The first camera motion category detection result is obtained by classification calculation according to the extracted video features, and the loss optimization of the camera motion detection model is performed, and then the camera motion category detection of the input video is performed by using the trained camera motion detection model, so that the efficient and accurate recognition of the camera motion category in the video can be realized, which is helpful to the performance improvement of the video generation, complex object motion trajectory analysis, video tracking analysis and other downstream tasks.

[0055] On the basis of the above-mentioned embodiment, the embodiment of the application continues to describe the model structure of the camera detection model.

[0056] Figure 3 A structural diagram of the camera motion detection model provided by the embodiment of the application is shown.

[0057] As shown in Figure 3 The camera motion detection model used in the embodiment of the application can integrate video preprocessing, feature extraction and classification into a unified framework, realize end-to-end training and reasoning, and thus realize an end-to-end video camera motion detection based on deep learning.

[0058] In the embodiment of the present application, the feature extraction of the continuous video frames of the video sample by using the camera motion detection model to obtain the first video feature and the cross-video frame feature extraction of the video sample can include: uniformly cropping the video frames of the video sample to obtain a plurality of video frame patches; flattening the video frame patches corresponding to one video frame into a first sequence, and performing attention encoding calculation on the first sequence and the position encoding of the video frame to obtain the video frame feature corresponding to the video frame; and performing feature extraction on the video frame features of the plurality of video frames of the video sample to obtain the first video feature and the second video feature.

[0059] Before the feature extraction of the video, the input video (video sample or input video to be detected) can be preprocessed.

[0060] The video preprocessing can include: frame extraction: for a video, a plurality of frames (such as 36 frames) can be extracted at equal intervals; cropping: all data is cropped around the center point to a size of 256*256; normalization: the pixel values (which can be represented by red, green and blue RGB) of all frames are normalized; each video frame is cropped according to a video frame patch size of 16*16, at this time, a plurality of video frame patches corresponding to the extracted video frames are obtained, and 256 video frame patches can be obtained for each video frame.

[0061] Then, the video frame patches of each video frame are input into the frame-level attention calculation network (Frame Transformer) of the camera motion detection model. In the embodiment of the present application, the frame-level attention calculation network can be stacked by a plurality of (for example, 6) Transformer encoders (Transformer Encoder), and each Transformer encoder can adopt a multi-head encoder, and the number of attention heads (Attention head) can be set to 8. The video frame patches corresponding to one video frame are flattened into a sequence, denoted as a first sequence, and the first sequence is input into the frame-level attention calculation network after adding the position encoding (position embedding) of the video frame, and the average of all video frame patches output by the frame-level attention calculation network is taken to obtain the video frame feature of the video frame , representing the number of the frame.

[0062] After the frame-level attention calculation network is calculated for each extracted video frame, the video frame feature corresponding to each video frame is output. The frame-level feature representation based on the video frame patch level, that is, the video frame is divided into video frame patches and encoded by using the Transformer, can capture the intra-frame spatial information in a more fine-grained manner, and provide more rich feature representation for subsequent time sequence modeling.

[0063] In the embodiment of the present application, the feature extraction of the continuous video frames of the video sample to obtain the first video feature can include: performing multi-round feature extraction of different scales on the video frame features of the plurality of video frames of the video sample, performing feature fusion on the extracted first local sequence features to obtain the first video feature; wherein the different scales correspond to different numbers of continuous video frames in the video sample.

[0064] After obtaining the video frame features of each video frame, the embodiment of the present application designs a slide window attention calculation module (Slide Window Transformer) to further extract video features. Since the target of the embodiment of the present application is to capture the changes of the video frames to detect the camera view, it is more meaningful to model the representation of the frames in a certain window (which can be denoted as a first window) than to observe the changes of the frames globally, therefore, the embodiment of the present application designs an attention (Attention) window based on the Transformer encoder, which enables the video frames to only consider the attention of the adjacent frames when calculating the context representation.

[0065] In the embodiment of the present application, the window size of the slide window attention calculation module can be set to 5, that is, for the frame , only the attention of needs to be considered (i.e. the two adjacent frames before and after it). Then, by stacking the attention, the receptive field (i.e. the size of the first window) can be gradually expanded to perform multi-round feature extraction of different scales.

[0066] In the embodiment of the present application, the slide window attention calculation module can be designed to stack 4 layers of Transformer encoders, and the number of attention heads can be set to 8.

[0067] In the embodiment of the present application, the feature extraction of the video sample across the video frames to obtain the second video feature can include: selecting a number of video frames corresponding to a second preset window size according to the second preset window size and a preset step length, and the interval between adjacent video frames in the video sample is the preset step length; performing feature extraction on the video frame features of the selected video frames to obtain the second video feature.

[0068] The embodiment of the present application also designs a cross window attention calculation module (Cross Window Transformer) to model larger scale video frame features. The main purpose of this module is to establish the representation context between the video frames in a large scale of time. By setting the window size (which can be denoted as a second window size) and the compensation between the video frames, the video frames take the video frames of the window size every certain step length to perform attention calculation. In the embodiment of the present application, the second window size can be set to 5, and the preset step length can be set to 3, that is, for the video frame , need to be considered attention.

[0069] In the embodiment of the application, the feature extraction is performed on the video frame features of the selected video frames to obtain second video features, which can include: performing multi-round feature extraction of different scales on the plurality of video frames of the video sample, and performing feature fusion on the extracted second local sequence features to obtain the second video features; wherein the different scales correspond to different numbers of video frames that are spaced apart by a preset step in the video sample. In actual application, the cross-window attention calculation module stacks 2 layers of Transformer encoders, and sets the number of attention heads to 8.

[0070] In the camera motion detection model, the sliding window attention calculation module and the cross-window attention calculation module are parallel modules.

[0071] The embodiment of the application can effectively capture the local and global dynamic characteristics of the camera motion in the video by combining the sliding window attention calculation module and the cross-window attention calculation module, which are two different scale time sequence modeling modules, effectively capture the inter-frame relationship in different time ranges, and thus more comprehensively depict the dynamic process of the camera motion.

[0072] Then, the output of the sliding window attention calculation module is averaged to obtain first video features, the output of the cross-window attention calculation module is averaged to obtain second video features, the first video features and the second video features are spliced to obtain a representation of the entire video, and the representation is input into a classifier of the camera motion detection model, so that the probabilities of a plurality of camera motion categories can be obtained. The camera motion category detection result is determined according to the size of the probability value.

[0073] In the embodiment of the application, the classifier can be composed of a multilayer perceptron (MLP) and an activation function (softmax function).

[0074] In the embodiment of the application, when training the camera motion detection model, an Adam weight decay optimizer (Adam Weight Decay Optimizer, AdamW) can be used for loss optimization calculation, the learning rate can be set to 0.00001, and the weight decay can be set to 0.001.

[0075] The video feature extraction method of the camera motion detection model introduced in the above embodiments is the overall video feature extraction of the video frame. In addition, the judgment of the camera motion category by the camera motion detection model can also be assisted according to the features of the target object in the video frame and / or the background features of the video frame. Then, the feature extraction of the continuous video frame of the video sample by the camera motion detection model to obtain the first video feature, and the feature extraction of the cross-video frame of the video sample to obtain the second video feature, can include: performing foreground object feature extraction and background object feature extraction on the video frames of the video sample to obtain the foreground object features and the background features of each of the plurality of video frames as the video frame features corresponding to the video frames; and performing feature extraction on the video frame features of the plurality of video frames of the video sample to obtain the first video feature and the second video feature.

[0076] Based on this, when the video frame features of the plurality of video frames of the video sample are extracted to obtain the first video feature and the second video feature, the foreground object features and the background features of the selected video frames can be extracted respectively to obtain the first video feature and the second video feature, and the first video features corresponding to the foreground object features and the background features respectively output by the sliding window attention calculation module are spliced, the second video features corresponding to the foreground object features and the background features respectively output by the cross-window attention calculation module are spliced, and then the spliced first video features and the spliced second video features are spliced, and the obtained result is input into the classifier for classification calculation.

[0077] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and the necessary general hardware platform, of course, it can also be realized by hardware, but in many cases the former is a better embodiment.

[0078] The embodiments of the present application also provide a camera motion detection device, which can include: an acquisition module configured to acquire a video sample with camera motion category labels; a model training module configured to perform feature extraction of continuous video frames of the video sample by a camera motion detection model to obtain a first video feature, perform feature extraction of cross-video frames of the video sample to obtain a second video feature, perform classification calculation according to the first video feature and the second video feature, and obtain a first camera motion category detection result; calculate a loss value according to the first camera motion category detection result and the corresponding camera motion category label, update the model parameters of the camera motion detection model by using the loss value, until a iteration training end condition is reached, and obtain a trained camera motion detection model; and a monitoring module configured to calculate a second camera motion category detection result according to an input video by using the trained camera motion detection model.

[0079] In the embodiment of the present application, the camera motion category of the video sample can include: camera perspective rotation around the target object, camera perspective rotation scanning, camera perspective rotation moving towards and tracking the target object, camera translation, camera static or micro-motion, and camera rotation.

[0080] In the embodiment of the present application, the camera motion category of the video sample can further include: non-natural scene video, and the non-natural scene video includes at least one of two-dimensional animation, superimposed video, marker video, multi-screen video, video with occluded subject screen, video obtained by screen capturing or screen recording, and video obtained by special effect production.

[0081] In the embodiment of the present application, the model training module can perform feature extraction on the continuous video frames of the video sample to obtain the first video feature, and perform feature extraction across the video frames of the video sample to obtain the second video feature, which can include: uniformly cropping the video frames of the video sample to obtain a plurality of video frame slices; flattening the video frame slices corresponding to one video frame into a first sequence, and performing attention coding calculation on the first sequence and the position coding of the video frame to obtain the video frame feature corresponding to the video frame; and performing feature extraction on the video frame features of the plurality of video frames of the video sample to obtain the first video feature and the second video feature.

[0082] In the embodiment of the present application, the model training module can perform feature extraction on the continuous video frames of the video sample to obtain the first video feature, which can include: performing multi-round feature extraction of different scales on the video frame features of the plurality of video frames of the video sample, performing feature fusion on the extracted first local sequence features to obtain the first video feature; wherein the different scales correspond to different numbers of continuous video frames in the video sample.

[0083] In the embodiment of the present application, the model training module can perform feature extraction across the video frames of the video sample to obtain the second video feature, which can include: selecting a number of video frames corresponding to the second preset window size according to the second preset window size and a preset step length, and the interval between adjacent video frames in the video sample is the preset step length; and performing feature extraction on the video frame features of the selected video frames to obtain the second video feature.

[0084] In the embodiment of the present application, the model training module can perform feature extraction on the video frame features of the selected video frames to obtain the second video feature, which can include: performing multi-round feature extraction of different scales on the plurality of video frames of the video sample, performing feature fusion on the extracted second local sequence features to obtain the second video feature; wherein the different scales correspond to different numbers of video frames in the video sample with a preset step length.

[0085] The features of the embodiments of the camera motion detection device can be referred to the related descriptions of the embodiments of the camera motion detection method, which will not be repeated here.

[0086] The embodiments of the present application also provide an electronic device, comprising a memory and a processor, the memory stores a computer program, and the processor is configured to execute the computer program to perform the steps in any of the above-mentioned embodiments of the camera motion detection method.

[0087] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program, wherein the computer program is configured to perform the steps in any of the above-mentioned embodiments of the camera motion detection method when executed.

[0088] In an example embodiment, the above-mentioned computer readable storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.

[0089] The embodiments of the present application also provide a computer program product, which comprises a computer program, and the computer program is executed by a processor to perform the steps in any of the above-mentioned embodiments of the camera motion detection method.

[0090] The embodiments of the present application also provide another computer program product, which comprises a non-volatile computer readable storage medium, and the non-volatile computer readable storage medium stores a computer program, and the computer program is executed by a processor to perform the steps in any of the above-mentioned embodiments of the camera motion detection method.

[0091] The skilled person can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in general terms in the above description. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0092] The camera motion detection method, device, equipment and computer readable storage medium provided by the present application are described in detail above. The principles and implementation modes of the present application are described by applying specific examples in this paper, and the above description of the embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that, for ordinary skilled persons in the technical field, some improvements and modifications can be made to the present application without departing from the principles of the present application, and these improvements and modifications also fall within the protection scope of the present application.

Claims

1. A camera motion detection method, characterized in that: include: Obtain video samples with camera motion category annotations; Using a camera motion detection model, extracting features of continuous video frames of the video sample to obtain a first video feature, extracting features across video frames of the video sample to obtain a second video feature, and performing classification calculation based on the first video feature and the second video feature to obtain a first camera motion category detection result; Calculating a loss value based on the first camera motion category detection result and the corresponding camera motion category label, and using the loss value to update model parameters of the camera motion detection model until an iterative training end condition is met, thereby obtaining the trained camera motion detection model; Utilizing the trained camera motion detection model to calculate a second camera motion category detection result based on the input video; The method of extracting features of continuous video frames of the video sample using a camera motion detection model to obtain a first video feature and extracting features across video frames of the video sample to obtain a second video feature includes: Performing foreground object feature extraction and background object feature extraction on the video frames of the video sample to obtain foreground object features and background features of the respective multiple video frames as video frame features corresponding to the video frames; extracting the first video features and the second video features from foreground object features and background features of a plurality of video frames of the video sample respectively; Performing feature extraction across video frames on the video sample to obtain a second video feature includes: According to a second preset window size and a preset step size, selecting a number of video frames corresponding to the second preset window size, and a distance between adjacent video frames in the video sample is the preset step size; performing multiple rounds of feature extraction at different scales on the plurality of video frames of the video sample, and performing feature fusion on the extracted second local sequence features to obtain the second video features; Different scales correspond to different numbers of video frames spaced by the preset step length in the video samples.

2. The camera motion detection method according to claim 1, wherein: The camera motion categories of the video samples include: camera view rotating around the target object, camera view rotating and sweeping, camera view moving and tracking toward the target object, camera translation, camera static or micro-motion, and camera rotation.

3. The camera motion detection method according to claim 2, wherein: The camera motion categories of the video samples also include: unnatural scene videos; The non-natural scene video includes at least one of a two-dimensional animation, a video that has been superimposed, a marker video, a multi-screen video, a video with the main screen blocked, a video obtained by screenshot or recording the screen, and a video produced by special effects.

4. The camera motion detection method according to claim 1, wherein: Using a camera motion detection model to perform feature extraction of continuous video frames on the video sample to obtain a first video feature, and performing feature extraction across video frames on the video sample to obtain a second video feature, including: Evenly cropping the video frames of the video sample to obtain a plurality of video frame slices; Flattening a video frame slice corresponding to the video frame into a first sequence, performing attention coding calculation on the first sequence and the position coding of the video frame to obtain a video frame feature corresponding to the video frame; Feature extraction is performed on video frame features of the plurality of video frames of the video sample to obtain the first video feature and the second video feature.

5. The camera motion detection method according to claim 1, wherein: Extracting features of continuous video frames of the video sample to obtain a first video feature includes: performing multiple rounds of feature extraction at different scales on the video frame features of the plurality of video frames of the video sample, and performing feature fusion on the extracted first local sequence features to obtain the first video features; Different scales correspond to different numbers of continuous video frames in the video sample.

6. A camera motion detection device, characterized in that: include: An acquisition module is used to obtain video samples with camera motion category annotations; A model training module is configured to use a camera motion detection model to perform feature extraction of continuous video frames on the video sample to obtain a first video feature, perform feature extraction across video frames on the video sample to obtain a second video feature, perform classification calculation based on the first video feature and the second video feature to obtain a first camera motion category detection result; calculate a loss value based on the first camera motion category detection result and the corresponding camera motion category label, and use the loss value to update model parameters of the camera motion detection model until an iterative training end condition is met, thereby obtaining the trained camera motion detection model; A monitoring module, configured to calculate a second camera motion category detection result based on the input video using the trained camera motion detection model; The method of extracting features of continuous video frames of the video sample using a camera motion detection model to obtain a first video feature and extracting features across video frames of the video sample to obtain a second video feature includes: Performing foreground object feature extraction and background object feature extraction on the video frames of the video sample to obtain foreground object features and background features of the respective multiple video frames as video frame features corresponding to the video frames; extracting the first video features and the second video features from foreground object features and background features of a plurality of video frames of the video sample respectively; Performing feature extraction across video frames on the video sample to obtain a second video feature includes: According to a second preset window size and a preset step size, selecting a number of video frames corresponding to the second preset window size, and a distance between adjacent video frames in the video sample is the preset step size; performing multiple rounds of feature extraction at different scales on the plurality of video frames of the video sample, and performing feature fusion on the extracted second local sequence features to obtain the second video features; Different scales correspond to different numbers of video frames spaced by the preset step length in the video samples.

7. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the camera motion detection method according to any one of claims 1 to 5 when executing the computer program.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the camera motion detection method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Abnormal data intelligent identification method, device and equipment based on video stream data

    CN116310985A