Scene detection model training method, scene detection method and device
By amplifying and training the camera posture data and using convolutional neural networks for scene category prediction, the problem of lagging responses for motion state changes in traditional methods is solved, and more accurate motion detection results are achieved.
Patent Information
- Application Number
- CN202510123166.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
Traditional motion scene detection methods are difficult to respond to changes in motion state in a timely manner, resulting in inaccurate detection results.
A scene detection model training method is provided, by obtaining camera pose data and annotating scene category labels, performing data amplification and model training, using convolutional neural network model for scene category prediction, and adjusting model parameters based on the loss function.
It improves the accuracy of motion detection results, can respond to changes in motion state in a timely manner, and optimizes the training process of scene detection model.
Smart Images

Figure CN120047771A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of motion detection technologies, and in particular, to a method for training a scene detection model, a scene detection method, and a device therefor. Background Art
[0002] With the development of short videos and self-media, cameras have now become an indispensable part of electronic devices, and people are also accustomed to using cameras to record their lives by shooting videos. To provide users with a better camera shooting experience, manufacturers have developed a series of algorithms for videos to make the captured videos smoother, more stable, and the operation process more "idiot-proof". The effects of these algorithms highly depend on the judgment of the motion scene, which is a common shooting scene. However, traditional solutions are difficult to respond promptly to changes in the motion state, resulting in inaccurate motion scene detection results. Summary of the Invention
[0003] Based on this, to address the above technical problems, it is necessary to provide a method for training a scene detection model, a scene detection method, and a device therefor that can ensure the accuracy of motion detection results.
[0004] In a first aspect, this application provides a method for training a scene detection model, the method comprising:
[0005] Obtaining camera pose data and the corresponding labeled scene category label for the camera pose data;
[0006] Performing data augmentation on the camera pose data to obtain a target data set;
[0007] Inputting the target data set into an initial scene detection model for scene category prediction to obtain a predicted scene category label;
[0008] Training the initial scene detection model based on a loss function according to the labeled scene category label and the predicted scene category label to obtain a target scene detection model.
[0009] In one embodiment, obtaining camera pose data and the corresponding labeled scene category label for the camera pose data comprises:
[0010] Obtaining the camera pose data collected when shooting video frames;
[0011] Using a decision tree to pre-label the camera pose data of consecutive multiple video frames to obtain a labeled scene category label.
[0012] In one embodiment, training the initial scene detection model based on a loss function according to the labeled scene category label and the predicted scene category label to obtain a target scene detection model comprises:
[0013] When the predicted scene category label of the current video frame is different from the labeled scene category label, obtain a first frame number and a second frame number; the first frame number is the frame number of the current video frame, and the second frame number is the frame number of the video frame corresponding to the labeled scene category label;
[0014] Wherein, the loss weight corresponding to the current video frame is positively correlated with the difference between the first frame number and the second frame number.
[0015] In one embodiment, the number of consecutive video frames is greater than a preset number, and the preset number is related to the frame rate.
[0016] In one embodiment, training an initial scene detection model based on a loss function according to the labeled scene category label and the predicted scene category label to obtain a target scene detection model, including:
[0017] When the predicted scene category label of the current video frame is different from the predicted scene category label of the previous video frame, regard the current video frame as the current mutation frame and obtain the frame number of the previous mutation frame;
[0018] If the difference between the frame number of the current mutation frame and the frame number of the previous mutation frame is less than the preset number, determine the loss weight between the predicted scene category label of the current mutation frame and the predicted scene category label of the previous mutation frame based on the loss function.
[0019] In one embodiment, the labeled scene category label represents the scene category of camera pose data; the scene categories include a motion scene and a panoramic shooting scene; the motion scene includes a walking scene, a running scene, and a static holding scene.
[0020] In one embodiment, performing data augmentation on the camera pose data to obtain a target data set, including:
[0021] Performing data augmentation on the camera pose data by using a data augmentation strategy to obtain a target data set; wherein, the data augmentation strategy includes multiple data augmentation methods executed in sequence; the data augmentation methods are related to the characteristics of the camera pose data and the prior information of the distribution of camera pose data in different scene categories.
[0022] In one embodiment, the multiple data augmentation methods include one or more of inter-axis data swapping, data noise addition, and scene continuation between data;
[0023] The execution order of inter-axis data swapping is before data noise addition, and the execution order of data noise addition is before scene continuation between data.
[0024] In one embodiment, inter-axis data random swapping includes random swapping of the X-axis data and the Y-axis data in the camera pose data;
[0025] Data noise addition includes adding noise to the camera pose data of the loop shooting scenario, walking scenario, and running scenario using random noise; among them, the scenario category of the camera pose data before and after noise addition remains unchanged;
[0026] Scene continuity between data means that the scenario category of the current camera pose data in the same video frame is the same as that of the adjacent camera pose data, and the adjacent camera pose data is the camera pose data adjacent to the current camera pose data.
[0027] In one embodiment, the timestamp of the adjacent camera pose data is related to the timestamp of the video frame to which it belongs and the sampling frequency.
[0028] In one embodiment, inputting the target data set into the initial scene detection model for scene category prediction to obtain the predicted scene category label includes:
[0029] Integrating the data in the target data set to obtain integrated data;
[0030] Inputting the integrated data into the initial scene detection model to obtain the predicted scene category label.
[0031] In one embodiment, the initial scene detection model is a convolutional neural network model.
[0032] In one embodiment, the initial scene detection model includes at least two convolutional layers; the activation function used in the at least two convolutional layers is the ReLU function.
[0033] In one embodiment, the camera pose data includes inertial measurement data.
[0034] In one embodiment, the inertial measurement data includes at least gyroscope data.
[0035] In a second aspect, the present application provides a scene detection method, and the method includes:
[0036] Obtaining the camera pose data collected when shooting a video;
[0037] Inputting the camera pose data into the target scene detection model obtained by the above-mentioned scene detection model training method, and obtaining the scene detection result output by the target scene detection model.
[0038] In one embodiment, inputting the camera pose data into the target scene detection model includes:
[0039] When the camera pose data is gyroscope data, performing integral processing on the gyroscope data and inputting the gyroscope data after integral processing into the target scene detection model.
[0040] In one embodiment, inputting the camera pose data into the target scene detection model includes:
[0041] Inputting the camera pose data of the current video frame in the video and the scene detection results of multiple video frames before the current video frame into the target scene detection model.
[0042] In one embodiment, the method further includes:
[0043] Obtaining a video anti-shake strategy corresponding to the scene detection result;
[0044] Performing anti-shake processing on the video according to the video anti-shake strategy to obtain a target video.
[0045] In a third aspect, the present application provides a scene detection model training device, which includes:
[0046] A data acquisition module, configured to acquire camera pose data and the labeled scene category label corresponding to the camera pose data;
[0047] A data augmentation module, configured to perform data augmentation on the camera pose data to obtain a target data set;
[0048] A scene prediction module, configured to input the target data set into an initial scene detection model for scene category prediction to obtain a predicted scene category label;
[0049] A model training module, configured to train the initial scene detection model based on the labeled scene category label and the predicted scene category label based on a loss function to obtain a target scene detection model.
[0050] In a fourth aspect, the present application provides a scene detection device, which includes:
[0051] A pose data acquisition module, configured to acquire camera pose data collected during video shooting;
[0052] A model input module, configured to input the camera pose data into the target scene detection model obtained by the above-mentioned scene detection model training method to obtain the scene detection result output by the target scene detection model.
[0053] In a fourth aspect, the present application provides an electronic device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the methods in the above aspects are implemented.
[0054] In a fifth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the methods in each aspect are implemented.
[0055] In a sixth aspect, the present application provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the methods in various aspects.
[0056] The above scene detection model training method, scene detection method, scene detection model training device, scene detection device, electronic device, computer-readable storage medium, and computer program product obtain camera pose data and the corresponding labeled scene category labels, then perform data augmentation on the camera pose data to obtain a target data set, input the target data set into an initial scene detection model for scene category prediction to obtain predicted scene category labels, and then train the initial scene detection model based on the labeled scene category labels and the predicted scene category labels using a loss function to obtain a target scene detection model, optimizing the training process of the target scene detection model and improving the accuracy of the scene detection results. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0058] Figure 1 It is an application environment diagram of the scene detection model training method and the scene detection method in an embodiment;
[0059] Figure 2 It is a flowchart of the scene detection model training method in an embodiment;
[0060] Figure 3 It is a flowchart of pre-labeling in an embodiment;
[0061] Figure 4 It is a flowchart of obtaining predicted scene category labels in an embodiment;
[0062] Figure 5 It is a flowchart of model training in an embodiment;
[0063] Figure 6 It is a network structure diagram of a convolutional neural network model in an embodiment;
[0064] Figure 7 It is a flowchart of obtaining loss weights in an embodiment;
[0065] Figure 8 It is a flowchart of a loss function for restricting prediction mutations in an embodiment;
[0066] Figure 9 It is a schematic flowchart of a scene detection method in an embodiment;
[0067] Figure 10 It is a structural block diagram of a scene detection model training device in an embodiment;
[0068] Figure 11 It is a structural block diagram of a scene detection device in an embodiment;
[0069] Figure 12 It is an internal structural diagram of an electronic device in an embodiment. Detailed implementation manners
[0070] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0071] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0072] It can be understood that terms such as "first" and "second" in the present application are only used to distinguish similar objects and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. It can be understood that "at least one" means one or more, and "a plurality" means two or more. For example, a plurality of video frames means two or more frames.
[0073] As used herein, the singular forms "a", "an" and "the" may also include the plural forms unless the context clearly dictates otherwise. It should also be understood that the terms "comprise / include" or "have" etc. specify the presence of the stated features, wholes, steps, operations, components, parts or combinations thereof, but do not exclude the possibility of the presence or addition of one or more other features, wholes, steps, operations, components, parts or combinations thereof. At the same time, the term "and / or" used in this specification includes any and all combinations of the related listed items.
[0074] Taking an electronic device as a mobile phone as an example, in order to provide users with a better video shooting experience, mobile phone manufacturers have developed a series of algorithms for videos, making the captured videos smoother, more stable, and the operation process more "idiot-proof". The effects of these algorithms generally highly depend on the judgment of the mobile phone's motion scenarios. For example, for the anti-shake algorithm, different motion scenarios of the mobile phone often mean different optimization goals. In scenarios of small-amplitude and high-frequency jitters, such as walking or running, users hope that the captured video should remain as stationary as possible. In some scenarios of large-amplitude and low-frequency, such as panning or camera movement, users often hope that the captured video can be as stable and uniform as possible to reflect the real movement. Therefore, in order to enable the mobile phone to capture the effects desired by users, mobile phone manufacturers generally first detect the current motion scenario of the camera, and then design different optimization goals for different motion scenarios. That is, the detection of the motion scenario is the cornerstone of the video shooting effect.
[0075] Current scene detection can generally be divided into three types of solutions: First, the decision-tree-based solution. The decision-tree method essentially detects the motion state by manually designing a decision tree and using the manually selected decision tree. Second, traditional machine learning solutions (such as support vector machines, k-nearest neighbors, naive Bayes, etc.). These algorithms manually extract features and then classify the scenes based on the manually extracted features. In recent years, there have also been many deep learning-based scene detection solutions, which can extract deeper implicit features and have generalization capabilities. However, the decision-tree-based solution, although simple to implement and deploy, has poor robustness and is prone to misjudgment of scenes. The traditional machine learning-based solutions have good robustness, but it is difficult to find features that can quickly and accurately distinguish different motion states. The deep learning-based solutions have good effects, but the design and training of the model and the production of data are all difficult. In addition, during real-time processing, it is difficult for the above algorithms to respond promptly to changes in the motion state.
[0076] Based on the above traditional technologies, the embodiments of this application propose a method for training a scene detection model related to camera motion scene detection and a motion scene detection method, which can respond promptly to changes in the motion state and ensure the accuracy of the motion detection results. It should be noted that the beneficial effects or the technical problems solved by the embodiments of this application are not limited to this one, and there may also be other implicit or related problems. For specific details, please refer to the description of the following embodiments.
[0077] The scene detection model training method and the motion scene detection method provided by the embodiments of this application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed on the cloud or other network servers. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted device can be a Virtual Reality (VR) device, an Augmented Reality (AR) device, smart glasses, etc. The server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.
[0078] Before introducing the specific embodiments of the present application, the professional terms involved in the present application are first explained:
[0079] IMU: Inertial Measurement Unit, Inertial Measurement Unit.
[0080] EIS: Electronic Image Stabilization, Electronic Image Stabilization.
[0081] Gyro: Gyroscope, Gyroscope.
[0082] CNN: Convolution Neural Network, Convolution Neural Network.
[0083] RNN: Recurrent Neural Network, Recurrent Neural Network.
[0084] The technical solution of the present application and how the technical solution of the present application solves the above technical problems are described in detail below with specific embodiments. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.
[0085] In an exemplary embodiment, as Figure 2 shown, a method for training a scene detection model is provided. Taking the application of this method to the terminal as an example, it can be understood that this method can also be applied to the server, and can also be applied to a system including the terminal and the server, and is implemented through the interaction between the terminal and the server. The method includes the following steps 202 to step 206.
[0086] Wherein:
[0087] Step 202: Obtain camera pose data and the corresponding labeled scene category label of the camera pose data.
[0088] Among them, the camera pose data may be data of the rotation state of the terminal's camera during shooting, and this exemplary embodiment is not limited thereto; for example, the camera pose data is data of the rotation state of the camera during video shooting.
[0089] Specifically, the embodiments of the present application can obtain camera pose data and the corresponding labeled scene category label of the camera pose data. Among them, the labeled scene category label corresponding to the camera pose data can be obtained through labeling. Exemplarily, the labeled scene category label can be understood as a scene detection label for indicating the scene category.
[0090] The scene detection in the embodiments of the present application can be understood as a temporal scene detection task. To address the problem of difficult data production for the temporal scene detection task, in a possible implementation, the present application can distinguish the camera pose data based on the data characteristics in the time series to obtain the corresponding labeled scene category label. It can be understood that the above-mentioned acquisition of the labeled scene category label can also adopt other forms, not limited to the forms already mentioned in the above embodiments, as long as it can achieve the function of determining the scene category corresponding to the camera pose data.
[0091] The present application realizes the acquisition of the data set by obtaining the camera pose data, and realizes the pre-labeling (such as initial labeling) of the scene categories of the data in the data set by obtaining the labeled scene category label corresponding to the camera pose data, optimizes the data production of the data set, and reduces the labeling time.
[0092] Step 204: Perform data augmentation on the camera pose data to obtain a target data set.
[0093] Specifically, in the case of obtaining the camera pose data, data augmentation can be performed on the camera pose data to obtain a target data set. Exemplarily, the target data set represents a data set for training a scene detection model.
[0094] Among them, data augmentation can be used to expand the data in the target data set. For example, without substantially adding data to the target data set, the existing camera pose data in the target data set is used to generate more valuable camera pose data, thereby expanding the data in the target data set and solving the problem of difficult data production.
[0095] Exemplarily, the data augmentation method includes, but is not limited to, transforming the camera pose data in the target dataset to increase the diversity of the data in the dataset. Optionally, multiple data augmentation methods can be designed according to prior knowledge, thereby reducing the dependence of scene detection on the amount of data. It can be understood that the above data augmentation method can also adopt other forms, not limited to the forms already mentioned in the above embodiments, as long as it can achieve the function of expanding the data in the target dataset.
[0096] In this application, a target dataset is obtained by performing data augmentation on camera pose data, reducing the demand for data in scene detection, improving the stability of scene detection model training, and being beneficial to the accuracy of scene detection results. Among them, data augmentation can reduce the dependence of scene detection on the amount of data and further improve the robustness and accuracy of the solution.
[0097] Step 206: Input the target dataset into the initial scene detection model for scene category prediction to obtain a predicted scene category label.
[0098] Specifically, the target dataset can be input into the initial scene detection model for scene category prediction to obtain a prediction result, where the prediction result can include a predicted scene category label.
[0099] Exemplarily, the initial scene detection model can be understood as a scene detection model to be trained. In some embodiments, the type of the initial scene detection model can be determined according to real-time requirements and the feasibility of mobile AI (Artificial Intelligence) deployment. For example, a model with a simple overall network structure can be determined as the initial scene detection model, or a model with a deployment-friendly operation and suitable for deployment on mobile devices can be determined as the initial scene detection model.
[0100] Optionally, the initial scene detection model can be a deep learning model. It can be understood that deep learning models include, but are not limited to: Deep Neural Network (DNN), Convolutional Neural Networks (CNN), Recurrent Neural Network (RNN), Dynamic Bayesian Network (DBN), and stacked auto-encoder network (SAE) models. Of course, in other embodiments of this application, other machine learning models or even models combined with multiple models can also be used.
[0101] Step 208: Based on the labeled scene category label and the predicted scene category label, train the initial scene detection model using a loss function to obtain the target scene detection model.
[0102] Specifically, the value of the loss function can be calculated according to the labeled scene category label and the predicted scene category label, and then the parameters of the initial scene detection model can be adjusted based on the value of the loss function to obtain the target scene detection model.
[0103] Among them, since it is difficult to collect data for each category evenly, this application proposes to use a loss function to supervise the training. In the embodiments of this application, the loss function can be used to calculate the degree of difference between the predicted value (predicted scene category label) and the true value (labeled scene category label). Exemplarily, the loss function can include a classification task loss function. Further, in the process of training the initial scene detection model based on the loss function, the mutation points can be processed, for example, corresponding constraints can be imposed in the loss function to avoid frequent jumps in the detected results.
[0104] The above scene detection model training method realizes pre-labeling by obtaining the labeled scene category label corresponding to the camera pose data, and reduces the requirement of scene detection for the data volume by data augmentation of the camera pose data, solves the problem of difficult data production, and improves the robustness and accuracy of the model. Further, based on the initial scene detection model and the loss function, etc., the model is trained, so that the target scene detection model trained in the embodiments of this application has a small computational overhead, can be combined with most camera algorithms, and can be applied to systems that require real-time processing, and can respond to changes in the motion state in a timely manner.
[0105] In one embodiment, the camera pose data includes inertial measurement data.
[0106] Specifically, the camera pose data in the embodiments of this application can include inertial measurement data. The embodiments of this application use inertial measurement data as the data source for model training and scene detection, which is beneficial to simple scene judgment and the detection results are relatively accurate.
[0107] Exemplarily, the camera pose data is inertial measurement data collected by the inertial measurement unit IMU of the terminal when shooting a video. The inertial measurement data can include the angular velocity data of the camera along three direction axes. Among them, the inertial measurement unit IMU is a sensor that can detect the pose data of the device in real time, outputs the angular velocities of the x, y, and z axes in the current reference coordinate system of the device, and the current pose of the device can be obtained by integrating the angular velocities. Further, taking the camera pose data as inertial measurement data as an example, the inertial measurement data can include accelerometer data (Acc) and gyroscope data (Gyro).
[0108] In one embodiment, the inertial measurement data includes at least gyroscope data.
[0109] Specifically, the present application proposes to use gyroscope data for scene detection, which helps to reduce the computational overhead. It should be noted that using accelerometer data consumes a large amount of power. The present application can use gyroscope data for data collection and production of the dataset, which is convenient for model training and applicable to terminal deployment, and can be combined with most camera algorithms.
[0110] Regarding the data collection and data production in the model training process of the present application, in an exemplary embodiment, as Figure 3 shown, step 202 may include steps 302 to 304. Among them:
[0111] Step 302, obtain the camera pose data collected when shooting video frames.
[0112] Specifically, the camera pose data may be the data of the rotation state of the terminal's camera when shooting video frames. Optionally, the camera pose data is the data of the rotation state of the terminal's camera when shooting the current video frame, and the current video frame may refer to the image frame corresponding to the current moment when the terminal is shooting a video, where the current video frame may be abbreviated as the current frame.
[0113] Step 304, use a decision tree to pre-label the camera pose data of multiple consecutive video frames to obtain a labeled scene category label.
[0114] Specifically, the present application designs a pre-labeling scheme in the data labeling process to reduce the labeling time, and the pre-labeling can be implemented through a decision tree. Optionally, the pre-labeling can be understood as an automatic pre-labeling based on the decision tree to perform an initial labeling of the data.
[0115] Exemplarily, a decision tree can be used to pre-label the camera pose data of multiple video frames. In practical applications, most frames in a group of videos are scenes that are easy to distinguish. The present application can reduce the labeling time by pre-labeling through a decision tree. In addition, after pre-labeling, the frame categories with incorrect labels can be adjusted to improve the accuracy of the labeled scene category labels.
[0116] In addition, the object of pre-labeling using a decision tree in the present application is the camera pose data of multiple consecutive video frames, that is, to label the scene category labels for multiple video frames with a long duration to avoid the problem of easy misdetection in the short-time window during scene detection. It should be noted that multiple consecutive video frames may refer to multiple video frames in the same time window, and the length of this time window can be set according to requirements.
[0117] It can be understood that taking the camera pose data including gyroscope data (Gyro) as an example, a 10-second 30-fps video generally has 120,000 groups of Gyro data. If all the collected Gyro data are manually labeled line by line, the labeling time will be unacceptable. In this regard, the present application proposes an automatic pre-labeling based on a decision tree to initially label the data, which can then quickly and accurately distinguish the time series with a long duration and obvious Gyro features in a video. After that, the frame categories with incorrect labels are manually adjusted.
[0118] The above scene detection model training method reduces the labeling time through a pre-labeling scheme in the data labeling process based on a decision tree, and pre-labels the camera pose data of multiple consecutive video frames at the same time. It not only solves the problem of difficult data production, but also ensures the accuracy of the detection results, and can be applied to systems that require real-time processing. It should be noted that the above pre-labeling scheme can be applied to scene detection based on deep learning.
[0119] Regarding the number of multiple consecutive video frames, in some embodiments, the number of multiple consecutive video frames is greater than a preset number, and the preset number is related to the frame rate.
[0120] Specifically, to ensure the accuracy of scene category detection, the present application proposes that the scene category can be determined only for the camera pose data that remains stable for a period of time (representing the same state), and the above period of time can be represented by a time window.
[0121] Exemplarily, the continuous time window can be ensured by the number of multiple consecutive video frames to avoid the problem of easy misdetection in the short-time window during scene detection. Among them, the number of multiple consecutive video frames can be greater than a preset number, and the preset number is related to the frame rate, that is, the preset number can be determined according to the frame rate. Optionally, the preset number can be 20, that is, the continuous time window greater than the state of 20 frames can be detected and marked with the scene category label.
[0122] Furthermore, regarding the labeled scene category label in the embodiments of the present application, in some embodiments, the labeled scene category label represents the scene category of the camera pose data; the scene category includes a motion scene and a panoramic shooting scene; the motion scene includes a walking scene, a running scene, and a static holding scene.
[0123] Specifically, the labeled scene category label in the embodiments of the present application can represent the scene category of the camera pose data. Optionally, the scene category can include a motion scene corresponding to the motion state and a panoramic shooting scene corresponding to the panoramic shooting state, where the motion scene can include a walking scene, a running scene, and a static holding scene.
[0124] It can be understood that the scene categories predicted in the embodiments of the present application are the motion state and the panoramic shooting state respectively. The motion state may include walking, running, and static holding. The panoramic shooting state may represent video panoramic shooting (for example, the user holds the device and rotates it in one direction to capture the surrounding images). Further, the panoramic shooting state in the embodiments of the present application refers to that the data of the three direction axes all represent panoramic shooting to ensure the detection accuracy of the panoramic shooting scene. Taking the camera attitude data using gyroscope data as an example, the panoramic shooting state includes that the X, Y, and Z axes are all in panoramic shooting.
[0125] It should be noted that corresponding to the labeled scene category tags, the predicted scene category tags in the embodiments of the present application are used to represent the predicted scene categories of the camera attitude data. Among them, the predicted scene categories may also include a motion scene and a panoramic shooting scene, and the motion scene includes a walking scene, a running scene, and a static holding scene.
[0126] As described above, the embodiments of the present application can detect multiple camera motion scenes, including three basic motion states: static holding, walking, and running, as well as the panoramic shooting states in three different directions.
[0127] In order to train a stable scene detection model, the present application also proposes a data augmentation strategy. Regarding the data augmentation in the embodiments of the present application, in one of the embodiments, the camera attitude data is augmented to obtain a target data set, including:
[0128] The camera attitude data is augmented using a data augmentation strategy to obtain a target data set; wherein, the data augmentation strategy includes multiple data augmentation methods executed in sequence; the data augmentation methods are related to the characteristics of the camera attitude data and the prior information of the distribution of the camera attitude data in different scene categories.
[0129] Specifically, the camera attitude data can be augmented using a data augmentation strategy to obtain a target data set; the data augmentation strategy in the embodiments of the present application may include multiple data augmentation methods executed in sequence. By executing these data augmentation methods in sequence, it is ensured that the augmented data is different from the original data, thereby not only reducing the dependence on the data volume for scene detection and model training, but also further improving the robustness and accuracy of scene detection.
[0130] Among them, the data augmentation methods are related to the characteristics of the camera attitude data and the prior information of the distribution of the camera attitude data in different scene categories. Taking the camera attitude data using gyroscope data as an example, various data augmentation methods can be determined according to the characteristics of the gyroscope data and the prior of the gyroscope data distribution in different motion states. It can be understood that the above method for determining the data augmentation methods can also adopt other forms, rather than being limited to the forms already mentioned in the above embodiments, as long as it can achieve the relevant functions of completing data augmentation.
[0131] In one embodiment, the multiple data augmentation methods include one or more of inter-axis data interchange, data noise addition, and scene continuity between data;
[0132] The execution order of the inter-axis data interchange is before the data noise addition, and the execution order of the data noise addition is before the scene continuity between data.
[0133] Specifically, the multiple data augmentation methods in the embodiments of the present application may include, but are not limited to, inter-axis data interchange, data noise addition, and scene continuity between data.
[0134] Among them, the inter-axis data interchange may refer to the interchange of data of each direction axis in the camera pose data. The data noise addition may refer to appropriately adding noise to the camera pose data. The scene continuity between data may refer to confirming that the scene categories corresponding to adjacent camera pose data are the same.
[0135] Further, the execution order of the inter-axis data interchange is before the data noise addition, and the execution order of the data noise addition is before the scene continuity between data, that is, in the data augmentation process, the inter-axis data interchange is first executed, then the data noise addition is executed, and finally the scene continuity between data is executed. It can be understood that the limitation of the execution order between the above various data augmentation methods in the present application is to ensure that the augmented data is different from the original data. In practical applications, the execution order between the various data augmentation methods can be adjusted according to requirements to ensure the training of a stable scene detection model.
[0136] In a possible implementation, the random inter-axis data interchange includes the random interchange of the X-axis data and the Y-axis data in the camera pose data;
[0137] The data noise addition includes adding noise to the camera pose data of the panoramic shooting scene, walking scene, and running scene by using random noise; among them, the scene category of the camera pose data before and after the noise addition remains unchanged;
[0138] The scene continuity between data includes that the scene category of the current camera pose data in the same video frame is the same as that of the adjacent camera pose data, and the adjacent camera pose data is the camera pose data adjacent to the current camera pose data.
[0139] Specifically, the random inter-axis data interchange in the present application may include the random interchange of the X-axis data and the Y-axis data in the camera pose data; taking the camera pose data using gyroscope data as an example, the random inter-axis data interchange may refer to the random interchange of the X data and the Y-axis data (i.e., X, Y-axis shuffle). The gyroscope data records the angular velocity of the three direction axes. When the user's terminal is a mobile phone and the video is shot in both landscape and portrait modes, the gyroscope data collected is that the X and Y axes are opposite. Therefore, for all the collected data, the X-axis and Y-axis data can be randomly interchanged.
[0140] Data noise addition may include adding noise to the camera pose data of non-static scenes using random noise, that is, adding noise to the camera pose data of panoramic shooting scenes, walking scenes, and running scenes. Among them, the scene category of the camera pose data before and after noise addition remains unchanged, and the noise amplitude of adding random noise can be set according to requirements.
[0141] Taking the camera pose data using gyroscope data as an example, data noise addition may refer to appropriately adding noise to the gyroscope data. For scene detection, the value of the gyroscope data has little significance for the detection of scene categories other than static. Therefore, this application proposes that the gyroscope data with a label other than static can be added with noise. For example, random noise of ±20% is added to the value of the gyroscope data. Among them, the state of the gyroscope data before adding noise remains the same after adding noise, that is, the scene category remains unchanged. It should be noted that the above 20% may refer to a limited amplitude of 20%, and this value can change (but not too abruptly). In addition, the way of adding noise in the embodiments of this application can be randomly added, and this example embodiment is not limited thereto.
[0142] Further, the scene continuation between data may include that the scene category of the current camera pose data in the same video frame is the same as that of the adjacent camera pose data, and the adjacent camera pose data is the camera pose data adjacent to the current camera pose data. It can be understood that in order to ensure that the states corresponding to the first row of data to the last row of data in the same video frame are the same (that is, the scene category is the same), this application proposes a moderately relaxed scene continuation strategy. For example, when the motion state of the current video frame is known, then the motion state of the adjacent camera pose data in this frame is consistent with that of the current camera pose data.
[0143] In one of the embodiments, the timestamp of the adjacent camera pose data is related to the timestamp of the video frame to which it belongs and the sampling frequency.
[0144] Specifically, in the data augmentation method of scene continuation between data, the camera pose data can be obtained using the timestamp to appropriately relax; among them, the timestamp of the adjacent camera pose data is related to the timestamp of the video frame to which it belongs and the sampling frequency.
[0145] Taking the camera pose data using gyroscope data as an example, the gyroscope data is obtained using the timestamp. When the motion state of a certain frame is known, then the state of the data corresponding to the timestamp of this frame ± sampling frequency / 2 is still the same motion state. It should be noted that the above timestamp ± sampling frequency / 2 is determined according to requirements, which can ensure that the states from the first row to the last row in the current video frame are the same.
[0146] It can be understood that when applied to scene detection based on deep learning, through the above data augmentation strategy, the present application can greatly reduce the data requirements of the deep learning-based scene detection solution. Among them, designing multiple data augmentation methods according to prior knowledge can reduce the dependence on the data volume of the deep learning solution, and at the same time, these data augmentation methods also further improve the robustness and accuracy of the algorithm.
[0147] In one embodiment, as Figure 4 shown, step 206 may include steps 402 to 404. Among them:
[0148] Step 402, integrate the data in the target dataset to obtain integrated data.
[0149] Specifically, before inputting the target dataset into the initial scene detection model, the data in the target dataset can be integrated to obtain integrated data. Among them, using the integrated data for model training is beneficial to the detection of the panoramic shooting scene. In addition, in practical applications, the integrated data needs to be used after anti-shake processing, that is, integrating the data is beneficial to actual scene detection.
[0150] Step 406, input the integrated data into the initial scene detection model to obtain the predicted scene category label.
[0151] Specifically, the integrated data can be input into the initial scene detection model to obtain the predicted scene category label. Taking the camera pose data using gyroscope data as an example, the motion state and panoramic shooting state of the integrated gyroscope data can be classified. It can be understood that the present application can also not integrate the data in the target dataset, but directly input the target dataset into the initial scene detection model for model training. In addition, during the process of using the target scene detection model for scene detection, the unintegrated acquired data can also be directly used.
[0152] To further illustrate the above data annotation and the data processing process in the training stage, a specific example is given below. As Figure 5 shown, taking the model training process applied to the terminal and the camera pose data including inertial measurement data (IMU data) as an example, in the data annotation stage, when the terminal acquires IMU data, the IMU data can be dumped (Dump IMU data), and then pre-annotated, and manually adjusted (that is, manually adjust the incorrectly annotated scene categories), so as to obtain a dataset containing scene category labels. In addition, the data processing in the training stage may include adding IMU data, and then performing data augmentation and integration, and inputting the processed data (target dataset) into the network. Among them, the network may refer to the initial scene detection model (such as a CNN model), and the IMU data in the training stage at least includes gyroscope data.
[0153] In summary, in the data annotation process of this application, a pre-annotation scheme is designed to reduce the annotation time, and various data augmentation methods are designed based on prior knowledge to reduce the dependence of scene detection on the amount of data. At the same time, these data augmentation methods also further improve the robustness and accuracy of the algorithm.
[0154] Regarding the type of the scene detection model in this application, in one embodiment, the initial scene detection model is a convolutional neural network model.
[0155] Specifically, this application uses a convolutional neural network model (CNN model) to perform scene detection, implementing a scene detection solution based on deep learning. Among them, the CNN model, as a network model with a simple network structure, is suitable for deployment on mobile devices. Exemplarily, the target data set can be input into the convolutional neural network model for scene category prediction, and then the predicted scene category label can be obtained. It can be understood that the model type of the target scene detection model in the embodiments of this application is the same as that of the initial scene detection model, and it can be a convolutional neural network model.
[0156] In one embodiment, the initial scene detection model includes at least two convolutional layers; the activation function used by the at least two convolutional layers is the ReLU function.
[0157] Specifically, the initial scene detection model in the embodiments of this application can include at least two convolutional layers (Conv), and the activation function used by the at least two convolutional layers is the ReLU function. Among them, after preprocessing the target data set (such as feature fusion and tensor expansion), it can be input into the convolutional layer for convolutional processing to generate an intermediate matrix, and then matrix multiplication processing and dimension rearrangement are performed on the intermediate matrix to generate a predicted scene category label. Among them, the activation function used in the above process is the ReLU function.
[0158] Exemplarily, the at least two convolutional layers include a first convolutional layer and a second convolutional layer, and the output of the first convolutional layer is connected to the input of the second convolutional layer. It can be understood that the model structure of the above initial scene detection model (such as the number of convolutional layers) and the type of activation function, etc. can also adopt other forms, rather than being limited to the forms mentioned in the above embodiments, as long as it can complete the corresponding processing process.
[0159] It should be noted that in the actual use of the target scene model, the input data of the target scene model includes the camera pose data collected at the current moment, and can also include the scene detection data (such as scene detection results) corresponding to the camera pose data before the current moment, so as to avoid sudden state changes in a small time interval and improve the accuracy of scene detection.
[0160] To further illustrate the network structure of the above convolutional neural network model, the following will be described with specific examples. Taking gyroscope data as an example for camera pose data, as Figure 6 shown, the convolutional neural network model may include two convolutional layers, and the activation function used in the two convolutional layers is the ReLU function. Among them, gyro0 represents the data input into the convolutional neural network model. Exemplarily, during the model training phase, gyro0 may represent the data in the target dataset, and during the model usage phase, gyro0 may be the collected gyroscope data (which may be gyroscope data that has been integrated). Further, Concat may refer to feature fusion, Unsqueeze represents inserting a new dimension to expand the tensor (adding a dimension to the matrix); Conv refers to the convolutional layer, Relu refers to the activation function, Gemm refers to the global matrix-to-matrix multiplication (multiplying two input matrices together to obtain an output matrix), such as General Matrix Multiplication (GEMM); Reshape refers to dimension rearrangement.
[0161] In addition, p_panning represents the probability of the panning state, and p_mode represents the probability of the motion state. It can be understood that the matrix representing the panning state in the model can be a 3*2 matrix, where 3 represents three direction axes and 2 represents two probabilities (representing panning and not panning respectively), and the matrix representing the motion state in the model can be a 1*3 matrix, where 3 represents three probabilities (walking, running, static holding). virtual_panning and virtual_mode represent the judgment results as probabilities (such as 0.8, 0.2), rather than the real judgment results (such as 1, 0). It can be understood that when the above CNN model is applied to actual scene detection, the above virtual_panning and virtual_mode may also include the scene detection data corresponding to the camera pose data before the current moment (such as the scene detection result). Taking the initialization data of the scene detection data of the first three frames before the current video frame as an example, it may include the data after the initialization of the scene detection data of the first three frames.
[0162] The overall network structure of this application is simple, and the operations used are all deployment-friendly and suitable for deployment on mobile devices. This means that this solution can be applied to systems that require real-time processing, and due to the small computational overhead of this solution, this solution can be combined with most camera algorithms.
[0163] Regarding the model training in this application, a loss function can be used to supervise the training. Among them, the Focal Loss function can be used as the loss function. The formula of the Focal Loss function is as follows: FL(p t ) = -αt (1 - p t ) γ log(p t ). Among them, p t is the predicted probability of the model for the correct category; α t is the balancing factor used to balance the influence of positive and negative samples; γ is the adjustment factor used to control the degree of attention to difficult-to-classify samples.
[0164] The output of the embodiment of this application is the scene category of the motion state and the panning state. Correspondingly, there are two Focal Losses, FL(p_mode) and FL(p_pan). Among them, FL(p_mode) corresponds to the motion state, and FL(p_pan) corresponds to the panning state. It can be understood that the above loss function can also adopt other forms, not limited to the forms mentioned in the above embodiments, as long as it can achieve the function of the function used to calculate the loss.
[0165] Furthermore, regarding the use of the loss function in the model training of this application, taking the camera pose data collected when shooting video frames as an example, in one of the embodiments, based on the labeled scene category label and the predicted scene category label, the initial scene detection model is trained based on the loss function to obtain the target scene detection model, including:
[0166] When the predicted scene category label of the current video frame is different from the labeled scene category label, obtain the first frame number and the second frame number; the first frame number is the frame number of the current video frame, and the second frame number is the frame number of the video frame corresponding to the labeled scene category label;
[0167] Among them, the loss weight corresponding to the current video frame is positively correlated with the difference between the first frame number and the second frame number.
[0168] Specifically, if the predicted scene category label of the current video frame is different from the labeled scene category label, it can indicate that the prediction result of the current video frame is incorrect. At this time, the frame number of the current video frame (the first frame number) can be recorded, and the frame number of the video frame corresponding to the labeled scene category label (the second frame number) can be recorded, and the loss weight corresponding to the current video frame can be configured accordingly to meet the situation where scene detection has high requirements for the delay of the detection result.
[0169] Exemplarily, the loss weight corresponding to the current video frame is positively correlated with the difference between the first frame number and the second frame number. That is, the greater the difference between the first frame number and the second frame number, the greater the weight given to the loss of the current video frame. Conversely, the smaller the difference, the smaller the corresponding loss weight, so as to avoid the situation where it takes many frames to recognize after the state switch. Taking the frame number of the current video frame as f_v and the frame number of the frame with a real state change in the label as f_r (i.e., the frame number of the video frame corresponding to the labeled scene category label) as an example, the greater the difference between f_v and f_r, the greater the weight given to the loss of the current video frame. Conversely, the smaller the difference, the smaller the corresponding loss weight. Through the above targeted constraints in the loss function, the model has characteristics such as high sensitivity and high accuracy.
[0170] Further, to limit the prediction of mutations, taking the camera pose data collected when shooting video frames and the number of consecutive video frames being greater than a preset number as an example, in one embodiment, as Figure 7 shown, step 208 may include steps 502 to 504. Among them:
[0171] Step 502, when the predicted scene category label of the current video frame is different from the predicted scene category label of the previous video frame, take the current video frame as the current mutation frame and obtain the frame number of the previous mutation frame.
[0172] Specifically, if the predicted scene category label of the current video frame is different from the predicted scene category label of the previous video frame, it can indicate that the state detected in the current video frame is different from the state detection result of the previous frame. In this application, the current video frame can be taken as the current mutation frame and the frame number of the previous mutation frame can be obtained.
[0173] Step 504, if the difference between the frame number of the current mutation frame and the frame number of the previous mutation frame is less than the preset number, determine the loss weight between the predicted scene category label of the current mutation frame and the predicted scene category label of the previous mutation frame based on the loss function.
[0174] Specifically, when the difference between the frame number of the current mutation frame and the frame number of the previous mutation frame is less than the preset number, the loss weight between the predicted scene category label of the current mutation frame and the predicted scene category label of the previous mutation frame can be determined based on the loss function, thereby being able to limit the frequent jump of the detected results and ensure that the duration window of the detected state is less than the preset number of video frames (for example, 20 frames).
[0175] To further illustrate the loss function process for restricting predicted mutations, the following provides a specific example. Taking the preset number as 20, in order to limit the frequent jumps in the detected results and avoid the duration window of the detected state being less than 20 frames, when the state result detected in the current video frame is different from the state result detected in the previous frame, the frame number of the current video frame can be recorded as crtDiff (and the corresponding label can also be recorded). Similarly, the frame number of the frame where the last state change occurred is denoted as lstDiff. When crtDiff - lstDiff < 20, it means that the time interval between two state mutations in the predicted result is too small, and a penalty needs to be imposed on this situation. At this time, the present application proposes to additionally calculate the loss between the current network output and the label corresponding to the last mutation frame lstDiff, and give a relatively large weight term.
[0176] As Figure 8 shown, Arg(p(t)) represents the state result (predicted scene category) detected in the video frame at time p(t), and Arg(p(t - 1)) represents the state result (predicted scene category) detected in the video frame at time p(t - 1). Arg(p(t))!= Arg(p(t - 1)) means that the state result detected in the video frame at time p(t) is different from the state result detected in the video frame at time p(t - 1). At this time, the frame number of the video frame at time p(t) can be denoted as crtDiff, and the corresponding label (predicted scene category label) can be recorded, and the frame number of the frame where the last state change occurred, lstDiff, can be obtained. When crtDiff - lstDiff < 20, the loss between the labels of the video frame at time p(t) and the video frame of IstDiff can be calculated, and a large weight term can be given.
[0177] Through the above data augmentation and the design for mutation points in the loss function, the scene detection results of the present application are less affected by noise data, the detection results are accurate, the detection results are smooth, and there will be no obvious jumps.
[0178] In the above scene detection model training method, a pre-annotation scheme is designed in the data annotation process to reduce the annotation time, and various data augmentation methods are designed according to prior knowledge to reduce the dependence of deep learning on the data volume. At the same time, these data augmentation methods also further improve the robustness and accuracy of the model. In addition, aiming at the two difficulties of easy misdetection in the short-time window and insensitive state detection in scene detection, two targeted constraints are proposed in the loss function of the model, so that the model has characteristics such as high sensitivity and high accuracy. Scene detection itself is a pre-algorithm for many camera algorithms. The overall network structure of the present application is simple, and the operations used are all deployment-friendly and suitable for deployment on mobile devices, and can respond promptly to changes in the motion state.
[0179] In an exemplary embodiment, as Figure 9 shown, a scene detection method is provided. Taking the application of this method to a terminal as an example, it includes the following steps 602 to step 604. Among them:
[0180] Step 602, obtain the camera pose data collected when shooting a video.
[0181] Specifically, the terminal can obtain the camera pose data collected when shooting a video. Among them, the camera pose data can be obtained through sensors on the terminal, and the type of sensor is not limited in this application.
[0182] Step 604, input the camera pose data into the target scene detection model obtained by the scene detection model training method, and obtain the scene detection result output by the target scene detection model.
[0183] Specifically, the camera pose data can be input into the target scene detection model obtained by the scene detection model training method, and the scene detection result output by the target scene detection model can be obtained.
[0184] In this application, the implementation solution of the scene detection model training in the scene detection method is similar to the implementation solution described in the above scene detection model training method. Therefore, the specific limitations in one or more scene detection method embodiments provided in this application embodiment can refer to the limitations of the foregoing scene detection model training method, and will not be elaborated here.
[0185] It can be understood that regarding the target scene detection model, considering the feasibility of mobile AI deployment, this application embodiment does not use sequence-related neural networks such as RNN.
[0186] In one of the embodiments, inputting the camera pose data into the target scene detection model may include: when the camera pose data is gyroscope data, performing integral processing on the gyroscope data, and inputting the gyroscope data after integral processing into the target scene detection model.
[0187] Specifically, taking the camera pose data as gyroscope data as an example, using gyroscope data for scene detection is essentially classifying one-dimensional data, and when scene detection is applied to a terminal, the requirement for real-time performance is very high. In this regard, based on the above-obtained target scene detection model, this application realizes classifying the gyroscope data after integration into scene categories through a simple network, such as classifying the motion state and the panoramic shooting state.
[0188] In one of the embodiments, inputting the camera pose data into the target scene detection model includes:
[0189] Input the camera pose data of the current video frame in the video and the scene detection results of multiple video frames before the current video frame into the target scene detection model.
[0190] Specifically, the number of multiple video frames before the current video frame can be set according to requirements. Exemplarily, the number of multiple video frames before the current video frame can be three, that is, the first three video frames of the current video frame.
[0191] In order to avoid sudden state changes in a small time interval as much as possible, this application can additionally input the scene detection data of multiple video frames before the current video frame as the input of the target scene detection model, so as to be able to handle the situation that the video EIS (electronic image stabilization) algorithm is sensitive to the results of state detection.
[0192] In one embodiment, the method further includes:
[0193] Obtain the video anti-shake strategy corresponding to the scene detection result;
[0194] Perform anti-shake processing on the video according to the video anti-shake strategy to obtain the target video.
[0195] Specifically, the embodiments of this application can be applied to the video anti-shake scenario. The terminal can obtain the video anti-shake strategy corresponding to the scene detection result, and then perform anti-shake processing on the video according to the video anti-shake strategy to obtain the target video.
[0196] As above, this application proposes a scene detection scheme based on deep learning. In the data annotation process, a pre-annotation scheme is designed to reduce the annotation time, and multiple data augmentation methods are designed according to prior knowledge to reduce the dependence of the deep learning scheme on the data volume. At the same time, these data augmentation methods also further improve the robustness and accuracy of the algorithm. Among them, a CNN model can be used as the scene detection model to distinguish the camera motion state. Aiming at the two difficulties of easy misdetection and insensitive state detection in the short-time window of scene detection, two targeted constraints are proposed in the loss function of the algorithm, so that the algorithm has the characteristics of high sensitivity and high accuracy. It can be understood that scene detection itself is a pre-algorithm for many camera algorithms. The network structure of the scene detection model in this application is overall simple, and the operations used are all deployment-friendly, suitable for deployment on mobile devices, and can respond to changes in the motion state in a timely manner.
[0197] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless specifically stated herein, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0198] Based on the same inventive concept, an embodiment of the present application further provides a scene detection model training device for implementing the above-mentioned scene detection model training method. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the scene detection model training device provided below can refer to the limitations on the scene detection model training method in the above text, and will not be repeated here.
[0199] In an exemplary embodiment, as Figure 10 shown, a scene detection model training device is provided, including:
[0200] A data acquisition module 801, configured to acquire camera pose data and the corresponding labeled scene category labels of the camera pose data;
[0201] A data augmentation module 802, configured to perform data augmentation on the camera pose data to obtain a target data set;
[0202] A scene prediction module 803, configured to input the target data set into an initial scene detection model for scene category prediction to obtain a predicted scene category label;
[0203] A model training module 804, configured to train the initial scene detection model based on the labeled scene category label and the predicted scene category label using a loss function to obtain a target scene detection model.
[0204] In one of the embodiments, the data acquisition module 801 includes:
[0205] A collected data acquisition module, configured to acquire the camera pose data collected when shooting video frames;
[0206] A pre-labeling module, configured to perform pre-labeling on the camera pose data of consecutive multiple video frames using a decision tree to obtain the labeled scene category labels.
[0207] In one embodiment, the model training module 804 is configured to obtain a first frame number and a second frame number when the predicted scene category label of the current video frame is different from the labeled scene category label; the first frame number is the frame number of the current video frame, and the second frame number is the frame number of the video frame corresponding to the labeled scene category label.
[0208] Wherein, the loss weight corresponding to the current video frame is positively correlated with the difference between the first frame number and the second frame number.
[0209] In one embodiment, the number of consecutive video frames is greater than a preset number, and the preset number is related to the frame rate.
[0210] In one embodiment, the model training module 804 includes:
[0211] The frame number acquisition module is configured to use the current video frame as the current mutation frame and obtain the frame number of the previous mutation frame when the predicted scene category label of the current video frame is different from the predicted scene category label of the previous video frame.
[0212] The weight determination module is configured to determine the loss weight between the predicted scene category label of the current mutation frame and the predicted scene category label of the previous mutation frame based on the loss function if the difference between the frame number of the current mutation frame and the frame number of the previous mutation frame is less than the preset number.
[0213] In one embodiment, the labeled scene category label represents the scene category of the camera pose data; the scene categories include a motion scene and a panoramic shooting scene; the motion scene includes a walking scene, a running scene, and a static holding scene.
[0214] In one embodiment, the data augmentation module 802 is configured to perform data augmentation on the camera pose data using a data augmentation strategy to obtain a target data set; wherein, the data augmentation strategy includes a plurality of data augmentation methods executed in sequence; the data augmentation methods are related to the characteristics of the camera pose data and the prior information of the distribution of the camera pose data in different scene categories.
[0215] In one embodiment, the plurality of data augmentation methods include one or more of inter-axis data swapping, data noise addition, and scene continuation between data.
[0216] The execution order of the inter-axis data swapping is before the data noise addition, and the execution order of the data noise addition is before the scene continuation between data.
[0217] In one embodiment, the random inter-axis data swapping includes the random swapping of the X-axis data and the Y-axis data in the camera pose data.
[0218] Data noise addition includes adding noise to the camera pose data of the loop shooting scenario, walking scenario, and running scenario using random noise; among them, the scenario category of the camera pose data before and after noise addition remains unchanged;
[0219] Scene continuity between data means that the scenario category of the current camera pose data in the same video frame is the same as that of the adjacent camera pose data, and the adjacent camera pose data is the camera pose data adjacent to the current camera pose data.
[0220] In one embodiment, the timestamp of the adjacent camera pose data is related to the timestamp of the video frame to which it belongs and the sampling frequency.
[0221] In one embodiment, the scene prediction module 803 includes:
[0222] An integration module for integrating the data in the target dataset to obtain integrated data;
[0223] A data input module for inputting the integrated data into the initial scene detection model to obtain a predicted scene category label.
[0224] In one embodiment, the initial scene detection model is a convolutional neural network model.
[0225] In one embodiment, the initial scene detection model includes at least two convolutional layers; the activation function used in the at least two convolutional layers is the ReLU function.
[0226] In one embodiment, the camera pose data includes inertial measurement data.
[0227] In one embodiment, the inertial measurement data includes at least gyroscope data.
[0228] Each module in the above scene detection model training device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the electronic device in hardware form or independent of it, or stored in the memory of the electronic device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0229] Based on the same inventive concept, the embodiments of the present application also provide a scene detection device for implementing the above-mentioned scene detection method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the scene detection device provided below can refer to the limitations on the scene detection method in the above text, and will not be repeated here.
[0230] In an exemplary embodiment, as Figure 11 shown, a scene detection device is provided, including:
[0231] An attitude data acquisition module 901 for acquiring camera attitude data collected during video shooting;
[0232] A model input module 902 for inputting the camera attitude data into the target scene detection model obtained by the above-mentioned scene detection model training method to obtain the scene detection result output by the target scene detection model.
[0233] In one embodiment, the model input module 902 includes:
[0234] An integration processing module for, when the camera attitude data is gyroscope data, performing integration processing on the gyroscope data and inputting the integrated gyroscope data into the target scene detection model.
[0235] In one embodiment, the model input module 902 is configured to input the camera attitude data of the current video frame in the video and the scene detection results of multiple video frames before the current video frame into the target scene detection model.
[0236] In one embodiment, the apparatus further includes:
[0237] An anti-shake module for obtaining a video anti-shake strategy corresponding to the scene detection result; performing anti-shake processing on the video according to the video anti-shake strategy to obtain a target video.
[0238] Each module in the above scene detection apparatus can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the electronic device in hardware form or be independent of the processor, or can be stored in the memory in the electronic device in software form so that the processor can call and execute the operations corresponding to the above modules.
[0239] In an exemplary embodiment, an electronic device is provided. The electronic device can be a terminal, and its internal structure diagram can be as Figure 12As shown in the figure. The electronic device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the electronic device is used to exchange information between the processor and external devices. The communication interface of the electronic device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a method for training a scene detection model and a scene detection method. The display unit of the electronic device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the electronic device, or an external keyboard, touchpad, or mouse, etc.
[0240] Those skilled in the art can understand that Figure 12 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the electronic device to which the solution of this application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0241] In one embodiment, an electronic device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0242] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0243] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0244] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0245] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include Read-Only Memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, Resistive Random Access Memory (ReRAM), Magnetoresistive Random Access Memory (MRAM), Ferroelectric Random Access Memory (FRAM), Phase Change Memory (PCM), graphene memory, etc. Volatile memory can include Random Access Memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, Artificial Intelligence (AI) processors, etc., and are not limited thereto.
[0246] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.
[0247] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application shall be subject to the appended claims.
Claims
1. A scene detection model training method, characterized in that: The method comprises: Obtaining camera pose data and annotated scene category labels corresponding to the camera pose data; Performing data augmentation on the camera posture data to obtain a target data set; Inputting the target data set into the initial scene detection model to perform scene category prediction to obtain a predicted scene category label; According to the annotated scene category label and the predicted scene category label, the initial scene detection model is trained based on a loss function to obtain a target scene detection model.
2. The method according to claim 1, characterized in that Obtaining camera pose data and the annotated scene category label corresponding to the camera pose data includes: Obtaining the camera posture data collected when shooting video frames; A decision tree is used to pre-label the camera posture data of a plurality of consecutive video frames to obtain the labeled scene category label.
3. The method according to claim 2, characterized in that According to the annotated scene category label and the predicted scene category label, the initial scene detection model is trained based on a loss function to obtain a target scene detection model, including: When the predicted scene category label of the current video frame is different from the annotated scene category label, obtaining a first frame number and a second frame number; the first frame number is the frame number of the current video frame, and the second frame number is the frame number of the video frame corresponding to the annotated scene category label; Among them, the loss weight corresponding to the current video frame is positively correlated with the difference between the first frame number and the second frame number.
4. The method according to claim 2, characterized in that: The number of the plurality of consecutive video frames is greater than a preset number, and the preset number is related to a frame rate.
5. The method according to claim 4, characterized in that According to the annotated scene category label and the predicted scene category label, the initial scene detection model is trained based on a loss function to obtain a target scene detection model, including: When the predicted scene category label of the current video frame is different from the predicted scene category label of the previous video frame, the current video frame is used as the current mutation frame, and the frame number of the previous mutation frame is obtained; If the difference between the frame number of the current mutation frame and the frame number of the previous mutation frame is less than the preset number, the loss weight between the predicted scene category label of the current mutation frame and the predicted scene category label of the previous mutation frame is determined based on the loss function.
6. The method according to claim 1, characterized in that The annotated scene category label represents the scene category of the camera posture data; the scene category includes motion scenes and panorama scenes; the motion scenes include walking scenes, running scenes and stationary scenes.
7. The method according to claim 6, characterized in that Performing data augmentation on the camera posture data to obtain a target data set includes: A data augmentation strategy is used to perform data augmentation on the camera pose data to obtain the target data set; wherein the data augmentation strategy includes a plurality of data augmentation methods that are executed sequentially; and the data augmentation methods are related to the characteristics of the camera pose data and prior information on the distribution of the camera pose data for different scene categories.
8. The method according to claim 7, characterized in that The multiple data amplification methods include one or more of inter-axis data exchange, data noise addition, and inter-data scene continuation; The execution order of the inter-axis data exchange is before the data noise addition, and the execution order of the data noise addition is before the inter-data scene continuation.
9. The method according to claim 8, characterized in that The random exchange of inter-axis data includes random exchange of X-axis data and Y-axis data in the camera posture data; The data noise addition includes adding noise to the camera posture data of the ring shooting scene, the walking scene, and the running scene using random noise; wherein the scene category of the camera posture data before and after the noise addition remains unchanged; The inter-data scene continuation includes that the scene category of the current camera pose data of the same video frame is the same as the scene category of the adjacent camera pose data, and the adjacent camera pose data is the camera pose data adjacent to the current camera pose data.
10. The method according to claim 9, characterized in that The timestamp of the neighboring camera posture data is related to the timestamp and sampling frequency of the corresponding video frame.
11. The method according to claim 1, characterized in that The target data set is input into the initial scene detection model to perform scene category prediction, and a predicted scene category label is obtained, including: Integrating the data in the target data set to obtain integrated data; The integral data is input into the initial scene detection model to obtain a predicted scene category label.
12. The method according to claim 1, characterized in that The initial scene detection model is a convolutional neural network model.
13. The method according to claim 12, characterized in that The initial scene detection model includes at least two convolutional layers; the activation function used by the at least two convolutional layers is a ReLU function.
14. The method according to any one of claims 1 to 13, characterized in that: The camera pose data includes inertial measurement data.
15. The method according to claim 14, characterized in that The inertial measurement data includes at least gyroscope data.
16. A scene detection method, characterized in that: The method comprises: Get the camera posture data collected when shooting the video; The camera posture data is input into a target scene detection model obtained based on the method described in any one of claims 1 to 15 to obtain a scene detection result output by the target scene detection model.
17. The method according to claim 16, characterized in that Inputting the camera posture data into the target scene detection model includes: When the camera posture data is gyroscope data, the gyroscope data is integrated and the integrated gyroscope data is input into the target scene detection model.
18. The method according to claim 16, characterized in that Inputting the camera posture data into the target scene detection model includes: The camera posture data of the current video frame in the video and the scene detection results of multiple video frames before the current video frame are input into the target scene detection model.
19. The method according to claim 16, characterized in that The method further comprises: Obtaining a video stabilization strategy corresponding to the scene detection result; The video is subjected to anti-shake processing according to the video anti-shake strategy to obtain a target video.
20. A scene detection model training device, characterized in that: The device comprises: A data acquisition module, used to acquire camera posture data and annotated scene category labels corresponding to the camera posture data; A data augmentation module, used to perform data augmentation on the camera posture data to obtain a target data set; A scene prediction module, used to input the target data set into an initial scene detection model to perform scene category prediction and obtain a predicted scene category label; The model training module is used to train the initial scene detection model based on the loss function according to the marked scene category label and the predicted scene category label to obtain a target scene detection model.
21. A scene detection device, characterized in that: The device comprises: A posture data acquisition module is used to obtain the camera posture data collected when shooting the video; A model input module is used to input the camera posture data into a target scene detection model obtained based on the method described in any one of claims 1 to 15, and obtain a scene detection result output by the target scene detection model.
22. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 19 are implemented.
23. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 19 are implemented.
24. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 19 are implemented.