A method and system for detecting camera movement type
By extracting the optical flow motion feature matrix of the video and using the recognition model to automatically identify the video mirror type, the problems of low manual recognition efficiency and poor accuracy are solved, and more efficient and accurate mirror type recognition is achieved.
Patent Information
- Application Number
- CN202510072760.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-17
AI Technical Summary
In the prior art, identifying video camera types depends on manual judgment, resulting in low recognition efficiency and poor accuracy.
By extracting the optical flow motion feature matrix of the storyboard video to be detected, the preset mirror type recognition model is used for identification, and the mirror type of the video is automatically identified.
It improves the recognition efficiency and recognition accuracy of the type of mirroring, reduces manual intervention, and improves video editing and shooting efficiency.
Smart Images

Figure CN119540657B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a method and system for detecting camera movement types. Background Art
[0002] Camera movement plays a vital role in video shooting. Choosing the appropriate camera movement method from existing videos can improve the efficiency of video shooting, but the prerequisite is to identify the camera movement type of the existing video.
[0003] Currently, manual discrimination is still used to identify the camera movement type of a video, that is, relevant personnel watch the video and identify the camera movement type of the video. However, this requires relevant personnel to have professional photography knowledge and takes a lot of time to watch the videos one by one. Errors are also prone to occur during the identification process, and the recognition efficiency and accuracy of the camera movement type are low. Summary of the invention
[0004] In view of this, an embodiment of the present invention provides a method and system for detecting camera movement types to solve the problems of low recognition efficiency and poor recognition accuracy of manual camera movement type recognition.
[0005] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions:
[0006] A first aspect of an embodiment of the present invention discloses a method for detecting a camera movement type, the method comprising:
[0007] Extracting a first sequence set of the storyboard video to be detected, wherein the first sequence set includes a coordinate sequence of multiple optical flow feature points, and the video screen of the storyboard video to be detected is divided into multiple main areas, and each of the main areas is divided into multiple sub-areas;
[0008] Dividing the coordinate sequence of the optical flow feature points in the first sequence set into a second sequence set matching the main area;
[0009] Calculate the optical flow motion feature vector of the sub-region based on the second sequence set;
[0010] splicing the optical flow motion feature vectors of the sub-regions to obtain an optical flow motion feature matrix of the storyboard video to be detected;
[0011] The optical flow motion feature matrix of the storyboard video to be detected is input into a preset camera movement type recognition model to perform camera movement type recognition, so as to obtain the camera movement type of the storyboard video to be detected.
[0012] Preferably, extracting a first sequence set of storyboard videos to be detected includes:
[0013] Calculate the optical flow information of each frame of the storyboard video to be detected;
[0014] Based on the optical flow information, extracting a coordinate sequence of optical flow feature points tracked in two adjacent frames of the video;
[0015] The coordinate sequence of the optical flow feature points tracked in two adjacent frames of the video is saved to obtain a first sequence set of the storyboard video to be detected.
[0016] Preferably, dividing the coordinate sequence of the optical flow feature points in the first sequence set into a second sequence set matching the main area includes:
[0017] According to a first coordinate point of the coordinate sequence of the optical flow feature points in the first sequence set, identifying the main area to which the optical flow feature point belongs;
[0018] A second sequence set matching the main area is constructed by using the coordinate sequence of the optical flow feature points belonging to the main area.
[0019] Preferably, calculating the optical flow motion feature vector of the sub-region based on the second sequence set includes:
[0020] For each sub-region in each main region, identifying a coordinate sequence of optical flow feature points in the second sequence set of the main region belonging to the sub-region;
[0021] Calculating the motion vector corresponding to the coordinate sequence of the optical flow feature points belonging to the sub-region, and calculating the standard deviation of the motion vector;
[0022] The motion vector with the largest standard deviation is determined as the optical flow motion feature vector of the sub-region.
[0023] Preferably, the optical flow motion feature vectors of the sub-regions are spliced to obtain an optical flow motion feature matrix of the storyboard video to be detected, including:
[0024] Adjusting the optical flow motion feature vector of the sub-region to a k-dimensional vector;
[0025] splicing the optical flow motion feature vectors of the sub-region adjusted to k-dimensional vectors into a feature matrix in the x-axis direction and a feature matrix in the y-axis direction;
[0026] The feature matrix in the x-axis direction and the feature matrix in the y-axis direction are superimposed in channel dimension to obtain the optical flow motion feature matrix of the storyboard video to be detected.
[0027] A second aspect of an embodiment of the present invention discloses a camera movement type detection system, the system comprising:
[0028] An extraction unit is used to extract a first sequence set of the storyboard video to be detected, wherein the first sequence set includes a coordinate sequence of multiple optical flow feature points, and the video screen of the storyboard video to be detected is divided into multiple main areas, and each of the main areas is divided into multiple sub-areas;
[0029] A division unit, configured to divide the coordinate sequence of the optical flow feature points in the first sequence set into a second sequence set matching the main area;
[0030] a calculation unit, configured to calculate an optical flow motion feature vector of the sub-region based on the second sequence set;
[0031] A splicing unit, used for splicing the optical flow motion feature vectors of the sub-regions to obtain an optical flow motion feature matrix of the storyboard video to be detected;
[0032] The recognition unit is used to input the optical flow motion feature matrix of the storyboard video to be detected into a preset camera movement type recognition model to perform camera movement type recognition to obtain the camera movement type of the storyboard video to be detected.
[0033] Preferably, the extraction unit comprises:
[0034] A calculation module, used to calculate the optical flow information of each frame of the storyboard video to be detected;
[0035] An extraction module, configured to extract, based on the optical flow information, a coordinate sequence of optical flow feature points tracked in two adjacent frames of the video;
[0036] The saving module is used to save the coordinate sequence of the optical flow feature points tracked in two adjacent frames of the video to obtain a first sequence set of the storyboard video to be detected.
[0037] Preferably, the division unit comprises:
[0038] an identification module, configured to identify the main area to which the optical flow feature point belongs according to the first coordinate point of the coordinate sequence of the optical flow feature point in the first sequence set;
[0039] A construction module is used to construct a second sequence set matching the main area by using the coordinate sequence of the optical flow feature points belonging to the main area.
[0040] Preferably, the calculation unit comprises:
[0041] An identification module, configured to identify, for each sub-region in each main region, a coordinate sequence of optical flow feature points in the second sequence set of the main region belonging to the sub-region;
[0042] A calculation module, used to calculate the motion vector corresponding to the coordinate sequence of the optical flow feature points belonging to the sub-region, and calculate the standard deviation of the motion vector;
[0043] The determination module is used to determine the motion vector with the largest standard deviation as the optical flow motion feature vector of the sub-region.
[0044] Preferably, the splicing unit comprises:
[0045] An adjustment module, used for adjusting the optical flow motion feature vector of the sub-region into a k-dimensional vector;
[0046] A splicing module, used for splicing the optical flow motion feature vectors of the sub-regions adjusted to k-dimensional vectors into a feature matrix in the x-axis direction and a feature matrix in the y-axis direction;
[0047] The superposition module is used to superimpose the feature matrix in the x-axis direction and the feature matrix in the y-axis direction in the channel dimension to obtain the optical flow motion feature matrix of the storyboard video to be detected.
[0048] A method and system for detecting camera movement types based on the above-mentioned embodiment of the present invention is provided, the method comprising: extracting a first sequence set of the camera storyboard video to be detected, the video screen of the camera storyboard video to be detected is divided into a plurality of main areas, each of the main areas is divided into a plurality of sub-areas; dividing the coordinate sequence of the optical flow feature points in the first sequence set into a second sequence set matching the main areas; calculating the optical flow motion feature vectors of the sub-areas based on the second sequence set; splicing the optical flow motion feature vectors of the sub-areas to obtain the optical flow motion feature matrix of the camera storyboard video to be detected; inputting the optical flow motion feature matrix of the camera storyboard video to be detected into a camera movement type recognition model for camera movement type recognition, so as to obtain the camera movement type of the camera storyboard video to be detected. In this scheme, the optical flow motion feature matrix of the camera storyboard video to be detected is extracted, and the camera movement type recognition model is used to identify the camera movement type of the camera storyboard video to be detected based on the optical flow motion feature matrix, without relying on manual recognition of the camera movement type, thereby improving the recognition efficiency and recognition accuracy of the camera movement type. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0050] Figure 1 A flowchart of a method for detecting a camera movement type provided by an embodiment of the present invention;
[0051] Figure 2An example diagram of dividing a video screen provided by an embodiment of the present invention;
[0052] Figure 3 A flowchart of extracting a first sequence set provided by an embodiment of the present invention;
[0053] Figure 4 Another flow chart of extracting a first sequence set provided by an embodiment of the present invention;
[0054] Figure 5 A flowchart of calculating the optical flow motion feature vector of a sub-region provided by an embodiment of the present invention;
[0055] Figure 6 Another flow chart of calculating the optical flow motion feature vector of a sub-region provided by an embodiment of the present invention;
[0056] Figure 7 A flow chart of obtaining an optical flow motion feature matrix of a storyboard video to be detected by splicing provided in an embodiment of the present invention;
[0057] Figure 8 A structural block diagram of a camera movement detection system provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0058] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0059] In this application, the terms "comprises", "comprising" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element.
[0060] Camera movement plays a vital role in video shooting. Choosing the appropriate camera movement method from existing videos can improve the efficiency of video shooting, but the prerequisite is to identify the camera movement type of the existing video.
[0061] Currently, manual discrimination is still used to identify the camera movement type of a video, that is, relevant personnel watch the video and identify the camera movement type of the video. However, this requires relevant personnel to have professional photography knowledge and takes a lot of time to watch the videos one by one. Errors are also prone to occur during the identification process, and the recognition efficiency and accuracy of the camera movement type are low.
[0062] In addition, currently, it is also possible to identify the camera movement type of a video by compiling feature rules. However, the scope of application of each feature rule is small, resulting in low accuracy of each feature rule in identifying the camera movement type, and conflicts are prone to occur between multiple feature rules.
[0063] In order to solve the above problems, this solution proposes a camera movement type detection method and system, which extracts the optical flow motion feature matrix of the storyboard video to be detected, and uses the camera movement type recognition model to identify the camera movement type of the storyboard video to be detected based on the optical flow motion feature matrix. It does not rely on manual identification of the camera movement type, thereby improving the recognition efficiency and accuracy of the camera movement type.
[0064] After applying this solution, one of the application scenarios is as follows: for enterprises with a large amount of video materials, the optical flow motion feature matrix of the storyboard video is input into the camera movement type recognition model to detect the camera movement type of the storyboard video; on the one hand, this can analyze the distribution of camera movement types of different storyboard videos, help editors better understand the motion characteristics of the storyboards, guide video editing, reduce the workload of editors, and improve video editing efficiency; on the other hand, the identified camera movement type can help shooting personnel choose appropriate camera movement methods to guide video shooting, thereby improving video shooting efficiency and effects.
[0065] The present solution is described in detail below through various embodiments.
[0066] See also Figure 1 , shows a flow chart of a method for detecting a camera movement type provided by an embodiment of the present invention, the detection method comprising:
[0067] Step S101: extracting a first sequence set of storyboard videos to be detected.
[0068] It should be noted that the video screen of the storyboard video to be detected is divided into a plurality of main areas, and each main area is divided into a plurality of sub-areas.
[0069] For example Figure 2As can be seen from the example diagram of the division of the video screen, each frame of the storyboard video to be detected (video screen image) is divided into 9 main areas, and these 9 main areas are recorded as MR1 to MR9; under the same camera movement mode, since the motion characteristics of objects in these 9 main areas are generally different, it is necessary to divide the video screen into 9 main areas (the number is only for example). After the main areas are divided, the motion characteristics of objects in the same main area have certain similarities.
[0070] After dividing into multiple main areas, each main area is divided into multiple sub-areas, for example Figure 2 As can be seen, after 9 main regions are divided, each main region is divided into 9 sub-regions. The 9 sub-regions divided from the main region MR1 are recorded as SR11 to SR19. The other main regions are divided into sub-regions in the same way, which will not be illustrated one by one here.
[0071] In the specific implementation of step S101, the optical flow information of each frame of the storyboard video to be detected is calculated using a feature point method, for example, the Lucas-Kanade method is used to calculate the optical flow information of each frame of the video.
[0072] It should be noted that the optical flow information refers to the optical flow information of a continuous video, that is, the optical flow information of the entire video screen, and the optical flow information includes the movement information of the optical flow feature points (represented by coordinates).
[0073] By utilizing the optical flow information of the video screen, the coordinate sequence of the optical flow feature points tracked (successfully tracked) in two adjacent frames of the video screen is saved in a feature point set, thereby obtaining a first sequence set, which includes the coordinate sequences of multiple optical flow feature points, and the coordinate sequence of a certain optical flow feature point includes multiple coordinates of the optical flow feature point.
[0074] Step S102: dividing the coordinate sequence of the optical flow feature points in the first sequence set into a second sequence set matching the main area.
[0075] In the specific implementation of step S102, after the first sequence set is extracted, the first sequence set contains the coordinate sequence of the optical flow feature points belonging to each main area, so the first sequence set is divided to obtain a second sequence set matching the main area.
[0076] The “second sequence set matching the main area” is composed of the following parts: the coordinate sequence of the optical flow feature points belonging to the main area in the first sequence set.
[0077] Specifically, according to the first coordinate of the coordinate sequence of the optical flow feature point in the first sequence set, the main area to which the optical flow feature point belongs is identified; and the coordinate sequence of the optical flow feature point belonging to the main area is used to construct a second sequence set matching the main area.
[0078] That is, according to the first coordinate of the coordinate sequence of the optical flow feature points in the first sequence set, first determine which main region each optical flow feature point belongs to. For each main region, the coordinate sequence of the optical flow feature points belonging to the main region is constructed into a second sequence set matching the main region.
[0079] For example: suppose that 9 main regions are divided, and these 9 main regions are recorded as MR1 to MR9; according to the above method, the coordinate sequence of the optical flow feature points in the first sequence set is divided into 9 second sequence sets, each second sequence set contains the coordinate sequence of the optical flow feature points belonging to a certain main region, and the 9 second sequence sets matching the 9 main regions MR1 to MR9 are recorded as MRP1 to MRP9 respectively.
[0080] Step S103: Calculate the optical flow motion feature vector of the sub-region based on the second sequence set.
[0081] In the process of specifically implementing step S103, after the first sequence set is divided into multiple second sequence sets matching each main area, for each sub-area of each main area, the optical flow motion feature vector of the sub-area is calculated using the coordinate sequence of the optical flow feature points belonging to the sub-area in the second sequence set of the main area.
[0082] By the above method, the optical flow motion feature vector of each sub-region in each main region can be calculated.
[0083] Step S104: splicing the optical flow motion feature vectors of the sub-regions to obtain an optical flow motion feature matrix of the storyboard video to be detected.
[0084] In the specific implementation of step S104, after the optical flow motion feature vector of each sub-region in each main region is calculated, the optical flow motion feature vectors of the sub-regions in each main region are spliced to obtain the optical flow motion feature matrix of the storyboard video to be detected.
[0085] Step S105: inputting the optical flow motion feature matrix of the storyboard video to be detected into a preset camera movement type recognition model to perform camera movement type recognition, so as to obtain the camera movement type of the storyboard video to be detected.
[0086] It should be noted that the camera movement type recognition model is obtained by training a neural network model based on sample data, wherein the sample data at least includes: a sample storyboard video carrying a real label (real camera movement type label); according to the above steps S101 to S104, the optical flow motion feature matrix of the sample storyboard video is first processed to obtain the optical flow motion feature matrix of the sample storyboard video, and then the neural network model (for example, a classification neural network model) is trained using the optical flow motion feature matrix of the sample storyboard video to obtain the camera movement type recognition model.
[0087] In the process of specifically implementing step S105, after the optical flow motion feature matrix of the storyboard video to be detected is obtained by splicing, the optical flow motion feature matrix stores the motion features of the storyboard video to be detected, and the optical flow motion feature matrix of the storyboard video to be detected is input into the camera movement type recognition model for camera movement type recognition, and the camera movement type of the storyboard video to be detected output by the camera movement type recognition model is obtained.
[0088] It should be noted that the number of camera movement types output by the fully connected layer of the camera movement type recognition model is the number of camera movement types that need to be judged; that is, the optical flow motion feature matrix of the storyboard video to be detected is input into the camera movement type recognition model, and the camera movement type recognition model outputs the probabilities of multiple camera movement types, and the camera movement type with the largest probability is taken as the camera movement type of the storyboard video to be detected.
[0089] For example, assuming that the camera movement type recognition model needs to judge the following nine camera movement types: push shot, pull shot, pan shot, horizontal movement, vertical movement, subjective shot, surround shot, lifting shot, and simulated aerial photography; the optical flow motion feature matrix of the storyboard video to be detected is input into the camera movement type recognition model, and the fully connected layer of the camera movement type recognition model will input the probabilities of the above nine camera movement types respectively: 0.1, 0.1, 0.3, 0.8, 0.6, 0.02, 0.09, 0.08, and 0.07; among them, 0.8 is the largest, so the camera movement type of the storyboard video to be detected is finally determined to be horizontal movement.
[0090] In an embodiment of the present invention, an optical flow motion feature matrix of a storyboard video to be detected is extracted, and a camera movement type recognition model is used to identify the camera movement type of the storyboard video to be detected based on the optical flow motion feature matrix, without relying on manual recognition of the camera movement type, thereby improving the recognition efficiency and recognition accuracy of the camera movement type.
[0091] For the above-mentioned embodiment of the present invention Figure 1 The first sequence set of extracting the storyboard video to be detected involved in step S101 is shown in FIG. Figure 3 , shows a flowchart of extracting a first sequence set provided by an embodiment of the present invention, including the following steps:
[0092] Step S301: Calculate the optical flow information of each frame of the storyboard video to be detected.
[0093] In the specific implementation of step S301, the optical flow information of each frame of the storyboard video to be detected is calculated using a feature point method, for example, the Lucas-Kanade method is used to calculate the optical flow information of each frame of the video.
[0094] Step S302: Based on the optical flow information, extract the coordinate sequence of the optical flow feature points tracked in two adjacent video frames.
[0095] In the specific implementation of step S302 , based on the optical flow information of each frame of video, a coordinate sequence of optical flow feature points tracked in two adjacent frames of video is extracted.
[0096] The “optical flow feature points tracked in two adjacent video frames” specifically refer to: optical flow feature points that are successfully tracked, or optical flow feature points that match in two adjacent video frames.
[0097] For example: assuming that the optical flow information of the first frame of video tracks optical flow feature points 1, 2, 3, 4 and 7, and the optical flow information of the second frame of video tracks optical flow feature points 1, 2, 3 and 4; then, the optical flow feature points successfully tracked are optical flow feature points 1, 2, 3 and 4, and optical flow feature point 7 is not a successfully tracked optical flow feature point.
[0098] Therefore, the coordinate sequences of the successfully tracked optical flow feature points 1, 2, 3 and 4 are saved.
[0099] Step S303: Save the coordinate sequence of the optical flow feature points tracked in two adjacent frames of video to obtain a first sequence set of the storyboard video to be detected.
[0100] In the specific implementation of step S303, the coordinate sequence of the optical flow feature points tracked in two adjacent frames of video is saved in a feature point set, thereby obtaining a first sequence set of the storyboard video to be detected.
[0101] That is to say, the coordinate sequence of the successfully tracked optical flow feature points is saved in the feature point set, thereby obtaining the first sequence set of the storyboard video to be detected.
[0102] For example, for a successfully tracked optical flow feature point a, assuming the number of cycles is m, then the coordinate sequence of the optical flow feature point a in the first sequence set is [(x 0 ,y 0 ),(x 1 ,y 1 ),…,(xm ,y m )], where x is the horizontal coordinate and y is the vertical coordinate; (x 0 ,y 0 ) is the coordinate of the optical flow feature point a obtained in the first cycle, (x m ,y m ) is the coordinate of the optical flow feature point a obtained in the mth cycle, so m can also represent the number of coordinates in the coordinate sequence of the optical flow feature point a.
[0103] It should be noted that the "number of cycles m" mentioned above is the number of cycles in the process of calculating the optical flow features by the optical flow method. For example, assuming that the number of cycles in the process of calculating the optical flow features by the optical flow method is 79, the number of cycles m is 79.
[0104] To further explain how to extract the first sequence set of the storyboard video to be detected, Figure 4 Another flowchart of extracting the first sequence set is shown for illustration. Figure 4 The following steps are involved:
[0105] Step S401: Input the storyboard video to be detected.
[0106] Step S402: Initialize the first frame of the storyboard video to be detected as FF.
[0107] In the specific implementation process of step S402, the first frame of the storyboard video to be detected is initialized, and the first frame of the video is recorded as FF during the initialization process.
[0108] Step S403: Get the next video frame SF of FF.
[0109] It should be noted that the next video frame of FF is recorded as SF.
[0110] Step S404: Calculate the optical flow information of FF and SF.
[0111] Step S405: Obtain the optical flow feature points that match FF and SF, and save the optical flow feature points that match FF and SF into a feature point set.
[0112] It should be noted that the optical flow feature points that match FF and SF are the optical flow feature points that are successfully tracked, and the coordinate sequences of all the successfully tracked optical flow feature points are saved in the feature point set to obtain the first sequence set.
[0113] Step S406: Determine whether there is a next video frame. If there is a next video frame, execute step S407; if there is no next video frame, end.
[0114] Step S407: Update FF=SF, and return to execute step S403.
[0115] above Figure 3 and Figure 4 , which is the relevant instructions on how to extract the first sequence set.
[0116] For the above-mentioned embodiment of the present invention Figure 1 The content of the optical flow motion feature vector of the sub-region involved in the calculation in step S103 is shown in Figure 5 , shows a flowchart of calculating the optical flow motion feature vector of a sub-region provided by an embodiment of the present invention, including the following steps:
[0117] Step S501: for each sub-region in each main region, identifying a coordinate sequence of optical flow feature points belonging to the sub-region in a second sequence set of the main region.
[0118] In the specific implementation of step S501, for each sub-region in each main region, a coordinate sequence of optical flow feature points belonging to the sub-region in the second sequence set of the main region is identified.
[0119] Specifically, the method for identifying the optical flow feature points belonging to the sub-region is: based on the first coordinate of the coordinate sequence of the optical flow feature point, determine which sub-region the optical flow feature point belongs to; through the above method, the coordinate sequence of the optical flow feature points belonging to the sub-region in the second sequence set can be obtained.
[0120] For example, for the main region MR1, the second sequence set of the main region MR1 is MRP1, MRP1-i is the coordinate sequence of the i-th optical flow feature point in MRP1, MRP1-i={(x j ,y j )|j=1~m},(x j ,y j ) is the jth coordinate in MRP1-i, and m is the number of coordinates in MRP1-i (that is, Figure 3 The number of loops m mentioned in step S303 in the above process); judging to which sub-region of the main region MR1 the i-th optical flow feature point belongs according to the first coordinate in MRP1-i.
[0121] Step S502: Calculate the motion vector corresponding to the coordinate sequence of the optical flow feature points belonging to the sub-region, and calculate the standard deviation of the motion vector.
[0122] In the process of specifically implementing step S502, for each sub-region in each main region, after identifying the coordinate sequence of the optical flow feature points belonging to the sub-region, the motion vectors corresponding to the coordinate sequences of each optical flow feature point belonging to the sub-region are calculated respectively, a motion vector is calculated for each coordinate sequence of the optical flow feature point, and the standard deviation of the motion vector is calculated.
[0123] Specifically, the motion vector corresponding to the coordinate sequence of the optical flow feature point is calculated using formula (1) and formula (2).
[0124] (1);
[0125] (2);
[0126] In formula (1) and formula (2), VMRP1-i is the motion vector of MRP1-i, MRP1-i is the coordinate sequence of the i-th optical flow feature point in the main region MRP1, c is any one of 1 to m-1, and m is the number of coordinates in MRP1-i;
[0127] x c is the horizontal coordinate of the cth coordinate in the coordinate sequence of the i-th optical flow feature point, y c vx is the ordinate of the cth coordinate in the coordinate sequence of the i-th optical flow feature point; c is the displacement between the abscissa of the c+1th coordinate and the abscissa of the cth coordinate in the coordinate sequence of the i-th optical flow feature point, vy c is the displacement between the ordinate of the c+1th coordinate and the ordinate of the cth coordinate in the coordinate sequence of the i-th optical flow feature point; vx in formula (1) 1 To vx m-1 ,vy 1 To vy m-1 It is calculated by formula (2).
[0128] The calculation method of the motion vector corresponding to the coordinate sequence of the optical flow feature points in the sub-areas of other main areas is similar, which will not be described here one by one.
[0129] Based on the above formulas (1) and (2), the standard deviation of the motion vector is calculated by formulas (3) to (5).
[0130] (3);
[0131] (4);
[0132] (5);
[0133] In formula (3) to formula (5), is the average value of the elements in the x-axis direction in the motion vector VMRP1-i, is the average value of the elements in the y-axis direction of the motion vector VMRP1-i, is the standard deviation of motion vector VMRP1-i.
[0134] Step S503: Determine the motion vector with the largest standard deviation as the optical flow motion feature vector of the sub-region.
[0135] In the specific implementation of step S503, for each sub-region of each main region, after calculating the standard deviation of the motion vector corresponding to the coordinate sequence of the optical flow feature points belonging to the sub-region, the motion vector with the largest standard deviation is determined as the optical flow motion feature vector of the sub-region.
[0136] For example, for the main region MR1, through the above steps S501 to S503, the optical flow motion feature vector of the sub-region to which MRP1-i belongs can be obtained as VSR1z.
[0137] To further explain how to obtain the optical flow motion feature vector of the sub-region, take the calculation of the optical flow motion feature vector of each sub-region in the main region MR1 as an example. Figure 6 Another flowchart of calculating the optical flow motion feature vector of the sub-region is shown as an example, wherein the optical flow motion feature vectors of the sub-regions SR11 to SR19 are recorded as VSR11 to VSR19, Figure 6 The steps include:
[0138] Step S601: Initialize the optical flow motion feature vectors VSR11-VSR19 of the sub-regions SR11-SR19 to zero vectors, and obtain the second sequence set MRP1 of the main region MR1.
[0139] Step S602: Acquire the next coordinate sequence MRP1-i from the second sequence set MRP1.
[0140] It should be noted that MRP1-i is the coordinate sequence of the i-th optical flow feature point in MRP1.
[0141] Step S603: Determine the sub-region to which MRP1-i belongs according to the position to which the first coordinate in MRP1-i belongs.
[0142] Step S604: Calculate the motion vector VMRP1-i of MRP1-i.
[0143] Step S605: Determine whether the optical flow motion feature vector VSR1z of the sub-region to which MRP1-i belongs is a zero vector. If VSR1z is a zero vector, execute step S606; if VSR1z is not a zero vector, execute step S607.
[0144] It should be noted that the optical flow motion feature vector of the sub-region to which MRP1-i belongs is recorded as VSR1z, and the value of z is any one of 1 to 9.
[0145] Step S606: VSR1z=VMRP1-i, execute step S608.
[0146] Step S607: Calculate the standard deviations of VSR1z and VMRP1-i respectively according to the standard deviation method, take the one with the larger standard deviation as the new VSR1z, and execute step S608.
[0147] It should be noted that the standard deviation of VSR1z can be calculated by the above formulas (3) to (5), which will not be described in detail here.
[0148] Step S608: determine whether the traversal of the second sequence set MRP1 is completed; if the traversal is not completed, return to execute step S602; if the traversal is completed, end, and obtain the optical flow motion feature vectors of the sub-regions SR11 to SR19.
[0149] above Figure 5 and Figure 6 , which is an explanation of how to calculate the optical flow motion feature vector of the sub-region.
[0150] For the above-mentioned embodiment of the present invention Figure 1 The optical flow motion feature matrix of the to-be-detected storyboard video involved in step S104 is shown in Figure 7 , shows a flow chart of obtaining an optical flow motion feature matrix of a storyboard video to be detected by splicing provided by an embodiment of the present invention, including the following steps:
[0151] Step S701: adjusting the optical flow motion feature vector of the sub-region to a k-dimensional vector.
[0152] In the specific implementation of step S701 , after obtaining the optical flow motion feature vector of each sub-region of each main region, for each sub-region of each main region, the optical flow motion feature vector of the sub-region is adjusted to a k-dimensional vector.
[0153] Specifically, the optical flow motion feature vector of the sub-region is filled or intercepted, so as to adjust the optical flow motion feature vector of the sub-region to a k-dimensional vector. The specific implementation method of filling and intercepting is as follows:
[0154] Filling: If the optical flow motion feature vector of the sub-region is less than k-dimensional (i.e., less than k-dimensional), fill the optical flow motion feature vector of the sub-region with 0 to adjust the optical flow motion feature vector of the sub-region to a k-dimensional vector.
[0155] Truncation: If the optical flow motion feature vector of the sub-region is greater than the k-dimensionality, the optical flow motion feature vector of the sub-region is processed in the order of "eliminating 0 elements" and "truncating at the end", so as to adjust the optical flow motion feature vector of the sub-region to a k-dimensional vector.
[0156] That is, if the optical flow motion feature vector of the sub-region is greater than the k dimension, the 0 elements in the optical flow motion feature vector of the sub-region are first removed. If the optical flow motion feature vector of the sub-region is still greater than the k dimension after all 0 elements are removed, the optical flow motion feature vector of the sub-region with 0 elements removed is truncated at the end, and the optical flow motion feature vector of the sub-region is adjusted to a k-dimensional vector.
[0157] For example: assuming k is 5, the optical flow motion feature vector of the sub-region is [1,2,3,6,0,1,9,8,9]; first remove the 0 element in the optical flow motion feature vector of the sub-region to obtain [1,2,3,6,1,9,8,9]. At this time, the dimension of the optical flow motion feature vector of the sub-region is still greater than 5, so [1,2,3,6,1,9,8,9] is truncated at the end to obtain a k-dimensional vector of [1,2,3,6,1].
[0158] It should be noted that the value of k in “adjusting the optical flow motion feature vector of the sub-region to a k-dimensional vector” can be set according to actual conditions, wherein two preferred ways of setting the value of k are as follows:
[0159] The first method of setting the value of k is to extract feature vectors from the sample data used in the trained camera movement type recognition model, and cluster the extracted feature vectors to obtain the vector dimension of the class center. The vector dimension of the class center is the actual value of k.
[0160] The second way to set the value of k: select k as a multiple of 8 according to actual conditions (to facilitate computer calculations, the specific value can be set by yourself), such as k is 64, 96 or 128.
[0161] Step S702: splicing the optical flow motion feature vectors of the sub-regions adjusted to k-dimensional vectors into a feature matrix in the x-axis direction and a feature matrix in the y-axis direction.
[0162] In the specific implementation process of step S702, after the optical flow motion feature vectors of each sub-region of each main region are adjusted to k-dimensional vectors, the optical flow motion feature vectors of the sub-regions adjusted to k-dimensional vectors are respectively spliced into feature matrices in the x-axis direction and feature matrices in the y-axis direction.
[0163] For example, assuming that the video screen is divided into 9 main areas, each main area is divided into 9 sub-areas, 81 optical flow motion feature vectors adjusted to k-dimensional vectors can be obtained (a total of 81 sub-areas); these 81 optical flow motion feature vectors adjusted to k-dimensional vectors are spliced into feature matrices in the x-axis direction and feature matrices in the y-axis direction, and the feature matrix in the x-axis direction is recorded as matrix x , the characteristic matrix in the y-axis direction is recorded as matrix y , matrix x and matrix y Both are matrices of 1*1*81*k.
[0164] Among them, matrix x and matrix y The specific content of is shown in formula (6).
[0165] (6);
[0166] In formula (6), VSR11 x VSR11 is the optical flow motion feature vector in the x-axis direction of the first sub-region SR11 of the first main region MR1. y is the optical flow motion feature vector of the first sub-region SR11 of the first main region MR1 in the y-axis direction. The other parameters in formula (6) are similar and will not be explained one by one here.
[0167] Step S703: superimpose the feature matrix in the x-axis direction and the feature matrix in the y-axis direction in the channel dimension to obtain the optical flow motion feature matrix of the storyboard video to be detected.
[0168] In the specific implementation of step S703, the characteristic matrix matrix in the x-axis direction is x And the characteristic matrix matrix in the y-axis direction y Channel dimensions are superimposed to obtain the optical flow motion feature matrix of the storyboard video to be detected. The optical flow motion feature matrix of the storyboard video to be detected is a matrix of 1*2*81*k.
[0169] above Figure 7 This is a description of the optical flow motion feature matrix obtained by splicing the storyboard video to be detected.
[0170] For the above-mentioned embodiment of the present invention Figure 1 The camera movement type recognition model involved in step S105, the process of training and obtaining the camera movement type recognition model is explained below.
[0171] The neural network model used to train the camera movement type recognition model includes at least a convolutional layer (denoted as conv2d), a backbone layer (denoted as backbone) and a fully connected layer, wherein the parameters of the convolutional layer are "conv2d (in_channel=2,out_channel,kernel=3*3)", and the backbone layer can collect the backbone layer of any classification network. The specific parameters and types of the convolutional layer and the backbone layer are not limited here, and the number of camera movement types output by the fully connected layer is the number of camera movement types that need to be judged.
[0172] Collect and organize sample data of the camera movement types that need to be identified. The sample data at least includes sample storyboard videos with real labels (real camera movement type labels). The camera movement types that need to be identified are the following 9 types: push shot, pull shot, pan shot, horizontal movement, vertical movement, subjective shot, surround shot, lifting shot, and simulated aerial photography. The aforementioned content about the camera movement types that need to be identified is only for illustration, and the camera movement types that need to be identified can be adjusted according to actual conditions.
[0173] Using the above embodiment of the present invention Figure 1 In the method of step S101 to step S104, the optical flow motion feature matrix of the sample storyboard video is extracted, and a training data set is constructed according to the optical flow motion feature matrix of the sample storyboard video. The specific form of the training data set is as follows: Dataset={data r ,r=1-p}, p is the amount of data in the training data set, data r =[optical flow motion feature matrix, number of camera movement types].
[0174] It should be noted that the same data in the training data set may have multiple camera movements, so r The number of camera movement types is from 1 to classes (for example, classes=9).
[0175] Model training is performed based on the training data set, the neural network model and the loss function to obtain a camera movement type recognition model. During the training process, the backbone layer of the neural network model can use pre-trained weights based on other classification data sets as the initial weights of the neural network model for fine-tuning.
[0176] All the above contents are related explanations about a method for detecting camera movement types proposed in this scheme. Overall, different camera movement methods will cause inconsistent regular movements of objects in the video screen. There is a certain consistency in the movement characteristics of objects in the video screen of the same camera movement type, so the camera movement type can be identified by analyzing the movement differences caused by the camera movement type.
[0177] The motion characteristics of objects in storyboard videos can be represented by optical flow motion characteristics. Considering that the motion characteristics of objects in different areas of the video screen under the same camera movement mode are generally different, it is necessary to consider the area where the object moves when expressing the camera movement through optical flow motion characteristics (this is the reason for dividing the main area and sub-area). Based on the above considerations, the obtained optical flow motion feature matrix can well characterize the camera movement characteristics of a storyboard video. At this time, the camera movement type recognition model is used to process the optical flow motion feature matrix to accurately identify the camera movement type of the storyboard video.
[0178] Corresponding to the detection method of a camera movement type provided in the above embodiment of the present invention, see Figure 8 , an embodiment of the present invention also provides a structural block diagram of a camera movement type detection system, the detection system comprising: an extraction unit 100, a division unit 200, a calculation unit 300, a splicing unit 400 and a recognition unit 500;
[0179] The extraction unit 100 is used to extract a first sequence set of the storyboard video to be detected, wherein the first sequence set includes a coordinate sequence of multiple optical flow feature points, and the video screen of the storyboard video to be detected is divided into multiple main areas, and each main area is divided into multiple sub-areas.
[0180] The division unit 200 is used to divide the coordinate sequence of the optical flow feature points in the first sequence set into a second sequence set matching the main area.
[0181] The calculation unit 300 is used to calculate the optical flow motion feature vector of the sub-region based on the second sequence set.
[0182] The splicing unit 400 is used to splice the optical flow motion feature vectors of the sub-regions to obtain the optical flow motion feature matrix of the storyboard video to be detected.
[0183] The recognition unit 500 is used to input the optical flow motion feature matrix of the storyboard video to be detected into a preset camera movement type recognition model to perform camera movement type recognition to obtain the camera movement type of the storyboard video to be detected, wherein the camera movement type recognition model is obtained based on a neural network model trained based on sample data.
[0184] In an embodiment of the present invention, an optical flow motion feature matrix of a storyboard video to be detected is extracted, and a camera movement type recognition model is used to identify the camera movement type of the storyboard video to be detected based on the optical flow motion feature matrix, without relying on manual recognition of the camera movement type, thereby improving the recognition efficiency and recognition accuracy of the camera movement type.
[0185] Preferably, combined Figure 8 The content shown, the extraction unit 100 includes a calculation module, an extraction module and a storage module, and the execution principle of each module is as follows:
[0186] The calculation module is used to calculate the optical flow information of each frame of the storyboard video to be detected.
[0187] The extraction module is used to extract the coordinate sequence of the optical flow feature points tracked in two adjacent frames of video images based on the optical flow information.
[0188] The saving module is used to save the coordinate sequence of the optical flow feature points tracked in two adjacent frames of video to obtain a first sequence set of the storyboard video to be detected.
[0189] Preferably, combined Figure 8 As shown in the figure, the division unit 200 includes an identification module and a construction module, and the execution principle of each module is as follows:
[0190] The identification module is used to identify the main area to which the optical flow feature point belongs according to the first coordinate point of the coordinate sequence of the optical flow feature point in the first sequence set.
[0191] The construction module is used to construct a second sequence set matching the main area by using the coordinate sequence of the optical flow feature points belonging to the main area.
[0192] Preferably, combined Figure 8 The content shown, the calculation unit 300 includes an identification module, a calculation module and a determination module, and the execution principle of each module is as follows:
[0193] The identification module is used to identify, for each sub-region in each main region, a coordinate sequence of optical flow feature points belonging to the sub-region in the second sequence set of the main region.
[0194] The calculation module is used to calculate the motion vector corresponding to the coordinate sequence of the optical flow feature points belonging to the sub-region, and calculate the standard deviation of the motion vector.
[0195] The determination module is used to determine the motion vector with the largest standard deviation as the optical flow motion feature vector of the sub-region.
[0196] Preferably, combined Figure 8 The splicing unit 400 includes an adjustment module, a splicing module and a superposition module. The execution principle of each module is as follows:
[0197] The adjustment module is used to adjust the optical flow motion feature vector of the sub-region into a k-dimensional vector.
[0198] The splicing module is used to splice the optical flow motion feature vectors of the sub-regions adjusted to k-dimensional vectors into a feature matrix in the x-axis direction and a feature matrix in the y-axis direction.
[0199] The superposition module is used to superimpose the feature matrix in the x-axis direction and the feature matrix in the y-axis direction in the channel dimension to obtain the optical flow motion feature matrix of the storyboard video to be detected.
[0200] In summary, an embodiment of the present invention provides a method and system for detecting camera movement types, which extracts an optical flow motion feature matrix of a storyboard video to be detected, and uses a camera movement type recognition model to identify the camera movement type of the storyboard video to be detected based on the optical flow motion feature matrix. This method does not rely on manual identification of the camera movement type, thereby improving the recognition efficiency and accuracy of the camera movement type.
[0201] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can refer to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without creative work.
[0202] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0203] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for detecting a camera movement type, characterized in that: The method comprises: Extracting a first sequence set of the storyboard video to be detected, including: calculating optical flow information of each frame of the storyboard video to be detected; extracting a coordinate sequence of optical flow feature points tracked in two adjacent frames of the video based on the optical flow information; saving the coordinate sequence of the optical flow feature points tracked in two adjacent frames of the video to obtain a first sequence set of the storyboard video to be detected, the first sequence set comprising a plurality of coordinate sequences of optical flow feature points, the video screen of the storyboard video to be detected being divided into a plurality of main regions, each of the main regions being divided into a plurality of sub-regions; Dividing the coordinate sequence of the optical flow feature points in the first sequence set into a second sequence set matching the main area; Calculating the optical flow motion feature vector of the sub-region based on the second sequence set, including: for each sub-region in each main region, identifying the coordinate sequence of optical flow feature points belonging to the sub-region in the second sequence set of the main region; calculating the motion vector corresponding to the coordinate sequence of the optical flow feature points belonging to the sub-region, and calculating the standard deviation of the motion vector; determining the motion vector with the largest standard deviation as the optical flow motion feature vector of the sub-region; splicing the optical flow motion feature vectors of the sub-regions to obtain an optical flow motion feature matrix of the storyboard video to be detected; The optical flow motion feature matrix of the storyboard video to be detected is input into a preset camera movement type recognition model to perform camera movement type recognition, so as to obtain the camera movement type of the storyboard video to be detected.
2. The method according to claim 1, characterized in that Dividing the coordinate sequence of the optical flow feature points in the first sequence set into a second sequence set matching the main area includes: According to a first coordinate point of the coordinate sequence of the optical flow feature points in the first sequence set, identifying the main area to which the optical flow feature point belongs; A second sequence set matching the main area is constructed by using the coordinate sequence of the optical flow feature points belonging to the main area.
3. The method according to claim 1 or 2, characterized in that: The optical flow motion feature vectors of the sub-regions are spliced to obtain an optical flow motion feature matrix of the to-be-detected storyboard video, including: Adjusting the optical flow motion feature vector of the sub-region to a k-dimensional vector; splicing the optical flow motion feature vectors of the sub-region adjusted to k-dimensional vectors into a feature matrix in the x-axis direction and a feature matrix in the y-axis direction; The feature matrix in the x-axis direction and the feature matrix in the y-axis direction are superimposed in channel dimension to obtain the optical flow motion feature matrix of the storyboard video to be detected.
4. A camera-moving type detection system, characterized in that: The system comprises: An extraction unit is used to extract a first sequence set of the storyboard video to be detected, wherein the first sequence set includes a coordinate sequence of multiple optical flow feature points, and the video screen of the storyboard video to be detected is divided into multiple main areas, and each of the main areas is divided into multiple sub-areas; A division unit, configured to divide the coordinate sequence of the optical flow feature points in the first sequence set into a second sequence set matching the main area; a calculation unit, configured to calculate an optical flow motion feature vector of the sub-region based on the second sequence set; A splicing unit, used for splicing the optical flow motion feature vectors of the sub-regions to obtain an optical flow motion feature matrix of the storyboard video to be detected; An identification unit, used for inputting the optical flow motion feature matrix of the to-be-detected storyboard video into a preset camera movement type identification model to perform camera movement type identification, so as to obtain the camera movement type of the to-be-detected storyboard video; Wherein, the extraction unit comprises: A calculation module, used to calculate the optical flow information of each frame of the storyboard video to be detected; An extraction module, configured to extract, based on the optical flow information, a coordinate sequence of optical flow feature points tracked in two adjacent frames of the video; A saving module, used for saving the coordinate sequence of the optical flow feature points tracked in two adjacent frames of the video to obtain a first sequence set of the storyboard video to be detected; Wherein, the computing unit comprises: An identification module, configured to identify, for each sub-region in each main region, a coordinate sequence of optical flow feature points in the second sequence set of the main region belonging to the sub-region; A calculation module, used to calculate the motion vector corresponding to the coordinate sequence of the optical flow feature points belonging to the sub-region, and calculate the standard deviation of the motion vector; The determination module is used to determine the motion vector with the largest standard deviation as the optical flow motion feature vector of the sub-region.
5. The system according to claim 4, characterized in that The division unit comprises: an identification module, configured to identify the main area to which the optical flow feature point belongs according to the first coordinate point of the coordinate sequence of the optical flow feature point in the first sequence set; A construction module is used to construct a second sequence set matching the main area by using the coordinate sequence of the optical flow feature points belonging to the main area.
6. The system according to claim 4 or 5, characterized in that: The splicing unit comprises: An adjustment module, used for adjusting the optical flow motion feature vector of the sub-region into a k-dimensional vector; A splicing module, used for splicing the optical flow motion feature vectors of the sub-regions adjusted to k-dimensional vectors into a feature matrix in the x-axis direction and a feature matrix in the y-axis direction; The superposition module is used to superimpose the feature matrix in the x-axis direction and the feature matrix in the y-axis direction in the channel dimension to obtain the optical flow motion feature matrix of the storyboard video to be detected.
Citation Information
Patent Citations
Video jitter detection method and device
CN110248048A
Micro-expression recognition method and device
CN111274978A