Video Semantic Recognition Method and Device
By using neural network methods to identify the spatial characteristics and semantics of video frames in video semantic recognition, divide time sequence fragments and prioritize the selection of exciting fragments for splicing, the problems of low excitement and computational complexity of video styling in the existing technology are solved, and efficient and exciting video styling effect is achieved.
Patent Information
- Application Number
- CN202011642456.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-24
- Filing Date
- 2020-12-31
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-12-31
AI Technical Summary
Existing video semantic recognition technology is difficult to effectively identify semantics at different levels in videos, resulting in low excitement during video stitching, and many parameters of three-dimensional convolutional neural networks, difficult model training, and long calculation time for video dual-stream networks.
A computer-executed neural network method is adopted, including input layer, airspace feature extraction layer, static semantic recognition layer, dynamic semantic recognition layer, timing segment division layer and output layer. By extracting the airspace features of video frames and identifying static and dynamic semantics, timing segments are divided, and timing segments with high excitement are preferred for splicing.
Effective recognition and video stitching of different levels of semantics in the video are realized, which improves the excitement of the stitching video and reduces the complexity and calculation time of model training.
Smart Images

Figure CN114078223B_ABST
Abstract
Description
[0001] This application claims the priority of Chinese patent applications filed with the China National Patent Office on August 17, 2020, with the application number 202010825602.6 and the application title "A Method and Device for Image Tag Processing in Videos", on August 31, 2020, with the application number 202010894732.5 and the application title "Video Semantic Extraction Method, Video Editing Method and Device", on November 30, 2020, with the application number 202011375148.5 and the application title "Method and Device for a Computer to Identify Video Semantics Using Neural Networks", on December 4, 2020, with the application number 202011405457.2 and the application title "A Method and Device for Determining Video Scenes", and on December 24, 2020, with the application number 202011554281.7 and the application title "A Method and Device for Video Processing". The entire contents of these applications are incorporated herein by reference in their entirety. Technical Field
[0002] This application relates to the field of terminal artificial intelligence (AI), particularly to the fields of video semantic understanding, video editing, video splicing, and video compression, and specifically to a method and device for video semantic recognition. Background Art
[0003] Since videos can display relatively rich content and provide a good audio-visual experience, watching videos has become part of users' daily entertainment activities. To meet users' needs for watching videos and further improve the audio-visual experience of videos, video publishers need to edit the videos to be published.
[0004] For video editing, it is usually necessary to identify the video semantics of multiple original videos, and then the user selects two or more videos from these multiple original videos according to a determined theme based on the video semantics for video splicing to obtain a spliced video that conforms to the theme and publishes it.
[0005] Currently, three-dimensional (3D) convolutional neural networks (CNNs) or video two-stream networks are usually used to recognize the semantics of videos. These two semantic recognition schemes usually classify videos according to preset themes. For example, the preset themes can be set to include two themes: sports and having a birthday. The above two semantic recognition schemes can analyze whether the content of a certain video belongs to sports or having a birthday. For a video that conforms to a certain theme, it may include both segments with a relatively low level of excitement and segments with a relatively high level of excitement. For example, for a video with the theme of sports, the excitement level of a segment that only contains a basketball is lower than that of a passing or dribbling segment; and the passing or dribbling segment is lower than that of a layup segment. Therefore, using the video classification results of the above two semantic recognition schemes for video splicing, the resulting spliced video is either too long or has a relatively low level of excitement.
[0006] Moreover, the three-dimensional convolutional neural network has more parameters, a larger model, and is difficult to train and converge. The video two-stream network uses optical flow information to extract the temporal information of the video, and the calculation takes a long time. Summary of the Invention
[0007] This application provides a video semantic recognition method and device.
[0008] Among them, the methods provided in some embodiments can extract different levels of semantics for different temporal segments in the same video; among them, for the same theme, its different levels of semantics respectively correspond to different levels of excitement. In some embodiments, when splicing videos, the temporal segments with a relatively high level of excitement can be preferentially selected. Thus, a video with a high level of excitement can be spliced conveniently and quickly.
[0009] In the first aspect, this application provides the following multiple method embodiments and device embodiments, including:
[0010] Embodiment 1. A computer-executed method for extracting video semantics using a neural network, the neural network including an input layer, a spatial domain feature extraction layer, a static semantics recognition layer, a dynamic semantics recognition layer, a temporal segment division layer, and an output layer; wherein, the static semantics recognition layer and the dynamic semantics recognition layer are arranged in parallel; the method includes: at the input layer, obtaining multiple video frames of a video; at the spatial domain feature extraction layer, extracting the spatial domain features of each of the multiple video frames; at the dynamic semantics recognition layer, determining the dynamic semantics of the Nth video frame among the N consecutive video frames according to the spatial domain features of the N consecutive video frames in the multiple video frames; N is a positive integer; at the static semantics recognition layer, determining the static semantics of the first video frame according to the spatial domain features of the first video frame in the multiple video frames; at the temporal segment division layer, when the number of consecutive video frames with the first dynamic semantics is greater than a first threshold, using the consecutive video frames with the first dynamic semantics to synthesize a first temporal segment, and determining the first dynamic semantics as the dynamic semantics of the first temporal segment; at the temporal segment division layer, when the number of consecutive video frames with the first static semantics is greater than a second threshold, using the consecutive video frames with the first static semantics to synthesize a second temporal segment, and determining the first static semantics as the static semantics of the second temporal segment; at the output layer, outputting the dynamic semantics of the first temporal segment and first position information; and outputting the static semantics of the second temporal segment and second position information; wherein, the first position information is represented by the positions of the first video frame and the last video frame in the first temporal segment in the video respectively; the second position information is represented by the positions of the first video frame and the last video frame in the second temporal segment in the video respectively.
[0011] Embodiment 2. According to the method of Embodiment 1, the neural network further includes an exciting temporal segment recognition layer; the method further includes: determining the spatial domain difference information between the first video frame and the second video frame according to the spatial domain features of the first video frame and the spatial domain features of the second video frame; the first video frame and the second video frame are adjacent in the multiple video frames; at the exciting temporal segment recognition layer, determining at least one exciting temporal segment according to the spatial domain difference information between each pair of adjacent video frames in the multiple video frames and the spatial domain features of the multiple video frames.
[0012] Embodiment 3. According to the method of Embodiment 2, the spatial domain features include RGB information, and the spatial domain difference information includes RGB difference information (RGB diff).
[0013] Embodiment 4. In the method according to Embodiment 2, the wonderful timing segment recognition layer includes a one-dimensional convolutional layer and a detailed dynamic semantic classification layer; the one-dimensional convolutional layer includes a first convolutional window, and the first convolutional window corresponds to a first detailed dynamic semantic; determining at least one wonderful timing segment according to the spatial domain difference information between each two adjacent video frames in the plurality of video frames and the spatial domain features of the plurality of video frames includes: in the one-dimensional convolutional layer, using the first convolutional window in the at least one convolutional window to perform convolutional processing on the spatial domain features of the plurality of video frames and the spatial domain differences between each two adjacent video frames in the plurality of video frames to obtain a plurality of convolutional results; in the detailed semantic classification layer, determining a wonderful timing segment having the first detailed semantic according to the plurality of convolutional results.
[0014] Embodiment 5. In the method according to Embodiment 2, the neural network further includes a joint logic judgment layer, and the joint logic judgment layer is the next layer of the timing segment division layer and the wonderful timing segment recognition layer; the first wonderful timing segment in the at least one wonderful timing segment is included in the first timing segment; the method further includes: in the joint logic judgment layer, judging whether the detailed dynamic semantic of the first wonderful timing segment matches the dynamic semantic of the first timing segment; when the detailed dynamic semantic of the first wonderful timing segment matches the dynamic semantic of the first timing segment, outputting the detailed dynamic semantic of the first wonderful timing segment and the third position information in the output layer; wherein, the third position information is represented by the positions of the first video frame and the last video frame in the first wonderful timing segment in the video respectively; when the detailed dynamic semantic of the first wonderful timing segment does not match the dynamic semantic of the first timing segment, the relevant information of the first wonderful timing segment is not output in the output layer.
[0015] Embodiment 6. In the method according to Embodiment 2, the neural network further includes a semantic smoothing layer, and the semantic smoothing layer is the upper layer of the timing segment division layer; the method further includes: in the semantic smoothing layer, smoothing the static semantics of the plurality of video frames according to the dependency relationship of the static semantics between consecutive video frames in the plurality of video frames.
[0016] Embodiment 7. In the method according to Embodiment 6, the smoothing process of the static semantics of the multiple video frames according to the dependency relationship of the static semantics between consecutive video frames in the multiple video frames includes: determining that the static semantics of the third video frame in P consecutive video frames are different from those of other video frames, and the static semantics of the other video frames are the same; P is greater than a third threshold; the other video frames are the video frames in the P consecutive video frames except the third video frame; and updating the static semantics of the third video frame according to the static semantics of the other video frames.
[0017] Embodiment 8. In the method according to Embodiment 1, the method further includes: at the temporal segment division layer, when the number of consecutive video frames with a second dynamic semantics is greater than the first threshold, using the consecutive video frames with the second dynamic semantics to synthesize a third temporal segment, and determining that the second dynamic semantics is the dynamic semantics of the third temporal segment; when the second dynamic semantics is the same as the first dynamic semantics and the number of video frames between the first temporal segment and the second temporal segment is less than a fourth threshold, merging the first temporal segment and the third temporal segment into the same temporal segment.
[0018] Embodiment 9. In the method according to Embodiment 1, the spatial domain features include multiple feature maps obtained by convolving the feature information of the video frames corresponding to the spatial domain features through a first convolutional layer, and the multiple feature maps correspond one by one to the multiple convolutional kernels of the first convolutional layer; the dynamic semantics recognition layer includes a second convolutional layer and a dynamic semantics classification layer; the determining of the dynamic semantics of the Nth video frame in the N consecutive video frames according to the spatial domain features of the N consecutive video frames in the multiple video frames includes: performing a feature map offset process on the N consecutive video frames to obtain the residual spatial domain features of the N consecutive video frames; wherein, the feature map offset process includes: using the first feature map of the kth video frame in the N consecutive video frames to replace the first feature map of the (k + 1)th video frame in the N consecutive video frames, where k sequentially takes integer values from 1 to N - 1; the first feature map of the kth video frame and the first feature map of the (k + 1)th video frame correspond to the same convolutional kernel of the first convolutional layer; in the second convolutional layer, convolving the residual spatial domain features of the N consecutive video frames to obtain the spatio-temporal features of the N consecutive video frames; and in the dynamic semantics classification layer, determining the dynamic semantics of the Nth video frame according to the spatio-temporal features of the N consecutive video frames.
[0019] Embodiment 10. A video editing method includes: obtaining a first theme of a target spliced video and a first duration of the target spliced video; determining a plurality of time-sequential segments whose semantics conform to the first theme; the plurality of time-sequential segments include time-sequential segments with dynamic semantics and time-sequential segments with static semantics; determining, according to the first duration, time-sequential segments for splicing the target spliced video from the plurality of time-sequential segments; wherein, when the total duration of the time-sequential segments with dynamic semantics is equal to or greater than the first duration, determining time-sequential segments for splicing the target spliced video from the time-sequential segments with dynamic semantics; when the total duration of the time-sequential segments with dynamic semantics is less than the first duration, determining that all of the time-sequential segments with dynamic semantics are used for splicing the target spliced video, and determining time-sequential segments for splicing the remaining video segments from the time-sequential segments with static semantics, the duration of the remaining video segments being equal to the difference between the first duration and the total duration.
[0020] Embodiment 11. A video editing method includes: obtaining a first theme of a target spliced video and a first duration of the target spliced video; determining a plurality of time-sequential segments whose semantics conform to the first theme; the plurality of time-sequential segments include time-sequential segments with detailed dynamic semantics and time-sequential segments with dynamic semantics; determining, according to the first duration, time-sequential segments for splicing the target spliced video from the plurality of time-sequential segments; wherein, when the total duration of the time-sequential segments with detailed dynamic semantics is equal to or greater than the first duration, determining time-sequential segments for splicing the target spliced video from the time-sequential segments with detailed dynamic semantics; when the total duration of the time-sequential segments with detailed dynamic semantics is less than the first duration, determining that all of the time-sequential segments with detailed dynamic semantics are used for splicing the target spliced video, and determining time-sequential segments for splicing the remaining video segments from the time-sequential segments with dynamic semantics, the duration of the remaining video segments being equal to the difference between the first duration and the total duration.
[0021] Embodiment 12. A video editing method includes: obtaining a first theme of a target spliced video and a first duration of the target spliced video; determining a plurality of temporal segments whose semantics conform to the first theme; the plurality of temporal segments include temporal segments with detailed dynamic semantics, temporal segments with dynamic semantics, and temporal segments with static semantics; determining, according to the first duration, temporal segments for splicing the target spliced video from the plurality of temporal segments; wherein, when the total duration of the temporal segments with detailed dynamic semantics and the temporal segments with dynamic semantics is less than the first duration, it is determined that all of the temporal segments with detailed dynamic semantics and the temporal segments with dynamic semantics are used for splicing the target spliced video, and temporal segments for splicing the remaining video segments are determined from the temporal segments with static semantics, and the duration of the remaining video segments is equal to the difference between the first duration and the total duration.
[0022] Embodiment 13. An apparatus for extracting video semantics includes: an input unit for obtaining a plurality of video frames of a video; an extraction unit for extracting spatial domain features of each of the plurality of video frames; a first recognition unit for determining dynamic semantics of the Nth video frame among the N consecutive video frames according to the spatial domain features of the N consecutive video frames in the plurality of video frames; N is a positive integer; a second recognition unit for determining static semantics of the first video frame according to the spatial domain features of the first video frame in the plurality of video frames; a partitioning unit for, when the number of consecutive video frames with first dynamic semantics is greater than a first threshold, synthesizing a first temporal segment using the consecutive video frames with first dynamic semantics, and determining the first dynamic semantics as the dynamic semantics of the first temporal segment; the partitioning unit is further configured to, when the number of consecutive video frames with first static semantics is greater than a second threshold, synthesize a second temporal segment using the consecutive video frames with first static semantics, and determine the first static semantics as the static semantics of the second temporal segment; an output unit for outputting the dynamic semantics of the first temporal segment and first position information; and outputting the static semantics of the second temporal segment and second position information; wherein, the first position information is represented by the positions of the first video frame and the last video frame in the first temporal segment in the video respectively; the second position information is represented by the positions of the first video frame and the last video frame in the second temporal segment in the video respectively.
[0023] Embodiment 14. The device according to Embodiment 13 further includes a third recognition unit; the third recognition unit is configured to: determine the spatial domain difference information between the first video frame and the second video frame according to the spatial domain features of the first video frame and the spatial domain features of the second video frame; the first video frame and the second video frame are adjacent in the plurality of video frames; determine at least one exciting timing segment according to the spatial domain difference information between each two adjacent video frames in the plurality of video frames and the spatial domain features of the plurality of video frames.
[0024] Embodiment 15. The device according to Embodiment 14, wherein the spatial domain features include RGB information, and the spatial domain difference information includes RGB difference information (RGB diff).
[0025] Embodiment 16. The device according to Embodiment 14, wherein the third recognition unit includes a convolution unit and a classification unit; the convolution unit is configured to perform convolution processing on the spatial domain features of the plurality of video frames and the spatial domain differences between each two adjacent video frames in the plurality of video frames by using a first convolution window in at least one convolution window to obtain a plurality of convolution results; the first convolution window corresponds to a first detailed dynamic semantics; the classification unit is configured to determine an exciting timing segment with the first detailed semantics according to the plurality of convolution results.
[0026] Embodiment 17. The first exciting timing segment in the at least one exciting timing segment is included in the first timing segment; the device further includes a judgment unit; the judgment unit is configured to judge whether the detailed dynamic semantics of the first exciting timing segment match the dynamic semantics of the first timing segment; the output unit is configured to output the detailed dynamic semantics of the first exciting timing segment and the third position information at the output layer when the detailed dynamic semantics of the first exciting timing segment match the dynamic semantics of the first timing segment; wherein, the third position information is represented by the positions of the first video frame and the last video frame in the first exciting timing segment in the video respectively; the output unit is configured to not output the relevant information of the first exciting timing segment at the output layer when the detailed dynamic semantics of the first exciting timing segment do not match the dynamic semantics of the first timing segment.
[0027] Embodiment 18. The device according to Embodiment 13 further includes a smoothing unit configured to smooth the static semantics of the plurality of video frames according to the dependency relationship of the static semantics between consecutive video frames in the plurality of video frames.
[0028] Example 19. For the device according to Example 18, the smoothing unit is further configured to: determine that the static semantics of the third video frame in P consecutive video frames are different from the static semantics of other video frames, and the static semantics of the other video frames are the same; P is greater than a third threshold; the other video frames are the video frames other than the third video frame in the P consecutive video frames; and update the static semantics of the third video frame according to the static semantics of the other video frames.
[0029] Example 20. For the device according to Example 13, the device further includes: a merging unit; when the number of consecutive video frames with a second dynamic semantics is greater than the first threshold, the merging unit is configured to synthesize a third time series segment using the consecutive video frames with the second dynamic semantics, and determine that the second dynamic semantics are the dynamic semantics of the third time series segment; when the second dynamic semantics are the same as the first dynamic semantics, and the number of video frames between the first time series segment and the second time series segment is less than a fourth threshold, the merging unit is further configured to merge the first time series segment and the third time series segment into the same time series segment.
[0030] Example 21. For the device according to the example, the spatial domain features include a plurality of feature maps obtained by convolving the feature information of the video frames corresponding to the spatial domain features through a first convolutional unit, and the plurality of feature maps correspond one-to-one to the plurality of convolutional kernels of the first convolutional unit; the first recognition unit includes an offset unit, a second convolutional unit, and a dynamic semantics classification unit; the offset unit is configured to perform feature map offset processing on the N consecutive video frames to obtain the residual spatial domain features of the N consecutive video frames; wherein, the feature map offset processing includes: using the first feature map of the kth video frame in the N consecutive video frames to replace the first feature map of the (k + 1)th video frame in the N consecutive video frames, where k sequentially takes integer values from 1 to N - 1; the first feature map of the kth video frame and the first feature map of the (k + 1)th video frame correspond to the same convolutional kernel of the first convolutional layer; the second convolutional unit is configured to convolve the residual spatial domain features of the N consecutive video frames to obtain the spatio-temporal features of the N consecutive video frames; the dynamic semantics classification unit is configured to determine the dynamic semantics of the Nth video frame according to the spatio-temporal features of the N consecutive video frames.
[0031] Embodiment 22. A video editing device, comprising: an acquisition unit configured to acquire a first theme of a target spliced video and a first duration of the target spliced video; a first determination unit configured to determine a plurality of temporal segments whose semantics conform to the first theme; the plurality of temporal segments include temporal segments with dynamic semantics and temporal segments with static semantics; a second determination unit configured to determine, according to the first duration, temporal segments for splicing the target spliced video from the plurality of temporal segments; wherein, when the total duration of the temporal segments with dynamic semantics is equal to or greater than the first duration, determine temporal segments for splicing the target spliced video from the temporal segments with dynamic semantics; when the total duration of the temporal segments with dynamic semantics is less than the first duration, determine that all of the temporal segments with dynamic semantics are used for splicing the target spliced video, and determine temporal segments for splicing the remaining video segments from the temporal segments with static semantics, the duration of the remaining video segments being equal to the difference between the first duration and the total duration.
[0032] Embodiment 23. A video editing device, comprising: an acquisition unit configured to acquire a first theme of a target spliced video and a first duration of the target spliced video; a first determination unit configured to determine a plurality of temporal segments whose semantics conform to the first theme; the plurality of temporal segments include temporal segments with detailed dynamic semantics and temporal segments with dynamic semantics; a second determination unit configured to determine, according to the first duration, temporal segments for splicing the target spliced video from the plurality of temporal segments; wherein, when the total duration of the temporal segments with detailed dynamic semantics is equal to or greater than the first duration, determine temporal segments for splicing the target spliced video from the temporal segments with detailed dynamic semantics; when the total duration of the temporal segments with detailed dynamic semantics is less than the first duration, determine that all of the temporal segments with detailed dynamic semantics are used for splicing the target spliced video, and determine temporal segments for splicing the remaining video segments from the temporal segments with dynamic semantics, the duration of the remaining video segments being equal to the difference between the first duration and the total duration.
[0033] Embodiment 24. A video editing device includes: an acquisition unit configured to acquire a first theme of a target spliced video and a first duration of the target spliced video; a first determination unit configured to determine a plurality of temporal segments whose semantics conform to the first theme; the plurality of temporal segments include temporal segments with detailed dynamic semantics, temporal segments with dynamic semantics, and temporal segments with static semantics; a second determination unit configured to determine, according to the first duration, temporal segments for splicing the target spliced video from the plurality of temporal segments; wherein, when the total duration of the temporal segments with detailed dynamic semantics and the temporal segments with dynamic semantics is less than the first duration, it is determined that all of the temporal segments with detailed dynamic semantics and the temporal segments with dynamic semantics are used for splicing the target spliced video, and temporal segments for splicing the remaining video segments are determined from the temporal segments with static semantics, and the duration of the remaining video segments is equal to the difference between the first duration and the total duration.
[0034] Embodiment 25. An electronic device includes: a processor and a memory; the memory is configured to store computer instructions; when the electronic device runs, the processor executes the computer instructions, so that the electronic device executes the method according to any one of Embodiments 1-9 or the method according to Embodiment 10 or the method according to Embodiment 11 or the method according to Embodiment 12.
[0035] Embodiment 26. A computer storage medium includes computer instructions, and when the computer instructions run on an electronic device, the electronic device is caused to execute the method according to any one of Embodiments 1-9 or the method according to Embodiment 10 or the method according to Embodiment 11 or the method according to Embodiment 12.
[0036] Embodiment 26. A computer program product, when the program code included in the computer program product is executed by a processor in an electronic device, implements the method according to any one of Embodiments 1-9 or the method according to Embodiment 10 or the method according to Embodiment 11 or the method according to Embodiment 12.
[0037] The video semantic recognition method provided by some of the above embodiments of the present application can identify temporal segments with static semantics and video segments with dynamic semantics in a video. Thus, when editing a video, temporal segments with static semantics or video segments with dynamic semantics can be selected as needed for video splicing, so that a more wonderful video can be obtained.
[0038] In a second aspect, the present application provides the following multiple method embodiments and device embodiments, including:
[0039] Embodiment 1. A method for processing video frame tags, comprising: obtaining a plurality of video frames of a video, each of the plurality of video frames being capable of carrying one or more category tags; the category tags being used to characterize the category of an object in the corresponding video frame; performing tag smoothing processing on the plurality of video frames according to a tag smoothing strategy corresponding to the video; wherein, when the tag smoothing strategy includes that a first category tag is continuous between adjacent video frames in the plurality of video frames, the performing tag smoothing processing on the plurality of video frames includes: when the first video frame and the second video frame in the plurality of video frames both carry or both do not carry the first category tag, determining that the scores of the first video frame and the second video frame under the first category tag are W; when one of the first video frame and the second video frame carries the first category tag and the other does not carry the first category tag, determining that the scores of the first video frame and the second video frame under the first category tag are W'; W' is less than W; the first video frame and the second video frame are two adjacent video frames in the plurality of video frames and form a pair of adjacent video frames; adding or deleting the first category tag of K video frames in the plurality of video frames so that the score of the plurality of video frames under the first category tag is maximized; the score of the plurality of video frames under the first category tag includes the sum of the scores of each pair of adjacent video frames in the plurality of video frames under the first category tag; K≥0 and is an integer; wherein, the K video frames include a third video frame, and the adding or deleting the first category tag of K video frames in the plurality of video frames includes: if the third video frame does not carry the first category tag when obtaining the plurality of video frames, adding the first category tag to the third video frame; if the third video frame carries the first category tag when obtaining the plurality of video frames, deleting the first category tag of the third video frame.
[0040] Embodiment 2. According to the method described in Embodiment 1, the score of the plurality of video frames under the first category tag further includes the sum of the tag consistency scores of each video frame in the plurality of video frames; the adding or deleting the first category tag of K video frames in the plurality of video frames so that the score of the plurality of video frames under the first category tag is maximized includes: determining that the tag consistency score of the video frames in the plurality of video frames except the K video frames is E; determining that the tag consistency score of the K video frames is E'; E' is less than E.
[0041] Example 3. The method according to Example 2, where the first category labels of K video frames among the multiple video frames are added or deleted to maximize the score of the multiple video frames under the first category label includes: solving the maximum sum of formula (I), and determining the K video frames according to the maximum sum, and adding or deleting the first category labels of the K video frames;
[0042] Where;
[0043] Max(N 1 +N 2 +N 3 ) (I);
[0044]
[0045]
[0046]
[0047] n is the number of the multiple video frames; w 1 、w 2 、w 3 are preset positive numbers; the absolute value of X 0,0 is 1; i corresponds to the first category label, and j corresponds to the j-th video frame among the multiple video frames; the multiple video frames = set C1 ∪ set C2; where, the elements in set C1 are the video frames carrying the first category label when obtaining the multiple video frames; the elements in set C2 are the video frames not carrying the first category label when obtaining the multiple video frames; during the solving process, X i,j and X i,j+1 take values between X 0,0 and -X 0,0 .
[0048] Example 4. The method according to Example 3, where the maximum sum of formula (I) is solved, and the K video frames are determined according to the maximum sum, and the first category labels of the K video frames are added or deleted includes:
[0049] Determine the column vector a = [X 0,0 ; X i,1 ;...; X i,j ;...; X i,n ;
[0050] Multiply the column vector a by the transpose vector of the column vector a to obtain the matrix M;
[0051] Taking the maximum of the summation result of formula (I) as the objective and the value of the element on the main diagonal of the matrix M being 1 as the constraint condition, solve using the ellipse method or the interior point method to obtain the positive semi - definite matrix M′;
[0052] According to the positive semi - definite matrix M′, determine the K video frames, and add or delete the first - type labels of the K video frames.
[0053] Example 5. According to the method described in Example 1, the first video frame and the second video frame are in the
[0054] where w 1 is a preset positive number; the absolute value of X 0,0 is 1; i corresponds to the first - type label, j corresponds to the first video frame, j + 1 corresponds to the second video frame; when the first video frame carries the first - type label, X i,j = X 0,0 ; when the first video frame does not carry the first - type label, X i,j = - X 0,0 ; when the second video frame carries the first - type label, X i,j+1 = X 0,0 ; when the second video frame does not carry the first - type label, X i,j+1 = - X 0,0 .
[0055] Example 6. According to the method described in Example 1, when the video - frame label smoothing strategy includes the co - existence of the second - type label and the third - type label on the same video frame, the label smoothing process for the multiple video frames includes: when the fourth video frame among the multiple video frames simultaneously carries or simultaneously does not carry the second - type label and the third - type label, determine the score of the fourth video frame under the second - type label and the third - type label as F; when the fourth video frame carries one of the second - type label and the third - type label and does not carry the other, determine the score of the fourth video frame under the second - type label and the third - type label as F′; F′ is less than F; add the third - type label to L video frames among the multiple video frames and / or delete the second - type label of L′ video frames among the multiple video frames, so that the scores of the multiple video frames under the second - type label and the third - type label are maximized; L≥0 and is an integer; L′≥0 and is an integer; the scores of the multiple video frames under the second - type label and the third - type label include the sum of the scores of each video frame among the multiple video frames under the second - type label and the third - type label.
[0056] Example 7. The method according to Example 6, wherein the fourth video frame is in the second category label and the
[0057] wherein, w 3 is a preset positive number; the absolute value of X 0,0 is 1; i * corresponds to the second category label, i corresponds to the third category label; j corresponds to the fourth video frame; when the fourth video frame carries the second category label, in X i*,j = X 0,0 ; when the fourth video frame carries the third category label, in X i,j = X 0,0 ; when the fourth video frame does not carry the third category label, X i,j = -X 0,0 .
[0058] Example 8. The method according to Example 1, when the label smoothing strategy includes that the fourth category label and the fifth category label do not coexist on the same video frame, the label smoothing process for the multiple video frames includes: when the fifth video frame among the multiple video frames does not carry both the fourth category label and the fifth category label at the same time, determining the score of the fifth video frame under the fourth category label and the fifth category label as H; when the fifth video frame carries both the fourth category label and the fifth category label at the same time, determining the score of the fifth video frame under the fourth category label and the fifth category label as H'; H' is less than H; deleting the fifth category labels of P video frames among the multiple video frames and / or deleting the fourth category labels of P' video frames among the multiple video frames, so that the scores of the multiple video frames under the fourth category label and the fifth category label are maximized; P≥0 and is an integer; P'≥0 and is an integer; the scores of the multiple video frames under the fourth category label and the fifth category label include the sum of the scores of each video frame among the multiple video frames under the fourth category label and the fifth category label.
[0059] Example 9. The method according to Example 8, the fifth video frame is in the fourth category label and the
[0060] wherein, w 4 is a preset positive number; the absolute value of X 0,0 is 1; i *Corresponding to the fourth category label, i corresponds to the fifth category label; j corresponds to the fifth video frame; when the fifth video frame carries the fourth category label, at X i*,j = X 0,0 ; when the fifth video frame carries the fifth category label, at X i,j = X 0,0 ; when the fifth video frame does not carry the fifth category label, X i,j = -X 0,0 .
[0061] Example 10. According to the method described in Example 1, when the label smoothing strategy includes that a sixth category label exists in the previous video frame among adjacent video frames, and a seventh category label exists in the subsequent video frame among the adjacent video frames; the previous video frame and the subsequent video frame are the front and rear video frames in the time sequence of the video; the label smoothing process for the multiple video frames includes: when the previous video frame in the first pair of adjacent video frame groups among the multiple video frames carries the sixth category label, and the subsequent video frame in the first pair of adjacent video frame groups carries the seventh category label; or, when the previous video frame in the first pair of adjacent video frame groups among the multiple video frames does not carry the sixth category label, and the subsequent video frame in the first pair of adjacent video frame groups does not carry the seventh category label; determining the score of the first pair of adjacent video frame groups under the sixth category label and the seventh category label as Z; when the previous video frame in the first pair of adjacent video frame groups carries the sixth category label, and the subsequent video frame in the first pair of adjacent video frame groups does not carry the seventh category label; or, when the previous video frame in the first pair of adjacent video frame groups does not carry the sixth category label, and the subsequent video frame in the first pair of adjacent video frame groups carries the seventh category label; determining the score of the first pair of adjacent video frame groups under the sixth category label and the seventh category label as Z'; Z' is less than Z; adding the seventh category label to Q video frames among the multiple video frames and / or deleting the sixth category of Q' video frames among the multiple video frames, so that the score of the multiple video frames under the sixth category label and the seventh category label is maximized; Q≥0 and is an integer; Q'≥0 and is an integer; the score of the multiple video frames under the sixth category label and the seventh category label includes the sum of the scores of each pair of adjacent video frame groups in the first category label among the multiple video frames.
[0062] Example 11. According to the method described in Example 10, the first pair of adjacent video frame groups under the sixth category label and the
[0063] wherein, w5 is a preset positive number; X 0,0 has an absolute value of 1; i * corresponds to the sixth category label, i corresponds to the seventh category label; j corresponds to the previous video frame in the first pair of adjacent video frames, and j + 1 corresponds to the subsequent video frame in the first pair of adjacent video frames; when the previous video frame in the first pair of adjacent video frames carries the sixth category label, at X i*,j = X 0,0 ; when the subsequent video frame in the first pair of adjacent video frames carries the seventh category label, X i,j+1 = X 0,0 ; when the subsequent video frame in the first pair of adjacent video frames carries the seventh category label, X i,j+1 = -X 0,0 .
[0064] Embodiment 12. A video frame label processing device, comprising:
[0065] An acquisition unit, configured to acquire a plurality of video frames of a video, and each video frame in the plurality of video frames may carry one or more category labels; the category labels are used to characterize the category of an object in the corresponding video frame;
[0066] A processing unit, configured to perform label smoothing processing on the plurality of video frames according to a label smoothing strategy corresponding to the video; wherein,
[0067] when the label smoothing strategy includes that the first category label is continuous between adjacent video frames in the plurality of video frames, the processing unit is configured to:
[0068] when the first video frame and the second video frame in the plurality of video frames both carry or both do not carry the first category label, determine that the scores of the first video frame and the second video frame under the first category label are W; when one of the first video frame and the second video frame carries the first category label and the other does not carry the first category label, determine that the scores of the first video frame and the second video frame under the first category label are W′; W′ is less than W; the first video frame and the second video frame are two adjacent video frames in the plurality of video frames and form a pair of adjacent video frames;
[0069] add or delete the first category labels of K video frames in the plurality of video frames so that the scores of the plurality of video frames under the first category label are maximized; the scores of the plurality of video frames under the first category label include the sum of the scores of each pair of adjacent video frames in the plurality of video frames under the first category label; K≥0 and is an integer;
[0070] Among them, the K video frames include a third video frame, and the processing unit is configured to:
[0071] If the third video frame does not carry the first category label when obtaining the multiple video frames, add the first category label to the third video frame;
[0072] If the third video frame carries the first category label when obtaining the multiple video frames, delete the first category label of the third video frame.
[0073] Embodiment 13. For the device according to Embodiment 12, the scores of the multiple video frames under the first category label further include the sum of the label consistency scores of each video frame in the multiple video frames; the processing unit is further configured to:
[0074] Determine that the label consistency score of the video frames other than the K video frames in the multiple video frames is E;
[0075] Determine that the label consistency score of the K video frames is E'; E' is less than E.
[0076] Embodiment 14. For the device according to Embodiment 13, the processing unit is further configured to: solve the maximum sum of formula (I), and determine the K video frames according to the maximum sum, and add and delete the first category labels of the K video frames;
[0077] Wherein;
[0078] Max(N 1 +N 2 +N 3 ) (I);
[0079]
[0080]
[0081]
[0082] n is the number of the multiple video frames; w 1 、w 2 、w 3 are preset positive numbers; the absolute value of X 0,0 is 1; i corresponds to the first category label, and j corresponds to the jth video frame in the multiple video frames; the multiple video frames = set C1 ∪ set C2; wherein, the elements in set C1 are the video frames that carry the first category label when obtaining the multiple video frames; the elements in set C2 are the video frames that do not carry the first category label when obtaining the multiple video frames; in the solving process, in Xi,j and X i,j+1 at X 0,0 and -X 0,0 to take values between them.
[0083] Example 15. For the apparatus according to Example 14, the processing unit is further configured to:
[0084] determine the column vector a = [X 0,0 ; X i,1 ;...; X i,j ;...; X i,n ;
[0085] multiply the column vector a by the transposed vector of the column vector a to obtain the matrix M;
[0086] aiming at maximizing the summation result of formula (I) and with the value of the element on the main diagonal of the matrix M being 1 as a constraint condition, solve using the ellipse method or the interior point method to obtain the positive semi - definite matrix M';
[0087] determine the K video frames according to the positive semi - definite matrix M', and add or delete the first - type labels of the K video frames.
[0088] Example 16. For the apparatus according to Example 12, the first video frame and the second video frame are in the
[0089] where w 1 is a preset positive number; the absolute value of X 0,0 is 1; i corresponds to the first - type label, j corresponds to the first video frame, and j + 1 corresponds to the second video frame; when the first video frame carries the first - type label, at X i,j = X 0,0 ; when the first video frame does not carry the first - type label, X i,j = -X 0,0 ; when the second video frame carries the first - type label, at X i,j+1 = X 0,0 ; when the second video frame does not carry the first - type label, X i,j+1 = -X 0,0 .
[0090] Example 17. For the apparatus according to Example 12, when the video frame label smoothing strategy includes co - existence of the second - type label and the third - type label on the same video frame, the processing unit is further configured to:
[0091] When the fourth video frame among the multiple video frames carries or does not carry the second category label and the third category label simultaneously, determine that the score of the fourth video frame under the second category label and the third category label is F; when the fourth video frame carries one of the second category label and the third category label and does not carry the other, determine that the score of the fourth video frame under the second category label and the third category label is F'; F' is less than F.
[0092] Add the third category label to L video frames among the multiple video frames and / or delete the second category label of L' video frames among the multiple video frames, so as to maximize the score of the multiple video frames under the second category label and the third category label; L≥0 and is an integer; L'≥0 and is an integer; the score of the multiple video frames under the second category label and the third category label includes the sum of the scores of each video frame among the multiple video frames under the second category label and the third category label.
[0093] Example 18. The device according to claim 17, wherein the fourth video frame is under the second category label and the
[0094] wherein, w 3 is a preset positive number; the absolute value of X 0,0 is 1; i * corresponds to the second category label, i corresponds to the third category label; j corresponds to the fourth video frame; when the fourth video frame carries the second category label, in X i*,j = X 0,0 ; when the fourth video frame carries the third category label, in X i,j = X 0,0 ; when the fourth video frame does not carry the third category label, X i,j = -X 0,0 .
[0095] Example 19. The device according to Example 12, when the label smoothing strategy includes that the fourth category label and the fifth category label do not coexist on the same video frame, the processing unit is further configured to:
[0096] When the fifth video frame among the multiple video frames does not carry the fourth category label and the fifth category label simultaneously, determine that the score of the fifth video frame under the fourth category label and the fifth category label is H; when the fifth video frame carries the fourth category label and the fifth category label simultaneously, determine that the score of the fifth video frame under the fourth category label and the fifth category label is H'; H' is less than H.
[0097] Delete the fifth category label of P video frames among the multiple video frames and / or delete the fourth category label of P' video frames among the multiple video frames, so as to maximize the scores of the multiple video frames under the fourth category label and the fifth category label; P≥0 and is an integer; P'≥0 and is an integer; the scores of the multiple video frames under the fourth category label and the fifth category label include the sum of the scores of each video frame in the multiple video frames under the fourth category label and the fifth category label.
[0098] Example 20. The device according to claim 19, wherein the fifth video frame is under the fourth category label and the
[0099] wherein, w 4 is a preset positive number; the absolute value of X 0,0 is 1; i * corresponds to the fourth category label, and i corresponds to the fifth category label; j corresponds to the fifth video frame; when the fifth video frame carries the fourth category label, X i*,j =X 0,0 ; when the fifth video frame carries the fifth category label, X i,j =X 0,0 ; when the fifth video frame does not carry the fifth category label, X i,j =-X 0,0 .
[0100] 21. The device according to claim 12, when the label smoothing strategy includes that the sixth category label exists in the previous video frame among adjacent video frames, and the seventh category label exists in the subsequent video frame among the adjacent video frames; the previous video frame and the subsequent video frame are the front and rear video frames in the time sequence of the video; the processing unit is further configured to:
[0101] When the previous video frame in the first pair of adjacent video frames among the multiple video frames carries the sixth category label, and the subsequent video frame in the first pair of adjacent video frames carries the seventh category label; or, when the previous video frame in the first pair of adjacent video frames among the multiple video frames does not carry the sixth category label, and the subsequent video frame in the first pair of adjacent video frames does not carry the seventh category label; determine that the score of the first pair of adjacent video frames under the sixth category label and the seventh category label is Z;
[0102] When the previous video frame in the first pair of adjacent video frame groups carries the sixth category label and the subsequent video frame in the first pair of adjacent video frame groups does not carry the seventh category label; or when the previous video frame in the first pair of adjacent video frame groups does not carry the sixth category label and the subsequent video frame in the first pair of adjacent video frame groups carries the seventh category label; determine that the score of the first pair of adjacent video frame groups under the sixth category label and the seventh category label is Z′; Z′ is less than Z;
[0103] Add the seventh category label to Q video frames among the multiple video frames and / or delete the sixth category of Q′ video frames among the multiple video frames, so that the score of the multiple video frames under the sixth category label and the seventh category label is maximized; Q≥0 and is an integer; Q′≥0 and is an integer; the score of the multiple video frames under the sixth category label and the seventh category label includes the sum of the scores of each pair of adjacent video frame groups among the multiple video frames under the first category label.
[0104] Example 22. The device according to Example 21, the first pair of adjacent video frame groups under the sixth category label and the
[0105] Wherein, w 5 is a preset positive number; the absolute value of X 0,0 is 1; i * corresponds to the sixth category label, i corresponds to the seventh category label; j corresponds to the previous video frame in the first pair of adjacent video frame groups, and j + 1 corresponds to the subsequent video frame in the first pair of adjacent video frame groups; when the previous video frame in the first pair of adjacent video frame groups carries the sixth category label, in X i*,j = X 0,0 ; when the subsequent video frame in the first pair of adjacent video frame groups carries the seventh category label, X i,j+1 = X 0,0 ; when the subsequent video frame in the first pair of adjacent video frame groups carries the seventh category label, X i,j+1 = -X 0,0 .
[0106] Example 23. An electronic device, comprising: a processor, a memory, and a transceiver; the memory is used to store computer instructions; when the electronic device runs, the processor executes the computer instructions, so that the electronic device executes the method according to any one of Examples 1-11.
[0107] Embodiment 24. A computer storage medium includes computer instructions that, when run on an electronic device, cause the electronic device to execute the method according to any one of Embodiments 1-11.
[0108] Embodiment 25. A computer program product, when the program code included in the computer program product is executed by a processor in an electronic device, implements the method according to any one of Embodiments 1-11.
[0109] In some embodiments provided in the second aspect, by providing the video frame tag processing method and device, the tags carried by the video frames in the video can be corrected according to the tag smoothing strategy corresponding to the video, so that the tag stream of the video is smooth or more in line with common sense.
[0110] In the third aspect, the present application provides the following multiple method embodiments and device embodiments, including:
[0111] Embodiment 1. A method for a computer to identify video semantics using a neural network, the neural network including an input layer, a spatial domain feature extraction layer, a plurality of offset layers, a classification layer, and an output layer; wherein, the spatial domain feature extraction layer includes at least one two-dimensional convolutional layer arranged in series; the plurality of offset layers include a first offset layer and a second offset layer arranged in parallel; the method includes:
[0112] In the input layer, obtain N video frames in a first video, where N is a positive integer greater than 1;
[0113] In the spatial domain feature extraction layer, extract the spatial domain features of each of the N video frames in a plurality of channels; wherein, different channels in the plurality of channels correspond to different convolutional kernels of the last convolutional layer in the at least one two-dimensional convolutional layer;
[0114] In the first offset layer, perform a first temporal offset on the spatial domain features of the N video frames in at least one of the plurality of channels to obtain the first spatial domain features of each of the N video frames;
[0115] In the second offset layer, perform a second temporal offset on the spatial domain features of the N video frames in at least one of the plurality of channels to obtain the second spatial domain features of each of the N video frames; wherein, the time offset amounts of the first temporal offset and the second temporal offset are different, or the offset channels are different;
[0116] In the classification layer, determine the semantics of the first video based at least on the first spatial domain features and the second spatial domain features of each of the N video frames;
[0117] At the output layer, output the semantics of the first video.
[0118] Example 2. The method according to Example 1, wherein the first temporal offset includes: at the first offset layer, performing a first temporal offset on the spatial domain features of each of the N video frames in at least one of the multiple channels, the at least one channel including a first channel, and a time offset corresponding to the first temporal offset being T, where T is a positive integer greater than or equal to 1 and less than N, or T is a negative integer greater than -N and less than or equal to -1, such that the spatial domain features of the k-th video frame among the N video frames in the first channel are offset to the spatial domain features of the (k + T)-th video frame in the first channel, and k sequentially takes positive integer values in the interval [1, N] to obtain the first spatial domain features of each of the N video frames;
[0119] The second temporal offset includes: at the second offset layer, performing a second temporal offset on the spatial domain features of each of the N video frames in at least one of the multiple channels, the at least one channel including a second channel, and a time offset corresponding to the second temporal offset being T′, where T′ is a positive integer greater than or equal to 1 and less than N, or T′ is a negative integer greater than -N and less than or equal to -1, such that the spatial domain features of the k-th video frame among the N video frames in the second channel are offset to the spatial domain features of the (k + T′)-th video frame in the second channel, and k sequentially takes positive integer values in the interval [1, N] to obtain the second spatial domain features of each of the N video frames.
[0120] Example 3. The method according to Example 2, the method further includes: T is not equal to T′, and the first channel and the second channel are the same channel; or,
[0121] T is equal to T′, and the first channel and the second channel are different channels; or,
[0122] T is not equal to T′, and the first channel and the second channel are different channels.
[0123] Example 4. The method according to Example 1, wherein the classification layer includes a plurality of two-dimensional convolutional layers arranged in parallel, and the convolutional layers in the plurality of two-dimensional convolutional layers correspond one-to-one with the offset layers in the plurality of offset layers;
[0124] Determining the semantics of the first video at least according to the first spatial domain features and the second spatial domain features of each of the N video frames includes:
[0125] In the first convolutional layer of the multiple two-dimensional convolutional layers, perform convolutional processing on the first spatial domain feature of the first video frame among the N video frames to obtain the first fused spatio-temporal feature of the first video frame; the first convolutional layer corresponds to the first offset layer;
[0126] In the second convolutional layer of the multiple two-dimensional convolutional layers, perform convolutional processing on the second spatial domain feature of the first video frame to obtain the second fused spatio-temporal feature of the first video frame;
[0127] Determine the semantics of the first video frame at least according to the first fused spatio-temporal feature and the second fused spatio-temporal feature of each video frame among the N video frames.
[0128] Example 5. According to the method described in claim 4, both the first fused spatio-temporal feature and the second fused spatio-temporal feature include features under M channels, and the features under each of the M channels are represented by matrices; M is a positive integer;
[0129] The determining the semantics of the first video frame at least according to the first fused spatio-temporal feature and the second fused spatio-temporal feature of each video frame among the N video frames includes:
[0130] Perform element-wise addition on the features under the i-th channel among the M channels included in the first fused spatio-temporal feature of the first video frame and the features under the i-th channel among the M channels included in the second fused spatio-temporal feature of the first video frame, where i takes integer values in the interval [1, M] in sequence, to obtain the third fused spatio-temporal feature of the first video frame;
[0131] Determine the semantics of the first video according to the third fused spatio-temporal feature of each video frame among the N video frames.
[0132] Example 6. According to the method described in claim 1, the determining the semantics of the first video at least according to the first spatial domain feature and the second spatial domain feature of each video frame among the N video frames includes:
[0133] Perform feature compensation on the first spatial domain feature of the first video frame according to the spatial domain features of the first video frame among the N video frames under the multiple channels to obtain a first residual spatial domain feature;
[0134] Perform feature compensation on the second spatial domain feature of the first video frame according to the spatial domain features of the first video frame among the N video frames under the multiple channels to obtain a second residual spatial domain feature;
[0135] Determine the semantics of the first video according to the first residual spatial domain feature and the second residual spatial domain feature.
[0136] Embodiment 7. A computer - executed method for identifying video semantics using a neural network, where the neural network includes an input layer, a spatial - domain feature extraction layer, at least one residual network layer arranged in series, a classification layer, and an output layer; wherein, the spatial - domain feature layer includes at least one two - dimensional convolutional layer arranged in series; each of the at least one residual network layer includes a plurality of spatio - temporal feature extraction layers and a spatial - domain feature compensation layer arranged in series; wherein, each spatio - temporal feature extraction layer in the plurality of spatio - temporal feature extraction layers includes an offset sub - layer and a convolutional sub - layer arranged in series;
[0137] The method includes:
[0138] At the input layer, obtain N video frames in a first video, where N is a positive integer greater than 1;
[0139] At the spatial - domain feature extraction layer, extract the spatial - domain features of each of the N video frames in multiple channels; wherein, different channels in the multiple channels correspond to different convolutional kernels of the last convolutional layer of the at least one two - dimensional convolutional layer;
[0140] At the first spatio - temporal feature extraction layer of the first residual network layer in the at least one residual network layer, extract the first fused spatio - temporal features of the N video frames; when the first spatio - temporal feature layer is the last one in the multiple spatio - temporal feature extraction layers of the first residual network layer, at the spatial - domain feature compensation layer of the first residual network layer, determine the residual spatial - domain features of the N video frames to be output by the first residual network layer according to the first fused spatio - temporal features of the N video frames and the spatial - domain features of the N videos in the multiple channels; wherein, extracting the first fused spatio - temporal features of the N video frames includes: at the offset sub - layer of the first spatio - temporal feature extraction layer, perform temporal offset on the spatial - domain features of the N video frames in at least one channel to obtain the first spatial - domain features of each of the N video frames, and the spatial - domain features in the at least one channel are a part of the spatial - domain features output by the previous layer of the first spatio - temporal feature extraction layer; at the convolutional sub - layer of the first spatio - temporal feature extraction layer, perform convolutional processing on the first spatial - domain features of each of the N video frames to obtain the first fused spatio - temporal features of the N video frames;
[0141] At the classification layer, determine the semantics of the first video according to the residual spatial - domain features of the N video frames output by the last residual network layer in the at least one residual network layer;
[0142] At the output layer, output the semantics of the first video.
[0143] Example 8. In the method according to 7, the temporal offset of the spatial domain features of the N video frames in at least one channel to obtain the first spatial domain features of each video frame in the N video frames includes:
[0144] Perform temporal offset on the spatial domain features of each video frame in the N video frames in the at least one channel. The at least one channel includes a first channel, and the time offset amount corresponding to the temporal offset is T, where T is a positive integer greater than or equal to 1 and less than N, or T is a negative integer greater than -N and less than or equal to -1, such that the spatial domain features of the k-th video frame in the N video frames in the first channel are offset to the spatial domain features of the (k + T)-th video frame in the first channel. k takes positive integer values in the interval [1, N] in sequence to obtain the first spatial domain features of each video frame in the N video frames.
[0145] Example 9. An apparatus for identifying video semantics includes: an input unit, a spatial domain feature extraction unit, a plurality of offset units, a classification unit, and an output unit; wherein, the spatial domain feature extraction unit includes at least one two-dimensional convolution unit arranged in series, and the plurality of offset units includes a first offset unit and a second offset unit arranged in parallel;
[0146] The input unit is configured to obtain N video frames of a first video, where N is a positive integer greater than 1;
[0147] The spatial domain feature extraction unit is configured to extract the spatial domain features of each video frame in the N video frames in a plurality of channels; wherein, different channels in the plurality of channels correspond to different convolution kernels of the last convolution unit in the at least one two-dimensional convolution unit;
[0148] The first offset unit is configured to perform a first temporal offset on the spatial domain features of the N video frames in at least one channel among the plurality of channels to obtain the first spatial domain features of each video frame in the N video frames;
[0149] The second offset unit is configured to perform a second temporal offset on the spatial domain features of the N video frames in at least one channel among the plurality of channels to obtain the second spatial domain features of each video frame in the N video frames; wherein, the time offset amounts of the first temporal offset and the second temporal offset are different, or the offset channels are different;
[0150] The classification unit is configured to determine the semantics of the first video at least according to the first spatial domain features and the second spatial domain features of each video frame in the N video frames;
[0151] The output unit is configured to output the semantics of the first video.
[0152] Embodiment 10. For the apparatus according to Embodiment 9, the first offset unit is further configured to: perform a first temporal offset on the spatial domain features of each of the N video frames in at least one of the multiple channels, where the at least one channel includes a first channel, and the time offset amount corresponding to the first temporal offset is T, and T is a positive integer greater than or equal to 1 and less than N, or T is a negative integer greater than -N and less than or equal to -1, so that the spatial domain features of the k-th video frame among the N video frames in the first channel are offset to the spatial domain features of the (k + T)-th video frame in the first channel, and k takes positive integer values in the interval [1, N] in sequence, to obtain the first spatial domain features of each of the N video frames;
[0153] The second offset unit is further configured to: perform a second temporal offset on the spatial domain features of each of the N video frames in at least one of the multiple channels, where the at least one channel includes a second channel, and the time offset amount corresponding to the second temporal offset is T′, and T′ is a positive integer greater than or equal to 1 and less than N, or T′ is a negative integer greater than -N and less than or equal to -1, so that the spatial domain features of the k-th video frame among the N video frames in the second channel are offset to the spatial domain features of the (k + T′)-th video frame in the second channel, and k takes positive integer values in the interval [1, N] in sequence, to obtain the second spatial domain features of each of the N video frames.
[0154] Embodiment 11. For the apparatus according to Embodiment 10, T is not equal to T′, and the first channel and the second channel are the same channel; or,
[0155] T is equal to T′, and the first channel and the second channel are different channels; or,
[0156] T is not equal to T′, and the first channel and the second channel are different channels.
[0157] Embodiment 12. For the apparatus according to 10, the classification unit includes a plurality of two-dimensional convolution units arranged in parallel, and the convolution units in the plurality of two-dimensional convolution units and the offset units in the plurality of offset units correspond one by one;
[0158] The first convolution unit in the plurality of two-dimensional convolution units is configured to perform convolution processing on the first spatial domain features of the first video frame among the N video frames to obtain the first fused spatio-temporal features of the first video frame; the first convolution unit corresponds to the first offset unit;
[0159] The second convolutional unit among the multiple two-dimensional convolutional units is used to perform convolutional processing on the second spatio-temporal feature of the first video frame to obtain the second fused spatio-temporal feature of the first video frame;
[0160] The classification unit is used to determine the semantics of the first video frame at least based on the first fused spatio-temporal feature and the second fused spatio-temporal feature of each video frame in N video frames.
[0161] Example 13. In the apparatus according to Example 12, both the first fused spatio-temporal feature and the second fused spatio-temporal feature include features under M channels, and the features under each of the M channels are represented by matrices; M is a positive integer;
[0162] The classification unit is used for:
[0163] Perform point-by-point addition on the features under the i-th channel among the M channels included in the first fused spatio-temporal feature of the first video frame and the features under the i-th channel among the M channels included in the second fused spatio-temporal feature of the first video frame, where i sequentially takes integer values in the interval [1, M], to obtain the third fused spatio-temporal feature of the first video frame;
[0164] Determine the semantics of the first video according to the third fused spatio-temporal features of each video frame in the N video frames.
[0165] Example 14. In the apparatus according to Example 9, the classification unit is used for:
[0166] Perform feature compensation on the first spatio-temporal feature of the first video frame according to the spatio-temporal features of the first video frame in the multiple channels in the N video frames to obtain a first residual spatio-temporal feature;
[0167] Perform feature compensation on the second spatio-temporal feature of the first video frame according to the spatio-temporal features of the first video frame in the multiple channels in the N video frames to obtain a second residual spatio-temporal feature;
[0168] Determine the semantics of the first video according to the first residual spatio-temporal feature and the second residual spatio-temporal feature.
[0169] Embodiment 15. An apparatus for identifying video semantics, comprising an input unit, a spatial domain feature extraction unit, at least one residual network unit arranged in series, a classification unit, and an output unit; wherein, the spatial domain feature extraction unit comprises at least one two-dimensional convolution unit arranged in series; each residual network unit in the at least one residual network unit comprises a plurality of spatio-temporal feature extraction units and a spatial domain feature compensation unit arranged in series; wherein, each spatio-temporal feature extraction unit in the plurality of spatio-temporal feature extraction units comprises an offset subunit and a convolution subunit arranged in series;
[0170] The input unit is configured to obtain N video frames in a first video, where N is a positive integer greater than 1;
[0171] The spatial domain feature extraction unit is configured to extract the spatial domain features of each of the N video frames in a plurality of channels; wherein, different channels in the plurality of channels correspond to different convolution kernels of the last convolution unit of the at least one two-dimensional convolution unit;
[0172] The first spatio-temporal feature extraction unit of the first residual network unit in the at least one residual network unit is configured to extract the first fused spatio-temporal features of the N video frames; when the first spatio-temporal feature unit is the last one of the plurality of spatio-temporal feature extraction units of the first residual network unit, the spatial domain feature compensation unit of the first residual network unit is configured to determine the residual spatial domain features of the N video frames to be output by the first residual network unit according to the first fused spatio-temporal features of the N video frames and the spatial domain features of the N videos in the plurality of channels; wherein, the offset subunit of the first spatio-temporal feature extraction layer is configured to perform temporal offset on the spatial domain features of the N video frames in at least one channel to obtain the first spatial domain features of each of the N video frames, and the spatial domain features in the at least one channel are a part of the spatial domain features output by the previous unit of the first spatio-temporal feature extraction unit; the convolution subunit of the first spatio-temporal feature extraction unit is configured to perform convolution processing on the first spatial domain features of each of the N video frames to obtain the first fused spatio-temporal features of the N video frames;
[0173] The classification unit is configured to determine the semantics of the first video according to the residual spatial domain features of the N video frames output by the last residual network unit in the at least one residual network unit;
[0174] The output unit is configured to output the semantics of the first video.
[0175] Embodiment 16. The apparatus according to Embodiment 15, wherein the offset subunit of the first spatio-temporal feature extraction layer is configured to:
[0176] Perform temporal offset on the spatial domain features of each of the N video frames in the at least one channel, where the at least one channel includes a first channel, and the time offset amount corresponding to the temporal offset is T, where T is a positive integer greater than or equal to 1 and less than N, or T is a negative integer greater than -N and less than or equal to -1, such that the spatial domain features of the k-th video frame among the N video frames in the first channel are offset to the spatial domain features of the (k + T)-th video frame in the first channel, and k sequentially takes positive integer values in the interval [1, N], so as to obtain the first spatial domain features of each video frame among the N video frames.
[0177] Example 17. A computing device, comprising: a processor, a memory;
[0178] The memory is used to store computer instructions;
[0179] When the computing device runs, the processor executes the computer instructions, so that the computing device executes the method according to any one of Examples 1-6.
[0180] Example 18. A computing device, comprising: a processor, a memory;
[0181] The memory is used to store computer instructions;
[0182] When the computing device runs, the processor executes the computer instructions, so that the computing device executes the method according to Example 7 or 8.
[0183] Example 19. A computer storage medium, the computer storage medium includes computer instructions, when the computer instructions run on a computing device, so that the computing device executes the method according to any one of Examples 1-6.
[0184] Example 20. A computer storage medium, the computer storage medium includes computer instructions, when the computer instructions run on a computing device, so that the computing device executes the method according to Example 7 or 8.
[0185] Example 21. A computer program product, when the program code included in the computer program product is executed by a processor in an electronic device, the method according to any one of Examples 1-8 is implemented.
[0186] In some embodiments of the method and apparatus for identifying video semantics provided in the third aspect of the present application, temporal offsets of different time amounts can be performed on the spatial domain features of video frames in a video, so as to capture information of objects with different motion frequencies in the video; in some embodiments, temporal offsets of spatial domain features in different channels can further be performed on the spatial domain features of video frames in the video, so as to ensure that information of potential object categories is captured, thereby improving the accuracy of video semantic classification results. Moreover, the video semantic recognition method provided in some embodiments of the present application uses a parallel method to perform temporal offsets of different time amounts or temporal offsets of spatial domain features in different channels on the spatial domain features of video frames in a video, avoiding the problem of poor interpretability or unexplainability caused by multiple mixing of the temporal information and spatial information of the video in a serial manner.
[0187] Fourthly, the present application provides the following multiple method embodiments and apparatus embodiments, including:
[0188] Embodiment 1. A method for determining scene type, the method comprising:
[0189] Obtain N video frames in a first video, the N video frames including a first video frame and a second video frame, the first video frame and the second video frame being adjacent in the N videos; N is a positive integer greater than 1;
[0190] When the first video frame includes a first object and the second video includes the first object, determine the position of the first object in the first video and determine the position of the first object in the video frame;
[0191] When the position of the first object in the first video is different from the position of the first object in the video frame, determine the scene type of the first video frame according to the shooting depth of the first object in the first video frame; wherein, the shooting depth is the distance between the camera and the first object when the camera shoots the first video frame.
[0192] Embodiment 2. The method according to Embodiment 1, the method further comprising:
[0193] Determine the size of the first object in the first video frame and determine the size of the first object in the second video frame;
[0194] When the size of the first object in the first video frame is different from the size of the first object in the second video frame, determine the scene type of the first video frame according to the shooting depth of the central region of the first video frame.
[0195] Example 3. The method according to Example 2, wherein the first object includes a plurality of objects; determining the scene type of the first video frame according to the shooting depth of the central region of the first video frame includes:
[0196] When the position of the first object in the first video is different from the position of the first object in the second video frame, and the size of the first object in the first video frame is different from the size of the first object in the second video frame, determine that the second object among the plurality of objects is closest to the central region of the first video frame;
[0197] Determine the scene type of the first video frame according to the shooting depth of the second object in the first video frame.
[0198] Example 4. The method according to Example 1, wherein the first object includes a plurality of objects, and different objects have different attention priorities;
[0199] Determining the scene type of the first video frame according to the shooting depth of the first object in the first video frame includes:
[0200] Determine the scene type of the first video frame according to the shooting depth of the object with the highest attention priority among the plurality of objects.
[0201] Example 5. The method according to Example 4, wherein the plurality of objects include people, animals, plants, and inanimate objects; among them, the attention priority of people is higher than that of animals, the attention priority of animals is higher than that of plants, and the attention priority of plants is higher than that of inanimate objects.
[0202] Example 6. The method according to Example 1, wherein determining the scene type of the first video frame according to the shooting depth of the first object in the first video frame includes:
[0203] When the shooting depth of the first object in the first video frame < the first distance, determine that the scene type of the first video frame is a close-up;
[0204] Or,
[0205] When the shooting depth of the first object in the first video frame > the second distance, determine that the scene type of the first video frame is a long shot;
[0206] Or,
[0207] When the shooting depth of the first object in the first video frame ≥ the first distance and the shooting depth of the first object in the first video frame ≤ the second distance, determine that the scene type of the first video frame is a medium shot.
[0208] Example 7. In the method according to Example 6, the first distance and / or the second distance is set by the user.
[0209] Example 8. In the method according to Example 7, before determining N frame images from the first video, the method further includes:
[0210] Displaying an input interface, the input interface including a first input box;
[0211] Determining a first length input by the user in the first input box as the first distance or the second distance.
[0212] Example 9. In the method according to 7, before determining N frame images from the first video, the method further includes:
[0213] Displaying a distance selection interface, the selection interface including a plurality of selection function areas, different selection functions in the plurality of selection functions corresponding to different lengths;
[0214] In response to an operation on a first selection function area among the plurality of selection function areas, determining the length corresponding to the first selection function area as the first distance or the second distance.
[0215] Example 10. In the method according to Examples 1-9, the first object is an object recognized from the first video frame by a visual saliency detection algorithm.
[0216] Example 11. A scene determination device, the device including:
[0217] An acquisition unit, configured to acquire N video frames in a first video, the N video frames including a first video frame and a second video frame, the first video frame and the second video frame being adjacent in the N videos; N is a positive integer greater than 1;
[0218] A first determination unit, configured to, when the first video frame includes a first object and the second video includes the first object, determine the position of the first object in the first video and determine the position of the first object in the video frame;
[0219] A second determination unit, configured to, when the position of the first object in the first video is different from the position of the first object in the video frame, determine the scene of the first video frame according to the shooting depth of the first object in the first video frame; wherein, the shooting depth is the distance between the camera and the first object when the camera shoots the first video frame.
[0220] Embodiment 12. For the device according to claim 11, the first determination unit is further configured to determine the size of the first object in the first video frame and determine the size of the first object in the second video frame;
[0221] The second determination unit is further configured to, when the size of the first object in the first video frame is different from the size of the first object in the second video frame, determine the scene type of the first video frame according to the shooting depth of the central region of the first video frame.
[0222] Embodiment 13. For the device according to Embodiment 12, the first object includes a plurality of objects; the second determination unit is further configured to:
[0223] When the position of the first object in the first video is different from the position of the first object in the second video frame, and the size of the first object in the first video frame is different from the size of the first object in the second video frame, determine that the second object among the plurality of objects is closest to the central region of the first video frame;
[0224] Determine the scene type of the first video frame according to the shooting depth of the second object in the first video frame.
[0225] Embodiment 14. For the device according to Embodiment 11, the first object includes a plurality of objects, where different objects have different attention priorities;
[0226] The second determination unit is further configured to determine the scene type of the first video frame according to the shooting depth of the object with the highest attention priority among the plurality of objects.
[0227] Embodiment 15. For the device according to Embodiment 14, the plurality of objects include people, animals, plants, and inanimate objects; among them, the attention priority of people is higher than that of animals, the attention priority of animals is higher than that of plants, and the attention priority of plants is higher than that of inanimate objects.
[0228] Embodiment 16. For the device according to Embodiment 11, the second determination unit is further configured to:
[0229] When the shooting depth of the first object in the first video frame < the first distance, determine that the scene type of the first video frame is a close-up;
[0230] Or,
[0231] When the shooting depth of the first object in the first video frame > the second distance, determine that the scene type of the first video frame is a long shot;
[0232] Alternatively,
[0233] When the shooting depth of the first object in the first video frame is ≥ the first distance and the shooting depth of the first object in the first video frame is ≤ the second distance, determine that the scene type of the first video frame is medium shot.
[0234] Embodiment 17. The apparatus according to Embodiment 16, wherein the first distance and / or the second distance is set by the user.
[0235] Embodiment 18. The apparatus according to Embodiment 17, wherein the apparatus further comprises:
[0236] A display unit, configured to display an input interface, and the input interface includes a first input box;
[0237] A third determination unit, configured to determine that the first length input by the user in the first input box is the first distance or the second distance.
[0238] Embodiment 19. The apparatus according to Embodiment 17, wherein the apparatus further comprises:
[0239] A display unit, configured to display a distance selection interface, and the selection interface includes a plurality of selection function areas, and different selection functions in the plurality of selection functions correspond to different lengths;
[0240] A third determination unit, configured to determine that the length corresponding to the first selection function area is the first distance or the second distance in response to an operation on the first selection function area in the plurality of selection function areas.
[0241] Embodiment 20. The apparatus according to Embodiments 11-19, wherein the first object is an object identified from the first video frame by a visual saliency detection algorithm.
[0242] Embodiment 21. A computing device, comprising: a processor and a memory;
[0243] The memory is configured to store computer instructions;
[0244] When the computing device runs, the processor executes the computer instructions, so that the computing device executes the method according to any one of Embodiments 1-10.
[0245] Embodiment 22. A computer storage medium, the computer storage medium includes computer instructions, and when the computer instructions run on a computing device, the computing device is caused to execute the method according to any one of Embodiments 1-10.
[0246] Embodiment 23. A computer program product, when the program code included in the computer program product is executed by a processor in an electronic device, implements the method described in any one of Embodiments 1-10.
[0247] In some embodiments of the scene determination method and apparatus provided in the fifth aspect of the present application, the attention of the user can be simulated to determine the object that the user is paying attention to. In other embodiments, the scene of the video can be further determined according to the shooting distance of the object to obtain a scene that better conforms to the subjective feeling of the user.
[0248] Fifth aspect, the present application provides the following multiple method embodiments and apparatus embodiments, including:
[0249] Embodiment 1. A video processing method, the method includes:
[0250] Using a neural network to extract the spatial domain features of each video frame in the first video under M channels, and to extract the spatial domain features of each video frame in the second video under the M channels; the neural network includes at least one two-dimensional convolutional layer, and different channels in the M channels correspond to different convolutional kernels of the last convolutional layer in the at least one two-dimensional convolutional layer; M is a positive integer greater than or equal to 1;
[0251] According to the spatial domain features of each video frame in the first video under the M channels and the spatial domain features of each video frame in the second video under the M channels, determine the pairwise video frame similarity between each video frame in the first video and each video frame in the second video;
[0252] According to the pairwise video frame similarity between each video frame in the first video and each video frame in the second video, determine the splicing point of the first video and the second video to splice the first video and the second video.
[0253] Embodiment 2. According to the method described in Embodiment 1, the neural network further includes a plurality of style transfer layers arranged in parallel, where different style transfer layers correspond to different object categories; the style transfer layer is obtained by training a generative adversarial network GAN with the spatial domain features of a reference object as training data; the reference object belongs to the object category corresponding to the style transfer layer;
[0254] The using a neural network to extract the spatial domain features of each video frame in the first video under M channels, and to extract the spatial domain features of each video frame in the second video under the M channels includes:
[0255] For each video frame in the first video and the second video, determine a style transfer layer corresponding to the video frame from the multiple style transfer layers according to the object category of the object included in the video frame;
[0256] Use the style transfer layer corresponding to the video frame to perform style transfer processing on the spatial domain features of the video frame in the M channels, and obtain first spatial domain features of the video frame in the M channels;
[0257] The determining of the pairwise video frame similarity between each video frame in the first video and each video frame in the second video according to the spatial domain features of each video frame in the first video in the M channels and the spatial domain features of each video frame in the second video in the M channels includes:
[0258] Determine the pairwise video frame similarity between each video frame in the first video and each video frame in the second video according to the first spatial domain features of each video frame in the first video in the M channels and the first spatial domain features of each video frame in the second video in the M channels.
[0259] Example 3. According to the method described in Example 1 or 2, the spatial domain features in each of the M channels are represented by a matrix with a length of K and a width of K';
[0260] The determining of the pairwise video frame similarity between each video frame in the first video and each video frame in the second video according to the spatial domain features of each video frame in the first video in the M channels and the spatial domain features of each video frame in the second video in the M channels includes:
[0261] Divide the spatial domain features of the first video frame in the M channels into K×K' vectors of the first video frame; where the i×j-th vector among the K×K' vectors of the first video frame is composed of elements with coordinates (i, j) in the spatial domain features of each channel in the M channels of the first video frame; i is a positive integer less than or equal to K, and j is a positive integer less than or equal to K'; the first video frame is a video frame in the first video;
[0262] Divide the spatial domain features of the second video frame in the M channels into K×K' vectors of the second video frame; where the i×j-th vector among the K×K' vectors of the second video frame is composed of elements with coordinates (i, j) in the spatial domain features of each channel in the M channels of the second video frame; the second video frame is a video frame in the second video;
[0263] Calculate the cosine distance between the \(i\times j\) -th vector of the first video frame and the \(i\times j\) -th vector of the second video frame;
[0264] Determine the similarity between the first video frame and the second video frame under the \(i\times j\) -th vector according to the cosine distance;
[0265] Determine the pairwise video frame similarity between the first video frame and the second video frame according to the similarities between the first video frame and the second video frame under each of the \(K\times K'\) vectors.
[0266] Example 4. According to the method described in Example 3, the method further includes:
[0267] Calculate the first average value of the element with coordinates \((i,j)\) in the spatial domain features of each channel in the \(M\) channels of the first video frame, and the second average value of the element with coordinates \((i,j)\) in the spatial domain features of each channel in the \(M\) channels of the second video frame;
[0268] The step of determining the similarity between the first video frame and the second video frame under the \(i\times j\) -th vector according to the cosine distance includes:
[0269] Calculate the first product of the cosine distance and the first average value, and the second product of the cosine distance and the second average value;
[0270] Determine the pairwise video frame similarity between the first video frame and the second video frame according to the first product and the second product.
[0271] Example 5. According to the method described in Example 1, the step of determining the splicing point between the first video and the second video according to the pairwise video frame similarity between each video frame in the first video and each video frame in the second video includes:
[0272] Determine the maximum pairwise video frame similarity among the pairwise video frame similarities between each video frame in the first video and each video frame in the second video;
[0273] Determine the two video frames corresponding to the maximum pairwise video frame similarity as the splicing point between the first video and the second video.
[0274] Example 6. According to the method described in Example 5, the two video frames are the third video frame in the first video and the fourth video frame in the second video; the method further includes:
[0275] Splice the first segment in the first video and the second segment in the second video into a third video;
[0276] Among them, the starting video frame of the first segment is the first video frame of the first video, and the ending video frame is the third video frame; the starting video frame of the second segment is the fourth video frame, and the ending video frame is the last video frame of the first video; in the third video, the first segment is located before the second segment; or,
[0277] the starting video frame of the first segment is the third video frame, and the ending video frame is the last video frame of the first video; the starting video frame of the second segment is the first video frame of the second video, and the ending video frame is the fourth video frame; in the third video, the second segment is located before the first segment.
[0278] Embodiment 7. The method according to Embodiment 1, the method further includes:
[0279] Obtain a fourth video;
[0280] Determine that the segment with the first detailed dynamic semantics in the fourth video is the first video, and determine that the segment with the second detailed dynamic semantics in the fourth video is the second video.
[0281] Embodiment 8. A video processing method, the method includes:
[0282] Perform frame extraction on a first video to obtain N video frames;
[0283] Use a neural network to extract the spatial domain features of each of the N video frames in M channels; the neural network includes at least one two-dimensional convolutional layer, and different channels among the M channels correspond to different convolutional kernels of the last convolutional layer in the at least one two-dimensional convolutional layer;
[0284] According to the spatial domain features of each of the N video frames in the M channels, determine the pairwise video frame similarity between adjacent video frames among the N video frames;
[0285] According to the pairwise video frame similarity between adjacent video frames among the N video frames and the time points of each video frame among the N video frames, determine a first curve;
[0286] Perform Fourier transform on the first curve to obtain a first transform result;
[0287] Determine the transform results less than a threshold from the first transform result;
[0288] According to the transform results less than the threshold and the time points corresponding to the transform results less than the threshold, determine a second curve;
[0289] Determine the video frame corresponding to the intersection of the first curve and the second curve as the key video frame;
[0290] Use the key video frames to synthesize a video to obtain the compressed first video.
[0291] Example 9. According to the method described in Example 8, the threshold is less than or equal to one-half of the maximum value in the first transformation result.
[0292] Example 10. According to the method described in 9, the threshold is one-fourth of the maximum value in the first transformation result.
[0293] Example 11. A video processing device, the device includes:
[0294] An extraction unit, configured to use a neural network to extract the spatial domain features of each video frame in the first video under M channels, and extract the spatial domain features of each video frame in the second video under the M channels; the neural network includes at least one two-dimensional convolutional layer, and different channels in the M channels correspond to different convolutional kernels of the last convolutional layer in the at least one two-dimensional convolutional layer; M is a positive integer greater than or equal to 1;
[0295] A first determination unit, configured to determine the pairwise video frame similarity between each video frame in the first video and each video frame in the second video according to the spatial domain features of each video frame in the first video under the M channels and the spatial domain features of each video frame in the second video under the M channels;
[0296] A second determination unit, configured to determine the splicing point of the first video and the second video according to the pairwise video frame similarity between each video frame in the first video and each video frame in the second video, so as to splice the first video and the second video.
[0297] Example 12. According to the device described in Example 11, the neural network further includes a plurality of style transfer layers arranged in parallel, where different style transfer layers correspond to different object categories; the style transfer layer is obtained by training a generative adversarial network GAN with the spatial domain features of a reference object as training data; the reference object belongs to the object category corresponding to the style transfer layer;
[0298] The extraction unit is further configured to:
[0299] For each video frame in the first video and the second video, determine the style transfer layer corresponding to the video frame from the plurality of style transfer layers according to the object category of the object included in the video frame;
[0300] Using the style transfer layer corresponding to the video frame, perform style transfer processing on the spatial domain features of the video frame in the M channels to obtain the first spatial domain features of the video frame in the M channels;
[0301] The first determining unit is further configured to: determine the pairwise video frame similarity between each video frame in the first video and each video frame in the second video according to the first spatial domain features of each video frame in the first video in the M channels and the first spatial domain features of each video frame in the second video in the M channels.
[0302] Example 13. For the apparatus according to Example 11 or 12, the spatial domain features in each of the M channels are represented by a matrix of length K and width K';
[0303] The first determining unit is further configured to:
[0304] Divide the spatial domain features of the first video frame in the M channels into K×K' vectors of the first video frame; wherein, the i×j-th vector among the K×K' vectors of the first video frame is composed of elements with coordinates (i, j) in the spatial domain features of each channel in the M channels of the first video frame; i is a positive integer less than or equal to K, and j is a positive integer less than or equal to K'; the first video frame is a video frame in the first video;
[0305] Divide the spatial domain features of the second video frame in the M channels into K×K' vectors of the second video frame; wherein, the i×j-th vector among the K×K' vectors of the second video frame is composed of elements with coordinates (i, j) in the spatial domain features of each channel in the M channels of the second video frame; the second video frame is a video frame in the second video;
[0306] Calculate the cosine distance between the i×j-th vector of the first video frame and the i×j-th vector of the second video frame;
[0307] Determine the similarity between the first video frame and the second video frame under the i×j-th vector according to the cosine distance;
[0308] Determine the pairwise video frame similarity between the first video frame and the second video frame according to the similarities between the first video frame and the second video frame under each of the K×K' vectors.
[0309] Example 14. The device according to Example 13, the device further includes a calculation unit, configured to calculate a first average value of an element with coordinates (i, j) in the spatial domain features of each channel among the M channels of the first video frame, and a second average value of an element with coordinates (i, j) in the spatial domain features of each channel among the M channels of the second video frame;
[0310] The first determination unit is further configured to:
[0311] calculate a first product of the cosine distance and the first average value, and a second product of the cosine distance and the second average value;
[0312] determine the pairwise video frame similarity between the first video frame and the second video frame according to the first product and the second product.
[0313] Example 15. The device according to Example 11, the second determination unit is further configured to:
[0314] determine the maximum pairwise video frame similarity among the pairwise video frame similarities between each video frame in the first video and each video frame in the second video;
[0315] determine the two video frames corresponding to the maximum pairwise video frame similarity as the splicing points of the first video and the second video.
[0316] Example 16. The device according to Example 15, the two video frames are the third video frame in the first video and the fourth video frame in the second video; the device further includes a splicing unit, configured to splice the first segment in the first video and the second segment in the second video into a third video;
[0317] wherein, the starting video frame of the first segment is the first video frame of the first video, and the ending video frame is the third video frame; the starting video frame of the second segment is the fourth video frame, and the ending video frame is the last video frame of the first video; in the third video, the first segment is located before the second segment; or,
[0318] the starting video frame of the first segment is the third video frame, and the ending video frame is the last video frame of the first video; the starting video frame of the second segment is the first video frame of the second video, and the ending video frame is the fourth video frame; in the third video, the second segment is located before the first segment.
[0319] Example 17. The device according to Example 11, the device further includes:
[0320] An acquisition unit, configured to acquire a fourth video;
[0321] A first determination unit, configured to determine a segment with a first detailed dynamic semantics in the fourth video as the first video, and determine a segment with a second detailed dynamic semantics in the fourth video as the second video.
[0322] Embodiment 18. A video processing apparatus, the apparatus comprising:
[0323] A frame extraction unit, configured to perform frame extraction on the first video to obtain N video frames;
[0324] An extraction unit, configured to use a neural network to extract spatial domain features of each of the N video frames in M channels; the neural network includes at least one two-dimensional convolutional layer, and different channels among the M channels correspond to different convolutional kernels of the last convolutional layer in the at least one two-dimensional convolutional layer;
[0325] A first determination unit, configured to determine pairwise video frame similarities between adjacent video frames among the N video frames according to the spatial domain features of each of the N video frames in the M channels;
[0326] A second determination unit, configured to determine a first curve according to the pairwise video frame similarities between adjacent video frames among the N video frames and the time points of each of the N video frames;
[0327] A transformation unit, configured to perform Fourier transform on the first curve to obtain a first transformation result;
[0328] A third determination unit, configured to determine transformation results less than a threshold from the first transformation result;
[0329] A fourth determination unit, configured to determine a second curve according to the transformation results less than the threshold and the time points corresponding to the transformation results less than the threshold;
[0330] A fifth determination unit, configured to determine a video frame corresponding to an intersection point of the first curve and the second curve as a key video frame;
[0331] A synthesis unit, configured to synthesize a video using the key video frames to obtain the compressed first video.
[0332] Embodiment 19. The apparatus according to Embodiment 18, wherein the threshold is less than or equal to one half of the maximum value in the first transformation result.
[0333] Embodiment 20. The apparatus according to Embodiment 19, wherein the threshold is one quarter of the maximum value in the first transformation result.
[0334] Example 21. A computing device, comprising: a processor and a memory;
[0335] The memory is used to store computer instructions;
[0336] When the computing device runs, the processor executes the computer instructions, so that the computing device executes the method according to any one of Examples 1-7.
[0337] Example 22. A computing device, comprising: a processor and a memory;
[0338] The memory is used to store computer instructions;
[0339] When the computing device runs, the processor executes the computer instructions, so that the computing device executes the method according to any one of Examples 8-10.
[0340] Example 23. A computer storage medium, the computer storage medium includes computer instructions, when the computer instructions run on a computing device, so that the computing device executes the method according to any one of Examples 1-7.
[0341] Example 24. A computer storage medium, the computer storage medium includes computer instructions, when the computer instructions run on a computing device, so that the computing device executes the method according to any one of Examples 8-10.
[0342] Example 25. A computer program product, when the program code included in the computer program product is executed by a processor in an electronic device, it implements the method according to any one of Examples 1-7 or the method according to any one of Examples 8-10.
[0343] In some embodiments of the video processing method and apparatus provided in the fifth aspect of the present application, the CNN features of video frames can be utilized to calculate the similarity of video frames in different videos or video segments, and according to the similarity of video frames, determine the splicing points for splicing different videos or video segments in the time dimension, so that the spliced video is smoother and the viewing effect of the video is improved. Moreover, directly using the CNN features of video frames extracted by the convolutional neural network to calculate the similarity improves the real-time performance of similarity calculation. The convolutional neural network can extract relatively rich features of video frames, thereby also improving the accuracy of similarity. Description of the Drawings
[0344] Figure 1 It is a schematic structural diagram of a neural network provided by an embodiment of the present application;
[0345] Figure 2 It is a flowchart of a video semantic recognition method provided by an embodiment of the present application;
[0346] Figure 3 It is a schematic structural diagram of a neural network provided by an embodiment of the present application;
[0347] Figure 4 It is a schematic structural diagram of a dynamic semantic recognition layer provided by an embodiment of the present application;
[0348] Figure 5 It is a flowchart of a video semantic recognition method provided by an embodiment of the present application;
[0349] Figure 6 It is a schematic structural diagram of a spatial domain feature extraction layer provided by an embodiment of the present application;
[0350] Figure 7 It is a schematic structural diagram of a dynamic semantic recognition layer provided by an embodiment of the present application;
[0351] Figure 8 It is a flowchart of a video semantic recognition method provided by an embodiment of the present application;
[0352] Figure 9 It is a flowchart of a video frame label processing method provided by an embodiment of the present application;
[0353] Figure 10 It is a schematic matrix diagram provided by an embodiment of the present application;
[0354] Figure 11 It is a schematic matrix diagram provided by an embodiment of the present application;
[0355] Figure 12 It is a block diagram provided by an embodiment of the present application for showing the smoothness of video frame labels in a video;
[0356] Figure 13 It is a schematic diagram of the logical expression of a scoring node provided by an embodiment of the present application;
[0357] Figure 14 It is a schematic diagram of a minimum cut algorithm provided by an embodiment of the present application;
[0358] Figure 15 It is a flowchart of a video frame label processing method in a video provided by an embodiment of the present application;
[0359] Figure 16A A block diagram provided by an embodiment of the present application for showing the smoothness of video frame labels in a video;
[0360] Figure 16B A block diagram provided by an embodiment of the present application for showing the smoothness of video frame labels in a video;
[0361] Figure 17AA block diagram provided by an embodiment of the present application for displaying the smoothness of video frame tags in a video;
[0362] Figure 17B A block diagram provided by an embodiment of the present application for displaying the smoothness of video frame tags in a video;
[0363] Figure 18A A block diagram provided by an embodiment of the present application for displaying the smoothness of video frame tags in a video;
[0364] Figure 18B A block diagram provided by an embodiment of the present application for displaying the smoothness of video frame tags in a video;
[0365] Figure 19A A block diagram provided by an embodiment of the present application for displaying the smoothness of video frame tags in a video;
[0366] Figure 19B A block diagram provided by an embodiment of the present application for displaying the smoothness of video frame tags in a video;
[0367] Figure 20 Is a flowchart of a scene determination method provided by an embodiment of the present application;
[0368] Figure 21 Is a flowchart of a scene determination method provided by an embodiment of the present application;
[0369] Figure 22 Is a flowchart of a scene determination method provided by an embodiment of the present application;
[0370] Figure 23 Is a flowchart of a scene determination method provided by an embodiment of the present application;
[0371] Figure 24 Is a flowchart of a video semantic recognition method provided by an embodiment of the present application;
[0372] Figure 25 Is an architecture diagram of a video semantic recognition provided by an embodiment of the present application;
[0373] Figure 26 Is a flowchart of a video editing method provided by an embodiment of the present application;
[0374] Figure 27 Is a flowchart of a video editing method provided by an embodiment of the present application;
[0375] Figure 28 Is a flowchart of a video editing method provided by an embodiment of the present application;
[0376] Figure 29It is a flowchart of a video editing method provided by an embodiment of the present application;
[0377] Figure 30 It is a flowchart of a video processing method provided by an embodiment of the present application;
[0378] Figure 31A It is a schematic diagram of a neural network structure provided by an embodiment of the present application;
[0379] Figure 31B It is a schematic diagram of a neural network structure provided by an embodiment of the present application;
[0380] Figure 32 It is a flowchart of a video processing method provided by an embodiment of the present application;
[0381] Figure 33 It is a similarity curve graph of two-by-two video frames provided by an embodiment of the present application;
[0382] Figure 34 It is a flowchart of a video processing method provided by an embodiment of the present application;
[0383] Figure 35 It is a flowchart of a video processing method provided by an embodiment of the present application;
[0384] Figure 36 It is a schematic diagram of a device for recognizing video semantics provided by an embodiment of the present application;
[0385] Figure 37 It is a schematic diagram of a video editing device provided by an embodiment of the present application;
[0386] Figure 38 It is a schematic diagram of the structure of a video processing device provided by an embodiment of the present application;
[0387] Figure 39 It is a schematic diagram of the structure of a video processing device provided by an embodiment of the present application;
[0388] Figure 40 It is a schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0389] Next, the technical solutions in the embodiments of the present invention will be described in conjunction with the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all of the embodiments.
[0390] In the description of this specification, "one embodiment" or "some embodiments" etc. mean that in one or more embodiments of this specification, the specific features, structures or characteristics described in connection with that embodiment are included. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways.
[0391] Among them, in the description of this specification, unless otherwise specified, " / " means "or". For example, A / B can mean A or B; "and / or" herein is merely a description of the relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of this specification, "a plurality of" means two or more than two.
[0392] In the description of this specification, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "include", "comprise", "have" and their variants all mean "include but not limited to", unless otherwise specifically emphasized in other ways.
[0393] In the embodiments of this application, a timing segment and three levels of semantics can be defined, specifically as follows.
[0394] A temporal segment refers to a video segment composed of multiple consecutive video frames in a video. "Consecutive" means adjacent in sequence, and multiple consecutive video frames mean that these multiple video frames are adjacent in sequence. A temporal segment can be represented by the position of the first video frame and the position of the last video frame among the multiple consecutive video frames that make up the temporal segment in the video. Among them, the first video frame can be called the starting frame of the temporal segment, and the last video frame can be called the ending frame of the temporal segment. The first video frame refers to the video frame with the most forward position in the video among these multiple consecutive video frames. That is to say, when playing this video, after playing the first video frame, other video frames among these multiple consecutive video frames are played in sequence. Correspondingly, the last video frame refers to the video frame with the most backward position in the video among these multiple consecutive video frames. Specifically, assume a video with a duration of 60 seconds, and multiple consecutive video frames from the 10th second to the 15th second form a temporal segment, then this temporal segment can be represented as the segment from the 10th second to the 15th second in this video. Among them, the video frame at the 10th second is the starting frame of this temporal segment, and the segment at the 15th second is the ending frame of this temporal segment.
[0395] Static semantics, which can also be called surface semantics, refers to the object semantics that can be recognized through a single video frame or at a glance. In other words, surface semantics is the subject, object, or scene presented by a single video frame. It should be noted that the scene here refers to the scene or situation. The subject refers to objects with life such as people and animals that can move autonomously. The object refers to objects without life such as basketballs, footballs, and birthday cakes that cannot move autonomously. Exemplarily, static semantics can refer to the object category. That is to say, for a certain video frame, it has static semantics, and its static semantics refers to the category of the objects contained in this video frame. It can be understood that the objects in the video frame are actually the images of the corresponding objects. For the convenience of description, the images of the objects in the video frame are called the categories of the objects in the video frame.
[0396] Dynamic semantics, which can also be called deep semantics, is information used to represent actions. It is difficult to see at a glance and requires multiple consecutive video frames to be combined to identify the semantics. Exemplarily, dynamic semantics can be information used to represent the actions of the subject, and this action is the subject action that requires multiple consecutive video frames to be combined to identify. For example, dynamic semantics can be playing basketball, playing football, etc.
[0397] Detail dynamic semantics, which can also be called detail semantics, is information used to represent specific actions, or information used to represent instant wonderful actions. For example, detail dynamic semantics can be basketball layups, football shots, etc. A temporal segment with detail dynamic semantics can be called a wonderful temporal segment.
[0398] Embodiments of the present application provide a video semantic recognition method and a video editing method. Among them, the video semantic recognition method can identify a temporal segment with dynamic semantics and a temporal segment with static semantics. When performing video editing, the video editing method can preferentially use the temporal segment with dynamic semantics for video splicing, so that a video with a relatively high level of excitement can be spliced.
[0399] For the video semantic recognition method and the video editing method provided by the embodiments of the present application, the technical solutions of the embodiments of the present application can be executed by a user device, which can be mobile or fixed. For example, the user device can be a mobile phone with video frame processing function, a tablet personal computer (TPC), a media player, a smart TV, a laptop computer (LC), a personal digital assistant (PDA), a personal computer (PC), a camera, a video camera, a smart watch, a wearable device (WD), etc. The embodiments of the present application do not make any limitations thereto.
[0400] Next, in different embodiments, the video semantic recognition method and the video editing method provided by the embodiments of the present application will be illustrated by examples.
[0401] In some embodiments, the video semantic recognition method provided by the embodiments of the present application can be implemented by Figure 1 the neural network shown. The neural network can include an input layer, a spatial domain feature extraction layer, a static semantic recognition layer, a dynamic semantic recognition layer, a temporal segment division layer, and an output layer. The static semantic recognition layer and the dynamic semantic recognition layer are arranged in parallel.
[0402] Referring to Figure 2 , the video semantic recognition method provided by the embodiments of the present application can include the following steps.
[0403] Step 201, in the input layer, obtain multiple video frames of the video.
[0404] Through the input layer, video A can be input into the neural network. That is to say, in the input layer, video A can be obtained. Video A is composed of a series of static images, and the images therein can be called frames or video frames. When obtaining video A or after that, video A can be video-framed. Video-framing can also be called video deframing, which is to decompose the video into a sequence of video frames or sequence of images. In other words, it is to decompose the video into multiple video frames or multiple images. Among them, the positional relationship between the video frames in the multiple video frames is the same as the positional relationship in the video.
[0405] Step 202, in the spatial domain feature extraction layer, extract the spatial domain features of each of the multiple video frames.
[0406] In the spatial domain feature extraction layer, extract the spatial domain features of each of the multiple video frames. The spatial domain features of a video frame refer to the spatial feature information of the video frame, or in other words, the feature information of the space or image represented by the video frame.
[0407] Exemplarily, the spatial domain features of a video frame can be obtained by convolution using a traditional two-dimensional (2D) convolutional neural network. For example, the spatial domain feature extraction layer can include at least one convolutional layer. The at least one convolutional layer can be connected in sequence such that the output of the previous layer can be used as the input of the next layer. The convolutional layer is used to extract the spatial domain features of the video frame. For one convolutional layer in the at least one convolutional layer, the convolutional layer can include multiple convolutional kernels. One of the multiple convolutional kernels can be used to perform a convolution operation on the video frame to obtain a feature map (the feature map can also be referred to as a channel). Multiple convolutional kernels can obtain multiple feature maps.
[0408] Among them, when the convolutional layer is not the last layer in the at least one convolutional layer, input the multiple feature maps obtained by this convolutional layer into the next convolutional layer. Each convolutional kernel of the next convolutional layer performs a convolution process on the multiple feature maps to obtain a feature map. When there are multiple convolutional kernels in the next layer, multiple feature maps can be obtained.
[0409] The multiple feature maps obtained in this convolutional layer can be stacked. Specifically, the feature map can be represented as a two-dimensional matrix, and the stacking of different feature maps can refer to the addition of elements at corresponding positions in different two-dimensional matrices. The stacked feature map can be used as the input of the next layer.
[0410] When the convolutional layer is the last layer in the at least one convolutional layer, this convolutional layer can directly output the multiple feature maps obtained in this convolutional layer. The multiple feature maps can be used as the spatial domain features of the corresponding video frame extracted in the spatial domain feature extraction layer.
[0411] Continue to refer to Figure 2, the video semantic recognition method provided by the embodiments of the present application further includes: Step 203a, in the dynamic semantic recognition layer, according to the spatial domain features of N consecutive video frames among the multiple video frames, determine the dynamic semantics of the Nth video frame among the N consecutive video frames; N is a positive integer. As described above, after video A is frame-divided, multiple video frames can be obtained. In the dynamic semantic recognition layer, the dynamic semantics of the Nth video frame among the N consecutive video frames can be determined according to the spatial domain features of the N consecutive video frames among the multiple video frames. N can be a set value, for example, it can be 5, or it can be 10, or it can be 15, and so on, which will not be elaborated here one by one.
[0412] Specifically, the spatial domain features of the N consecutive video frames can be combined to determine whether the N consecutive video frames as a whole have or represent dynamic semantics. If the N consecutive video frames as a whole have or represent dynamic semantics, then determine the dynamic semantics as the dynamic semantics of the Nth video frame among the N consecutive video frames. If the N consecutive video frames as a whole do not have or do not represent dynamic semantics, then determine that the Nth video frame among the N consecutive video frames does not have dynamic semantics.
[0413] It can be understood that for the 1st video frame to the (N - 1)th video frame among the multiple video frames, since it is difficult to determine whether they have dynamic semantics according to the above scheme, therefore, in the embodiments of the present application, it can be defaulted that the 1st video frame to the (N - 1)th video frame do not have dynamic semantics.
[0414] In some embodiments, in the dynamic semantic recognition layer, based on the temporal shift module (TSM) scheme, the dynamic semantics represented by the N consecutive video frames can be determined, and then the dynamic semantics of the Nth video frame among the N consecutive video frames can be determined. In the TSM scheme, the spatial domain features of the N consecutive video frames stacked together are temporally shifted on one or some channels (feature maps) to obtain the residual spatial domain features of the N consecutive video frames. Thus, through the fusion on different video frame channels, the fusion in the time domain of different video frames is realized. Then, the residual spatial domain features of the N consecutive video frames are input into a convolutional neural network for feature extraction, so that while further extracting the spatial domain features, the temporal domain features between the N consecutive video frames can be extracted. The spatial domain features of the N consecutive video frames after channel fusion can be input into a convolutional neural network for extraction, and the extraction result is called the spatio-temporal features of the N consecutive video frames. Using the spatio-temporal features of the N consecutive video frames to classify the N consecutive video frames, the dynamic semantics that the N consecutive video frames as a whole have or represent can be obtained.
[0415] Next, an example of the timing offset module solution will be given. For a video frame, it can be set that the spatial feature extraction layer extracts L channels. The L channels can include channel B1, channel B2, channel B3, channel B4, etc. It can be understood that different channels among the L channels are obtained by convolution processing with different convolution kernels of the spatial feature extraction layer. In the dynamic semantic recognition layer, the B1 channel of the first video frame among the N consecutive video frames can be used to replace the B1 channel of the second video frame among the N consecutive video frames, the B1 channel of the second video frame among the N consecutive video frames can be used to replace the B1 channel of the third video frame among the N consecutive video frames,..., the B1 channel of the (N - 1)-th video frame among the N consecutive video frames can be used to replace the B1 channel of the N-th video frame among the N consecutive video frames. Thus, the offset processing of channel B1 is completed. Similarly, the offset processing of channel B2 can also be performed. Exemplarily, the number of channels for which offset is performed is one-fourth of L.
[0416] Through the above channel offset processing, the residual spatial features of N consecutive video frames are obtained. Inputting the residual spatial features of N consecutive video frames into a convolutional neural network for convolution processing can obtain the spatio-temporal features of the N consecutive video frames. In addition, it should be noted that if the convolutional neural network has multiple convolutional layers and each convolutional layer has multiple convolutional kernels. Then channel offset layers can be set between different convolutional layers of the multiple convolutional layers. In the channel offset layer, the channels of the N consecutive video frames output by the upper convolutional layer can be offset processed with reference to the above channel offset processing scheme, and then used as the input of the lower convolutional layer. Among them, the convolution processing of the residual spatial features of N consecutive video frames by the convolutional neural network can refer to the convolution processing process of the stacked channels of an image by a 2D convolutional neural network, which will not be elaborated here.
[0417] Through the above scheme, the spatio-temporal features of N consecutive video frames can be extracted, and then the whole of the N consecutive video frames can be classified (for example, classified by the softmax function) to obtain the dynamic semantics possessed or represented by the whole of the N consecutive video frames.
[0418] In some embodiments, another scheme for determining the dynamic semantics of a video frame is provided. Refer to Figure 3, this solution can determine the dynamic semantics possessed or represented by the whole of the N video frames based on the spatial domain features of each of the N consecutive video frames in multiple channels, and further determine the dynamic semantics of the Nth video frame among the N video frames. Among them, the N video frames can be N consecutive video frames, that is to say, the N video frames are adjacent to each other in sequence in video A. It can also be N video frames obtained by frame extraction processing on video A. For two adjacent video frames among the N video frames, they may not be adjacent in video A. Generally speaking, performing offsets of different time amounts of the spatial domain features in the same channel among the N video frames, and / or performing offsets of the spatial domain features in different channels among the N video frames, so that the spatial domain features of video frame A1 among the N video frames are mixed with the spatial domain features of the video frames before video frame A1 and / or the spatial domain features of the video frames after video frame A1. For the convenience of description, the spatial domain features of video frame A1 obtained after mixing the spatial domain features of other video frames in one or more channels in the spatial domain features of video frame A1 in multiple channels can be called the mixed spatial domain features of video frame A1. In other words, for video frame A1 in video A1, its mixed spatial domain features refer to the spatial domain features obtained after mixing the spatial domain features of the video frames before it in at least one channel and / or the spatial domain features of the video frames after it in at least one channel in its spatial domain features. Among them, in the embodiments of the present application, unless otherwise specified, the front and back refer to the front and back in the time sequence or time domain of the video. For example, the video frame before video frame A1 refers to the video frame before video frame A1 in the time sequence or time domain of video A. In other words, when video A is played, the video frame before video frame A1 is played first, and then video frame A1 is played. Similarly, the video frame after video frame A1 refers to the video frame after video frame A1 in the time sequence or time domain of video A.
[0419] More specifically, Figure 4 shows a specific structure of the dynamic semantics recognition layer. Among them, it includes multiple offset layers, a classification layer, and an output layer. The multiple offset layers can be arranged in parallel.
[0420] Based on Figure 4 the shown dynamic semantics recognition layer, the dynamic semantics of the video frame can be determined through the steps as shown in Figure 5 shown.
[0421] Among them, before performing the steps shown in Figure 5 the dynamic semantics recognition layer can obtain the spatial domain features of the N video frames output by the spatial domain feature extraction layer. The spatial domain features of each of the N video frames can be the spatial domain features in multiple channels. In one example, refer to Figure 6, the spatio-temporal feature extraction layer may include at least one two-dimensional convolutional layer. The two-dimensional convolutional layers in the at least one two-dimensional convolutional layer may be connected in sequence, and the output of the previous layer may be used as the input of the next layer. The two-dimensional convolutional layer is used for spatio-temporal feature extraction of video frames. For one two-dimensional convolutional layer in the at least one two-dimensional convolutional layer, it may include multiple convolutional kernels. A convolutional kernel may also be referred to as a convolution operator and can be understood as a filter for extracting specific information in a video frame. Essentially, a convolutional kernel may be a weight matrix, and the parameters therein may be obtained through supervised training. For example, the initial parameters in the convolutional kernel may be preset values or randomized values. During the supervised training of the neural network, the loss between the result output by the neural network and the training set label is calculated, and the parameters in the convolutional kernel are updated in the direction of reducing the loss. In this way, the parameters in the convolutional kernel can be obtained after training. In addition, for the convenience of description, hereinafter, unless otherwise specified, the convolutional layer refers to the two-dimensional convolutional layer.
[0422] For the convolutional layer in the spatio-temporal feature extraction layer, a convolutional kernel in the multiple convolutional kernels of the convolutional layer can be used to perform a convolution operation on the video frame, and spatio-temporal features can be extracted. The spatio-temporal features obtained by the convolution operation of the convolutional kernel can be referred to as spatio-temporal features under a channel (channel) or as a feature map. Multiple convolutional kernels of a convolutional layer can obtain spatio-temporal features under multiple channels, and the channels in the multiple channels correspond to the convolutional kernels in the multiple convolutional kernels one by one.
[0423] Among them, when a certain convolutional layer is not the last layer in the at least one convolutional layer, the multiple feature maps obtained by the convolutional layer are input into the next convolutional layer. Each convolutional kernel of the next convolutional layer performs a convolution process on the multiple feature maps to obtain a feature map. In one example, the specific process of each convolutional kernel performing a convolution process on the multiple feature maps to obtain a feature map may be: a convolutional kernel can perform a convolution process on each of the multiple feature maps to obtain a new feature map. In this way, multiple feature maps can obtain multiple new feature maps under the convolution process of one convolutional kernel. The multiple new feature maps are superimposed to obtain a feature map. Specifically, a feature map can be represented as a two-dimensional matrix, and the superposition of different feature maps may refer to the addition of elements at corresponding positions in different two-dimensional matrices.
[0424] Referring to the above method, through the convolution process of the multiple convolutional kernels of the next layer, spatio-temporal features under multiple channels of the next layer can be obtained. The multiple channel kernels correspond to the multiple convolutional kernels one by one.
[0425] Refer to Figure 6, when the convolutional layer is the last convolutional layer in the spatial domain feature extraction layer, through channel stacking, the spatial domain features of each video frame among the N video frames extracted by the last convolutional layer under multiple channels can be stacked. The channels in the multiple channels correspond one-to-one with the convolutional kernels of the last convolutional layer.
[0426] In one example, as Figure 6 shown, it can be set that the N video frames are video frame A1, video frame A2, video frame A3, and video frame A4. The last convolutional layer in the spatial domain feature extraction layer has 8 convolutional kernels, and the 8 convolutional kernels are the convolutional kernel corresponding to channel C1, the convolutional kernel corresponding to channel C2, the convolutional kernel corresponding to channel C3, the convolutional kernel corresponding to channel C4, the convolutional kernel corresponding to channel C5, the convolutional kernel corresponding to channel C6, the convolutional kernel corresponding to channel C7, and the convolutional kernel corresponding to channel C8. In this example, through the spatial domain feature extraction layer, the spatial domain features of video frame A1, video frame A2, video frame A3, and video frame A4 under channel C1, under channel C2, under channel C3, under channel C4, under channel C5, under channel C6, under channel C7, and under channel C8 can be obtained respectively. Then, according to the temporal relationship among video frame A1, video frame A2, video frame A3, and video frame A4, channel stacking is performed, and the stacked features as shown in Figure 6 can be obtained.
[0427] The spatial domain feature extraction layer can output the spatial domain features of the stacked N video frames under multiple channels to multiple offset layers.
[0428] Back to Figure 4 , the multiple offset layers can include a first offset layer and a second offset layer.
[0429] Back to Figure 5, the user equipment may perform step 501a to perform a first temporal offset on the spatial domain features of the N video frames in at least one of the multiple channels in the first offset layer to obtain the first spatial domain features of each video frame in the N video frames. The computer may perform step 501b to perform a second temporal offset on the spatial domain features of the N video frames in at least one of the multiple channels in the second offset layer to obtain the second spatial domain features of each video frame in the N video frames; wherein, the time offset amounts of the first temporal offset and the second temporal offset are different, or the channels to be offset are different. The temporal offset of the spatial domain features under a channel may refer to the offset of the spatial domain features under the channel between different video frames. Exemplarily, the spatial domain feature offset may be understood as a spatial domain feature replacement. For example, the spatial domain feature of video frame A1 in channel C1 is offset to the spatial domain feature of video frame A2 in channel C1, which may be understood as that the spatial domain feature of video frame A2 in channel C1 is replaced by the spatial domain feature of video frame A1 in channel C1. The different time offset amounts of the first temporal offset and the second temporal offset may refer to the different numbers of video frames with spatial domain feature offset under the channel in the first temporal offset and the second temporal offset. The different channels to be offset by the first temporal offset and the second temporal offset may refer to the different channels corresponding to the spatial domain features offset by the first offset and the spatial domain features offset by the second offset.
[0430] Next, the temporal offset of the spatial domain features under a channel, the different time offset amounts, and the different channels to be offset will be specifically described.
[0431] In an illustrative example, as Figure 4 shown, the multiple offset layers may include offset layer B1 and offset layer B2. Among them, the time amount offset by the channel temporal offset performed in offset layer B1 is different from the time amount offset by the channel temporal offset performed in offset layer B2.
[0432] Specifically, in the offset layer B1, the spatial domain features under the channel Cx can be offset by T video frames. The channel Cx can include one or more channels, and T is a positive integer greater than or equal to 1 and less than N, or T is a negative integer greater than -N and less than or equal to -1. That is, the spatial domain features of the current video frame under the channel Cx are offset to the video frame that is T - 1 video frames apart from the current video frame, so as to replace the spatial domain features of the video frame that is T - 1 video frames apart from the current video frame under the channel Cx. In other words, in the offset layer B1, the temporal offset is performed on the spatial domain features of each of the N video frames under the channel Cx, and the corresponding time offset amount of this temporal offset is T, where T is a positive integer greater than or equal to 1 and less than N, or T is a negative integer greater than -N and less than or equal to -1, such that the spatial domain features of the k-th video frame among the N video frames under the channel Cx are offset to the spatial domain features of the (k + T)-th video frame under the channel Cx, and k successively takes positive integer values in the interval [1, N], so as to obtain the mixed spatial domain features of each of the N video frames. Exemplarily, the k-th can refer to the k-th in the positive order in the sequence of the N video frames arranged in the time order in video A, or can refer to the k-th in the reverse order.
[0433] For a video frame among the N video frames, after the temporal offset of the spatial domain features under the channel Cx, the spatial domain features in multiple channels thereof are mixed with the spatial domain features of other video frames under the channel Cx (for example, the spatial domain features of the current video frame under the channel Cx are replaced by the spatial domain features of other video frames under the channel Cx). Therefore, after the channel temporal offset of the spatial domain features under the channel Cx performed in the offset layer B1, the spatial domain features of the video frame in multiple channels can be referred to as the mixed spatial domain features B11 of this video frame. That is, the mixed spatial domain features B11 of the N video frames include the mixed spatial domain features B11 of each of the N video frames.
[0434] In the offset layer B2, the spatial domain feature under the channel Cx can be offset by T' video frames, where T' is a positive integer greater than or equal to 1 and less than N, or T' is a negative integer greater than -N and less than or equal to -1. That is, the spatial domain feature of the current video frame under the channel Cx is used to replace the spatial domain feature of the video frame that is T'-1 video frames apart from the current video frame under the channel Cx. In other words, in the offset layer B2, the temporal sequence of the spatial domain features of each of the N video frames under the channel Cx is offset, and the corresponding temporal offset is T', where T' is a positive integer greater than or equal to 1 and less than N, or T is a negative integer greater than -N and less than or equal to -1, such that the spatial domain feature of the k-th video frame among the N video frames under the channel Cx is offset to the spatial domain feature of the (k + T)-th video frame under the channel Cx, and k takes positive integer values in the interval [1, N] in sequence, so as to obtain the mixed spatial domain features of each of the N video frames. Exemplarily, the k-th may refer to the k-th in the positive sequence in the sequence of the N video frames arranged in the time order in video A, or may refer to the k-th in the reverse order.
[0435] The offset result obtained by temporally offsetting the spatial domain features of the N video frames under the channel Cx in the offset layer B2 can be referred to as the mixed spatial domain feature B21 of the N video frames. After a video frame among the N video frames undergoes temporal offset of the spatial domain feature under the channel Cx of the offset layer B2, the spatial domain features of this video frame under multiple channels can be referred to as the mixed spatial domain feature B21 of this video frame. That is, the mixed spatial domain feature B21 of the N video frames includes the mixed spatial domain features B21 of each of the N video frames.
[0436] It should be noted that in the embodiments of the present application, the fact that two video frames are separated by T - 1 (or T' - 1) video frames means that there are T - 1 (or T' - 1) video frames between these two video frames in the sequence of the N video frames arranged in the time order in video A.
[0437] In a specific example, taking the N video frames as video frame A1, video frame A2, video frame A3, video frame A4, and the channel Cx as channel C1 and channel C2 as an example for illustration. In this example, video frame A1, video frame A2, video frame A3, and video frame A4 are adjacent in sequence in the sequence arranged in the time order in video A. The offset is performed in the positive sequence of the time order of the N video frames in video A.
[0438] As Figure 4As shown, T is 1. Then the specific channel timing offset performed in the offset layer B1 is as follows: the spatial domain features of video frame A1 in channel C1 are offset to the spatial domain features of video frame A2 in channel C1; the spatial domain features of video frame A2 in channel C1 are offset to the spatial domain features of video frame A3 in channel C1; the spatial domain features of video frame A3 in channel C1 are offset to the spatial domain features of video frame A4 in channel C1; the spatial domain features of video frame A1 in channel C2 are offset to the spatial domain features of video frame A2 in channel C2; the spatial domain features of video frame A2 in channel C2 are offset to the spatial domain features of video frame A3 in channel C2; the spatial domain features of video frame A3 in channel C2 are offset to the spatial domain features of video frame A4 in channel C2.
[0439] As Figure 4 shown, T' is 2. Then the specific timing offset of the spatial domain features in channel Cx performed in the offset layer B2 is as follows: the spatial domain features of video frame A1 in channel C1 are offset to the spatial domain features of video frame A3 in channel C1; the spatial domain features of video frame A2 in channel C1 are offset to the spatial domain features of video frame A4 in channel C1; the spatial domain features of video frame A1 in channel C2 are offset to the spatial domain features of video frame A3 in channel C2; the spatial domain features of video frame A2 in channel C2 are offset to the spatial domain features of video frame A4 in channel C2.
[0440] In another specific example, taking N video frames as video frame A1, video frame A2, video frame A3, video frame A4, and channel Cx as channel C3 and channel C4 as an example for illustration. In this example, in the sequence arranged according to the time order in video A, video frame A1, video frame A2, video frame A3, and video frame A4 are adjacent in sequence. The offset is performed in the reverse order of the time order of the N video frames in video A.
[0441] As Figure 4 shown, T is 1. Then the specific timing offset of the spatial domain features in channel Cx performed in the offset layer B1 is as follows: the spatial domain features of video frame A4 in channel C3 are offset to the spatial domain features of video frame A3 in channel C3; the spatial domain features of video frame A3 in channel C3 are offset to the spatial domain features of video frame A2 in channel C3; the spatial domain features of video frame A2 in channel C3 are offset to the spatial domain features of video frame A1 in channel C3; the spatial domain features of video frame A4 in channel C4 are offset to the spatial domain features of video frame A3 in channel C4; the spatial domain features of video frame A3 in channel C4 are offset to the spatial domain features of video frame A2 in channel C4; the spatial domain features of video frame A2 in channel C4 are offset to the spatial domain features of video frame A1 in channel C4.
[0442] AsFigure 4 As shown, T’ is 2. Then the temporal offset of the spatial domain features under channel Cx in the offset layer B2 is specifically as follows: offset the spatial domain features of video frame A4 under channel C3 to the spatial domain features of video frame A2 under channel C3; offset the spatial domain features of video frame A3 under channel C3 to the spatial domain features of video frame A1 under channel C3; offset the spatial domain features of video frame A4 under channel C4 to the spatial domain features of video frame A2 under channel C4; offset the spatial domain features of video frame A3 under channel C4 to the spatial domain features of video frame A1 under channel C4.
[0443] In an illustrative example, as Figure 4 shown, the multiple offset layers may include offset layer B1 and offset layer B3. Among them, the temporal offset of the spatial domain features under channel Cx that can be performed in offset layer B1. For details, reference can be made to the above introduction and will not be elaborated here. In offset layer B3, the spatial domain features under channel Cy can be offset by T video frames. T is a positive integer greater than or equal to 1 and less than N, or T is a negative integer greater than -N and less than or equal to -1. That is, offset the spatial domain features of the current video frame under channel Cy to the spatial domain features of the video frame that is T - 1 video frames away from the current video frame under channel Cy. In other words, offset layer B3 performs temporal offset on the spatial domain features of each of the N video frames under channel Cy, and the corresponding time offset amount of this temporal offset is T, where T is a positive integer greater than or equal to 1 and less than N, or T is a negative integer greater than -N and less than or equal to -1, so that the spatial domain features of the kth video frame among the N video frames under channel Cy are offset to the spatial domain features of the (k + T)th video frame under channel Cy, and k successively takes positive integer values in the interval [1, N] to obtain the mixed spatial domain features of each video frame among the N video frames. Exemplarily, the kth can refer to the kth in the positive order in the sequence of the N video frames arranged in the time order in video A, or it can refer to the kth in the reverse order. The offset result obtained after performing temporal offset on the spatial domain features of the N video frames under channel Cy in offset layer B2 can be called the mixed spatial domain feature B31 of the N video frames. After the spatial domain features of a video frame among the N video frames are temporally offset under channel Cy of offset layer B3, the spatial domain features of this video frame under multiple channels can be called the mixed spatial domain feature B31 of this video frame. That is to say, the mixed spatial domain feature B31 of the N video frames includes the mixed spatial domain features B31 of each video frame among the N video frames.
[0444] Channel Cy may include one or more channels, and the one or more channels are different from channel Cx or not included in channel Cx.
[0445] In a specific example, taking N video frames as video frame A1, video frame A2, video frame A3, video frame A4, channel Cx as channel C1 and channel C2, and channel Cy as channel C8 as an example, an illustrative description is given. In this example, in the sequence arranged according to the time order in video A, video frame A1, video frame A2, video frame A3, and video frame A4 are adjacent in sequence. Offset is performed in the positive order of the time order of the N video frames in video A. Among them, in this example, the channel time sequence offset performed in offset layer B1 can refer to the above introduction. Regarding the channel time sequence offset performed in offset layer B3, as shown in Figure 4, it can be set that T is 1, then the spatial domain feature of video frame A1 under channel C8 is offset to the spatial domain feature of video frame A2 under channel C8; the spatial domain feature of video frame A2 under channel C8 is offset to the spatial domain feature of video frame A3 under channel C8; the spatial domain feature of video frame A3 under channel C8 is offset to the spatial domain feature of video frame A4 under channel C8.
[0446] In another specific example, taking N video frames as video frame A1, video frame A2, video frame A3, video frame A4, channel Cx as channel C1 and channel C2, and channel Cy as channel C5 as an example, an illustrative description is given. In this example, in the sequence arranged according to the time order in video A, video frame A1, video frame A2, video frame A3, and video frame A4 are adjacent in sequence. Offset is performed in the reverse order of the time order of the N video frames in video A. Among them, in this example, the channel time sequence offset performed in offset layer B1 can refer to the above introduction. Regarding the channel time sequence offset performed in offset layer B3, as shown in Figure 4, it can be set that T is 1, then the spatial domain feature of video frame A4 under channel C5 is offset to the spatial domain feature of video frame A3 under channel C5; the spatial domain feature of video frame A3 under channel C5 is offset to the spatial domain feature of video frame A2 under channel C5; the spatial domain feature of video frame A2 under channel C5 is offset to the spatial domain feature of video frame A1 under channel C5.
[0447] In an illustrative example, as Figure 4 shown, multiple offset layers can simultaneously include offset layer B1 and offset layer B3. The channel time sequence offset performed in each offset layer can refer to the above description.
[0448] In an illustrative example, as Figure 4 shown, multiple offset layers can also simultaneously include offset layers such as offset layer B1, offset layer B2, and offset layer B3. The channel time sequence offset performed in each offset layer can refer to the above description.
[0449] In an illustrative example, when the spatial domain features of a video frame, such as video frame A1, under channel Cx (or channel Cy) are offset to the spatial domain features of other video frames under Cx (or channel Cy), if no spatial domain features are offset to the spatial domain features of video frame A1 under channel Cx (or channel Cy), the spatial domain features of video frame A1 under channel Cx (or channel Cy) can be replaced with zeros. Specifically, the spatial domain features can be in matrix form, and replacing the spatial domain features with zeros means filling the matrix with zeros, that is, the elements in the matrix are zeros.
[0450] In an illustrative example, for a video frame, the spatial domain features of this video frame output by the spatial domain feature extraction layer under multiple channels can be used to perform feature compensation on the spatial domain features of this video frame output by the offset layer under multiple channels. Specifically, the spatial domain features of this video frame output by the spatial domain feature extraction layer under multiple channels and the spatial domain features of this video frame output by the offset layer under multiple channels can be accumulated at corresponding positions. It can be understood that the spatial domain features can be represented by a matrix, or rather, the spatial domain features are in matrix form. For a video frame, the spatial domain features of this video frame output by the spatial domain feature extraction layer under channel Cj and the spatial domain features of this video frame output by the offset layer under channel Cj can be added point by point. The result after addition can be called the residual spatial domain features of this video frame under channel Cj. Among them, channel Cj is one or more channels among multiple channels. For the convenience of expression, hereinafter, the residual spatial domain features can also be called the mixed spatial domain features.
[0451] Through the above offset processing of the spatial domain features, the mixed spatial domain features of N video frames are obtained. The mixed spatial domain features of N video frames are input into the classification layer.
[0452] The computer can execute step 502, and in the classification layer, at least determine the semantics of the N video frames according to the first spatial domain features and the second spatial domain features of each video frame in the N video frames. Among them, the semantics of the N video frames refer to the specific or expressed dynamic semantics of the whole of the N video frames.
[0453] In an illustrative example, the first spatial domain feature in step 502 can refer to the mixed spatial domain feature B11 described above, and the second spatial domain feature can refer to the mixed spatial domain feature B21 described above. In some embodiments, the first spatial domain feature in step 502 can refer to the mixed spatial domain feature B11 described above, and the second spatial domain feature can refer to the mixed spatial domain feature B31 described above. In some embodiments, the first spatial domain feature in step 502 can refer to the mixed spatial domain feature B21 described above, and the second spatial domain feature can refer to the mixed spatial domain feature B31 described above.
[0454] In an illustrative example, the classification layer may include multiple two-dimensional convolutional layers, such as convolutional layer b1, convolutional layer b2, convolutional layer b3, etc. The convolutional layers in the multiple two-dimensional convolutional layers correspond one-to-one with the offset layers in the multiple offset layers. For example, convolutional layer b1 corresponds to offset layer B1, convolutional layer b2 corresponds to offset layer B2, and convolutional layer b3 corresponds to offset layer B3.
[0455] The offset layer can output the obtained mixed spatial domain features to the corresponding convolutional layer, and perform convolutional processing in the convolutional layer to obtain fused spatio-temporal features. It can be understood that the mixed spatial domain features obtained by performing temporal offset on the spatial domain features in the offset layer are only a simple or mechanical mixture of the spatial domain features of different video frames, and it is still difficult to reflect the temporal information between different video frames. Further convolutional processing of the mixed spatial domain features can fuse the spatial domain features of the video frames to obtain information that can reflect the temporal features between different video frames. In the embodiments of the present application, the information obtained after convolutional processing of the mixed spatial domain features can be referred to as fused spatio-temporal features.
[0456] The different convolutional layers in the multiple two-dimensional convolutional layers in the classification layer have the same type and number of convolutional kernels. It can be set that each convolutional layer includes M convolutional kernels, where the i-th convolutional kernel among the M convolutional kernels of different convolutional layers is the same, and i takes positive integer values in the interval [1, M]. It can be understood that one convolutional kernel is used to extract a certain feature of the image. The sameness of two convolutional kernels means that the features extracted by the two convolutional kernels are the same. For example, the convolutional kernel of convolutional layer b1 for extracting the red pixel feature of the image (feature under the R channel) and the convolutional kernel of convolutional layer b2 for extracting the red pixel feature of the image (feature under the R channel) are the same convolutional kernel. The convolutional kernel of convolutional layer b1 for extracting the green pixel feature of the image (feature under the G channel) and the convolutional kernel of convolutional layer b2 for extracting the green pixel feature of the image (feature under the G channel) are the same convolutional kernel. The convolutional kernel of convolutional layer b1 for extracting the blue pixel feature of the image (feature under the B channel) and the convolutional kernel of convolutional layer b2 for extracting the blue pixel feature of the image (feature under the B channel) are the same convolutional kernel.
[0457] In the first example, the multiple offset layers include an offset layer B1 and an offset layer B2, and the classification layer includes a convolutional layer b1 corresponding to the offset layer B1 and a convolutional layer b2 corresponding to the offset layer B2. In the convolutional layer b1, M convolutional kernels can be used to perform convolutional processing on the mixed spatial domain feature B11 of each of the N video frames output by the offset layer B1, so as to obtain the fused spatio-temporal feature b11 of each video frame in M channels. In the convolutional layer b2, M convolutional kernels can be used to perform convolutional processing on the mixed spatial domain feature B21 of each of the N video frames output by the offset layer B2, so as to obtain the fused spatio-temporal feature b21 of each video frame in M channels. As can be seen from the above, the i-th convolutional kernel among the M convolutional kernels of the offset layer b1 and the i-th convolutional kernel among the M convolutional kernels of the offset layer b2 are the same convolutional kernel. The fused spatio-temporal feature under the i-th channel among the M channels is obtained by performing convolutional processing on the mixed spatial domain feature by the i-th convolutional kernel among the M convolutional kernels. For a video frame, the fused spatio-temporal feature b11 of this video frame output by the convolutional layer b1 under the i-th channel among the M channels and the fused spatio-temporal feature b21 of this video frame output by the convolutional layer b2 under the i-th channel among the M channels can be subjected to cumulative summation of corresponding positions, where i takes a positive integer value in the interval [1, M]. The fused spatio-temporal feature under each channel can be represented by a matrix, and the summation of corresponding positions of the fused spatio-temporal feature under one channel and the fused spatio-temporal feature under another channel can be the point-to-point addition of matrices and matrices.
[0458] In the second example, the multiple offset layers include an offset layer B1 and an offset layer B3, and the classification layer includes a convolutional layer b1 corresponding to the offset layer B1 and a convolutional layer b3 corresponding to the offset layer B3. In the convolutional layer b3, M convolutional kernels can be used to perform convolutional processing on the mixed spatial domain feature B31 of each of the N video frames output by the offset layer B3, so as to obtain the fused spatio-temporal feature b31 of each video frame in M channels. For a video frame, the fused spatio-temporal feature b11 of this video frame output by the convolutional layer b1 under the i-th channel among the M channels and the fused spatio-temporal feature b31 of this video frame output by the convolutional layer b3 under the i-th channel among the M channels can be subjected to cumulative summation of corresponding positions, where i takes a positive integer value in the interval [1, M]. The fused spatio-temporal feature under each channel can be represented by a matrix, and the summation of corresponding positions of the fused spatio-temporal feature under one channel and the fused spatio-temporal feature under another channel can be the point-to-point addition of matrices and matrices.
[0459] In the third example, the multiple offset layers include an offset layer B2 and an offset layer B3, and the classification layer includes a convolutional layer b2 corresponding to the offset layer B2 and a convolutional layer b3 corresponding to the offset layer B3. For a video frame, the fused spatio-temporal feature b21 of the video frame output by the convolutional layer b2 in the i-th channel among the M channels and the fused spatio-temporal feature b31 of the video frame output by the convolutional layer b3 in the i-th channel among the M channels can be subjected to cumulative summation at corresponding positions, where i takes positive integer values in the range [1, M]. The fused spatio-temporal feature under each channel can be represented by a matrix, and the cumulative summation of the corresponding positions of the fused spatio-temporal feature under one channel and the fused spatio-temporal feature under another channel can be the element-wise addition of the matrices.
[0460] In the fourth example, the multiple offset layers include an offset layer B1, an offset layer B2, and an offset layer B3, and the classification layer includes a convolutional layer b1 corresponding to the offset layer B1, a convolutional layer b2 corresponding to the offset layer B2, and a convolutional layer b3 corresponding to the offset layer B3. For a video frame, the fused spatio-temporal feature b11 of the video frame output by the convolutional layer b1 in the i-th channel among the M channels, the fused spatio-temporal feature b21 of the video frame output by the convolutional layer b2 in the i-th channel among the M channels, and the fused spatio-temporal feature b31 of the video frame output by the convolutional layer b3 in the i-th channel among the M channels can be subjected to cumulative summation at corresponding positions, where i takes positive integer values in the range [1, M]. The fused spatio-temporal feature under each channel can be represented by a matrix, and the cumulative summation of the corresponding positions of the fused spatio-temporal feature under one channel and the fused spatio-temporal feature under another channel can be the element-wise addition of the matrices.
[0461] In the fifth example, the mixed spatial domain feature of the video frame processed by the convolutional layer in the classification layer is the residual spatial domain feature of the video frame, that is, the mixed spatial domain feature obtained by the convolutional layer is the spatial domain feature obtained after compensating the spatial domain feature output by the offset layer with the spatial domain feature output by the spatial domain feature extraction layer. A more detailed introduction to the residual spatial domain feature can be referred to the above description and will not be elaborated here.
[0462] For ease of description, the result of cumulative summation of the corresponding positions of the fused spatio-temporal features of the video frame output by different convolutional layers in the i-th channel among the M channels can be referred to as the cumulative fused spatio-temporal feature of the video frame in the i-th channel. That is to say, the result of cumulative summation at corresponding positions in the first example, the result of cumulative summation at corresponding positions in the second example, the result of cumulative summation at corresponding positions in the third example, and the result of cumulative summation at corresponding positions in the fourth example described above can all be referred to as the cumulative fused spatio-temporal feature in the i-th channel. Referring to the determination method of the cumulative fused spatio-temporal feature in the i-th channel, the cumulative fused spatio-temporal features in each of the M channels can be obtained.
[0463] In the classification layer, the accumulated fusion spatio-temporal features of each video frame among N video frames under M channels can be utilized to classify the semantics of video A, or in other words, to recognize the semantics of video A. The specific solution can be as follows.
[0464] In an illustrative example, for a video frame A1 among N video frames, the accumulated fusion spatio-temporal feature of it under the i-th channel among M channels can be mapped into a feature value. Exemplarily, the classification layer can also include a fully connected layer. In the fully connected layer, the accumulated fusion spatio-temporal feature of video frame A1 under the i-th channel among M channels can be mapped into a feature value. For example, the accumulated fusion spatio-temporal feature of video frame A1 under the i-th channel among M channels can be convolved to obtain a feature value. Among them, the parameters in the convolution kernel used for this convolution process can be obtained through supervised training. Referring to the determination method of the feature value corresponding to the i-th channel, the feature values corresponding to each channel among M channels can be obtained. The feature values corresponding to each channel among the M channels of video frame A1 can be accumulated to obtain an accumulated feature value.
[0465] The accumulated feature value of video frame A1 is multiplied by the weight coefficient q1 among Q weight coefficients, and the bias corresponding to the weight coefficient q1 is added to obtain the value q11 corresponding to video frame A1. Referring to the determination method of the value q11 corresponding to video frame A1, the values q11 corresponding to each video frame among N video frames can be obtained. The accumulated sums obtained by accumulating the values q11 corresponding to each video frame are used as elements in a column vector. Using each of the Q weight coefficients and the bias corresponding to the weight coefficient, the foregoing calculations are performed on the accumulated feature values of the video frames among N video frames, and a column vector containing Q elements can be obtained. Among them, different weight coefficients among the Q weight coefficients correspond to preset different semantics or classification results. For example, the weight coefficient q1 among the Q weight coefficients corresponds to playing basketball, the weight coefficient q2 corresponds to playing football, and so on. The weight coefficients and their corresponding biases can be obtained through supervised training.
[0466] The softmax function can be used to map the column vector determined above into Q probability values, where each probability value corresponds to a classification result. The classification result corresponding to the maximum probability value among the Q probability values can be used as the semantics of video A.
[0467] Back to Figure 5 , the computer can execute step 503, and at the output layer, output the semantics of N video frames.
[0468] The semantics of N video frames can be used as the dynamic semantics of the N-th video frame among the N video frames.
[0469] Figure 5 The dynamic semantic recognition method shown can perform temporal offsets of different time amounts on the spatial domain features of video frames in a video, so as to capture information of objects with different motion frequencies in the video; it can also perform temporal offsets of spatial domain features in different channels on the spatial domain features of video frames in the video, so as to ensure that information of potential object categories is captured, thereby improving the accuracy of video semantic classification results. Moreover, the video semantic recognition method provided in the embodiments of the present application uses a parallel method to perform temporal offsets of different time amounts or temporal offsets of spatial domain features in different channels on the spatial domain features of video frames in the video, avoiding the problem of unexplainability or poor interpretability caused by multiple hybridizations of the temporal information and spatial information of the video in a serial manner.
[0470] In some embodiments, another scheme for determining the dynamic semantics of video frames is provided. Next, in combination with Figure 7 、 Figure 8 this scheme will be illustrated by examples.
[0471] Figure 7 FIG. shows a specific structure of the dynamic semantic recognition layer. It includes at least one residual network layer, a classification layer, and an output layer arranged in series; wherein, the spatial domain feature layer includes at least one two-dimensional convolutional layer arranged in series; each residual network layer in the at least one residual network layer includes a plurality of spatio-temporal feature extraction layers and a spatial domain feature compensation layer arranged in series; wherein, each spatio-temporal feature extraction layer in the plurality of spatio-temporal feature extraction layers includes an offset sub-layer and a convolutional sub-layer arranged in series.
[0472] Based on Figure 7 the dynamic semantic recognition layer shown, the dynamic semantics of video frames can be determined through the steps shown in Figure 8 FIG.
[0473] Before performing Figure 8 the steps shown, the spatial domain features of each of the N video frames in multiple channels can be obtained from the dynamic semantic recognition layer. The spatial domain features can refer to the introduction of the embodiments shown in Figure 6 above and will not be elaborated here.
[0474] Step 801, in the first spatio-temporal feature extraction layer of the first residual network layer in the at least one residual network layer, extract the first fused spatio-temporal features of the N video frames; when the first spatio-temporal feature layer is the last one of the multiple spatio-temporal feature extraction layers in the first residual network layer, in the spatial domain feature compensation layer of the first residual network layer, determine the residual spatial domain features of the N video frames to be output by the first residual network layer according to the first fused spatio-temporal features of the N video frames and the spatial domain features of the N videos in the multiple channels.
[0475] The first fused spatio-temporal feature extraction of the N video frames includes: in the offset sub-layer of the first spatio-temporal feature extraction layer, performing temporal offset on the spatial domain features of the N video frames in at least one channel to obtain the first spatial domain features of each video frame in the N video frames, where the spatial domain features in the at least one channel are a part of the spatial domain features output by the previous layer of the first spatio-temporal feature extraction layer; in the convolution sub-layer of the first spatio-temporal feature extraction layer, performing convolution processing on the first spatial domain features of each video frame in the N video frames to obtain the first fused spatio-temporal feature of the N video frames.
[0476] In an illustrative example, as Figure 7 shown, at least one residual network may include residual network layer 1......residual network layer m arranged in series. Among them, residual network layer 1 is the first residual network layer in the at least one residual network layer. Residual network layer 1 includes spatio-temporal feature extraction layer 1......spatio-temporal feature extraction layer n arranged in series. In spatio-temporal feature extraction layer 1, temporal offset can be performed on the spatial domain features of each video frame in the N video frames output by the spatial domain feature extraction layer in at least one channel among multiple channels. The temporal offset of the spatial domain features can specifically refer to the introduction of steps 501a and 501b above, and will not be elaborated here. In the convolution sub-layer of spatio-temporal feature extraction layer 1, convolution processing is performed on the output of spatio-temporal feature extraction layer 1. Specifically, it can refer to the introduction of step 502 above, and will not be elaborated here. The offset sub-layer of each spatio-temporal feature extraction layer in residual network layer 1 performs temporal offset on the output of the previous layer of this offset layer for the spatial domain features, and the convolution layer of this spatio-temporal feature extraction layer performs convolution processing on the output of the previous layer of this convolution layer. The spatial domain feature compensation layer in residual network layer 1 can use the spatial domain features of the N video frames extracted by the spatial domain feature extraction layer in multiple channels to perform spatial domain feature compensation on the output of the convolution sub-layer of spatio-temporal feature extraction layer n (i.e., the last spatio-temporal feature extraction layer among the multiple spatio-temporal feature extraction layers of residual network layer 1). The spatial domain feature compensation can specifically refer to the introduction of Figure 5 the method embodiment shown above, and will not be elaborated here.
[0477] The spatial domain feature compensation layer of residual network layer 1 can output its processing result to the next residual network layer, and the offset sub-layer of the next residual network can perform temporal offset on the output result of this spatial domain feature compensation layer for the spatial domain features in at least one channel, and so on. The feature extraction method of the next residual network layer can refer to the introduction of the feature extraction method of residual network layer 1 above, and will not be elaborated here.
[0478] When the next residual network layer is not the last one among at least one residual network layer, the spatial domain feature compensation layer of the next residual network layer can output its processing result to its next residual network layer.
[0479] When the next residual network is the last one among at least one residual network layer, for example, when the next residual network is Figure 7 the residual network layer m shown, the spatial domain feature compensation layer of the residual network layer m can output its processing result to the classification layer.
[0480] Step 802, in the classification layer, determine the semantics of the N video frames according to the residual spatial domain features of the N video frames output by the last residual network layer among the at least one residual network layer. Specifically, reference can be made to the introduction of step 502 above, which will not be elaborated here.
[0481] Step 803, in the output layer, output the semantics of the N video frames. The semantics of the N video frames can be determined as the semantics of the Nth video frame among the N video frames.
[0482] Figure 8 The dynamic semantics determination method shown can perform multiple temporal offsets on the spatial domain features of video frames in a video. During the multiple temporal offset processes, feature compensation is performed on the spatially domain features after temporal offset, achieving the guarantee that spatial domain information is not lost or lost less while extracting temporal information multiple times, and providing the accuracy of video semantic recognition.
[0483] Through the above solution, video frames with static semantics and video frames with dynamic semantics can be obtained.
[0484] Returning to Figure 2 , the video semantic recognition method provided by the embodiment of the present application further includes: Step 203b, in the static semantic recognition layer, determine the static semantics of the first video frame according to the spatial domain features of the first video frame among the multiple video frames.
[0485] In the static semantic recognition layer, the static semantics of a video frame can be recognized according to the spatial domain features corresponding to the video frame. In one example, the spatial domain features of the video frame can be input into the softmax function, and then calculated to obtain the softmax value. The softmax value can be used to classify the video frame. Thus, the static semantics of the video frame can be obtained.
[0486] Referring to Figure 1 , the neural network provided by the embodiment of the present application further includes a temporal segment division layer. In the temporal segment division layer, temporal segments with static semantics or dynamic semantics can be divided according to the static semantics or dynamic semantics of video frames. Next, an example is introduced.
[0487] Refer to Figure 2 Figure 2 , the video semantic recognition method provided by the embodiment of the present application further includes: Step 204a, in the timing segment division layer, when the number of consecutive video frames with the first dynamic semantics is greater than the first threshold, use the consecutive video frames with the first dynamic semantics to synthesize the first timing segment, and determine the first dynamic semantics as the dynamic semantics of the first timing segment
[0488] In the timing segment division layer, consecutive video frames with the same dynamic semantics can be screened out from the multiple video frames obtained by video frame division of video A1. For example, the dynamic semantics can be set as dynamic semantics F1 (such as playing football), dynamic semantics F2 (such as washing an apple), etc. It can be set that each of the h-th to (h'+s')-th video frames in the multiple frames of video has the dynamic semantics F1, while the (h'-1)-th video frame and the (h'+s'+1)-th video frame do not have the dynamic semantics F1. Then, the h'-th to (h'+s')-th video frames are consecutive video frames with the dynamic semantics F1. For convenience of description, consecutive video frames with the same dynamic semantics can be referred to as the set of dynamic video frames to be integrated
[0489] It can be understood that one or more sets of dynamic video frames to be integrated can be screened out from the multiple video frames. For a set of dynamic video frames to be integrated, it can be determined whether the number of video frames included therein is greater than the threshold D1. The threshold D1 can be a preset value, such as 10, or 15, etc., which will not be listed one by one here
[0490] If the number of video frames included in the set of dynamic video frames to be integrated is greater than the threshold D1, the set of dynamic video frames to be integrated can be synthesized into a timing segment, and the dynamic semantics of the video frames in the set of dynamic video frames to be integrated can be used as the dynamic semantics of the timing segment. It should be noted that the meaning of "synthesis" here is not to synthesize multiple images or video frames into one image, but to form a sequence of video frames from multiple images or video frames, and this sequence is the timing segment
[0491] If the number of video frames included in the set of dynamic video frames to be integrated is less than or equal to the threshold D1, the video frames in the set of dynamic video frames to be integrated are not used to synthesize the timing segment
[0492] Through the above solution, one or more timing segments with dynamic semantics are obtained
[0493] In some embodiments, if the timing segment E1 has a dynamic semantics F1, the timing segment E2 also has the dynamic semantics F1. And if the number of video frames between the timing segment E1 and the timing segment E2 is less than a threshold D2, the timing segment E2, the timing segment E2, and the video frames between the two can be combined into one timing segment, and the dynamic semantics F1 is used as the dynamic semantics of the combined timing segment. The threshold D2 can be a preset value. For example, it can be 2, it can be 3, and so on. They are not listed one by one here.
[0494] Refer to Figure 2 , the video semantics recognition method provided by the embodiments of the present application further includes: Step 204b, at the timing segment division layer, when the number of consecutive video frames with the first static semantics is greater than a second threshold, use the consecutive video frames with the first static semantics to synthesize a second timing segment, and determine the first static semantics as the static semantics of the second timing segment.
[0495] At the timing segment division layer, consecutive video frames with the same static semantics can be screened out from the multiple video frames obtained by video frame division of video A1. For example, it can be set that the static semantics include static semantics C1 (such as football), static semantics C2 (such as apple), etc. It can be set that each of the hth to the (h + s)th video frames in the multi-frame video has the static semantics C1, while the (h - 1)th video frame and the (h + s + 1)th video frame do not have the static semantics C1. Then the hth to the (h + s)th video frames are consecutive video frames with the static semantics C1. For ease of description, consecutive video frames with the same static semantics can be referred to as a set of static video frames to be integrated.
[0496] It can be understood that one or more sets of static video frames to be integrated can be screened out from the multiple video frames. For a set of static video frames to be integrated, it can be determined whether the number of video frames it includes is greater than a threshold D3. The threshold D3 can be a preset value. For example, it can be 10, or it can be 15, and so on. They are not listed one by one here.
[0497] If the number of video frames included in the set of static video frames to be integrated is greater than the threshold D3, the set of static video frames to be integrated can be combined into one timing segment, and the static semantics of the video frames in the set of static video frames to be integrated are used as the static semantics of the timing segment. It should be noted that the meaning of "synthesis" here is not to synthesize multiple images or video frames into one image, but to form a sequence of video frames from multiple images or video frames, and this sequence is the timing segment.
[0498] If the number of video frames included in the set of static video frames to be integrated is less than or equal to the threshold D3, the video frames in the set of static video frames to be integrated are not used to synthesize a temporal segment.
[0499] Through the above solution, one or more temporal segments with static semantics are obtained.
[0500] In some embodiments, if the temporal segment E3 has the static semantics C1 and the temporal segment E4 also has the static semantics C1. And the number of video frames between the temporal segment E3 and the temporal segment E4 is less than the threshold D4, the temporal segment E3, the temporal segment E4, and the video frames between the two can be synthesized into one temporal segment, and the static semantics C1 is used as the static semantics of the synthesized temporal segment. The threshold D4 can be a preset value, for example, it can be 2, it can be 3, and so on, which will not be listed one by one here.
[0501] In some embodiments, continue to refer to Figure 1 , the neural network provided by the embodiment of the present application may further include a semantic smoothing layer. The semantic smoothing layer is the upper layer of the temporal segment division layer and the lower layer of the static semantic recognition layer and the dynamic semantic recognition layer.
[0502] In the semantic smoothing layer, the static semantics (or dynamic semantics) of the multiple video frames can be smoothed according to the dependence relationship of the static semantics (or dynamic semantics) between consecutive video frames in the multiple video frames.
[0503] In an illustrative example, it can be determined that the static semantics of the video frame G1 in P consecutive video frames is different from the static semantics of other video frames, and the static semantics of other video frames are the same; the other video frames are the video frames in the P consecutive video frames except the video frame G1, and P is greater than the threshold D5. The threshold D5 can be a preset value, for example, it can be 10, or it can be 12, and so on, which will not be listed one by one here. Update (or replace) the static semantics of the video frame G1 according to the static semantics of the other video frames. In one example, the video frame G1 is not the edge video frame in the P consecutive videos. Among them, the edge video frames of the P consecutive video frames can be defined as the first to the pth video frames and the (P - p)th to the Pth video frames in the P consecutive video frames. p can be a preset value and less than half of P.
[0504] In an illustrative example, it can be determined that the dynamic semantics of video frame G2 in Q consecutive video frames are different from those of other video frames, and the dynamic semantics of the other video frames are the same; the other video frames are the video frames in the Q consecutive video frames except video frame G2, and Q is greater than threshold D6. Threshold D6 can be a preset value, for example, it can be 10, or it can be 12, etc., which will not be listed one by one here. Update (or replace) the dynamic semantics of video frame G2 according to the dynamic semantics of the other video frames. In one example, video frame G2 is not an edge video frame in the Q consecutive videos. Among them, the edge video frames of the Q consecutive video frames can be defined as the first to the q-th video frames and the (Q - q)-th to the Q-th video frames in the Q consecutive video frames. q can be a preset value and is less than one-half of Q.
[0505] In some embodiments, another solution is provided for smoothing the static semantics (or dynamic semantics) of multiple video frames according to the dependency relationship of the static semantics (or dynamic semantics) between consecutive video frames in the multiple video frames.
[0506] As described above, the static semantics refer to the categories of the objects included in the video frame. Through the static semantics recognition layer, the static semantics of each video frame can be determined, so that category labels can be carried or added to the video frames. The category labels can represent the objects included in the video frames. That is to say, the category labels are used to characterize the information of the categories of the objects in the video frames (it can be understood that the objects in the video frames are actually the images of the corresponding objects. For the convenience of description, the images of the objects in the video frames are called the objects in the video frames). Generally speaking, one category label corresponds to one object category. For example, the subway label is a category label, which corresponds to the object category of the subway. For another example, the bus label is another category label, which corresponds to the object category of the bus. For another example, the cat label is another category label, which corresponds to the object category of the cat. Etc., which will not be listed one by one here.
[0507] For a video frame A1 in a video, when analyzing the object category B1 included in it through the video frame or video semantics recognition model, the category label corresponding to the object category B1 can be carried for video frame A1; when analyzing that it includes the object category B2 through the video frame or video semantics recognition model, the category label corresponding to the object category B2 can be carried for video frame A1; when analyzing that it does not include an object category through the video frame or video semantics recognition model, no category label can be carried for it. That is to say, video frame A1 can carry one or more category labels, or can carry no category label, which can be specifically determined by the analysis result of the video frame or video semantics recognition model.
[0508] The smoothness of the video frame labels in the video means that the category labels carried by the video frames in the video are consistent with common sense or experience. For example, the previous video frame carries the subway label, and the next video frame also carries the subway label. For another example, the previous video frame carries the cat label, and the next video frame carries the cat food label.
[0509] In contrast to smooth video frame labels in a video, non-smooth video frame labels in a video means that one or some of the category labels carried in the video do not conform to common sense or experience. For example, a video frame carries both a subway label and an airplane label. For another example, the previous video frame carries a cat label, and the next video frame carries a dog label.
[0510] The uneven video frame labels in the video are often caused by errors in the video frame or video semantic recognition model, whether the camera direction changes suddenly during video shooting (such as shaking), etc. The uneven video frame labels in the video are not good for video editing. For example, they do not meet the requirements of certain video editing algorithms. For another example, they affect the viewing quality of the edited video.
[0511] The embodiment provides a dependency relationship between static semantics (or dynamic semantics) of consecutive video frames in multiple video frames, and a solution for smoothing the static semantics (or dynamic semantics) of the multiple video frames. According to the label smoothing strategy, the category label carried by at least one video frame in the multiple video frames is added or deleted to achieve the category label smoothing of the multiple video frames. Among them, adding or deleting the category label carried by at least one video frame in the multiple video frames can be called label smoothing of the multiple video frames or the video.
[0512] It should be noted that in the embodiment of the present application, the meaning of "adding and deleting category labels of video frames" is: when acquiring the multiple video frames, that is, before the label smoothing processing is performed, if the video frame does not carry the category label, then add the category label to the video frame; if the video frame carries the category label, then delete the category label carried by the video frame.
[0513] Next, combine Figures 9 - 19B , the dependency relationship between static semantics (or dynamic semantics) of consecutive video frames in the multiple video frames provided in this embodiment, and the scheme for smoothing the static semantics (or dynamic semantics) of the multiple video frames are illustrated.
[0514] In step 901, multiple video frames in a video can be obtained. It can be understood that a video frame in a video is an image, multiple video frames can also be called multiple images, and video frames can also be called images. Specifically, multiple video frames can be obtained from a static semantic recognition layer, or multiple video frames can be obtained from a dynamic semantic recognition layer.
[0515] Exemplarily, the multiple video frames can be multiple sequentially adjacent video frames in the video. For example, it can be set that the video has a total of 100 video frames, and these 100 video frames can be used as the multiple video frames.
[0516] Exemplarily, the multiple video frames can be multiple video frames obtained by frame extraction processing on the video. For example, in the video, one can be extracted every m (m is an integer greater than 1), and after such extraction multiple times, the multiple video frames can be obtained. That is to say, in this example, for two adjacent video frames in the multiple video frames, such as video frame A1 and video frame A2, their positions in the video are not necessarily adjacent, and there can be m - 1 video frames between them.
[0517] The video frames in the multiple video frames can carry one or more class labels, where the class label carried by each video frame can be a label added to the video frame after object recognition of the video frame by the video frame or the video semantic recognition model. For specific details, reference can be made to the above introduction, and details will not be elaborated here.
[0518] Step 902: Perform label smoothing processing on the multiple video frames according to the label smoothing strategy corresponding to the video.
[0519] Next, the label smoothing strategy will be introduced by way of example first.
[0520] The label smoothing strategy can be understood as a strategy for expecting or controlling what label categories different video frames or the same video frame should carry. In the embodiments of the present application, the label smoothing strategy can also be referred to as prior knowledge.
[0521] In an illustrative example, it can be understood that, according to common sense or experience, a certain object category tends to or should have continuity in multiple video frames of the same video. The continuity of the object category can refer to that, according to common sense or experience, the object category tends to or should appear or not appear simultaneously in adjacent video frames before and after in the video. For example, under normal circumstances, if video frame A1 in the video includes the object category of subway, the next video frame of video frame A1 (i.e., the video frame located after video frame A1 in the time sequence of the video) should also include the object category of subway. If video frame A in the video does not include the object category of subway, the next video frame of video frame A1 should also not include the object category of subway. Thus, in the embodiments of the present application, for an object category with continuity, when its corresponding category label is carried by video frame A1, it is expected that the next video frame of video frame A1 also carries this category label; when its corresponding category label is not carried by video frame A1, it is expected that the next video frame of video frame A1 also does not carry this category label. That is to say, for an object category with continuity, its corresponding category label should also have continuity between adjacent video frames before and after in the video. In this way, the video is smooth under this category label. For example, the object category of subway has continuity. If video frame A1 carries the subway category label, it is expected that the next video frame of video frame A1 also carries the subway category label; if video frame A1 does not carry the subway category label, it is expected that the next video frame of video frame A1 also does not carry the subway category label. That is, the subway category label should also have continuity.
[0522] Thus, for the above-mentioned multiple video frames, the label smoothing strategy can be or include that the category label C1 is continuous between adjacent video frames in these multiple video frames. The object category corresponding to the category label C has continuity in the multiple video frames. Among them, for the convenience of the following description, it can be considered that the continuity of one or more category labels between adjacent video frames in these multiple video frames is a smoothing sub-strategy.
[0523] In an illustrative example, it can be understood that, based on common sense or experience, two or more object categories tend to or should coexist in the same video frame, which can be referred to as the coexistence of object categories. For example, the object category of subway and the object category of people have coexistence, and they tend to or should coexist in the same video frame. Thus, in the embodiments of the present application, for two or more object categories with coexistence, when the category label corresponding to one object category is carried by video frame A1, it is expected that video frame A1 also carries the category labels corresponding to other object categories. That is to say, the category labels corresponding to the coexistent object categories should also coexist on the same video frame. For example, if video frame A1 carries the category label corresponding to the object category of subway, it is expected that video frame A1 also carries the category label corresponding to the object category of people.
[0524] Thus, for the above-mentioned multiple video frames, the label smoothing strategy can be or include the coexistence of category label C2 and category label C3 on the same video frame among the multiple video frames. The object category corresponding to category label C2 and the object category corresponding to category label C3 have coexistence. Among them, for the convenience of the following description, the coexistence of category label C2 and category label C3 on the same video frame among the multiple video frames can be regarded as a smoothing sub-strategy.
[0525] In an illustrative example, it can be understood that, based on common sense or experience, two or more object categories tend not to or should not coexist in the same video frame, which can be referred to as the non-coexistence of object categories. For example, the object category of subway and the object category of airplane have non-coexistence, and they tend not to or should not coexist in the same video frame. Thus, in the embodiments of the present application, for two or more object categories with non-coexistence, when the category label corresponding to one object category is carried by video frame A1, it can be expected that video frame A1 does not carry the labels corresponding to other object categories. That is to say, the category labels corresponding to the non-coexistent object categories should also have non-coexistence on the same video frame. For example, if video frame A1 carries the category label corresponding to the object category of subway, it can be expected that video frame A1 does not carry the category label corresponding to the object category of airplane.
[0526] Thus, for the above-mentioned multiple video frames, the label smoothing strategy can be or include the non-coexistence of category label C4 and category label C5 on the same video frame among the multiple video frames. The object category corresponding to category label C4 and the object category corresponding to category label C5 have non-coexistence. Among them, for the convenience of the following description, the non-coexistence of category label C4 and category label C5 on the same video frame among the multiple video frames can be regarded as a smoothing sub-strategy.
[0527] In an illustrative example, it can be understood that, based on common sense or experience, when a video frame A1 in a video contains an object category B1, the next video frame of video frame A1 often or should contain an object category B2, which can be referred to as the derivability between object category B1 and object category B2. That is to say, video frame A1 containing object category B1 implies that the next video frame of video frame A1 should contain object category B2. For example, video frame A1 contains the object category of subway security gate, which implies that the next video frame of video frame A1 should contain the object category of subway turnstile. Thus, in the embodiment of the present application, when video frame A1 carries the category label corresponding to object category B1, it can be expected that the next video frame of video frame A1 should carry the category label corresponding to object category B2.
[0528] Thus, for the above-mentioned multiple video frames, the label smoothing strategy can be or include that the label category corresponding to object category B1 exists in the previous video frame among adjacent video frames in the multiple video frames, and the label category corresponding to object category B2 exists in the subsequent video frame among adjacent video frames.
[0529] For the convenience of the following description, it can be considered that the label category corresponding to object category B1 exists in the previous video frame among adjacent video frames in the multiple video frames, and the label category corresponding to object category B2 exists in the subsequent video frame among adjacent video frames as a smoothing sub-strategy.
[0530] In an illustrative example, it can be understood that, based on common sense or experience, when a video frame A1 in a video contains an object category B3, the next video frame of video frame A1 often does not or should not contain an object category B4, which can also be referred to as the derivability between object category B3 and object category B4. That is to say, video frame A1 containing object category B3 implies that the next video frame of video frame A1 should not contain object category B4. For example, video frame A1 contains the object category of grassland, which implies that the next video frame of video frame A1 should not contain the object category of desert. Thus, in the embodiment of the present application, when video frame A1 carries the category label corresponding to object category B3, it can be expected that the next video frame of video frame A1 should not carry the category label corresponding to object category B4.
[0531] Thus, for the above-mentioned multiple video frames, the label smoothing strategy can be or include that the label category corresponding to object category B3 exists in the previous video frame among adjacent video frames in the multiple video frames, and the label category corresponding to object category B4 does not exist in the subsequent video frame among the adjacent video frames. For the convenience of the following description, it can be considered that the label category corresponding to object category B3 exists in the previous video frame among adjacent video frames in the multiple video frames, and the label category corresponding to object category B4 does not exist in the subsequent video frame among the adjacent video frames as a kind of smoothing sub-strategy.
[0532] In an illustrative example, the label smoothing strategy can be a pre-set strategy corresponding to the video to which the multiple video frames belong. Exemplarily, a user or a business person (such as a business person in the video online beautification service) can input video processing requirements for a certain type of video, so that the electronic device can generate a label smoothing strategy according to the video processing requirements. For example, for a video for subway promotion, relevant personnel can input "subways are continuous, subways coexist with people, and subways do not coexist with airplanes". Thus, the set label smoothing strategy can include that subway category labels are continuous between adjacent video frames, subway labels and human labels coexist in the same video frame, and subway labels and airplane labels do not coexist in the same video frame. Referring to the foregoing method, relevant personnel can input different video processing requirements for different types of videos to enable the electronic device to generate different label smoothing strategies corresponding to different types of videos. When label smoothing processing needs to be performed on a certain video, the video category to which the video belongs can be determined according to the semantics of the video, and then the video can be subjected to label smoothing processing according to the label smoothing strategy corresponding to the video category to which the video belongs.
[0533] Among them, in one example, the semantics of the video can be obtained by analyzing the video through a three-dimensional convolutional neural network (3D CNN).
[0534] Next, an example is introduced to describe the process of performing label smoothing processing on the multiple video frames according to the label smoothing strategy.
[0535] In an illustrative example, it can be set that the label smoothing strategy includes that the class label C1 is continuous between adjacent video frames among the multiple video frames. Then, performing label smoothing processing on the multiple video frames can include: determining whether the video frames among the multiple video frames carry the class label C1, and determining whether two video frames in each pair of adjacent video frame groups carry or do not carry the class label C1 simultaneously. If two video frames in a pair of adjacent video frame groups carry or do not carry the class label C1 simultaneously, a higher score W is assigned to this pair of adjacent video frame groups; if one of the two video frames in a pair of adjacent video frame groups carries the class label C1 and the other does not carry the class label C1, a lower score W' is assigned to this pair of adjacent video frame groups. For convenience of description, the score W and the score W' here can be referred to as the scores of the pair of adjacent video frames under the class label C1.
[0536] Among them, a pair of video frame groups refers to a combination composed of two adjacent video frames among the multiple video frames. For example, it can be set that the multiple video frames include video frame A1, video frame A2,..., video frame An, where video frame A1, video frame A2,..., video frame An are adjacent in sequence among the multiple video frames. Then, video frame A1 and video frame A2 form a pair of video frame groups, video frame A2 and video frame A3 form a pair of video frame groups,..., and video frame An - 1 and video frame An form a pair of video frame groups.
[0537] It can be targeted at maximizing the sum of the scores of each pair of adjacent video frame groups among the multiple video frames under the class label C1, and adding or deleting the class label C1 of K video frames among the multiple video frames. Or rather, by adding or deleting the class label C1 of K video frames among the multiple video frames, the sum of the scores of each pair of adjacent video frame groups among the multiple video frames under the class label C1 is maximized. The specific meaning of adding or deleting the class label of the video frame can refer to the above introduction and will not be elaborated here. K ≥ 0 and is an integer.
[0538] Among them, which specific K video frames' class label C1 is added or deleted can be determined in the following way.
[0539] In an illustrative example, it can be attempted to add or delete the class label C1 of any one or more video frames, and then calculate the sum of the scores of each pair of adjacent video frame groups among the multiple video frames under the class label C1. The sum of the scores can also be referred to as the total score. By continuously attempting, multiple total scores can be obtained. The video frames with the class label C1 added or deleted corresponding to the highest score among the multiple total scores are determined as the above-mentioned K video frames.
[0540] In an illustrative example, it can be set that X i,jIndicates the situation where the j-th video frame among the multiple video frames carries the class label i (specifically divided into two situations: carrying and not carrying). It is also possible to set X 0,0 to indicate carrying, -X 0,0 to indicate not carrying. That is, when X i,j = X 0,0 , it indicates that the j-th video frame carries the class label i; when X i,j = -X 0,0 , it indicates that the j-th video frame does not carry the class label i.
[0541] For a group of two adjacent video frames under the class label C1
[0542] When the label smoothing strategy includes the sub-strategy that the class label C1 is continuous between adjacent video frames among the multiple video frames, formula (1) can be set to calculate the sum of the scores of each group of two adjacent video frames under the class label C1 among the multiple video frames.
[0543] Max(N 1 ) (1)
[0544] Where
[0545] w1 is a preset positive number. The absolute value of X 0,0 is 1. j corresponds to the j-th video frame among the multiple video frames, j + 1 corresponds to the j + 1-th video frame among the multiple video frames, and i corresponds to the class label C1. That is, X i,j indicates whether the j-th video frame among the multiple video frames carries the class label C1. If X i,j = X 0,0 , it indicates that the j-th video frame carries the class label C1; if X i,j = -X 0,0 , it indicates that the j-th video frame does not carry the class label C1. n is the number of video frames among the multiple video frames; during the solution process, X i,j and X i,j+1 take values between X 0,0 and -X 0,0 .
[0546] By calculating using formula (1), under the condition that the value of formula (1) is the largest, the situation of each video frame among the multiple video frames carrying the class label C1 can be obtained. Thus, the smoothing result is obtained. This smoothing result is the smoothing result when the label smoothing strategy is that the class label C1 is continuous between adjacent video frames among the multiple video frames.
[0547] In an illustrative example, it can be understood that all video frames among multiple video frames do not carry the category label C1, or all video frames carry the category label C1, and the score of each pair of adjacent video frame groups among the multiple video frames is the highest under the category label C1. The video frame label smoothing scheme provided by the embodiments of the present application is used to eliminate the non-smoothness (not conforming to common sense) of the category labels of some video frames in the video caused by factors such as camera jitter, video frame or video semantic recognition model recognition errors. It can be understood that generally, camera jitter, video frames or video semantic recognition models do not occur commonly. Correspondingly, the proportion of video frames with non-smooth labels in the entire video is relatively small. To avoid the imbalance of the category labels of video frames with a relatively small proportion, and adding or deleting the category labels of most video frames in the video, in this embodiment, on the basis of including one or more of the above-mentioned smoothing sub-strategies, the label smoothing strategy may further include minimizing the number of video frames with added or deleted category labels. That is to say, before and after the label smoothing operation, the number of video frames with added or deleted category labels among the multiple video frames is as small as possible. Or, it can be said that the situation where the multiple video frames carry category labels before the label smoothing operation (when obtaining the multiple video frames in step 901) can be referred to as the input, and the processing result of step 902 can be referred to as the output. Then, the smoothing sub-strategy of minimizing the number of video frames with added or deleted category labels can also be referred to as making the input and output as consistent as possible.
[0548] In an illustrative example, the smoothing sub-strategy of minimizing the number of video frames with added or deleted category labels can be characterized by formula (2).
[0549] Max(N 2 +N 3 ) (2)
[0550] Wherein,
[0551]
[0552] w2 is a preset positive number; set C1 ∪ set C2 = the multiple video frames, the elements in set C1 are the video frames that carry the first category label when obtaining the multiple video frames; the elements in set C2 are the video frames that do not carry the first category label when obtaining the multiple video frames; in the solving process, it takes values between X i,j In X 0,0 and -X 0,0 The values of X i,j , X 0,0 , -X 0,0 can be referred to the above introduction and will not be elaborated here.
[0553] Thus, when the label smoothing strategy includes two sub-strategies: the class label C1 is continuous between adjacent video frames among the multiple video frames and the number of video frames with class label addition or deletion is minimized, formula (3) can be set up to calculate the sum of the scores of each pair of adjacent video frame groups among the multiple video frames under the class label C1.
[0554] Max(N 1 +N 2 +N 3 ) (3).
[0555] In an illustrative example, it can be understood that 3 = 1 + 1 + 1. Thus, formula (3) can be represented in the form of formula (3').
[0556]
[0557] The form of formula (3') is more general. In actual use, the meaning of X ij and X pq can be set according to specific situations. For example, X ij X pq in formula (3') can represent X 0,0 X 0,0 mentioned above. At this time, X ij = X 00 , X pq = X 00 . Again, for example, X ij X pq in formula (3') can represent X i,j X i,j+1 mentioned above. At this time, X ij = X i,j , X pq = X i,j+1 . And so on. Here, no more examples will be listed one by one.
[0558] In the process of solving formula (3'), X ij takes values between 1 and -1, and X pq also takes values between 1 and -1. w ij,pq is a preset positive number.
[0559] In addition, it can be understood that 1 is a constant value and it will not affect the maximum sum solving process of formula (3'). Therefore, formula (3') can be simplified to formula (3").
[0560]
[0561] Thus, the maximum sum solving problem of formula (3) can be converted into the maximum sum solving problem of formula (3").
[0562] There may be an NP-hard problem in the computer field for formula (3”), making it impossible to quickly solve formula (3) or rather formula (3’). To improve the efficiency of smooth processing of video frame labels in a video, in this embodiment, it can be represented by a semi-definite matrix, and the maximum sum solution problem of formula (3’) is converted into a semi-definite programming problem to achieve fast solution. Among them, semi-definite programming is a mathematical optimization term, which refers to a type of mathematical programming with a semi-definite matrix as the feasible region, linear constraints, and a linear objective function. The specific solution is as follows.
[0563] The column vector a can be set as a = [X 0,0 ; X i,1 ;...; X i,j ;...; X i,n . Among them, i corresponds to the category label C1, and j corresponds to the j-th video frame among n video frames (the above-mentioned multiple video frames). The column vector a has n + 1 columns.
[0564] Multiply the column vector a by the transpose vector of the column vector a (a T ), to obtain a matrix Y. And X ij X pq in formula (3”) is an element in the matrix Y. Thus, formula (3”) can be represented or converted into formula (4).
[0565]
[0566] Y is a semi-definite matrix; (4)
[0567] Rank(Y) = 1;
[0568] Y(i, i) = 1.
[0569] That is to say, the maximum sum solution problem of formula (3”) is converted into a problem of semi-definite programming with the objective of maximizing the sum result of formula (3”) (or rather formula (3)), rank(Y) = 1, and the value of the element on the main diagonal of the matrix Y being 1 as the constraints.
[0570] Among them, rank(Y) = 1 is a non-convex condition, making the process of semi-definite programming with the objective of maximizing the sum result of formula (3”) (or rather formula (3)), rank(Y) = 1, and the value of the element on the main diagonal of the matrix Y being 1 as the constraints cannot be quickly solved.
[0571] Exemplarily, the problem of semi - definite optimization with the objective of maximizing the summation result of formula (3”) (or formula (3)), rank(Y) = 1, and the elements on the main diagonal of matrix Y being 1 as constraints can be relaxed, removing or ignoring the constraint of rank(Y) = 1. Here, relaxation is a term in mathematical optimization, which means removing part of the constraints of a problem to reduce its difficulty and obtain an approximate solution.
[0572] Thus, formula (4) can be expressed or transformed into formula (4’).
[0573]
[0574] Y is a semi - definite matrix; (4’)
[0575] Y(1, 1) = 1.
[0576] That is to say, the problem of maximizing the sum of formula (3”) (or formula (3)) can be transformed into a semi - definite optimization problem with the objective of maximizing the summation result of formula (3”) (or formula (3)) and the elements on the main diagonal of matrix Y being 1 as constraints. This problem can be solved using the interior point method or the ellipsoid method. Both the ellipsoid method and the interior point method are terms in mathematical optimization and are both mathematical optimization algorithms that can solve semi - definite optimization problems. For specific introductions to the interior point method or the ellipsoid method in the prior art, please refer to the relevant content and will not be elaborated here.
[0577] Thus, a relaxed solution S can be obtained. S is a semi - definite matrix and the elements on the main diagonal are 1. The relaxed solution can also be understood as an approximate solution. Since the condition of rank(Y) = 1 is removed or ignored during the calculation process, the calculation result obtained is an approximate solution. It is necessary to recover the relaxed solution S to obtain the true answer. Recovery is a term in mathematical optimization, which refers to the process of converting the optimal solution of the relaxed problem into a feasible solution of the original problem after obtaining the optimal solution of the relaxed problem. The specific recovery process is as follows.
[0578] The semi - definite matrix S can be decomposed by eigenvalues to obtain S = Y′·Y T . Where Y′ is a matrix with n + 1 columns, and n is the number of video frames in the above - mentioned multiple video frames. The columns in Y′ correspond one - to - one with the columns in the above - mentioned column vector a.
[0579] A unit vector r can be randomly selected. If r TMultiply the first column of Y’ and if the product ≥ 0, then set the first column of Y’ to correspond to 1. Correspondingly, set the first column of column vector a, i.e., X 00 = 1. If r T Multiply the first column of Y’ and if the product < 0, then set the first column of Y’ to correspond to 1. Correspondingly, set the first column of column vector a, i.e., X 00 = -1.
[0580] r T Multiply the i-th column of Y’ and then multiply by X 0,0 , to obtain the product result. If the product result ≥ 0, then the i-th column in column vector a is 1; if the product result < 0, then the i-th column in column vector a is -1. i = 2, 3...... Thus, the values of each element in column vector a can be obtained, and combined with the value of X 00 , add or delete the class label C1 of the video frames in multiple video frames, so that the situation of the video frames carrying the class label C1 in multiple video frames is consistent with the elements in the column vector. Specifically, if in the column vector, X 0,0 = 1, X i,j = 1, then the j-th video frame carries the class label C1 (if the j-th video frame originally carried the class label C1, no deletion is made; if the j-th video frame did not originally carry the class label C1, then add the class label C1 to the j-th video frame). If in the column vector, X 0,0 = -1, X i,j = -1, then the j-th video frame carries the class label C1 (if the j-th video frame originally carried the class label C1, no deletion is made; if the j-th video frame did not originally carry the class label C1, then add the class label C1 to the j-th video frame). That is to say, when X i,j = X 0,0 in the column vector, then in the label smoothing processing result, the j-th video frame carries the class label C1. If X i,j = -X 0,0 in the column vector, then in the label smoothing processing result, the j-th video frame does not carry the class label C1.
[0581] The above takes the example that the label smoothing strategy includes the continuity of the class label C1 between adjacent video frames in the multiple video frames, and takes the example of the two sub-strategies that the label smoothing strategy includes the continuity of the class label C1 between adjacent video frames in the multiple video frames and minimizing the number of video frames for adding or deleting class labels, to introduce the process of label smoothing processing for multiple video frames according to the label smoothing strategy.
[0582] In an illustrative example, it can be set that the label smoothing strategy includes the coexistence of class label C2 and class label C3 on the same video frame among multiple video frames. Then, the label smoothing process for the multiple video frames can include: It can be determined whether any one of the multiple video frames, for example, video frame A1, carries both class label C2 and class label C3. When video frame A1 carries both or does not carry both class label C2 and class label C3, a higher score F is assigned to video frame A1. When video frame A1 carries class label C2 but does not carry class label C3, a lower score F' (F' is less than F) is assigned to video frame A1; when video frame A1 does not carry class label C2 but carries class label C3, a lower score F' is also assigned to video frame A1. For ease of description, the score F and the score F' here can be referred to as the scores of the video frame under class label C2 and class label C3.
[0583] It can be targeted at maximizing the sum of the scores of each video frame in the multiple video frames under class label C2 and class label C3, and add class label C3 to L video frames (the L video frames did not originally carry class label C3) in the multiple video frames and / or delete the class label C2 of L' video frames (the L' video frames did not originally carry class label C2) in the multiple video frames. L≥0 and is an integer; L'≥0 and is an integer. In other words, it can be targeted at maximizing the sum of the scores of each video frame in the multiple video frames under class label C2 and class label C3, and add or delete the class label C2 or class label C3 of L'' video frames in the multiple video frames. Or rather, by adding or deleting the class label C2 or class label C3 of L'' video frames in the multiple video frames, the sum of the scores of each video frame in the multiple video frames under class label C2 and class label C3 is maximized. L''≥0 and is an integer. The specific meaning of adding or deleting the class label of the video frame can refer to the above introduction and will not be elaborated here.
[0584] Among them, which specific L'' video frames' class label C2 or class label C3 to add or delete can be determined by the following method.
[0585] In an illustrative example, it can be attempted to add or delete the class label C2 or class label C3 of any one or more video frames, and then, calculate the sum of the scores of each video frame in the multiple video frames under class label C2 and class label C3. The sum of the scores can be called the total score. By continuously attempting, multiple total scores can be obtained. The video frames whose class label C2 or class label C3 has been added or deleted corresponding to the highest total score among the multiple total scores are determined as the above L'' video frames.
[0586] In an illustrative example, X can be set i*,jIndicates the situation where the j-th video frame among the multiple video frames carries the class label i* (specifically divided into two situations: carrying and not carrying). X can also be set 0,0 Indicates carrying, -X 0,0 Indicates not carrying. That is, when X i*,j = X 0,0 it indicates that the j-th video frame carries the class label i*; when X i*,j = -X 0,0 it indicates that the j-th video frame does not carry the class label i*. X can be set i,j Indicates the situation where the j-th video frame among the multiple video frames carries the class label i (specifically divided into two situations: carrying and not carrying). When X i,j = X 0,0 it indicates that the j-th video frame carries the class label i; when X i,j = -X 0,0 it indicates that the j-th video frame does not carry the class label i. Among them, the class label i* is the class label C2, and the class label i is the class label C3.
[0587] A video frame is in
[0588] When the label smoothing strategy includes the sub-strategy that the class labels C2 and C3 coexist on the same video frame among multiple video frames, formula (5) can be set to calculate the maximum sum of the scores of each video frame among the multiple video frames under the class labels C2 and C3.
[0589] Max(N 4 ) (5)
[0590] Among them,
[0591] Among them, n is the number of video frames among the multiple video frames; w 3 is a preset positive number; the absolute value of X 0,0 is 1; i* corresponds to the class label C2, i corresponds to the class label C3; j corresponds to the j-th video frame among the multiple video frames; when the j-th video frame carries the class label C2, in X i*,j = X 0,0 ; when the j-th video frame does not carry the class label C2, in X i*,j = -X 0,0 ; when the j-th video frame carries the class label C3, in X i,j = X 0,0 ; when the j-th video frame does not carry the class label C3, X i,j = -X 0,0 .
[0592] By performing calculations using formula (5), it is possible to obtain, under the condition that the value of formula (5) is maximized, the situations of each video frame among multiple video frames carrying class label C2 and carrying class label C3. Thus, a smoothing result is obtained. This smoothing result is the smoothing result of the label smoothing strategy for the coexistence of class label C2 and class label C3 on the same video frame among the multiple video frames.
[0593] In an illustrative example, the label smoothing strategy includes the coexistence of class label C2 and class label C3 on the same video frame among the multiple video frames, and also includes minimizing the number of video frames with added or deleted class labels. Formula (6) can be set to calculate the sum of the scores of each video frame among the multiple video frames under class label C2 and class label C3.
[0594] Max(N 4 +N 2 +N 3 )(6).
[0595] Wherein, N 4 , N 2 , N 3 can be referred to the above introduction and will not be elaborated here.
[0596] In some embodiments, formula (6) can also be converted into the forms of the above formula (3') and formula (3"). Similarly, a vector a = [X 0,0 ; X i*,1 ;...; X i*,j ;...; X i*,n ; X i,1 ;...; X i,j ;...; X i,n can be set. Wherein, i* corresponds to class label C2, i corresponds to class label C3, and j corresponds to the j-th video frame among n video frames (the above multiple video frames). The column vector a has 2n + 1 columns. Similarly, the maximum sum solution problem of formula (6) can be converted into the semi-definite matrix optimization problems shown in the above formula (4) and formula (4'), and under the corresponding constraint conditions, the interior point method or the ellipsoid method can be used to solve it. Then, the obtained relaxed answer can be restored to obtain the label smoothing result, and based on the label smoothing result, the class label C2 or class label C3 of the video frames among the multiple video frames can be added or deleted. Specifically, it can be referred to the above introduction and will not be elaborated here.
[0597] The above takes the case where the label smoothing strategy includes the co - existence of class label C2 and class label C3 on the same video frame in multiple video frames, and the case where the label smoothing strategy includes the co - existence of class label C2 and class label C3 on the same video frame in multiple video frames and minimizing the number of video frames with class labels added or deleted as two sub - strategies, and gives an example of the process of performing label smoothing on multiple video frames according to the label smoothing strategy.
[0598] In an illustrative example, it can be set that the label smoothing strategy includes that class label C4 and class label C5 do not co - exist on the same video frame in multiple video frames. Then, performing label smoothing on multiple video frames can include: determining whether any one of the multiple video frames, such as video frame A1, carries both class label C4 and class label C5 at the same time. When video frame A1 does not carry class label C4 and class label C5 at the same time; or when video frame A1 carries only one of class label C4 and class label C5 and does not carry the other, a higher score H is assigned to video frame A1. When video frame A1 carries both class label C4 and class label C5 at the same time, a lower score H' (H' is less than H) is assigned to video frame A1. For the convenience of description, the score H and score H' here can be referred to as the scores of the video frame under class label C4 and class label C5.
[0599] It can be aimed at maximizing the sum of the scores of each video frame in the multiple video frames under class label C4 and class label C5, and deleting the class label C5 of P video frames (the P video frames originally carried class label C5) in the multiple video frames and / or deleting the class label C4 of P' video frames (the P' video frames originally carried class label C4) in the multiple video frames; P≥0 and is an integer; P'≥0 and is an integer.
[0600] In other words, it can be aimed at maximizing the sum of the scores of each video frame in the multiple video frames under class label C4 and class label C5, and adding or deleting the class label C4 or class label C5 of P'' video frames in the multiple video frames. Or rather, by adding or deleting the class label C4 or class label C5 of P'' video frames in the multiple video frames, the sum of the scores of each video frame in the multiple video frames under class label C4 and class label C5 is maximized. P''≥0 and is an integer. The specific meaning of adding or deleting the class label of the video frame can refer to the above introduction and will not be elaborated here.
[0601] Among them, which specific P'' video frames' class label C4 or class label C5 to add or delete can be determined by the following method.
[0602] In an illustrative example, it is possible to attempt to add or delete the class label C4 or the class label C5 of any one or more video frames, and then calculate the sum of the scores of each video frame among the multiple video frames under the class label C4 and the class label C5. The sum of the scores can be referred to as the total score. By continuously attempting, multiple total scores can be obtained. The video frame with the highest total score among the multiple total scores, for which the class label C4 or the class label C5 has been added or deleted, is determined as the above-mentioned "P" video frames.
[0603] In an illustrative example, X can be set i*,j to represent the situation where the j-th video frame among the multiple video frames carries the class label i* (specifically divided into two situations: carrying and not carrying). X can also be set 0,0 to represent carrying, and -X 0,0 to represent not carrying. That is to say, when X i*,j = X 0,0 , it means that the j-th video frame carries the class label i*; when X i*,j = -X 0,0 , it means that the j-th video frame does not carry the class label i*. X can be set i,j to represent the situation where the j-th video frame among the multiple video frames carries the class label i (specifically divided into two situations: carrying and not carrying). When X i,j = X 0,0 , it means that the j-th video frame carries the class label i; when X i,j = -X 0,0 , it means that the j-th video frame does not carry the class label i. Among them, the class label i* is the class label C4, and the class label i is the class label C5.
[0604] A video frame in
[0605] When the label smoothing strategy includes the sub-strategy that the class label C4 and the class label C5 do not coexist on the same video frame among multiple video frames, formula (7) can be set to calculate the scores of each video frame among the multiple video frames under the class label C4 and the class label C5.
[0606] Max(N 5 ) (7)
[0607] Among them,
[0608] Among them, n is the number of video frames among the multiple video frames; w 4 is a preset positive number; the absolute value of X 0,0 is 1; i* corresponds to the class label C4, i corresponds to the class label C5; j corresponds to the j-th video frame among the multiple video frames; when the j-th video frame carries the class label C4, in Xi*,j = X 0,0 ; When the j-th video frame does not carry the class label C4, at X i*,j = -X 0,0 ; When the j-th video frame carries the class label C5, at X i,j = X 0,0 ; When the j-th video frame does not carry the class label C5, X i,j = -X 0,0 .
[0609] By calculating using formula (7), it is possible to obtain, under the condition that the value of formula (7) is the largest, the situations of each video frame among multiple video frames carrying the class label C4 and carrying the class label C5. Thus, a smoothing result is obtained. This smoothing result is a smoothing result where the label smoothing strategy is that the class label C4 and the class label C5 do not coexist on the same video frame among these multiple video frames.
[0610] In an illustrative example, the label smoothing strategy includes that the class label C4 and the class label C5 do not coexist on the same video frame among these multiple video frames, and also includes minimizing the number of video frames with class labels added or deleted. Formula (8) can be set to calculate the sum of the scores of each video frame among multiple video frames under the class label C4 and the class label C5.
[0611] Max(N 5 + N 2 + N 3 ) (8).
[0612] Among them, N 5 , N 2 , N 3 can be referred to the above introduction and will not be elaborated here.
[0613] In some embodiments, formula (8) can also be converted into the forms of the above formula (3') and formula (3"). Similarly, a vector a = [X 0,0 ; X i*,1 ;...; X i*,j ;...; X i*,n ; X i,1 ;...; X i,j ;...; X i,nAmong them, i* corresponds to the category label C4, i corresponds to the category label C5, and j corresponds to the j-th video frame among n video frames (the above-mentioned multiple video frames). The column vector a has 2n + 1 columns. Similarly, the maximum sum solution problem of formula (8) can be transformed into the semi-definite matrix optimization problem shown in formula (4) and formula (4’), and under the corresponding constraints, the interior point method or the ellipsoid method can be used to solve it. Then, the obtained relaxed answer can be restored to obtain the label smoothing result, and based on the label smoothing result, the category label C4 or the category label C5 of the video frames in the multiple video frames can be added or deleted. For specific reference, please refer to the above introduction and will not be elaborated here.
[0614] The above takes the example that the label smoothing strategy includes that the category labels C4 and C5 do not coexist on the same video frame among multiple video frames, and takes the two sub-strategies that the label smoothing strategy includes that the category labels C4 and C5 do not coexist on the same video frame among multiple video frames and the number of video frames for adding or deleting category labels is minimized as examples to introduce the process of label smoothing for multiple video frames according to the label smoothing strategy.
[0615] In an illustrative example, it can be set that the label smoothing strategy includes that the category label C6 exists in the previous video frame among adjacent video frames, and the category label C7 exists in the subsequent video frame among the adjacent video frames. Then, the label smoothing for multiple video frames can include: it can be determined whether the previous video frame in each pair of adjacent video frame groups among the multiple video frames carries the category label C6, and whether the subsequent video frame carries the category label C7. The meaning of each pair of adjacent video frame groups can be referred to the above introduction and will not be elaborated here. For a pair of adjacent video frame groups, if its previous video frame carries the category label C6 and its subsequent video frame carries the category label C7, a higher score Z is assigned to this pair of adjacent video frame groups. Or, if its previous video frame does not carry the category label C6 and its subsequent video frame does not carry the category label C7, a higher score Z is assigned to this pair of adjacent video frame groups. If its previous video frame carries the category label C6 and its subsequent video frame does not carry the category label C7, a lower score Z’ (Z’ is less than Z) is assigned to this pair of adjacent video frame groups. Or, if its previous video frame does not carry the category label C6 and its subsequent video frame carries the category label C7, a higher score Z’ is assigned to this pair of adjacent video frame groups. For ease of description, the scores Z and Z’ here can be referred to as the scores of the video frames under the category labels C6 and C7.
[0616] The maximization of the scores of the multiple video frames under class label C6 and class label C7 can be used as the objective to add class label C7 to Q video frames among the multiple video frames (the Q video frames originally do not carry class label C7) and / or delete class label C6 of Q' video frames among the multiple video frames (the Q' video frames originally carry class label C6); Q≥0 and is an integer; Q'≥0 and is an integer.
[0617] In other words, the maximization of the scores of the multiple video frames under class label C6 and class label C7 can be used as the objective to add or delete class label C6 or class label C7 of Q'' video frames among the multiple video frames. Or rather, by adding or deleting class label C6 or class label C7 of Q'' video frames among the multiple video frames, the sum of the scores of each pair of adjacent video frame groups among the multiple video frames under class label C6 and class label C7 is maximized. Q''≥0 and is an integer. The specific meaning of adding or deleting the class labels of video frames can refer to the above introduction and will not be elaborated here.
[0618] Among them, which specific Q'' video frames' class label C6 or class label C7 to add or delete can be determined by the following method.
[0619] In an illustrative example, it is possible to attempt to add or delete class label C6 or class label C7 of any one or more video frames, and then calculate the sum of the scores of each pair of adjacent video frame groups among the multiple video frames under class label C6 and class label C7. The sum of the scores can be referred to as the total score. By continuously attempting, multiple total scores can be obtained. The video frames whose class label C6 or class label C7 has been added or deleted corresponding to the highest total score among the multiple total scores are determined as the above-mentioned Q'' video frames.
[0620] In an illustrative example, X can be set i*,j to represent the situation where the j-th video frame among the multiple video frames carries class label i* (specifically divided into two situations: carrying and not carrying). It is also possible to set X 0,0 to represent carrying, and -X 0,0 to represent not carrying. That is to say, when X i*,j =X 0,0 , it means that the j-th video frame carries class label i*; when X i*,j =-X 0,0 , it means that the j-th video frame does not carry class label i*. It is possible to set X i,j+1 to represent the situation where the (j + 1)-th video frame among the multiple video frames carries class label i (specifically divided into two situations: carrying and not carrying). When X i,j+1 =X 0,0 , it means that the (j + 1)-th video frame carries class label i; when X i,j+1 =-X0,0 When it indicates that the (j + 1)-th video frame does not carry the class label i. Among them, the class label i* is the class label C6, and the class label i is the class label C7.
[0621] The score of a pair of adjacent video frame groups under the class labels C6 and C7 =
[0622] When the label smoothing strategy includes the existence of the class label C6 in the previous video frame among adjacent video frames and the existence of the class label C7 in the subsequent video frame among the adjacent video frames, the formula (9) can be set to calculate the maximum sum of the scores of pairs of adjacent video frame groups among multiple video frames under the class labels C6 and C7.
[0623] Max(N 6 ) (9)
[0624] Among them,
[0625] Among them, n is the number of video frames among multiple video frames; w 5 is a preset positive number; the absolute value of X 0,0 is 1; i* corresponds to the class label C6, i corresponds to the class label C7; j corresponds to the j-th video frame among multiple video frames; when the j-th video frame carries the class label C6, in X i*,j = X 0,0 ; when the j-th video frame does not carry the class label C6, in X i*,j = -X 0,0 ; when the (j + 1)-th video frame carries the class label C7, in X i,j+1 = X 0,0 ; when the (j + 1)-th video frame does not carry the class label C7, X i,j+1 = -X 0,0 . The j-th video frame is the previous video frame in the pair of adjacent video frame groups, and the (j + 1)-th video frame is the subsequent video frame in the pair of adjacent video frame groups.
[0626] By using the formula (9) for calculation, under the condition that the value of the formula (9) is the largest, the situations of each video frame among multiple video frames carrying the class label C6 and carrying the class label C7 can be obtained. Thus, the smoothing result is obtained. This smoothing result is the smoothing result when the label smoothing strategy is that the class label C6 exists in the previous video frame among adjacent video frames and the class label C7 exists in the subsequent video frame among the adjacent video frames.
[0627] In an illustrative example, the label smoothing strategy includes that the category label C6 exists in the previous video frame among adjacent video frames, and the category label C7 exists in the subsequent video frame among the adjacent video frames, and the label smoothing strategy further includes minimizing the number of video frames with category labels added or deleted. Formula (10) can be set to calculate the sum of scores of each pair of video frame groups among multiple video frames under the category label C6 and the category label C7.
[0628] Max(N 6 +N 2 +N 3 ) (6).
[0629] Wherein, N 6 、N 2 、N 3 can be referred to the above introduction and will not be elaborated here.
[0630] In an illustrative example, formula (10) can also be converted into the forms of the above formula (3) and formula (3”). Similarly, the vector a = [X 0,0 ; X i*,1 ;...; X i*,j ;...; X i*,n ; X i,1 ;...; X i,j ;...; X i,n can be set. Wherein, i* corresponds to the category label C6, i corresponds to the category label C7, and j corresponds to the j-th video frame among n video frames (the above multiple video frames). The column vector a has 2n + 1 columns. Similarly, the maximum sum solving problem of formula (10) can be converted into the semi-definite matrix optimization problems shown in the above formula (4) and formula (4’), and under the corresponding constraints, the interior point method or the ellipsoid method can be used to solve it. Then, the obtained relaxed answer can be restored to obtain the label smoothing processing result, and according to the label smoothing processing result, the category label C6 or the category label C7 of the video frames among multiple video frames can be added or deleted. Specifically, it can be referred to the above introduction and will not be elaborated here.
[0631] The above takes the label smoothing strategy including that the category label C6 exists in the previous video frame among adjacent video frames and the category label C7 exists in the subsequent video frame among the adjacent video frames as an example to introduce the process of performing label smoothing processing on multiple video frames according to the label smoothing strategy.
[0632] In an illustrative example, it can be set that the label smoothing strategy includes that the class label C8 exists in the previous video frame among adjacent video frames, and the class label C9 does not exist in the subsequent video frame among the adjacent video frames. Then, performing label smoothing processing on multiple video frames can include: determining whether the previous video frame in each pair of adjacent video frame groups among the multiple video frames carries the class label C8, and whether the subsequent video frame carries the class label C9. If the previous video frame in a pair of adjacent video frame groups carries the class label C8 and the subsequent video frame does not carry the class label C9, a higher score D is assigned to this pair of adjacent video frame groups. Or, if the previous video frame in this pair of adjacent video frame groups does not carry the class label C8 and the subsequent video frame carries the class label C9, a higher score D is assigned to this pair of adjacent video frame groups. Or, if the previous video frame in this pair of adjacent video frame groups does not carry the class label C8 and the subsequent video frame also does not carry the class label C9, a higher score D is assigned to this pair of adjacent video frames. If the previous video frame in this pair of adjacent video frame groups carries the class label C8 and the subsequent video frame carries the class label C9, a lower score D' (D' is less than D) is assigned to this pair of adjacent video frames. For ease of description, the score D and the score D' here can be referred to as the scores of the pair of adjacent video frame groups under the class label C8 and the class label C9.
[0633] It can be targeted at maximizing the sum of the scores of each video frame in the multiple video frames under the class label C8 and the class label C9, and deleting the class label C9 of V video frames (the V video frames originally carried the class label C9) in the multiple video frames and / or deleting the class label C8 of V' video frames (the V' video frames originally carried the class label C8) in the multiple video frames; V≥0 and is an integer; V'≥0 and is an integer.
[0634] In other words, it can be targeted at maximizing the sum of the scores of each pair of adjacent video frame groups in the multiple video frames under the class label C8 and the class label C9, and adding or deleting the class label C8 or the class label C9 of V'' video frames in the multiple video frames. Or rather, by adding or deleting the class label C8 or the class label C9 of V'' video frames in the multiple video frames, the sum of the scores of each pair of adjacent video frame groups in the multiple video frames under the class label C8 and the class label C9 is maximized. V''≥0 and is an integer. The specific meaning of adding or deleting the class label of the video frame, and the specific meaning of the pair of adjacent video frame groups can refer to the above introduction and will not be elaborated here.
[0635] Among them, specifically which V'' video frames' class label C8 or class label C9 to add or delete can be determined by the following method.
[0636] In an illustrative example, it is possible to attempt to add or delete the class label C9 or the class label C8 of any one or more video frames. Then, calculate the sum of scores of each pair of adjacent video frame groups among the multiple video frames under the class label C8 and the class label C9. The sum of scores can be referred to as the total score. By continuously attempting, multiple total scores of the multiple video frames under the class label C8 and the class label C9 can be obtained. Determine the video frames with the class label C8 or the class label C9 added or deleted corresponding to the highest total score among the multiple total scores as the above-mentioned "V" video frames.
[0637] In an illustrative example, X can be set i*,j to represent the situation where the j-th video frame among the multiple video frames carries the class label i* (specifically divided into two situations: carrying and not carrying). X can also be set 0,0 to represent carrying, and -X 0,0 to represent not carrying. That is to say, when X i*,j = X 0,0 , it means that the j-th video frame carries the class label i*; when X i*,j = -X 0,0 , it means that the j-th video frame does not carry the class label i*. X can be set i,j+1 to represent the situation where the (j + 1)-th video frame among the multiple video frames carries the class label i (specifically divided into two situations: carrying and not carrying). When X i,j+1 = X 0,0 , it means that the (j + 1)-th video frame carries the class label i; when X i,j+1 = -X 0,0 , it means that the (j + 1)-th video frame does not carry the class label i. Among them, the class label i* is the class label C8, and the class label i is the class label C9.
[0638] The score of a pair of video frames under the class label C8 and the class label C9 =
[0639] When the label smoothing strategy includes the sub-strategy that the class label C8 exists in the previous video frame among adjacent video frames and the class label C9 does not exist in the subsequent video frame among the adjacent video frames, formula (11) can be set to calculate the maximum sum of scores of each pair of video frame groups among the multiple video frames under the class label C8 and the class label C9.
[0640] Max(N 7 ) (11)
[0641] Among them,
[0642] Among them, n is the number of video frames among the multiple video frames; w 6 is a preset positive number; X0,0 The absolute value of is 1; i* corresponds to the class label C8, and i corresponds to the class label C9; j corresponds to the j-th video frame among multiple video frames; when the j-th video frame carries the class label C8, at X i*,j = X 0,0 ; when the j-th video frame does not carry the class label C8, at X i*,j = -X 0,0 ; when the (j + 1)-th video frame carries the class label C9, at X i,j+1 = X 0,0 ; when the (j + 1)-th video frame does not carry the class label C5, X i,j+1 = -X 0,0 . The j-th video frame is the previous video frame in a pair of adjacent video frames, and the (j + 1)-th video frame is the subsequent video frame in this pair of adjacent video frames.
[0643] By calculating using formula (11), under the condition that the value of formula (11) is the largest, the situations where each video frame among multiple video frames carries the class label C8 and the class label C9 can be obtained. Thus, a smoothing result is obtained. This smoothing result is a smoothing result where the label smoothing strategy is that the class label C8 exists in the previous video frame among adjacent video frames, and the class label C9 does not exist in the subsequent video frame among these adjacent video frames.
[0644] In an illustrative example, the label smoothing strategy includes that the class label C8 exists in the previous video frame among adjacent video frames, the class label C9 does not exist in the subsequent video frame among these adjacent video frames, and also includes minimizing the number of video frames with class labels added or deleted. Formula (12) can be set to calculate the sum of the scores of each video frame among multiple video frames under the class label C4 and the class label C5.
[0645] Max(N 7 + N 2 + N 3 ) (12).
[0646] Among them, N 5 , N 2 , N 3 can be referred to the above introduction and will not be elaborated here.
[0647] In some embodiments, formula (12) can also be converted into the forms of the above formula (3') and formula (3"). Similarly, the vector a = [X 0,0 ; X i*,1 ;...; X i*,j ;...; X i*,n ; X i,1 ;...; X i,j ;...; Xi,n .
[0648] Among them, i* corresponds to the category label C8, i corresponds to the category label C9, and j corresponds to the j-th video frame among n video frames (the above-mentioned multiple video frames). The column vector a has 2n + 1 columns. Similarly, the maximum sum solution problem of formula (12) can be transformed into the semi-definite matrix optimization problem shown in formula (4) and formula (4’), and under the corresponding constraints, the interior point method or the ellipse method can be used to solve it. Then, the obtained relaxed answer can be restored to obtain the label smoothing result, and according to the label smoothing result, the category label C8 or the category label C9 of the video frames in multiple video frames can be added or deleted. For specific reference, please refer to the above introduction and will not be elaborated here.
[0649] The above takes the label smoothing strategy including that the category label C8 exists in the previous video frame among adjacent video frames and the category label C9 does not exist in the subsequent video frame among adjacent video frames as an example to introduce the process of performing label smoothing processing on multiple video frames according to the label smoothing strategy.
[0650] Through Figure 9 the video frame label processing method shown, the label carried by the video frames in the video can be corrected according to the label smoothing strategy corresponding to the video, so that the label stream of the video is smooth or conforms to common sense.
[0651] Next, in a specific example, the video frame label processing method provided by the embodiment of the present application will be introduced by way of example.
[0652] It can be set that a video frame can carry at most m category labels, which are described as L1, L2,..., Lm. It can be set that X 0,0 represents the label exists (carried by the video frame), and -X 0,0 represents the label does not exist (not carried by the video frame). The absolute value of X 0,0 is 1. If the video has n video frames, then the label stream of the video can be represented as Figure 10 the matrix shown. Among them, X i,j represents the value of the i-th row and the -th column. The label stream of the video is composed of the labels carried by the n video frames.
[0653] For the convenience of description, it can be set that m≥7 and n≥7. In this way, the matrix used to represent the label stream of the video can include at least the matrix shown in Figure 11 as shown. As Figure 11As shown, the class label represented by L1 can be set as the class label corresponding to a bus, the class label represented by L2 can be set as the class label corresponding to an airplane, the class label represented by L3 can be set as the class label corresponding to a subway, the class label represented by L4 can be set as the class label corresponding to a handrail, the class label represented by L5 can be set as the class label corresponding to a cat, the class label represented by L6 can be set as the class label corresponding to a chair, and the class label represented by L7 can be set as the class label corresponding to a door.
[0654] In the embodiment of the present application, a block diagram as shown in Figure 12 is drawn to represent the situation where each video frame carries a class label. Among them, the white block represents X 0,0 , and the black block represents -X 0,0 .
[0655] In the embodiment of the present application, score nodes such as consecutive nodes, coexistence nodes, derivation nodes, and consistency nodes are defined. The score node represents a logical relationship, and a Boolean expression corresponds to it. The Boolean expression can also be called Boolean logic, which is a mathematical term and a logical expression based on Boolean algebra. Specifically as follows.
[0656] A consecutive node is a node used to reflect the continuity of labels of the same category. According to common sense or experience, the same object category should be continuous between adjacent video frames. For example, if the current video frame carries the subway class label, it is expected that the next one also carries the subway class label. Another example is that if the current video frame does not carry the subway class label, it is expected that the next one also does not carry the subway class label. The Boolean expression corresponding to the consecutive node is
[0657] A coexistence node is a node used to reflect the coexistence of different class labels in the same video frame. It can be understood that according to common sense or experience, some object categories often exist simultaneously in the same [context], while some object categories can hardly exist simultaneously in the same [context]. For example, a subway and a person often exist simultaneously in the same [context], while a subway and an airplane can hardly exist simultaneously in the same [context]. Therefore, it is expected that the subway class label and the human class label are carried by the same video frame simultaneously, and it is not expected that the subway class label and the airplane class label are carried by the same video frame simultaneously. The Boolean expression corresponding to the coexistence node is X i*,j →X i,j Or,
[0658] A derivation node is a node used to reflect the derivation between object categories that appear at different times in the video. It can be understood that according to common sense or experience, some object categories that exist currently imply what object categories should exist or not exist next. The Boolean expression corresponding to the derivation node is X i*,j →Xi,j+1 Alternatively,
[0659] A consistent node is a node that reflects the consistency of the class labels carried by the video frames before and after. In this example, it is expected that the class labels carried by the video frames at input and output are as consistent as possible. The Boolean expression corresponding to the consistent node is X 0,0 →X i,j Alternatively, The former expectation is that the video frame carries the corresponding class label at output. Or the expectation is that the video frame does not carry the corresponding class label at output.
[0660] In this example, mathematical expressions corresponding to various Boolean expressions are set, as specifically Figure 13 shown.
[0661] For a given video label stream, the corresponding label smoothing strategy for the video (which can also be called prior knowledge) (taking the Figure 11 matrix shown as an example, the label smoothing strategy can include that all class labels are continuous between adjacent video frames, and the class labels of airplanes and subways do not coexist, etc.) can be used to smooth the labels of the video frames in the video. Among them, the label smoothing strategy can be preset, and it can include Figure 11 or Figure 10 the weight value w between the element units of the matrix shown (for example, when the weight value between two element units is 0, it means that there is no dependency relationship between these two element units. For another example, when the weight array between two element units is infinite, it means that there is an absolute dependency relationship between these two element units).
[0662] Corresponding scoring nodes can be created according to the label smoothing strategy. The objective function is the sum of all scoring nodes. In this way, the problem can be expressed by formula (3’).
[0663]
[0664] Among them, X ij takes values between 1 and -1, and X pq also takes values between 1 and -1. w ij,pq is a preset positive number.
[0665] Secondly, X ij X pq can be represented by a positive semi-definite matrix, and the original problem is relaxed to convert the original problem into a positive semi-definite optimization problem. The method is as follows.
[0666] Flatten the Figure 10 matrix shown into a column vector a = [X 0,0 ; X1,1 X 1,2 ;...; X i,n ;...; X 2,1 X 2,2 ...; X m,n . Multiply the column vector a by the transpose vector of the column vector a (a T ), to obtain a matrix Y. And X ij X pq is an element in the matrix Y. Thus, the formula (3’) can be expressed or converted into formula (4).
[0667]
[0668] Y is a positive semi - definite matrix; (4)
[0669] Rank(Y) = 1;
[0670] Y(i, i) = 1.
[0671] That is to say, convert the maximum - sum problem of formula (3”) into a positive semi - definite optimization problem with the objective of maximizing the summation result of formula (3”), rank(Y) = 1, and the values of the elements on the main diagonal of the matrix Y being 1 as constraints.
[0672] Among them, rank(Y) = 1 is a non - convex condition, making the process of positive semi - definite optimization with the objective of maximizing the summation result of formula (3”), rank(Y) = 1, and the values of the elements on the main diagonal of the matrix Y being 1 as constraints, cannot be solved quickly.
[0673] Exemplarily, for the positive semi - definite optimization problem with the objective of maximizing the summation result of formula (3”), rank(Y) = 1, and the values of the elements on the main diagonal of the matrix Y being 1 as constraints, it can be relaxed by removing or ignoring the constraint condition of rank(Y) = 1. Among them, relaxation is a mathematical optimization term, which means removing a part of the constraint conditions of the problem, thereby reducing the difficulty of the problem and obtaining an approximate solution.
[0674] Thus, the formula (4) can be expressed or converted into formula (4’).
[0675]
[0676] Y is a positive semi - definite matrix; (4’)
[0677] Y(i, i) = 1.
[0678] That is to say, the problem of maximizing the sum in formula (3”) can be transformed into a problem of semi - definite optimization with the goal of maximizing the sum result of formula (3”) and the elements on the main diagonal of matrix Y being 1 as constraints. This problem can be solved using the interior - point method or the ellipsoid method. Both the ellipsoid method and the interior - point method are terms in mathematical optimization, and both are mathematical optimization algorithms that can solve semi - definite optimization problems. For specific details, reference can be made to the introductions of the interior - point method or the ellipsoid method in the prior art, which will not be elaborated here.
[0679] Thus, a relaxed answer S can be obtained. S is a semi - definite matrix, and the elements on its main diagonal are 1. The relaxed answer can also be understood as an estimated solution. Since the condition rank(Y) = 1 is removed or ignored during the calculation process, the calculation result obtained is an estimated solution.
[0680] It is necessary to restore the relaxed answer S to obtain the true answer. Restoration is a term in mathematical optimization, which refers to the process of converting the optimal solution of the relaxed problem into a feasible solution of the original problem after obtaining the optimal solution of the relaxed problem. The specific restoration process is as follows.
[0681] The semi - definite matrix S can be eigen - value decomposed to obtain S = Y′·y T . Among them, Y′ is a matrix with n + 1 columns, where n is the number of video frames in the above - mentioned multiple video frames. The columns in Y′ correspond one - to - one with the columns in the above - mentioned column vector a.
[0682] A unit vector r can be randomly selected. If r T multiplies the first column of Y’, and the resulting product ≥ 0, then the first column of Y’ is set to correspond to 1. Correspondingly, the first column of column vector a, that is, X 00 = 1. If r T multiplies the first column of Y’, and the resulting product < 0, then the first column of Y’ is set to correspond to 1. Correspondingly, the first column of column vector a, that is, X 00 = - 1.
[0683] r T multiplies the i - th column of Y’, and then multiplies by X 0,0 , to obtain a product result. If the product result ≥ 0, then the i - th column in column vector a is 1; if the product result < 0, then the i - th column in column vector a is - 1. i = 2, 3...... In this way, the values of each element in column vector a can be obtained, and combined with the value of X 00 , it can be determined whether the values of each element in vector a are X 00 , or - X 00 . In this way, the class labels carried by each video frame in the video can be obtained when the value in formula (3’) is the largest, realizing the smoothness of the label stream of this video.
[0684] In an illustrative example, when the space requirement for calculating Equation (4’) exceeds the maximum limit of the computer, a divide-and-conquer algorithm can be adopted to estimate the above problem. Divide-and-conquer is a term in computer science, literally meaning “divide and conquer”, that is, dividing a complex problem into two or more identical or similar sub-problems until the final sub-problems can be simply and directly solved, and the solution of the original problem is the combination of the solutions of the sub-problems.
[0685] For a given estimation degree k, its k-order algorithm is as follows.
[0686] The original label stream (the class labels carried by video frames in the video when input, such as Figure 10 the matrix shown) can be divided into two sub-parts. If k - 1 is not 0, the k - 1 order estimation algorithm is respectively adopted for these two sub-parts and then divided again. In this way, after k divisions, 2 k sub-parts can be obtained. For each sub-part, the semi-definite optimization algorithm can be adopted in the form of Equation (4’) to obtain 2 k sub-results. Since the original label stream is divided into 2 k sub-parts, and for each sub-part, label smoothing is performed according to the label smoothing strategy corresponding to this sub-part, thereby reducing the space complexity of the calculation to 2 k one of that before divide-and-conquer. Correspondingly, the time complexity of the calculation is 2 k one of that before divide-and-conquer. Space complexity is a term in computer science, and its theoretical expression is the space required by an algorithm, reflecting the relationship between the algorithm time and the input size. After obtaining the label smoothing results of each of the 2 k sub-parts, the label smoothing results of the 2 k sub-parts can be combined. And according to the combined result, the class labels carried by each video frame in the video are set to achieve label stream smoothing of the video.
[0687] Next, the implementation process of the division is introduced.
[0688] When performing sub-part division, the minimum cut algorithm can be first adopted for division. Minimum cut, a term in graph theory: A cut is a partition of the vertices in a network that divides all the vertices in the network into two vertex sets S and T, where the source point s ∈ S and the sink point t ∈ T. It is denoted as CUT(S, T), and the minimum cut (min cut) from S to T that satisfies the condition.
[0689] Next, combined with Figure 14 , for division Figure 10Taking the element unit in the matrix shown as an example, the minimum cut algorithm is introduced. Among them, an element unit, which can also be called a node, refers to a position in the matrix. For example, the position indicated by the i-th row and the j-th column in the matrix can be called an element unit.
[0690] Refer to Figure 14 , B1, B2, B3, B4, B5, and B6 each represent Figure 10 an element unit in the matrix shown. Among them, B1 and B2 are in the same row, B3 and B4 are in the same row, and B5 and B6 are in the same row. As Figure 14 shown, the weight value between the element units in the same row is set to infinity. The weight value between different rows of elements is consistent with the label smoothing strategy. For example, it is set that there is coexistence between B1 and B3, that is, the class labels corresponding to B1 and B3 coexist in the same video frame, then the weight value between B1 and B3 is w in the above formula (5) 3 . For another example, it is set that there is a derivability between B3 and B6. The class label corresponding to B3 exists in the previous video frame of two adjacent video frames, and the class label corresponding to B6 exists (or does not exist) in the latter video frame of this adjacent video frame. Then the weight value between B3 and B6 is w in the above formula (9) 5 (or w in the above formula (11) 6 ).
[0691] In the above manner, the weight values are set between the element units in the matrix shown in Figure 10 . Then, using the minimum cut algorithm, the matrix is segmented according to the weight values between the element units, and two sub-parts can be obtained. For example, if w 5 (or w 6 ) is less than w 3 , then the element units in the row where B6 is located and the element units in the row where B3 is located are separated. Thus, the matrix shown in Figure 10 is divided into two sub-parts.
[0692] If the two sub-parts obtained after segmentation using the minimum cut algorithm are unbalanced (the number of element units contained in the two sub-parts differs greatly. For example, the proportion of element units contained in one of the two sub-parts is greater than 60% of the total number of element units contained in the two sub-parts, or less than 40% of the total number of element units contained in the two sub-parts.), then the minimum bisection algorithm is used for re-segmentation. Minimum bisection is a graph theory term, similar to the minimum cut, but additionally requires that the sizes of S and T are the same. The minimum bisection algorithm adopted in this embodiment can be called the Kernighan-Lin algorithm.
[0693] It should be noted that since the Kernighan-Lin algorithm also performs partitioning based on the weight values between element units in the matrix, before using the Kernighan-Lin algorithm for partitioning, it is also necessary to set the weight values between different element units in the same row to infinity, or in other words, fuse different element units in the same row into one node (i.e., regard different element units in the same row as one element unit). The weight values between element units in different rows are set according to the corresponding label smoothing strategy. Then, the set matrix is input into the Kernighan-Lin algorithm for partitioning. Thus, the matrix can be partitioned into two sub-parts, and the number of element units contained in these two sub-parts is equal.
[0694] Next, in an example, in combination with Figure 15 , the execution process of the video frame label processing method shown in Figure 9 will be introduced by way of example.
[0695] First, input the corresponding logical expression, weight, label stream, and optimization method. Among them, the logical expression and weight are determined by the label balancing strategy. The label stream is the class labels carried by multiple video frames in the video. The optimization method can be the interior point method or the ellipsoid method.
[0696] Second, create the corresponding logical node, weight matrix, and semi-definite relaxation expression. Among them, the logical node is the scoring node mentioned above. The weight matrix is a matrix with corresponding weights between element units.
[0697] Then, determine whether the space complexity of the algorithm exceeds the maximum limit.
[0698] If the space complexity of the algorithm does not exceed the maximum limit, use the interior point method or the ellipsoid method for semi-definite optimization. Then restore the optimization result and output it.
[0699] If the space complexity of the algorithm exceeds the maximum limit, input the estimation degree k. k ≥ 0 and k is a positive number. Determine whether k is equal to 0. If k is not equal to 0, then use the minimum cut to partition each sub-part (for the first partition, there is only one sub-part, that is, the matrix that has not been partitioned) respectively to obtain two new sub-parts. If the two new sub-parts obtained by partitioning the same sub-part are balanced, then let k = k - 1. If the two new sub-parts obtained by partitioning the same sub-part are not balanced, then use the Kernighan-Lin algorithm to re-partition the same sub-part and let k = k - 1. In this way, iterate until k = 0. Thus, 2 k sub-parts can be obtained. For these 2k For one of the sub - parts, label smoothing is independently performed according to the label smoothing strategy corresponding to that sub - part to obtain the label smoothing result of that sub - part. The specific process of label smoothing can refer to the above description and will not be elaborated here. The processes of using min - cut for segmentation and using the Kernighan - Lin algorithm for segmentation can refer to the above introduction and will not be elaborated here.
[0700] Two k sub - parts' label smoothing results can be merged (or synthesized) to obtain the final result. Thus, according to the final result, the values of the elements in the matrix can be set to achieve smooth label flow of the video.
[0701] In the embodiments of the present application, prior knowledge (continuity, co - existence, consistency, derivation) of the smoothing problem is expressed using Boolean logic. This expression method constitutes the basic framework of the solution, which can effectively reflect the internal relationship between data. Its corresponding mathematical expression can be relaxed and expressed by semi - definite programming for rapid solution. In addition, compared with dynamic programming and sliding window algorithms, this expression method is more general and can effectively reflect the connection between data, such as the dependencies between the current one and the current one, the current one and the next one, and the current one and the next two. This ability allows the method to combine prior knowledge of Markov of any order and autoregressive models of any order to obtain excellent answers.
[0702] Moreover, the embodiments of the present application use X 00 →X i,j or to reflect the consistency between input and output instead of using the common (X i,j - 1) 2 or (X i,j + 1) 2 . This design method not only reveals the relationship between consistency and Boolean logic but also ensures that the problem can be relaxed and expressed by semi - definite programming.
[0703] In addition, when the algorithm space is too large, this method improves the existing two algorithms and uses the divide - and - conquer algorithm to perform a secondary estimation of the original problem. This method exponentially reduces the space requirement while sacrificing some accuracy.
[0704] Next, combined with Figure 11 the matrix shown and different label smoothing strategies, the technical effects of the video frame label processing method provided in the embodiments of the present application are demonstrated.
[0705] In some embodiments, the matrix shown in Figure 11 can be used as the input, and the adopted label smoothing strategy is that various category labels are continuous between adjacent video frames. Among them, Figure 11The matrix shown corresponds to the block diagram as shown in Figure 16A According to this label smoothing strategy, label smoothing is performed. After setting the values of the elements in the matrix according to this label smoothing process, the block diagram corresponding to the obtained matrix is as shown in Figure 16B shown.
[0706] In an illustrative example, the matrix shown in Figure 11 can be used as the input, and the label smoothing strategy adopted is that label category L5 and label category L6 coexist in the same video frame. Among them, Figure 11 the matrix shown corresponds to the block diagram as shown in Figure 17A According to this label smoothing strategy, label smoothing is performed. After setting the values of the elements in the matrix according to this label smoothing process, the block diagram corresponding to the obtained matrix is as shown in Figure 17B shown.
[0707] In an illustrative example, the matrix shown in Figure 11 can be used as the input, and the label smoothing strategy adopted is that class label L1 and label category L2 do not coexist in the same video frame, class label L1 and label category L2 do not coexist in the same video frame, and class label L2 and label category L3 do not coexist in the same video frame. That is, class label L1, class label L1, and class label L1 do not coexist. Among them, Figure 11 the matrix shown corresponds to the block diagram as shown in Figure 18A According to this label smoothing strategy, label smoothing is performed. After setting the values of the elements in the matrix according to this label smoothing process, the block diagram corresponding to the obtained matrix is as shown in Figure 18B shown.
[0708] In an illustrative example, the matrix shown in Figure 11 can be used as the input, and the label smoothing strategy adopted is that label category L4 exists in the previous video frame of the adjacent video frames, and label category 5 exists in the next video frame of the adjacent video frames. Among them, Figure 11 the matrix shown corresponds to the block diagram as shown in Figure 19A According to this label smoothing strategy, label smoothing is performed. After setting the values of the elements in the matrix according to this label smoothing process, the block diagram corresponding to the obtained matrix is as shown in Figure 19B shown.
[0709] Among them, taking the example of smoothing the static semantic labels of video frames in the above embodiments, an example description of the label smoothing scheme is given. Similarly, this label smoothing scheme can also be used to smooth the labels of the dynamic semantics of video frames. Among them, the labels of dynamic semantics are information used to describe actions, such as playing basketball, playing football, etc. It can be understood that the dynamic semantics of different video frames should also conform to common sense or experience. For example, if the dynamic semantics of the previous video frame is playing basketball, then the dynamic semantics of the subsequent video frame should also be playing basketball, rather than playing football. Thus, when smoothing the labels of the dynamic semantics of video frames, a label smoothing strategy or prior knowledge can also be set. Then, using the label smoothing strategy, the above scheme is used to smooth the labels of the dynamic semantics. The specific smoothing process can refer to the above description of Figures 10 - 19B and will not be elaborated here one by one.
[0710] Returning to Figure 1 , in some embodiments, the neural network provided by the embodiments of the present application may further include an exciting time series segment recognition layer. In the exciting time series segment recognition layer, exciting time series segments can be recognized. An exciting time series segment refers to a time series segment synthesized or composed of video frames with the same detailed dynamic semantics.
[0711] In an illustrative example, the exciting time series segment recognition layer can be the next layer of the spatial domain feature extraction layer. Thus, in the exciting time series segment layer, the spatial domain features of each video frame of video A1 output by the spatial domain feature extraction layer can be used to determine the exciting time series segments in video A1.
[0712] Exemplarily, in the exciting time series segment recognition layer, the spatial domain difference information between two adjacent video frames can be calculated according to the spatial domain features of the two adjacent video frames. In an example, the spatial domain features may include the RGB information of the corresponding image. Among them, R in RGB represents red, G represents green, and B represents blue. The RGB difference information (RGB diff) between two adjacent video frames can be calculated according to the RGB information of the two adjacent video frames. The RGB diff between two adjacent video frames can be used as the spatial domain difference information between the two adjacent videos.
[0713] One or more exciting time series segments can be determined according to the spatial domain difference information between each pair of adjacent video frames among multiple video frames and the spatial domain features of each video frame.
[0714] In one example, the exciting time sequence segment recognition layer may include a one-dimensional convolutional layer and a detailed semantic recognition layer. The spatial domain difference information between each pair of adjacent video frames and the spatial domain features of each video frame among multiple video frames can be input into the one-dimensional convolutional layer for convolutional processing. Among them, the one-dimensional convolutional layer may include one or more convolutional windows, and each convolutional window has a certain coverage range. The coverage range can be understood as the width, which is represented by the number of video frames. For example, the coverage range of a convolutional window is Z video frames. When convolutional processing is performed using this convolutional window, convolutional processing can be performed on the Z videos covered by this convolutional window. Exemplarily, a convolutional window can also correspond to a kind of detailed dynamic semantics. It can be understood that when training the one-dimensional convolutional neural network, time sequence segments marked with detailed dynamic semantics can be used as training samples for training. The weight parameters corresponding to the convolutional windows of the trained one-dimensional convolutional neural network can use this detailed dynamic semantics as the detailed dynamic semantics corresponding to this convolutional window.
[0715] It can be set that the one-dimensional convolutional layer has a convolutional window R, the coverage range of the convolutional window R is Z, and it corresponds to the detailed dynamic semantics r (such as playing basketball). This convolutional window can be used to perform a convolutional operation on the spatial domain fe...
Claims
1. A computer - implemented method for identifying video semantics using a neural network, characterized in that, the neural network includes an input layer, a spatial - domain feature extraction layer, a static semantics recognition layer, a dynamic semantics recognition layer, a temporal segment division layer, and an output layer; wherein, the static semantics recognition layer and the dynamic semantics recognition layer are arranged in parallel; the method includes: at the input layer, obtaining multiple video frames of a video; at the spatial - domain feature extraction layer, extracting the spatial - domain features of each of the multiple video frames; at the dynamic semantics recognition layer, according to the spatial - domain features of N consecutive video frames among the multiple video frames, determining the dynamic semantics of the Nth video frame among the N consecutive video frames; N is a positive integer; at the static semantics recognition layer, according to the spatial - domain features of the first video frame among the multiple video frames, determining the static semantics of the first video frame; at the temporal segment division layer, when the number of consecutive video frames with a first dynamic semantics is greater than a first threshold, using the consecutive video frames with the first dynamic semantics to synthesize a first temporal segment, and determining the first dynamic semantics as the dynamic semantics of the first temporal segment; at the temporal segment division layer, when the number of consecutive video frames with a first static semantics is greater than a second threshold, using the consecutive video frames with the first static semantics to synthesize a second temporal segment, and determining the first static semantics as the static semantics of the second temporal segment; at the output layer, outputting the dynamic semantics of the first temporal segment and first position information; and outputting the static semantics of the second temporal segment and second position information; wherein, the first position information is represented by the positions of the first video frame and the last video frame in the first temporal segment in the video respectively; the second position information is represented by the positions of the first video frame and the last video frame in the second temporal segment in the video respectively.
2. The method according to claim 1, characterized in that, the neural network further includes an exciting temporal segment recognition layer; the method further includes: determining the spatial - domain difference information between the first video frame and the second video frame according to the spatial - domain features of the first video frame and the spatial - domain features of the second video frame; the first video frame and the second video frame are adjacent among the multiple video frames; at the exciting temporal segment recognition layer, determining at least one exciting temporal segment according to the spatial - domain difference information between each pair of adjacent video frames among the multiple video frames and the spatial - domain features of the multiple video frames.
3. The method according to claim 2, characterized in that, the spatial - domain features include RGB information, and the spatial - domain difference information includes RGB difference information (RGB diff).
4. The method according to claim 2, characterized in that, the exciting temporal segment recognition layer includes a one - dimensional convolutional layer and a detailed dynamic semantics classification layer; the one - dimensional convolutional layer includes a first convolutional window, and the first convolutional window corresponds to a first detailed dynamic semantics; Determining at least one wonderful time sequence segment according to the spatial domain difference information of each pair of adjacent video frames among the multiple video frames and the spatial domain features of the multiple video frames includes: In the one-dimensional convolutional layer, using a first convolutional window among the at least one convolutional window to perform convolutional processing on the spatial domain features of the multiple video frames and the spatial domain differences of each pair of adjacent video frames among the multiple video frames, to obtain a plurality of convolutional results; In the detailed dynamic semantic classification layer, determining a wonderful time sequence segment with a first detailed dynamic semantic according to the plurality of convolutional results.
5. The method according to claim 2, wherein, the neural network further includes a joint logic judgment layer, and the joint logic judgment layer is the next layer of the time sequence segment division layer and the wonderful time sequence segment recognition layer; a first wonderful time sequence segment among the at least one wonderful time sequence segment is included in the first time sequence segment; the method further includes: in the joint logic judgment layer, judging whether the detailed dynamic semantic of the first wonderful time sequence segment matches the dynamic semantic of the first time sequence segment; when the detailed dynamic semantic of the first wonderful time sequence segment matches the dynamic semantic of the first time sequence segment, outputting the detailed dynamic semantic of the first wonderful time sequence segment and third position information in the output layer; wherein, the third position information is represented by the positions of the first video frame and the last video frame in the first wonderful time sequence segment in the video respectively; when the detailed dynamic semantic of the first wonderful time sequence segment does not match the dynamic semantic of the first time sequence segment, not outputting the relevant information of the first wonderful time sequence segment in the output layer.
6. The method according to claim 1, wherein, the neural network further includes a semantic smoothing layer, and the semantic smoothing layer is the upper layer of the time sequence segment division layer; the method further includes: in the semantic smoothing layer, smoothing the static semantics of the multiple video frames according to the dependence relationship of the static semantics between consecutive video frames among the multiple video frames.
7. The method according to claim 6, wherein, smoothing the static semantics of the multiple video frames according to the dependence relationship of the static semantics between consecutive video frames among the multiple video frames includes: determining that the static semantic of the third video frame among P consecutive video frames is different from the static semantics of other video frames, and the static semantics of the other video frames are the same; P is greater than a third threshold; the other video frames are the video frames other than the third video frame among the P consecutive video frames; updating the static semantic of the third video frame according to the static semantics of the other video frames.
8. The method according to claim 1, wherein, the method further includes: in the time sequence segment division layer, when the number of consecutive video frames with a second dynamic semantic is greater than the first threshold, using the consecutive video frames with the second dynamic semantic to synthesize a third time sequence segment, and determining the second dynamic semantic as the dynamic semantic of the third time sequence segment. When the second dynamic semantics is the same as the first dynamic semantics, and the number of video frames between the first time sequence segment and the second time sequence segment is less than a fourth threshold, merge the first time sequence segment and the third time sequence segment into the same time sequence segment.
9. The method according to claim 1, wherein, the spatial domain features include a plurality of feature maps obtained by convolving the feature information of the video frames corresponding to the spatial domain features through a first convolutional layer, and the plurality of feature maps correspond one-to-one to the plurality of convolutional kernels of the first convolutional layer; the dynamic semantics recognition layer includes a second convolutional layer and a dynamic semantics classification layer; the determining of the dynamic semantics of the Nth video frame among the N consecutive video frames according to the spatial domain features of the N consecutive video frames in the plurality of video frames includes: performing feature map offset processing on the N consecutive video frames to obtain residual spatial domain features of the N consecutive video frames; wherein, the feature map offset processing includes: using the first feature map of the kth video frame among the N consecutive video frames to replace the first feature map of the (k + 1)th video frame among the N consecutive video frames, where k sequentially takes integer values from 1 to N - 1; the first feature map of the kth video frame and the first feature map of the (k + 1)th video frame correspond to the same convolutional kernel of the first convolutional layer; in the second convolutional layer, convolving the residual spatial domain features of the N consecutive video frames to obtain spatio-temporal features of the N consecutive video frames; in the dynamic semantics classification layer, determining the dynamic semantics of the Nth video frame according to the spatio-temporal features of the N consecutive video frames.
10. A video editing method, wherein, based on the semantics recognition method according to any one of claims 1-9 to obtain different semantics, the method includes: acquiring a first theme of a target spliced video and a first duration of the target spliced video; determining a plurality of time sequence segments whose semantics conform to the first theme; the plurality of time sequence segments include time sequence segments with dynamic semantics and time sequence segments with static semantics; determining, according to the first duration, time sequence segments for splicing the target spliced video from the plurality of time sequence segments; wherein, when the total duration of the time sequence segments with dynamic semantics is equal to or greater than the first duration, determining time sequence segments for splicing the target spliced video from the time sequence segments with dynamic semantics; when the total duration of the time sequence segments with dynamic semantics is less than the first duration, determining that all the time sequence segments with dynamic semantics are used for splicing the target spliced video, and determining time sequence segments for splicing the remaining video segments from the time sequence segments with static semantics, the duration of the remaining video segments being equal to the difference between the first duration and the total duration.
11. A video editing method, wherein, based on the semantics recognition method according to any one of claims 1-9 to obtain different semantics, the method includes: acquiring a first theme of a target spliced video and a first duration of the target spliced video; Determine multiple temporal segments whose semantics conform to the first theme; the multiple temporal segments include temporal segments with detailed dynamic semantics and temporal segments with dynamic semantics; Determine, according to the first duration, the temporal segments for splicing the target spliced video from the multiple temporal segments; Wherein, when the total duration of the temporal segments with detailed dynamic semantics is equal to or greater than the first duration, determine the temporal segments for splicing the target spliced video from the temporal segments with detailed dynamic semantics; When the total duration of the temporal segments with detailed dynamic semantics is less than the first duration, determine that all the temporal segments with detailed dynamic semantics are used for splicing the target spliced video, and determine the temporal segments for splicing the remaining video segments from the temporal segments with dynamic semantics, where the duration of the remaining video segments is equal to the difference between the first duration and the total duration.
12. A video editing method, Characterized in that, Based on the semantic recognition method according to any one of claims 1-9, different semantics are obtained, and the method includes: Obtain the first theme of the target spliced video and the first duration of the target spliced video; Determine multiple temporal segments whose semantics conform to the first theme; the multiple temporal segments include temporal segments with detailed dynamic semantics, temporal segments with dynamic semantics, and temporal segments with static semantics; Determine, according to the first duration, the temporal segments for splicing the target spliced video from the multiple temporal segments; Wherein, when the total duration of the temporal segments with detailed dynamic semantics and the temporal segments with dynamic semantics is less than the first duration, determine that all the temporal segments with detailed dynamic semantics and the temporal segments with dynamic semantics are used for splicing the target spliced video, and determine the temporal segments for splicing the remaining video segments from the temporal segments with static semantics, where the duration of the remaining video segments is equal to the difference between the first duration and the total duration.
13. An electronic device, Characterized in that, Comprising: A processor and a memory; The memory is used to store computer instructions; When the electronic device runs, the processor executes the computer instructions, so that the electronic device executes the method according to any one of claims 1-9 or the method according to claim 10 or the method according to claim 11 or the method according to claim 12.
14. A computer storage medium, Characterized in that, The computer storage medium includes computer instructions, and when the computer instructions run on an electronic device, the electronic device is caused to execute the method according to any one of claims 1-9 or the method according to claim 10 or the method according to claim 11 or the method according to claim 12.
Citation Information
Patent Citations
Video positioning method and device and electronic equipment
CN110225368A
Method for extracting wonderful clip of badminton event video based on machine learning
CN111291617A