A method for video data recognition, an electronic device, and a storage medium
By dividing the video data into multiple video units, using the feature fusion of reference frames and predicted frames, the problems of low video data recognition efficiency and low accuracy are solved, and efficient video data recognition is achieved.
Patent Information
- Application Number
- CN202510281997.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-03-11
AI Technical Summary
The prior art has problems with low recognition efficiency and low accuracy in video data recognition, especially when processing a large number of video frames, the resource occupation is increased or partial frame processing leads to information loss.
The video data is divided into multiple video units, and the fusion characteristics of the video unit are determined by fusion of reference frames and predicted frames, and the video data is identified using the team characteristics of the frame team to reduce input content, and improve recognition efficiency and accuracy.
By dividing the video data into multiple video units and fusion of features, the recognition efficiency and accuracy of the video data are improved, and long video data can be effectively processed.
Smart Images

Figure CN119785275B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of video data recognition, and in particular, to a video data recognition method, an electronic device, and a storage medium. Background Art
[0002] With the explosive growth of computer applications, the recognition of video data has entered a stage of rapid development. Video data generally contains a large amount of redundant information. How to efficiently understand the content of video data and extract valuable information from the vast amount of data has become the focus of attention.
[0003] In the current technology, for the recognition of video data, usually a deep neural network model is constructed, and the video spatial domain image sequence is used as the input. The extracted video features are mapped to the semantic space to generate semantic information. However, using the frame sequence of the complete video data as the input, as the number of video frames increases, the resource occupancy increases linearly, resulting in low recognition efficiency of the recognition data; or a large model is used to calculate the similarity of the form of image-text pairs, but only some video frames are extracted for processing, which will lose a large amount of content information, resulting in a decrease in the recognition accuracy, and further affecting the recognition accuracy of video data. Summary of the Invention
[0004] The main technical problem to be solved by this application is to provide a video data recognition method, an electronic device, and a storage medium. By dividing the recognition data into multiple video units, and then using the reference frame features of the reference frames and the temporal dynamic features of the predicted frames in the video units to determine the first fusion feature of each video unit. Then, using the first temporal feature, the first motion information, and the first fusion feature corresponding to at least one video unit to determine the team features of multiple frame teams and the frame teams, and then using the team features of multiple teams to determine the video fusion feature of the video data, and recognizing the video fusion feature to obtain the recognition result, which can effectively improve the recognition efficiency and the recognition accuracy.
[0005] To solve the above technical problems, a technical solution adopted in this application is: to provide a video data recognition method, including: dividing the video data to be processed to obtain multiple video units, where each of the video units includes a reference frame and multiple prediction frames; obtaining the reference frame features corresponding to the reference frame, and obtaining the temporal dynamic features corresponding to each of the prediction frames, and using the reference frame features and the multiple temporal dynamic features for fusion to determine the first fusion feature of each of the video units; using the first temporal feature, the first motion information, and the first fusion feature corresponding to each of the video units to determine multiple frame teams and the team features of each of the frame teams, where the frame team includes the first temporal feature, the first motion information, and the first fusion feature corresponding to at least one of the video units; using the multiple team features to determine the video fusion feature of the video data to be processed, and then recognizing the video fusion feature to obtain the recognition result of the video data to be processed.
[0006] In some embodiments, the obtaining the reference frame features corresponding to the reference frame, and obtaining the temporal dynamic features corresponding to each of the prediction frames includes: obtaining the reference frame features and the first position information corresponding to the reference frame in each of the video units; obtaining the first motion vector of each of the prediction frames, and then using the first position information and the first motion vector to determine the second position information of each of the prediction frames; using the reference frame features, the second position information, and the multiple prediction frames to determine the temporal dynamic features.
[0007] In some embodiments, the using the reference frame features and the multiple temporal dynamic features for fusion to determine the first fusion feature of each of the video units includes: obtaining the second temporal features corresponding to the reference frame and each of the prediction frames of each of the video units; using the second temporal features, the reference frame features, and the temporal dynamic features corresponding to the multiple prediction frames for fusion to determine the first fusion feature of each of the video units.
[0008] In some embodiments, the using the first temporal feature, the first motion information, and the first fusion feature corresponding to each of the video units to determine multiple frame teams and the team features of each of the frame teams includes: obtaining the first temporal feature of each of the video units and the reference frame in each of the video units, and obtaining the first motion information corresponding to each of the video units; using the first temporal feature and the first fusion feature of at least one of the video units to form a frame team with a preset length; fusing the first temporal feature, the first fusion feature, and the first motion information corresponding to at least one of the video units in each of the frame teams to determine the team feature of each of the frame teams.
[0009] In some embodiments, obtaining the first motion information corresponding to each of the video units includes: obtaining second motion vectors of each macroblock region in the video unit, and then determining the change amount of each macroblock region; using the change amount to determine the first motion information corresponding to the video unit.
[0010] In some embodiments, the preset length is less than or equal to the length of the video data to be processed, or the preset length is greater than or equal to the length of the video unit.
[0011] In some embodiments, using a plurality of the team features to determine the video fusion feature of the video data to be processed, and then identifying the video fusion feature to obtain the identification result of the video data to be processed includes: performing weighted processing on the plurality of the team features to obtain the video fusion feature of the video data to be processed; before the inverse quantization operation in the video decoding process, identifying the video fusion feature to obtain the identification result of the video data to be processed.
[0012] In some embodiments, the video data to be processed is incompletely decoded video frame data.
[0013] To solve the above technical problems, another technical solution adopted by this application is: to provide an electronic device, the electronic device includes a memory and a processor coupled to the memory, the memory stores at least one computer program, and when the at least one computer program is loaded and executed by the processor, it is used to implement the video data identification method as described above.
[0014] To solve the above technical problems, another technical solution adopted by this application is: to provide a computer-readable storage medium, the computer-readable storage medium has at least one segment of program, and when the at least one segment of program is loaded and executed by a processor, it is used to implement the video data identification method as described above.
[0015] Different from the current technology, the video data recognition method provided by this application includes: dividing the video data to be processed to obtain multiple video units, where each video unit includes a reference frame and multiple prediction frames; obtaining the reference frame features corresponding to the reference frame, and obtaining the temporal dynamic features corresponding to each prediction frame, and fusing the reference frame features and multiple temporal dynamic features to determine the first fusion feature of each video unit; using the first temporal feature, the first motion information, and the first fusion feature corresponding to each video unit to determine multiple frame groups and the group features of each frame group, where the frame group includes the first temporal feature, the first motion information, and the first fusion feature corresponding to at least one video unit; using multiple group features to determine the video fusion feature of the video data to be processed, and then recognizing the video fusion feature to obtain the recognition result of the video data to be processed; that is, in this application, by dividing the recognition data into multiple video units, and then using the reference frame features of the reference frames in the video units and the temporal dynamic features of the prediction frames to determine the first fusion feature of each video unit, and then using the first temporal feature, the first motion information, and the first fusion feature corresponding to at least one video unit to determine multiple frame groups and the group features of the frame groups, and then using multiple group features to determine the video fusion feature of the video data, and recognizing the video fusion feature to obtain the recognition result, the recognition accuracy can be effectively improved, and long videos can be processed, and the recognition efficiency can be effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. Among them:
[0017] Figure 1 is a schematic flowchart of an embodiment of the video data recognition method in this application;
[0018] Figure 2 is a schematic structural diagram of an embodiment of the video unit processing module in this application;
[0019] Figure 3 is a schematic flowchart of an embodiment of obtaining motion information in this application;
[0020] Figure 4 is a schematic structural diagram of an embodiment of obtaining the video fusion feature in this application;
[0021] Figure 5 is a schematic structural diagram of an embodiment of the video data recognition system in this application;
[0022] Figure 6It is a schematic structural diagram of an embodiment of an electronic device in the present application;
[0023] Figure 7 It is a schematic structural diagram of an embodiment of a computer-readable storage medium in the present application. Detailed implementation manners
[0024] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be specifically noted that the following embodiments are only used to illustrate the present invention, but do not limit the scope of the present invention. Similarly, the following embodiments are only partial embodiments of the present invention rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0025] Referring to "embodiments" herein means that the specific features, structures, or characteristics described in connection with the embodiments may be included in at least one embodiment of the present invention. The phrase appears in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0026] In traditional identification data recognition methods, usually, a deep neural network model is constructed, and with the video spatial domain image sequence as the input, the extracted video features are mapped to the semantic space to generate semantic information. However, using the frame sequence of the complete video data as the input, as the number of video frames increases, the resource occupancy increases linearly, resulting in low recognition efficiency of the identification data; or the similarity calculation is performed on the form of image-text pairs using a large model, but only part of the video frames are extracted for processing, which will lose a large amount of content information, leading to a decrease in the recognition accuracy, and further affecting the recognition accuracy of video data.
[0027] Therefore, a video data recognition method is provided. Taking the video unit as the processing unit, feature fusion is performed, and then at least one video unit is used to determine the frame team, reducing the input content and improving the video data recognition efficiency; furthermore, the video fusion feature of the video data is determined using the team feature of the frame team, and then the video fusion feature is recognized to obtain the recognition result. Combining all the features can effectively improve the recognition accuracy of video data.
[0028] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of an embodiment of the video data recognition method in the present application; it should be noted that if there are substantial results, the method of the present application is not limited to Figure 1 the process sequence shown.
[0029] As Figure 1As shown, the video data recognition method of the present application may include the following steps.
[0030] S10. Divide the video data to be processed to obtain a plurality of video units, where each video unit includes a reference frame and a plurality of predicted frames.
[0031] Among them, the video data to be processed refers to the video data that has not been processed yet; the video unit refers to a group of frames, that is, GOP data (Group of Pictures), generally the distance between two IDR frames (Instantaneous Decoder Refresh Frame), which is used to describe the video sequence structure and can be a picture group.
[0032] Specifically, use the acquisition module to acquire the video data to be processed, and then divide the video data to be processed in the GOP form to obtain a plurality of video units, and use the first frame in the video unit as the reference frame, and the subsequent other frames as the predicted frames.
[0033] In some embodiments, the video data to be processed refers to the frequency-domain data of an incompletely decoded frame sequence.
[0034] S20. Obtain the reference frame feature corresponding to the reference frame, and obtain the temporal dynamic feature corresponding to each predicted frame, and use the reference frame feature and a plurality of temporal dynamic features for fusion to determine the first fusion feature of each video unit.
[0035] Among them, the reference frame feature refers to the feature extracted from the reference frame that can characterize the reference frame; the temporal dynamic feature refers to the feature extracted from the predicted frame that can characterize the dynamics of the predicted frame; the fusion refers to the feature formed by superimposing the two features together.
[0036] Specifically, after obtaining the reference frame and a plurality of predicted frames of each video unit, use the reference frame processing module to process the reference frame, input the reference frame into the reference frame processing module to extract the reference frame feature corresponding to the reference frame; and use the predicted frame processing module to process the predicted frame, that is, input the predicted frame into the predicted frame processing module to extract the temporal dynamic feature corresponding to each predicted frame; then use the intra-group feature fusion module to perform fusion processing on the features of each video unit, that is, input the reference frame feature and a plurality of dynamic temporal features in the video unit into the intra-group feature fusion module to obtain the first fusion feature corresponding to each video unit.
[0037] S30. Use the first temporal feature, the first motion information, and the first fusion feature corresponding to each video unit to determine a plurality of frame teams and the team feature of each frame team, where the frame team includes the first temporal feature, the first motion information, and the first fusion feature corresponding to at least one video unit.
[0038] Among them, the first timing feature refers to the time series corresponding to the first fusion feature; the first motion information refers to the motion information corresponding to the video unit, representing the motion information of the target object in the video unit; the frame team refers to a queue composed of one or more video units; the team feature refers to the fusion of the features of all video units in the frame team.
[0039] Specifically, after obtaining the first fusion feature of each video unit, it is also necessary to use the frame group motion information acquisition module to obtain the first timing feature and the first motion information corresponding to each video unit, and then divide the first fusion feature and the timing feature into frame teams with a fixed length. Furthermore, the time series, the first fusion feature, and the first motion information corresponding to one or more video units within the frame team are fused to obtain the team feature of each frame team.
[0040] In some embodiments, the fusion here can be to splice the first fusion feature, the first timing feature, and the first motion information of each video unit to obtain the team feature of each frame team.
[0041] S40. Use multiple team features to determine the video fusion feature of the video data to be processed, and then identify the video fusion feature to obtain the recognition result of the video data to be processed.
[0042] Among them, the video fusion feature refers to the fusion of the team features of all frame teams of the entire video data to be processed; the recognition refers to the understanding and analysis result of the content of the video data to be processed by identifying the video fusion feature, that is, the recognition result.
[0043] Specifically, after obtaining multiple frame teams and the team feature of each frame team, the team features of all frame teams in the video data to be processed are fused to obtain the video fusion feature corresponding to the video data to be processed, and then the recognition module is used to identify the video fusion feature to obtain the recognition result of the video data to be processed.
[0044] In this embodiment, by dividing the recognition data into multiple video units, and then using the reference frame feature of the reference frame and the timing dynamic feature of the prediction frame in the video unit to determine the first fusion feature of each video unit, and then using the first timing feature, the first motion information, and the first fusion feature corresponding to at least one video unit to determine multiple frame teams and the team feature of the frame team, and then using multiple team features to determine the video fusion feature of the video data, and identifying the video fusion feature to obtain the recognition result, it can effectively improve the recognition accuracy, and can process long videos, effectively improving the recognition efficiency.
[0045] In some embodiments, the video unit processing module may include sub-modules such as a reference frame processing module, a predicted frame processing module, and an intra-group-of-pictures fusion module.
[0046] Refer to Figure 2 , Figure 2 which is a schematic structural diagram of an embodiment of the video unit processing module in the present application.
[0047] As Figure 2 shown, the video unit processing module includes a reference frame processing module N I , a predicted frame processing module N p , and an intra-group-of-pictures feature fusion module N G . Among them, the reference frame processing module N I takes the reference frame I I as input and extracts the reference frame feature F I ; the predicted frame processing module N p takes I Pi as input and extracts the temporal dynamic feature F Pi of the predicted frame, and F P1 is the temporal dynamic feature of the first predicted frame, F I is the reference frame feature, I P1 is the first predicted frame, F Pj is the temporal dynamic feature of the j-th predicted frame, I Pj is the j-th predicted frame, F Pm-1 is the temporal dynamic feature of the (m - 1)-th predicted frame, I Pm-1 is the P m-1 -th predicted frame; the P m-1 -th predicted frame is input into the predicted frame processing module N p to obtain the corresponding temporal dynamic feature F Pm-1 , and then the first fusion feature G G of the video unit is obtained through the intra-group-of-pictures feature fusion module N i ; N P1 is the first processing unit of the intra-group-of-pictures feature fusion module N P , S P1 is the first motion information of the first predicted frame, N Pj is the j-th processing unit of the intra-group-of-pictures feature fusion module N P , S Pj is the first motion information of the j-th predicted frame, N Pm-1 is the (m - 1)-th processing unit of the intra-group-of-pictures feature fusion module N P , S Pm-1 is the first motion information of the (m - 1)-th predicted frame, and the first motion information of all predicted frames within the video unit constitutes the first motion information of the entire video unit; T 0 refers to the temporal feature of the reference frame, T 1is the temporal feature of the first predicted frame, T j is the temporal feature of the j-th predicted frame, T m-1 is the temporal feature of the m-1 predicted frame.
[0048] In some embodiments, S10 divides the video data to be processed to obtain multiple video units, where each video unit includes a reference frame and multiple predicted frames, and may include the following operations.
[0049] The acquisition module acquires the incompletely decoded video data as the video data to be processed, and then divides the video data to be processed in the GOP format to obtain multiple video units, and uses the first frame in the video unit as the reference frame, and the subsequent other frames as the predicted frames.
[0050] It can be understood that the size of the video data to be processed can be selected according to the actual situation, and the size is not limited here; the format of the video data to be processed can be a playable video data format, such as mp4, avi, ts, wmv, etc.; the compression standard of the video data to be processed can be MPEG, H.263, H.264, etc.; a group of video frames that can be played independently is used as a GOP, that is, a video unit. Each GOP contains a key frame, multiple predicted frames, and may also include different types of video frame data such as bidirectional predicted frames. And the differences in pixels, brightness, chrominance, etc. between video frames are small, and they have high correlation. The key frame can display a complete image independently, while the predicted frame only contains the macroblocks and motion vectors of the difference part between the current frame and other frames. Therefore, the complete information of the current frame needs to be obtained according to the data of the key frame and its previous and subsequent frames, and combined with the motion vector.
[0051] In some embodiments, obtaining the reference frame features corresponding to the reference frame in S20, and obtaining the temporal dynamic features corresponding to each predicted frame may include the following operations.
[0052] First, obtain the reference frame features and the first position information corresponding to the reference frame in each video unit.
[0053] Among them, each video unit may include a reference frame; the first position information refers to the position information of the reference frame in the video unit.
[0054] Specifically, after using the acquisition module to obtain the reference frame in each video unit, the reference frame processing module processes the reference frame to obtain the reference frame features corresponding to the reference frame, and determines the first position information of the reference frame in the video unit.
[0055] For example, a video data to be processed includes n video units, each video unit includes m video frames, and the first video frame of each video unit is denoted as reference frame II , other video frames are denoted as predicted frame I Pi , the reference frame processing module is N 1 , then use the reference frame processing module N 1 to process the reference frames that can be independently decoded and displayed. The specific processing is as follows:
[0056] .
[0057] Among them, the reference frame processing module N 1 takes the reference frame I within the video unit I as the input, and then extracts the reference frame feature F I . It can be understood that the position information of the reference frame macroblock is S I , that is, the first position information of the reference frame can be S I .
[0058] Then, obtain the first motion vector of each predicted frame, and then use the first position information and the first motion vector to determine the second position information of each predicted frame.
[0059] Among them, the first motion vector refers to the motion vector of the current predicted frame, and the second position information refers to the position information of the predicted frame in the video unit.
[0060] Specifically, after obtaining multiple predicted frames of the video unit, for each predicted frame, obtain the first motion vector of each predicted frame, and then combine the first position information of the reference frame and the first motion vector of the current predicted frame to determine the second position information of each predicted frame.
[0061] Then, use the reference frame feature, the second position information, and multiple predicted frames to determine the temporal dynamic feature.
[0062] Specifically, after obtaining the second position information of each predicted frame, then the temporal dynamic feature of each predicted frame can be determined by using the reference frame feature, multiple predicted frames, and the second position information of each predicted frame.
[0063] For example, the first position information of the reference frame is S I , the first motion vector of the predicted frame is M i , , m is the number of video frames in a video unit. And set the second position information of the predicted frame to S Pi , and depends on the position information of the previous video frame and the motion vector of the current video frame.
[0064] Specifically, when the current predicted frame is the first predicted frame, then there is:
[0065] .
[0066] Among them, S I is the first position information of the reference frame, M 1 is the first motion vector of the first predicted frame, and S P1 is the second position information of the first predicted frame.
[0067] When the current predicted frame is other predicted frames except the first predicted frame, then there is:
[0068] .
[0069] Among them, S Pi-1 is the second position information of the (i - 1)-th predicted frame, M i is the first motion vector of the i-th predicted frame, and S Pi is the second position information of the i-th predicted frame.
[0070] Then there is:
[0071] .
[0072] .
[0073] Among them, F P1 is the temporal dynamic feature of the first predicted frame, F I is the reference frame feature, I P1 is the first predicted frame, F Pi is the temporal dynamic feature of the i-th predicted frame, I Pi is the i-th predicted frame, F Pi-1 is the temporal dynamic feature of the (i - 1)-th predicted frame, and the value range of i is from 2 to m -1, and N P is the predicted frame processing module.
[0074] In some embodiments, the operation of using the reference frame feature and multiple temporal dynamic features for fusion in S20 to determine the first fusion feature of each video unit may include the following operations.
[0075] First, obtain the second temporal features corresponding to the reference frame and each predicted frame in each video unit.
[0076] Among them, since the video frames included in the video data to be processed also have a temporal order, it is necessary to add temporal information to the frame features within the video unit to represent the temporal differences of the video frames.
[0077] Specifically, obtain the reference frame and multiple predicted frames of each video unit, and determine the temporal order of each predicted frame, that is, determine the second temporal feature corresponding to each predicted frame, and then add the second temporal feature to the temporal dynamic feature.
[0078] For example, the second temporal feature is denoted as {T 1 , T P1… , T Pm-1}, where the second temporal feature can be obtained by using relative position encoding or methods such as neural network learning, which is not limited here. Then:
[0079] .
[0080] .
[0081] Among them, T I is the second temporal feature corresponding to the first prediction frame, F I is the reference frame feature, F Pi is the temporal dynamic feature of the i-th prediction frame, T Pi is the second temporal feature of the i-th prediction frame, F Pi+1 is the temporal dynamic feature of the (i + 1)-th prediction frame.
[0082] Next, the temporal feature, the reference frame feature, and the second temporal dynamic features corresponding to multiple prediction frames are fused to determine the first fusion feature of each video unit.
[0083] Among them, since each video unit has a corresponding reference frame and multiple prediction frames, there are multiple frame features, and the first fusion feature is the fusion feature of the multiple frame features in the video unit.
[0084] Specifically, after obtaining the temporal dynamic feature of each prediction frame in the video unit, the frame feature fusion module within the frame group is used to fuse the frame features in the video unit with the added second temporal feature, so as to obtain the first fusion feature of each video unit.
[0085] For example, it is set that the frame feature fusion module within the frame group is N G , the first fusion feature of the video unit is G j , and the feature sequence within the video unit with the added second temporal feature is {F I , F P0… , F Pm-1}, then:
[0086] .
[0087] Among them, F I is the reference frame feature, F P0 is the temporal dynamic feature of the first prediction frame, F Pm-1 is the temporal dynamic feature of the m prediction frame.
[0088] In this embodiment, by obtaining the reference frame features corresponding to the reference frame and obtaining a plurality of prediction frames and the second position information of each prediction frame, the temporal dynamic features of each prediction frame are determined, and then the first fusion features of each video unit are obtained by fusing the second temporal features, reference frame features and a plurality of temporal dynamic features of each prediction frame, effectively combining the temporal information in the video unit, so that the first fusion features can clearly represent the data frame features in the video unit.
[0089] In some embodiments, S30 uses the first temporal features, first motion information, and first fusion features corresponding to each video unit to determine a plurality of frame groups and the group features of each frame group, which may include the following operations.
[0090] First, obtain the first temporal features of each video unit and the reference frame in each video unit, and obtain the first motion information corresponding to each video unit.
[0091] Among them, the first temporal feature refers to the temporal feature of each video unit; the first motion information refers to the set of motion information of multiple prediction frames in the video unit and the previous frame, that is, the set of change amounts of motion vectors.
[0092] Specifically, obtain the reference frame and a plurality of prediction frames in each video unit, and obtain the first temporal features corresponding to each video unit, and then use the frame group motion information acquisition module to obtain the first motion information corresponding to the video unit.
[0093] Next, use the first temporal features and first fusion features of at least one video unit to form a frame group with a preset length.
[0094] Among them, the first temporal feature refers to the set of second temporal features of each prediction frame in the video unit, that is, the temporal feature sequence; the preset length refers to the length of the frame group set in advance, which can be a fixed value or can be set according to the actual situation, that is, the frame group can be a video unit or a set of multiple video units.
[0095] Specifically, after obtaining the second temporal features of each prediction frame, determine the set of second temporal features of multiple prediction frames in the video unit as the first temporal features of the video unit; then determine the length of each video unit, and then use the first temporal features and first fusion features of at least one video unit to form a frame group with a preset length.
[0096] Then, fuse the first temporal features, first fusion features, and first motion information corresponding to at least one video unit in each frame group to determine the group features of each frame group.
[0097] Among them, the team feature refers to the combined feature of the first temporal feature, the first fusion feature, and the first motion information of at least one video unit in the frame team, and can also be a splicing feature.
[0098] Specifically, after obtaining the frame team, determine the video units included in the frame team, and determine the corresponding first temporal feature, first fusion feature, and first motion information for each video unit. Then, fuse the first temporal feature, first fusion feature, and first motion information included in the frame team to obtain the team feature of each frame team.
[0099] In some embodiments, the preset length is less than the length of the video data to be processed, and the preset length is greater than or equal to the length of the video unit.
[0100] Specifically, the video data to be processed may include multiple frame teams. Therefore, the preset length of the frame team should be less than the length of the video data to be processed; and the frame team may include one or more video units. Therefore, the preset length of the frame team should be greater than or equal to the length of the video unit.
[0101] Further, obtaining the first motion information corresponding to each video unit may include the following operations.
[0102] Obtain the second motion vector of each macroblock region in the video unit, and then determine the change amount of each macroblock region.
[0103] Among them, the motion intensity of the target object in the video data can be represented by the motion vector of the macroblock. A macroblock refers to a basic processing unit in video coding, and each prediction frame may include multiple macroblocks.
[0104] Specifically, use the frame group motion information acquisition module to obtain the data frames in the video unit. The data frames include reference frames and multiple prediction frames. Then, obtain the second motion vector of each macroblock region in each data frame. Use multiple second motion vectors to determine the third motion vector of each data frame, and then use multiple third motion vectors to determine the fourth motion vector of the video unit; and the change amount of the motion vector is determined by the difference between the motion vector of the previous frame and the motion vector of the current frame, and then determine the change amount of each macroblock region.
[0105] Use the change amount to determine the first motion information corresponding to the video unit.
[0106] Among them, after determining the change amount of each macroblock region, the change amount of each data frame can be determined, and then the change amount of each video unit can be further determined. Then, use the change amount of each video unit to determine the first motion information corresponding to the video unit, that is, the motion information of the target object in the video unit.
[0107] For example, in the video data to be processed, the motion states of the target object can be divided into three types: stationary, regular motion, and irregular motion. The target object with a large change in motion state contains more important information. Therefore, the motion vector of the i-th predicted frame is set as M i , then the motion vector of the video unit is M j ={M i}, where the value range of i is [1, m -1], and the value range of j is [0, n -1]; then, the change amount of the motion vector is:
[0108] V j ={M i -M i-1}.
[0109] Among them, V j is the change amount of the motion vector of the video unit.
[0110] It can be understood that for a still image or a macroblock area that has not changed, the corresponding change amount of the motion vector is a 0 vector.
[0111] Then, the first motion information of the video unit is calculated as follows:
[0112] .
[0113] Among them, S j is the first motion information of the video unit, and N m is the frame group motion information acquisition module.
[0114] In addition, the sequence of the first fusion features included in the video data to be processed is set as:
[0115] .
[0116] Among them, G 0 is the first fusion feature of the first data frame, and G n-1 is the first fusion feature of the (n-1)-th data frame.
[0117] And the corresponding time feature sequence is:
[0118] .
[0119] Among them, T Go is the time feature of the first data frame, and T Gn-1 is the time feature of the (n-1)-th data frame.
[0120] Refer to Figure 3 , Figure 3 which is the flowchart of an embodiment for obtaining motion information in this application.
[0121] As shown Figure 3 in the figure, set the motion vector of each video unit as M j , then the change amount of the motion vector is V j ; furthermore, through the group-of-frames motion information acquisition module N m obtain the first motion information of each video unit. Among them, M1 refers to the motion vector of the first video frame in the video unit, V2 refers to the change amount of the motion vector between the first video frame and the second video frame, and S refers to the first motion information of the entire video unit.
[0122] Then, divide the first fusion feature and the time feature sequence of the video unit into frame groups with a preset length, and the preset length is w , and then for the w video units in the frame group are fused to obtain the team feature:
[0123] .
[0124] Among them, W k is the team feature, Concat(*) is the feature concatenation operation, that is, concatenate the first fusion features of different video units in the frame group; j The value range of is kw , and the value range of k is:
[0125] .
[0126] In this embodiment, by obtaining the second motion vector of each macroblock region in the video data, determining the change amount of each macroblock region, and then determining the change amount of the whole video unit, finally determining the first motion information corresponding to the video unit, and then combining the video frame with the motion information to determine the corresponding video fusion feature, the scale of the data can be effectively reduced, providing a basis for improving the recognition efficiency in the follow-up.
[0127] In some embodiments, S40 uses multiple team features to determine the video fusion feature of the video data to be processed, and then recognizes the video fusion feature to obtain the recognition result of the video data to be processed, which may include the following operations.
[0128] First, perform weighted processing on multiple team features to obtain the video fusion feature of the video data to be processed.
[0129] Specifically, the video data to be processed may include multiple frame groups, so there are multiple team features, and then perform weighted processing on multiple team features to obtain the video fusion feature of the video data to be processed.
[0130] For example, by performing weighted processing on multiple team features, obtain the video fusion feature of the video data to be processed:
[0131]
[0132] Among them, O is the video fusion feature.
[0133] Refer to Figure 4 , Figure 4 which is a schematic structural diagram of an embodiment for obtaining the video fusion feature in this application.
[0134] As Figure 4 shown, G is the first fusion feature of the video unit, and G0 is the first fusion feature of the first video unit in the frame team; here, three frame teams are set, that is, the three large frames in the second step. Each frame team has corresponding multiple video units, and T G0 refers to the first timing feature of the first video unit, and S 0 refers to the first motion information of the first video unit, and then the team information W of the frame team is determined k , W 0 refers to the team feature of the first frame team, and then the team features of multiple frame teams are spliced to obtain the video fusion feature O.
[0135] Next, before the inverse quantization operation in the video decoding process, the video fusion feature is recognized to obtain the recognition result of the video data to be processed.
[0136] Among them, the video decoding process includes: inputting a video signal; entropy encoding a video frame; frequency-domain compressing a video frame, then performing an inverse quantization operation to obtain an unquantized video frame, and then performing an inverse transform to obtain a decoded video frame.
[0137] Specifically, before the inverse quantization operation in the video encryption process, the obtained video fusion feature is recognized, and then the recognition result of the video data to be processed is obtained. There is no need to perform an inverse quantization operation, which can save the computing power overhead of video decoding, reduce the input data scale at the same time, improve the training speed of the neural network, and shorten the model training time.
[0138] In this application, a video data recognition system is also provided.
[0139] Refer to Figure 5 , Figure 5 which is a schematic structural diagram of an embodiment of the video data recognition system in this application.
[0140] As Figure 5As shown in the figure, the video data recognition system 400 includes: a first acquisition module 410, a second acquisition module 420, a determination module 430, and an identification module 440. Among them, the first acquisition module 410 divides the video data to be processed to obtain a plurality of video units, where each video unit includes a reference frame and a plurality of prediction frames. The second acquisition module 420 acquires the reference frame features corresponding to the reference frame, and acquires the temporal dynamic features corresponding to each prediction frame, and fuses the reference frame features and the plurality of temporal dynamic features to determine the first fusion feature of each video unit. The determination module 430 uses the first temporal feature, the first motion information, and the first fusion feature corresponding to each video unit to determine a plurality of frame groups and the team features of each frame group, where the frame group includes the first temporal feature, the first motion information, and the first fusion feature corresponding to at least one video unit. The identification module 440 uses the plurality of team features to determine the video fusion feature of the video data to be processed, and then identifies the video fusion feature to obtain the recognition result of the video data to be processed.
[0141] In this application, an electronic device is also provided.
[0142] Refer to Figure 6 , Figure 6 which is a schematic structural diagram of an embodiment of the electronic device in this application. This electronic device can execute the steps of the video data recognition method in the above method.
[0143] The electronic device 500 includes a memory 520, a processor 510 coupled to the memory, and at least one computer program stored in the memory 520 and executable on the processor 510. When the processor 510 loads and executes the at least one computer program, it is used to implement the steps of the video data recognition method in the above method. For the relevant content, please refer to the detailed description in the above method, and details will not be repeated here.
[0144] In this application, a computer-readable storage medium is also included.
[0145] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of an embodiment of the computer-readable storage medium in this application.
[0146] The computer-readable storage medium 600 stores at least one segment of program 610. When the at least one segment of program 610 is loaded and executed by the processor, it is used to implement the steps of the video data recognition method in the above method. For the relevant content, please refer to the detailed description in the above method, and details will not be repeated here.
[0147] In the above solution, by dividing the recognition data into multiple video units, and then using the reference frame features of the reference frames and the temporal dynamic features of the predicted frames in the video units to determine the first fusion feature of each video unit. Further, using the first temporal feature, the first motion information, and the first fusion feature corresponding to at least one video unit to determine the team features of multiple frame teams and the frame teams, and then using the multiple team features to determine the video fusion feature of the video data, and performing recognition on the video fusion feature to obtain the recognition result, the recognition efficiency can be effectively improved, and the recognition accuracy can be improved.
[0148] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling, or communication connection can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.
[0149] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0150] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0151] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0152] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A video data recognition method, characterized in that: include: Dividing the video data to be processed to obtain a plurality of video units, wherein each of the video units includes a reference frame and a plurality of prediction frames; Acquire a reference frame feature corresponding to the reference frame, and acquire a temporal dynamic feature corresponding to each of the predicted frames, and fuse the reference frame feature and a plurality of the temporal dynamic features to determine a first fusion feature of each of the video units; Determine a plurality of frame teams and a team feature of each of the frame teams by using the first timing feature, the first motion information, and the first fusion feature corresponding to each of the video units, wherein the frame team includes the first timing feature, the first motion information, and the first fusion feature corresponding to at least one of the video units, and the first timing feature refers to a time series corresponding to the first fusion feature; The video fusion features of the video data to be processed are determined by using a plurality of the team features, and then the video fusion features are identified to obtain the identification result of the video data to be processed.
2. The method according to claim 1, characterized in that The obtaining of reference frame features corresponding to the reference frame, and obtaining temporal dynamic features corresponding to each of the predicted frames, includes: Acquire reference frame features and first position information corresponding to the reference frame in each of the video units; Acquire a first motion vector of each of the predicted frames, and then determine second position information of each of the predicted frames using the first position information and the first motion vector; The temporal dynamic characteristics are determined using the reference frame characteristics, the second position information and the plurality of predicted frames.
3. The method according to claim 2, characterized in that The step of fusing the reference frame feature with the plurality of temporal dynamic features to determine a first fusion feature of each video unit includes: Acquire a second temporal feature corresponding to a reference frame and each predicted frame of each of the video units; The second temporal feature, the reference frame feature and the temporal dynamic features corresponding to a plurality of the predicted frames are fused to determine the first fusion feature of each of the video units.
4. The method according to claim 1, characterized in that: The determining of a plurality of frame teams and a team feature of each of the frame teams by using the first timing feature, the first motion information and the first fusion feature corresponding to each of the video units includes: Acquire a first timing feature of each of the video units and the reference frame in each of the video units, and acquire the first motion information corresponding to each of the video units; Using the first timing feature and the first fusion feature in at least one of the video units, forming a frame team with a preset length; The first timing feature, the first fusion feature, and the first motion information corresponding to at least one of the video units in each of the frame teams are fused to determine the team feature of each of the frame teams.
5. The method according to claim 4, characterized in that The obtaining the first motion information corresponding to each of the video units includes: Obtain a second motion vector of each macroblock region in the video unit, and then determine the change amount of each macroblock region The first motion information corresponding to the video unit is determined using the change amount.
6. The method according to claim 4, characterized in that The preset length is smaller than the length of the video data to be processed, and the preset length is greater than or equal to the length of the video unit.
7. The method according to claim 1, characterized in that The step of using a plurality of the team features to determine the video fusion features of the video data to be processed, and then identifying the video fusion features to obtain the identification result of the video data to be processed, includes: Performing weighted processing on the plurality of team features to obtain video fusion features of the video data to be processed; Before the inverse quantization operation in the video decoding process, the video fusion feature is identified to obtain an identification result of the video data to be processed.
8. The method according to claim 1, characterized in that The video data to be processed is incompletely decoded video frame data.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor coupled to the memory, the memory stores at least one computer program, and when the at least one computer program is loaded and executed by the processor, it is used to implement the method according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium has at least one program, and when the at least one program is loaded and executed by the processor, it is used to implement the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Dynamic gesture recognition method and system based on video coding data multi-feature fusion
CN113489958A
Depth feature fusion video super-resolution method based on time sequence grouping
CN115760565A