Video data coding method and device, electronic equipment and storage medium

By performing scene detection and multimodal feature extraction on video data, regions of interest are identified and classified, solving the problem of insufficient adaptability in traditional video coding methods and achieving high-quality video coding results.

CN121125992AActive Publication Date: 2025-12-12HEBEI XIONGXIN TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511387755.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2025-12-12
Estimated Expiration
2045-09-26

AI Technical Summary

Technical Problem

Traditional video coding methods struggle to achieve high-quality adaptive coding, failing to intelligently perceive video content and adaptively adjust coding strategies, leading to resource waste or loss of important information. This is especially true when computing resources are limited on edge AI devices, making high-quality coding difficult to achieve.

Method used

By performing scene detection on the target video data, dividing it into sub-target video data, extracting multimodal features, identifying the target region of interest in the video frame, classifying the importance of the features, and finally encoding the video data.

Benefits of technology

It enables differentiated encoding based on video content, improves the adaptability and quality of video encoding, optimizes resource allocation, and enhances the encoding efficiency of edge devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121125992A_ABST
    Figure CN121125992A_ABST
Patent Text Reader

Abstract

The invention provides a video data coding method and device, electronic equipment and a storage medium, and relates to the technical field of video coding. The method comprises the following steps: performing scene detection on target video data, and dividing the target video data into sub-target video data according to a target scene detection result; performing multi-modal feature extraction on each segment of sub-target video data to obtain a sub-target motion feature, a sub-target visual attention feature and a sub-target scene semantic feature; identifying each target region of interest of each video frame in each segment of sub-target video data according to the sub-target motion features and the sub-target visual attention features; according to the semantic features of the sub-target scenes, importance grading is carried out on the corresponding target regions of interest; and performing video data coding according to the importance grading result of each target region of interest of each video frame in each segment of sub-target video data. According to the invention, high-quality adaptive coding based on intelligent perception of video content can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video encoding technology, and in particular to a video data encoding method, apparatus, electronic device, and storage medium. Background Technology

[0002] In today's era of rapid development in digital media, video content has become a primary carrier of information dissemination. From short videos on social media to high-definition live streaming, from online video-on-demand to virtual reality applications, efficient video encoding technology is indispensable.

[0003] However, traditional video coding methods have many limitations, such as: fixed bitrate coding cannot adjust coding resources according to content importance, leading to resource waste or loss of important information; inaccurate background modeling can cause edge inconsistencies and artifacts; or, coding methods based solely on motion detection struggle to distinguish the differences in importance between different moving objects. Furthermore, existing coding schemes, when running on edge AI devices, struggle to achieve high-quality adaptive content coding due to the limited computing resources of these devices.

[0004] Therefore, there is an urgent need for a video coding method that can intelligently perceive video content and adaptively adjust coding strategies. Summary of the Invention

[0005] This invention provides a video data encoding method, apparatus, electronic device, and storage medium to solve the problem that traditional video encoding methods are unable to achieve high-quality adaptive encoding.

[0006] In a first aspect, embodiments of the present invention provide a video data encoding method, including: Scene detection is performed on the target video data to obtain the target scene detection result, and the target video data is divided into sub-target video data according to the target scene detection result; Multimodal feature extraction is performed on each segment of sub-target video data to obtain sub-target motion features, sub-target visual attention features, and sub-target scene semantic features; Based on the sub-target motion features and the sub-target visual attention features, identify the region of interest for each target in each video frame of each sub-target video data segment; The importance of each region of interest in each video frame of each sub-target video data segment is classified according to the semantic features of the sub-target scene. Video data is encoded based on the importance classification results of each target region of interest in each video frame of each sub-target video data segment.

[0007] In one possible implementation, the step of performing scene detection on the target video data to obtain the target scene detection result includes: Visual and audio features are extracted from the target video data to obtain the target's visual and audio features. Scene detection is performed based on the target visual features to obtain a first scene detection result; Scene detection is performed based on the target audio features to obtain a second scene detection result; The first scene detection result is corrected based on the second scene detection result to obtain the target scene detection result.

[0008] In one possible implementation, the first scene detection result is corrected based on the second scene detection result to obtain the target scene detection result, including: Compare whether the detection result of the first scene is consistent with the detection result of the second scene; If the first scene detection result is inconsistent with the second scene detection result, the inconsistent scene detection result in the first scene detection result is recorded as the first suspicious scene detection result, and the inconsistent scene detection result in the second scene detection result is recorded as the second suspicious scene detection result. For each pair of first suspicious scene detection results and second suspicious scene detection results, the inter-frame distance between the first video frame corresponding to the first suspicious scene detection result in the target video data and the second video frame corresponding to the second suspicious scene detection result in the target video data is calculated. In the target video data, the first video frame and the second video frame are extended according to the inter-frame distance to obtain a corrected video. Visual features are extracted from the corrected video to obtain corrected visual features. Scene detection is performed based on the corrected visual features to obtain the corrected scene detection result corresponding to the first suspicious scene detection result. The target scene detection result is obtained based on the corrected scene detection result and the consistent scene detection result in the first scene detection result.

[0009] In one possible implementation, identifying each region of interest (ROI) in each video frame of each sub-target video data segment based on the sub-target motion features and the sub-target visual attention features includes: Based on the motion features of the sub-target, identify each first region of interest in each video frame of each sub-target video data segment; Based on the visual attention features of the sub-target, identify each second region of interest in each video frame of each sub-target video data segment; Based on each of the first region of interest and each of the second region of interest, determine each target region of interest in each video frame of each sub-target video data segment.

[0010] In one possible implementation, determining each target region of interest (ROI) of each video frame in each sub-target video data segment based on each of the first ROI and each of the second ROI includes: The union of each first region of interest and each second region of interest is calculated to determine each target region of interest in each video frame of each sub-target video data segment.

[0011] In one possible implementation, the importance of each target region of interest in each video frame of each sub-target video data segment is classified according to the semantic features of the sub-target scene, including: Obtain the basic importance classification results of each target region of interest in each video frame in each sub-target video data; Calculate the correlation between the semantic features of the sub-target scene and the region of interest of each target in each video frame of each sub-target video data segment; Based on the correlation, the basic importance classification results of each target region of interest in each video frame in each sub-target video data are corrected to obtain the importance classification results of each target region of interest in each video frame in each sub-target video data.

[0012] In one possible implementation, obtaining the basic importance ranking results of each target region of interest in each video frame of each sub-target video data segment includes: Based on the motion features of the sub-target, identify each first region of interest in each video frame of each sub-target video data segment; Based on the visual attention features of the sub-target, identify each second region of interest in each video frame of each sub-target video data segment; Determine whether each target region of interest in each video frame of each segment of target video data contains the overlapping area of ​​each first region of interest and each second region of interest in each video frame of the segment of target video data; If a target region of interest does not contain the overlapping region, then the basic importance classification result of the target region of interest is determined as the first preset value; If a target region of interest contains the overlapping region, then the sum of the proportion of the overlapping region in the target region of interest and the first preset value is used as the basic importance classification result of the target region of interest.

[0013] In a second aspect, embodiments of the present invention provide a video data encoding apparatus, comprising: The first processing module is used to perform scene detection on the target video data, obtain the target scene detection result, and divide the target video data into sub-target video data according to the target scene detection result; The second processing module is used to extract multimodal features for each segment of sub-target video data to obtain sub-target motion features, sub-target visual attention features and sub-target scene semantic features; The third processing module is used to identify the region of interest of each target in each video frame of each segment of sub-target video data based on the sub-target motion features and the sub-target visual attention features; The fourth processing module is used to classify the importance of each target region of interest in each video frame in each segment of sub-target video data according to the semantic features of the sub-target scene; The encoding module is used to encode video data based on the importance classification results of each target region of interest in each video frame of each sub-target video data segment.

[0014] Thirdly, embodiments of the present invention provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect or any possible implementation thereof.

[0015] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect or any possible implementation thereof.

[0016] In this embodiment of the invention, scene detection is first performed on the target video data to obtain target scene detection results. Based on these results, the target video data is divided into sub-target video data. Then, multimodal feature extraction is performed on each sub-target video data segment to obtain sub-target motion features, sub-target visual attention features, and sub-target scene semantic features. Based on the sub-target motion features and sub-target visual attention features, the regions of interest (ROIs) of each video frame in each sub-target video data segment are identified. Based on the sub-target scene semantic features, the importance of each ROI in each video frame of each sub-target video data segment is graded. Finally, video data encoding is performed based on the importance grading results of each ROI in each video frame of each sub-target video data segment. This approach first uses scene detection to distinguish different content within the target video data, then distinguishes the importance of different ROIs within different content based on the sub-target motion features, sub-target visual attention features, and sub-target scene semantic features corresponding to different content, thereby achieving high-quality adaptive encoding based on the importance of different ROIs within different content. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating the implementation of the video data encoding method provided in this embodiment of the invention; Figure 2This is a schematic diagram of the structure of the video data encoding device provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0018] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0019] See Figure 1 The document illustrates a flowchart of the video data encoding method provided in an embodiment of the present invention, which is described in detail below: Step 101: Perform scene detection on the target video data to obtain the target scene detection results, and divide the target video data into sub-target video data based on the target scene detection results.

[0020] In this embodiment, by performing scene detection on the target video data and dividing the target video data into sub-target video data according to the target scene detection results, the target video data can be divided based on video content, so as to facilitate differentiated encoding for different content in the future.

[0021] In one embodiment, step 101 includes: Visual and audio features are extracted from the target video data to obtain the target visual and audio features.

[0022] Scene detection is performed based on the target's visual features to obtain the first scene detection result.

[0023] Scene detection is performed based on the target audio features to obtain the second scene detection result.

[0024] The detection results of the first scene are corrected based on the detection results of the second scene to obtain the detection results of the target scene.

[0025] For example, visual feature extraction of target video data can be performed by calculating the differences in color histograms, texture features, or optical flow (representing motion) between adjacent video frames, or by using a pre-trained neural network model. The specific visual feature extraction method can be determined according to actual needs.

[0026] For example, audio feature extraction from target video data can be achieved by detecting silence intervals (usually indicating topic transitions), or by first converting speech into text and then using a Large Language Model (LLM) to analyze the semantic coherence of the text. The specific audio feature extraction method can also be determined according to actual needs.

[0027] For example, scene detection based on target visual features to obtain a first scene detection result, and scene detection based on target audio features to obtain a second scene detection result, can be achieved by setting corresponding thresholds for the differences in visual features and audio features, respectively. Alternatively, to improve adaptability, the target visual features and target audio features can be analyzed separately using machine learning models or deep learning models.

[0028] In this embodiment, after scene detection is performed based on the target visual features to obtain a first scene detection result, and scene detection is performed based on the target audio features to obtain a second scene detection result, the first scene detection result is corrected based on the second scene detection result, so as to obtain a more accurate target scene detection result.

[0029] In one embodiment, correcting the first scene detection result based on the second scene detection result to obtain the target scene detection result includes: Compare whether the detection results of the first scene are consistent with the detection results of the second scene.

[0030] If the detection results of the first scene and the detection results of the second scene are inconsistent, the inconsistent scene detection results in the first scene detection results shall be recorded as the first suspicious scene detection results, and the inconsistent scene detection results in the second scene detection results shall be recorded as the second suspicious scene detection results.

[0031] For each pair of first suspicious scene detection results and second suspicious scene detection results, calculate the inter-frame distance between the first video frame corresponding to the first suspicious scene detection result in the target video data and the second video frame corresponding to the second suspicious scene detection result in the target video data; in the target video data, extend the first video frame and the second video frame according to the inter-frame distance to obtain the corrected video; extract visual features from the corrected video to obtain corrected visual features; perform scene detection based on the corrected visual features to obtain the corrected scene detection result corresponding to the first suspicious scene detection result.

[0032] The target scene detection result is obtained based on the consistent scene detection results in the corrected scene detection result and the first scene detection result.

[0033] For example, if the first scene detection result includes 5 scenes and the second scene detection result includes 4 scenes, then the first scene detection result and the second scene detection result can be considered inconsistent. Alternatively, if both the first scene detection result and the second scene detection result include 5 scenes, but the third scene in the first scene detection result differs significantly from the third scene in the second scene detection result, then the first scene detection result and the second scene detection result can also be considered inconsistent.

[0034] For example, the difference between the third scene in the first scene detection result and the third scene in the second scene detection result can be determined based on the difference between the video frames corresponding to the third scene in the first scene detection result and the video frames corresponding to the third scene in the second scene detection result. For instance, it can be determined by calculating the frame difference between the video frames corresponding to the third scene in the first scene detection result and the video frames corresponding to the third scene in the second scene detection result, or by calculating the similarity between the video frames corresponding to the third scene in the first scene detection result and the video frames corresponding to the third scene in the second scene detection result.

[0035] For example, when determining the corrected video, assuming the first video frame is the 100th frame in the target video data and the second video frame is the 136th frame in the target video data, the inter-frame distance between the first video frame and the second video frame is 35 frames. Then, by extending the first video frame forward by 35 frames, we get the 65th frame, and by extending the second video frame backward by 35 frames, we get the 171st frame. The corrected video is then frames 65 to 171 in the target video data.

[0036] In this embodiment, after determining that the first scene detection result and the second scene detection result are inconsistent, the first video frame and the second video frame are extended in the target video data according to the inter-frame distance to obtain the corrected video. This allows for a more comprehensive extraction of video frames that may affect the scene detection result from the target video data, thereby more accurately determining the corrected scene detection result corresponding to the first suspicious scene detection result based on the corrected video.

[0037] Step 102: Perform multimodal feature extraction for each sub-target video data segment to obtain sub-target motion features, sub-target visual attention features, and sub-target scene semantic features.

[0038] In this embodiment, motion features (corresponding to sub-target motion features) are powerful cues for attracting attention in videos. They can be obtained primarily by capturing changes between video frames. For example, optical flow can be used to represent the instantaneous motion vector of a pixel, describing the apparent motion from one frame to the next. Then, the scene motion intensity distribution of each segment of sub-target video data can be analyzed using optical flow to obtain sub-target motion features. Alternatively, motion history or motion energy can be used to represent the region of motion over a period of time as sub-target motion features. Or, inter-frame difference can be used to detect motion regions by comparing pixel values ​​of consecutive frames to obtain sub-target motion features.

[0039] In this embodiment, visual attention features (corresponding to sub-target visual attention features) simulate the way the human eye observes the world, that is, it always focuses on the salient parts of the scene first. Therefore, the visual saliency model is used to detect the human eye's attention region in each segment of sub-target video data, obtaining the sub-target visual attention features. Visual saliency can be divided into two types: Bottom-up: Driven by the low-level visual features of the image itself (such as color, brightness, texture, and directional contrast), it is a rapid, pre-attentive, and unconscious process. For example, a single red dot in a sea of ​​green makes the red area prominent.

[0040] Top-down: Driven by high-level cognitive tasks (such as target-specific search and semantic understanding), it is a conscious, task- and context-dependent process. For example, finding a person wearing red in a crowd.

[0041] In this embodiment, the scene semantic features (corresponding to the sub-target scene semantic features) can be obtained by extracting the foreground target and its category information in each frame of each sub-target video data through a semantic segmentation network, and generating a semantic segmentation map.

[0042] In this embodiment, by extracting multimodal features from each segment of sub-target video data, the region of interest and importance classification under different video content can be accurately identified by combining sub-target motion features, sub-target visual attention features, and sub-target scene semantic features, thereby helping to achieve high-quality adaptive encoding of video data.

[0043] Step 103: Identify the region of interest of each target in each video frame of each sub-target video data segment based on the sub-target motion characteristics and sub-target visual attention characteristics.

[0044] In one embodiment, step 103 includes: Identify the first region of interest (ROI) for each video frame in each segment of sub-target video data based on the motion characteristics of the sub-target.

[0045] Identify the second region of interest (ROI) for each video frame in each segment of sub-target video data based on the visual attention features of the sub-target.

[0046] Based on each first region of interest and each second region of interest, determine each target region of interest in each video frame of each sub-target video data segment.

[0047] For example, based on each first region of interest and each second region of interest, the target region of interest for each video frame in each segment of sub-target video data is determined, including: Find the union of each first region of interest and each second region of interest to determine the target region of interest for each video frame in each segment of the target video data.

[0048] In this embodiment, each first region of interest (ROI) in each video frame of each sub-target video data segment is identified based on the sub-target motion characteristics, i.e., the motion salient region of each video frame in each sub-target video data segment is identified based on the sub-target motion characteristics. Each second region of interest (ROI) in each video frame of each sub-target video data segment is identified based on the sub-target visual attention characteristics, i.e., the spatial salient region of each video frame in each sub-target video data segment is identified based on the sub-target visual attention characteristics. Therefore, the salient regions of each video frame in each sub-target video data segment are identified as the target ROI from both the motion and spatial dimensions.

[0049] Step 104: Classify the importance of each target region of interest in each video frame in each segment of sub-target video data according to the semantic features of the sub-target scene.

[0050] In one embodiment, step 104 includes: Obtain the basic importance classification results of each target region of interest in each video frame of each sub-target video data.

[0051] Calculate the correlation between the semantic features of the sub-target scene and the region of interest of each target in each video frame of each sub-target video data segment.

[0052] The basic importance classification results of each target region of interest in each video frame in each sub-target video data are corrected based on the correlation degree, so as to obtain the importance classification results of each target region of interest in each video frame in each sub-target video data.

[0053] In one embodiment, obtaining the basic importance ranking results of each target region of interest in each video frame of each sub-target video data includes: Identify the first region of interest (ROI) for each video frame in each segment of sub-target video data based on the motion characteristics of the sub-target.

[0054] Identify the second region of interest (ROI) for each video frame in each segment of sub-target video data based on the visual attention features of the sub-target.

[0055] Determine whether the region of interest (ROI) of each video frame in each segment of target video data contains the overlapping region of each first ROI and each second ROI of each video frame in that segment of target video data.

[0056] If a target region of interest does not contain overlapping regions, then the basic importance classification result of the target region of interest is determined as the first preset value.

[0057] If a target region of interest contains overlapping regions, the sum of the proportion of the overlapping regions in the target region of interest and a first preset value is used as the basic importance classification result of the target region of interest.

[0058] In this embodiment, the basic importance classification results of each target region of interest in each video frame in each segment of sub-target video data are first obtained based on the sub-target motion features and sub-target visual attention features. Then, the basic importance classification results of each target region of interest in each video frame in each segment of sub-target video data are corrected based on the sub-target scene semantic features. This allows for more accurate importance classification of each target region of interest in each video frame in each segment of sub-target video data from the dimensions of scene semantics, motion, and space, thereby improving both the adaptive capability and quality of video encoding.

[0059] For example, calculating the correlation between the semantic features of the sub-target scene and the region of interest of each target in each video frame of each sub-target video data segment can include: Calculate the cosine similarity and Euclidean distance between the semantic features of the sub-target scene and the semantic features of the scene in each target region of interest in each video frame of each sub-target video data segment. Then, perform convolution extraction on the semantic features of the sub-target scene and the semantic features of the scene in each target region of interest in each video frame of each sub-target video data segment to obtain the local information similarity between the semantic features of the sub-target scene and the semantic features of the scene in each target region of interest in each video frame of each sub-target video data segment.

[0060] The cosine similarity, Euclidean distance, and local information similarity corresponding to each target region of interest are weighted and summed to obtain the correlation between the semantic features of the sub-target scene and each target region of interest in each video frame of each sub-target video data segment.

[0061] In this embodiment, to accurately determine the correlation between the semantic features of the sub-target scene and the regions of interest (ROIs) of each video frame in each sub-target video data segment, multiple perspectives are considered regarding the semantic features of the sub-target scene and the semantic features of the ROIs of each video frame in each sub-target video data segment. The cosine similarity and Euclidean distance between the semantic features of the sub-target scene and the semantic features of the ROIs of each video frame in each sub-target video data segment are calculated. Furthermore, convolution extraction is performed on the semantic features of the sub-target scene and the semantic features of the ROIs of each video frame in each sub-target video data segment to obtain the local information similarity between the semantic features of the sub-target scene and the semantic features of the ROIs of each video frame in each sub-target video data segment. Therefore, based on the cosine similarity, Euclidean distance, and local information similarity corresponding to each ROI, the correlation between the semantic features of the sub-target scene and the ROIs of each video frame in each sub-target video data segment is obtained.

[0062] Step 105: Encode the video data based on the importance classification results of each target region of interest in each video frame of each sub-target video data segment.

[0063] For example, the quantization parameters of each target region of interest can be determined based on the importance classification results of each target region of interest in each video frame in each sub-target video data segment. Based on the quantization parameters of each target region of interest, each target region of interest in each video frame in each sub-target video data segment is encoded to obtain the compressed bitstream.

[0064] This invention first performs scene detection on the target video data to obtain target scene detection results, and then divides the target video data into sub-target video data based on the target scene detection results. Next, multimodal feature extraction is performed on each sub-target video data segment to obtain sub-target motion features, sub-target visual attention features, and sub-target scene semantic features. Based on the sub-target motion features and sub-target visual attention features, each target region of interest in each video frame of each sub-target video data segment is identified. Based on the sub-target scene semantic features, the importance of each target region of interest in each video frame of each sub-target video data segment is graded. Finally, video data encoding is performed based on the importance grading results of each target region of interest in each video frame of each sub-target video data segment. This allows for the initial use of scene detection to distinguish different content within the target video data, followed by the differentiation of the importance of different target regions of interest within different content based on the sub-target motion features, sub-target visual attention features, and sub-target scene semantic features. This enables high-quality adaptive encoding based on the importance of different target regions of interest within different content segments.

[0065] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0066] The following are device embodiments of the present invention. For details not described in detail, please refer to the corresponding method embodiments described above.

[0067] Figure 2 A schematic diagram of the video data encoding apparatus provided in an embodiment of the present invention is shown. For ease of explanation, only the parts related to the embodiment of the present invention are shown, and are described in detail below: like Figure 2 As shown, the video data encoding device includes: The first processing module 21 is used to perform scene detection on the target video data, obtain the target scene detection result, and divide the target video data into sub-target video data according to the target scene detection result.

[0068] The second processing module 22 is used to extract multimodal features for each segment of sub-target video data to obtain sub-target motion features, sub-target visual attention features and sub-target scene semantic features.

[0069] The third processing module 23 is used to identify the region of interest of each target in each video frame of each segment of sub-target video data based on the sub-target motion characteristics and sub-target visual attention characteristics.

[0070] The fourth processing module 24 is used to classify the importance of each target region of interest in each video frame in each segment of sub-target video data according to the semantic features of the sub-target scene.

[0071] The encoding module 25 is used to encode video data based on the importance classification results of each target region of interest in each video frame in each segment of sub-target video data.

[0072] In one possible implementation, the first processing module 21 is specifically used for: Visual and audio features are extracted from the target video data to obtain the target visual and audio features.

[0073] Scene detection is performed based on the target's visual features to obtain the first scene detection result.

[0074] Scene detection is performed based on the target audio features to obtain the second scene detection result.

[0075] The detection results of the first scene are corrected based on the detection results of the second scene to obtain the detection results of the target scene.

[0076] In one possible implementation, the first processing module 21 is specifically used for: Compare whether the detection results of the first scene are consistent with the detection results of the second scene.

[0077] If the detection results of the first scene and the detection results of the second scene are inconsistent, the inconsistent scene detection results in the first scene detection results shall be recorded as the first suspicious scene detection results, and the inconsistent scene detection results in the second scene detection results shall be recorded as the second suspicious scene detection results.

[0078] For each pair of first suspicious scene detection results and second suspicious scene detection results, calculate the inter-frame distance between the first video frame corresponding to the first suspicious scene detection result in the target video data and the second video frame corresponding to the second suspicious scene detection result in the target video data; in the target video data, extend the first video frame and the second video frame according to the inter-frame distance to obtain the corrected video; extract visual features from the corrected video to obtain corrected visual features; perform scene detection based on the corrected visual features to obtain the corrected scene detection result corresponding to the first suspicious scene detection result.

[0079] The target scene detection result is obtained based on the consistent scene detection results in the corrected scene detection result and the first scene detection result.

[0080] In one possible implementation, the third processing module 23 is specifically used for: Identify the first region of interest (ROI) for each video frame in each segment of sub-target video data based on the motion characteristics of the sub-target.

[0081] Identify the second region of interest (ROI) for each video frame in each segment of sub-target video data based on the visual attention features of the sub-target.

[0082] Based on each first region of interest and each second region of interest, determine each target region of interest in each video frame of each sub-target video data segment.

[0083] In one possible implementation, the third processing module 23 is specifically used for: Find the union of each first region of interest and each second region of interest to determine the target region of interest for each video frame in each segment of the target video data.

[0084] In one possible implementation, the fourth processing module 24 is specifically used for: Obtain the basic importance classification results of each target region of interest in each video frame of each sub-target video data.

[0085] Calculate the correlation between the semantic features of the sub-target scene and the region of interest of each target in each video frame of each sub-target video data segment.

[0086] The basic importance classification results of each target region of interest in each video frame in each sub-target video data are corrected based on the correlation degree, so as to obtain the importance classification results of each target region of interest in each video frame in each sub-target video data.

[0087] In one possible implementation, the fourth processing module 24 is specifically used for: Identify the first region of interest (ROI) for each video frame in each segment of sub-target video data based on the motion characteristics of the sub-target.

[0088] Identify the second region of interest (ROI) for each video frame in each segment of sub-target video data based on the visual attention features of the sub-target.

[0089] Determine whether the region of interest (ROI) of each video frame in each segment of target video data contains the overlapping region of each first ROI and each second ROI of each video frame in that segment of target video data.

[0090] If a target region of interest does not contain overlapping regions, then the basic importance classification result of the target region of interest is determined as the first preset value.

[0091] If a target region of interest contains overlapping regions, the sum of the proportion of the overlapping regions in the target region of interest and a first preset value is used as the basic importance classification result of the target region of interest.

[0092] Figure 3 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. For example... Figure 3 As shown, the electronic device 3 of this embodiment includes a processor 30 and a memory 31. The memory 31 stores a computer program 32. When the processor 30 executes the computer program 32, it implements the steps in the various method embodiments described above. Alternatively, when the processor 30 executes the computer program 32, it implements the functions of each module / unit in the various device embodiments described above.

[0093] For example, computer program 32 may be divided into one or more modules / units, which are stored in memory 31 and executed by processor 30 to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of computer program 32 in electronic device 3.

[0094] Electronic device 3 may include, but is not limited to, processor 30 and memory 31. Those skilled in the art will understand that... Figure 3 This is merely an example of electronic device 3 and does not constitute a limitation on electronic device 3. It may include more or fewer components than shown, or combine certain components, or different components. For example, electronic device 3 may also include input / output devices, network access devices, buses, etc.

[0095] The processor 30 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0096] The memory 31 can be an internal storage unit of the electronic device 3, such as a hard disk or memory of the electronic device 3. The memory 31 can also be an external storage device of the electronic device 3, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device 3. Furthermore, the memory 31 can include both internal and external storage units of the electronic device 3. The memory 31 is used to store the computer program 32 and other programs and data required by the electronic device 3. The memory 31 can also be used to temporarily store data that has been output or will be output.

[0097] For the sake of simplicity and clarity, only the above-described functional modules / units are used as examples. In practical applications, the functions described above can be assigned to different functional modules / units as needed. These modules / units can be implemented in hardware, software, or a combination of both.

[0098] This invention also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the methods described in the above-described method embodiments.

[0099] Computer programs include computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. Computer-readable media can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.

[0100] In the above embodiments, the descriptions of each embodiment have their own emphasis. Parts not detailed or described in a particular embodiment can be referred to in the relevant descriptions of other embodiments. Unless otherwise specified or in conflict with logic, the terminology and / or descriptions between different embodiments are consistent and can be referenced interchangeably. Technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships.

[0101] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A video data encoding method, characterized in that, include: Scene detection is performed on the target video data to obtain the target scene detection result, and the target video data is divided into sub-target video data according to the target scene detection result; Multimodal feature extraction is performed on each segment of sub-target video data to obtain sub-target motion features, sub-target visual attention features, and sub-target scene semantic features; Based on the sub-target motion features and the sub-target visual attention features, identify the region of interest for each target in each video frame of each sub-target video data segment; The importance of each region of interest in each video frame of each sub-target video data segment is classified according to the semantic features of the sub-target scene. Video data is encoded based on the importance classification results of each target region of interest in each video frame of each sub-target video data segment.

2. The video data encoding method according to claim 1, characterized in that, The process of performing scene detection on the target video data to obtain the target scene detection result includes: Visual and audio features are extracted from the target video data to obtain the target's visual and audio features. Scene detection is performed based on the target visual features to obtain a first scene detection result; Scene detection is performed based on the target audio features to obtain a second scene detection result; The first scene detection result is corrected based on the second scene detection result to obtain the target scene detection result.

3. The video data encoding method according to claim 2, characterized in that, The first scene detection result is corrected based on the second scene detection result to obtain the target scene detection result, including: Compare whether the detection result of the first scene is consistent with the detection result of the second scene; If the first scene detection result is inconsistent with the second scene detection result, the inconsistent scene detection result in the first scene detection result is recorded as the first suspicious scene detection result, and the inconsistent scene detection result in the second scene detection result is recorded as the second suspicious scene detection result. For each pair of first suspicious scene detection results and second suspicious scene detection results, the inter-frame distance between the first video frame corresponding to the first suspicious scene detection result in the target video data and the second video frame corresponding to the second suspicious scene detection result in the target video data is calculated. In the target video data, the first video frame and the second video frame are extended according to the inter-frame distance to obtain a corrected video. Visual features are extracted from the corrected video to obtain corrected visual features. Scene detection is performed based on the corrected visual features to obtain the corrected scene detection result corresponding to the first suspicious scene detection result. The target scene detection result is obtained based on the corrected scene detection result and the consistent scene detection result in the first scene detection result.

4. The video data encoding method according to claim 1, characterized in that, Identify the region of interest (ROI) for each video frame in each segment of sub-target video data based on the sub-target motion features and the sub-target visual attention features, including: Based on the motion features of the sub-target, identify each first region of interest in each video frame of each sub-target video data segment; Based on the visual attention features of the sub-target, identify each second region of interest in each video frame of each sub-target video data segment; Based on each of the first region of interest and each of the second region of interest, determine each target region of interest in each video frame of each sub-target video data segment.

5. The video data encoding method according to claim 4, characterized in that, Based on each of the first region of interest and each of the second region of interest, determine each target region of interest in each video frame of each sub-target video data segment, including: The union of each first region of interest and each second region of interest is calculated to determine each target region of interest in each video frame of each sub-target video data segment.

6. The video data encoding method according to claim 1, characterized in that, Based on the semantic features of the sub-target scene, the importance of each target region of interest in each video frame of each sub-target video data segment is classified, including: Obtain the basic importance classification results of each target region of interest in each video frame in each sub-target video data; Calculate the correlation between the semantic features of the sub-target scene and the region of interest of each target in each video frame of each sub-target video data segment; Based on the correlation, the basic importance classification results of each target region of interest in each video frame in each sub-target video data are corrected to obtain the importance classification results of each target region of interest in each video frame in each sub-target video data.

7. The video data encoding method according to claim 6, characterized in that, The basic importance ranking results of each target region of interest in each video frame of each sub-target video data segment include: Based on the motion features of the sub-target, identify each first region of interest in each video frame of each sub-target video data segment; Based on the visual attention features of the sub-target, identify each second region of interest in each video frame of each sub-target video data segment; Determine whether each target region of interest in each video frame of each segment of target video data contains the overlapping area of ​​each first region of interest and each second region of interest in each video frame of the segment of target video data; If a target region of interest does not contain the overlapping region, then the basic importance classification result of the target region of interest is determined as the first preset value; If a target region of interest contains the overlapping region, then the sum of the proportion of the overlapping region in the target region of interest and the first preset value is used as the basic importance classification result of the target region of interest.

8. A video data encoding device, characterized in that, include: The first processing module is used to perform scene detection on the target video data, obtain the target scene detection result, and divide the target video data into sub-target video data according to the target scene detection result; The second processing module is used to extract multimodal features for each segment of sub-target video data to obtain sub-target motion features, sub-target visual attention features and sub-target scene semantic features; The third processing module is used to identify the region of interest of each target in each video frame of each segment of sub-target video data based on the sub-target motion features and the sub-target visual attention features; The fourth processing module is used to classify the importance of each target region of interest in each video frame in each segment of sub-target video data according to the semantic features of the sub-target scene; The encoding module is used to encode video data based on the importance classification results of each target region of interest in each video frame of each sub-target video data segment.

9. An electronic device, characterized in that, It includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method as described in any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • ROI-based video coding method and system and video transmission and coding system

    CN111447449A

  • Video abstract generation method and device, equipment and storage medium

    CN120236231A

  • Video target identification method and device based on artificial intelligence, and storage medium

    CN120431318A

  • Complex scene understanding method and system based on mixed attention dynamic feedback adjustment

    CN120472460A

  • Content summarization leveraging systems and processes for key moment identification and extraction

    US20200372066A1