Video processing method and device, computer equipment and storage medium

Through the video processing method of multi-dimensional feature fusion and a two-layer decision-making mechanism, the problems of long training cycles and high cost of AI models are solved, efficient and stable video scene recognition and transcoding are achieved, and resource costs and project cycles are reduced.

CN120281937APending Publication Date: 2025-07-08中央广播电视总台
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510555639.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the prior art, video scene recognition problems with long training cycles, high cost and large output errors.

Method used

By dividing the video to be transcoded into multiple sub-videos, extracting feature values from multiple video dimensions, using a multi-dimensional feature fusion strategy to determine the scene category, and determining the transcoding parameters based on the scene category for transcoding, avoiding relying on NVIDIA hardware, and simplifying content classification annotation and model training.

Benefits of technology

It realizes efficient and stable video content classification and identification, reduces resource cost investment, shortens project cycle, and improves the stability and accuracy of scene classification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281937A_ABST
    Figure CN120281937A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video processing method, a video processing device, computer equipment and a computer storage medium, and relates to the technical field of video processing. The method comprises the following steps: acquiring a to-be-transcoded video, and dividing the to-be-transcoded video to obtain a plurality of sub-videos; for each target sub-video, respectively extracting feature values of the target sub-video from the plurality of video dimensions to obtain a plurality of video feature values; determining a scene category of the target sub-video based on the plurality of video feature values to obtain an initial scene category of each sub-video, and determining a target scene category corresponding to the to-be-transcoded video according to the initial scene category of each sub-video; and determining a scene coding parameter corresponding to the to-be-transcoded video according to the target scene category, and transcoding the to-be-transcoded video based on the scene coding parameter to obtain a transcoded target video. According to the method, high-efficiency and high-accuracy identification of video scene categories can be realized under the condition of low cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of video processing. Specifically, it relates to a video processing method, a video processing device, a computer device, and a computer storage medium. Background Art

[0002] In the current continuous evolution of new media, scene coding technology can optimize transcoding parameters according to the scene to which the content picture belongs, so as to meet the requirement of providing a better picture experience without increasing the bandwidth / traffic cost, or reducing the video bit rate while keeping the picture quality unchanged.

[0003] In related technical solutions, artificial intelligence (AI) technology is usually adopted to train an AI model so that the AI model has different scene recognition capabilities, and then output a higher-quality visual effect. Then, corresponding scene coding parameters are called respectively according to the recognition result to present a higher-quality visual effect.

[0004] However, the training period of the above method for the AI model is long and the cost is high. Moreover, for video content containing a large amount of complex information, due to the difference in data distribution characteristics, the error output by the AI model will also increase significantly, affecting the accuracy of video scene recognition. Summary of the Invention

[0005] In the embodiments of this application, a video processing method, a video processing device, a computer device, and a computer storage medium are provided, which can at least overcome to a certain extent the technical problems of long model training period, high cost, and large error in the output result due to the limitations and defects of related technologies.

[0006] In the first aspect of the embodiments of this application, a video processing method is provided. The method includes: obtaining a video to be transcoded, and dividing the video to be transcoded to obtain multiple sub-videos; for each target sub-video among the multiple sub-videos, extracting eigenvalue of the target sub-video from multiple video dimensions respectively to obtain multiple video eigenvalues; determining the scene category of the target sub-video based on the multiple video eigenvalues to obtain the initial scene category of each sub-video, and determining the target scene category corresponding to the video to be transcoded according to the initial scene category of each sub-video; determining the scene coding parameters corresponding to the video to be transcoded according to the target scene category, and transcoding the video to be transcoded based on the scene coding parameters to obtain the transcoded target video.

[0007] In an optional embodiment of this application, dividing the video to be transcoded to obtain multiple sub-videos includes: obtaining the group of pictures (GOP) interval of the video to be transcoded; and segmenting the video to be transcoded according to the GOP interval to obtain multiple sub-videos.

[0008] In an optional embodiment of the present application, the multiple video dimensions include two or more of the video information complexity, the color complexity of the video picture, and the difference in inter-frame information.

[0009] In an optional embodiment of the present application, the multiple video dimensions at least include the video information complexity. Feature values of the target sub-video are respectively extracted from the multiple video dimensions to obtain multiple video feature values, including: obtaining the number of first video frames included in the target sub-video; for each first sub-video frame in the target sub-video, determining the variance value of the discrete cosine transform (DCT) coefficients of each first sub-video frame; based on the DCT coefficient variance values of each first sub-video frame and the number of first video frames, determining the average value of the DCT coefficient variance value corresponding to the target sub-video to obtain the video feature value of the target sub-video in the video information complexity dimension.

[0010] In an optional embodiment of the present application, the multiple video dimensions at least include the color complexity of the video picture. Feature values of the target sub-video are respectively extracted from the multiple video dimensions to obtain multiple video feature values, including: obtaining multiple color channels corresponding to the target sub-video, and determining the histogram variance value of each color channel; determining the average value of the histogram variance values of the multiple color channels to obtain the color complexity of the target sub-video.

[0011] In an optional embodiment of the present application, the multiple video dimensions at least include the difference in inter-frame information. Feature values of the target sub-video are respectively extracted from the multiple video dimensions to obtain multiple video feature values, including: obtaining the number of second video frames included in the target sub-video; for each second sub-video frame in the target sub-video, determining the Manhattan norm of the motion vector in each second sub-video frame; based on the Manhattan norm of the motion vector in each second sub-video frame and the number of second video frames, determining the average value of the Manhattan norm of the motion vector in the target sub-video to obtain the video feature value of the target sub-video in the difference in inter-frame information dimension.

[0012] In an alternative embodiment of the present application, determining the scene category of the target sub-video based on multiple video feature values includes: determining that the scene category of the target sub-video is the first scene category in response to the video information complexity of the target sub-video being greater than or equal to a first threshold and the color complexity of the video frame being greater than or equal to a second threshold; determining that the scene category of the target sub-video is the second scene category in response to the video information complexity of the target sub-video being greater than or equal to the first threshold and the inter-frame information difference being greater than or equal to a third threshold; determining that the scene category of the target sub-video is the third scene category in response to the video information complexity of the target sub-video being greater than or equal to the first threshold and the inter-frame information difference being less than the third threshold; determining that the scene category of the target sub-video is the fourth scene category in response to the video information complexity of the target sub-video being greater than or equal to the first threshold and the color complexity of the video frame being less than the second threshold, or in response to the video information complexity of the target sub-video being less than the first threshold.

[0013] In an alternative embodiment of the present application, determining the target scene category corresponding to the video to be transcoded according to the initial scene categories of each sub-video includes: counting the occupancy ratios of each scene category, and determining the scene category corresponding to the highest occupancy ratio as the target scene category corresponding to the video to be transcoded.

[0014] The second aspect of the embodiments of the present application provides a video processing apparatus, which includes: a video segmentation module configured to obtain the video to be transcoded and divide the video to be transcoded into multiple sub-videos; a feature extraction module configured to extract the feature values of each target sub-video from multiple video dimensions for each target sub-video among the multiple sub-videos, to obtain multiple video feature values; a scene category determination module configured to determine the scene category of the target sub-video based on the multiple video feature values, to obtain the initial scene categories of each sub-video, and determine the target scene category corresponding to the video to be transcoded according to the initial scene categories of each sub-video; a video transcoding module configured to determine the scene coding parameters corresponding to the video to be transcoded according to the target scene category, and transcoding the video to be transcoded based on the scene coding parameters to obtain the transcoded target video.

[0015] In an alternative embodiment of the present application, the video segmentation module is specifically configured to obtain the group of pictures (GOP) interval of the video to be transcoded; and segment the video to be transcoded according to the GOP interval to obtain multiple sub-videos.

[0016] In an alternative embodiment of the present application, the multiple video dimensions include two or more of video information complexity, color complexity of the video frame, and inter-frame information difference.

[0017] In an optional embodiment of the present application, among multiple video dimensions, at least the video information complexity is included. The feature extraction module is configured to execute the steps of obtaining the number of first video frames included in the target sub-video; for each first sub-video frame in the target sub-video, determining the variance value of the discrete cosine transform (DCT) coefficients of each first sub-video frame; based on the DCT coefficient variance values of each first sub-video frame and the number of first video frames, determining the average value of the DCT coefficient variance values corresponding to the target sub-video, so as to obtain the video feature value of the target sub-video in the video information complexity dimension.

[0018] In an optional embodiment of the present application, among multiple video dimensions, at least the color complexity of the video picture is included. The feature extraction module is configured to execute the steps of obtaining multiple color channels corresponding to the target sub-video, determining the histogram variance value of each color channel; determining the average value of the histogram variance values of multiple color channels, so as to obtain the color complexity of the target sub-video.

[0019] In an optional embodiment of the present application, among multiple video dimensions, at least the inter-frame information difference is included. The feature extraction module is configured to execute the steps of obtaining the number of second video frames included in the target sub-video; for each second sub-video frame in the target sub-video, determining the Manhattan norm of the motion vector in each second sub-video frame; based on the Manhattan norms of the motion vectors in each second sub-video frame and the number of second video frames, determining the average value of the Manhattan norms of the motion vectors in the target sub-video, so as to obtain the video feature value of the target video in the inter-frame information difference dimension.

[0020] In an optional embodiment of the present application, the scene category determination module is configured to execute the step of determining that the scene category of the target sub-video is the first scene category in response to the video information complexity of the target sub-video being greater than or equal to a first threshold and the color complexity of the video picture being greater than or equal to a second threshold; the scene category determination module is configured to execute the step of determining that the scene category of the target sub-video is the second scene category in response to the video information complexity of the target sub-video being greater than or equal to the first threshold and the inter-frame information difference being greater than or equal to a third threshold; the scene category determination module is configured to execute the step of determining that the scene category of the target sub-video is the third scene category in response to the video information complexity of the target sub-video being greater than or equal to the first threshold and the inter-frame information difference being less than the third threshold; the scene category determination module is configured to execute the step of determining that the scene category of the target sub-video is the fourth scene category in response to the video information complexity of the target sub-video being greater than or equal to the first threshold and the color complexity of the video picture being less than the second threshold, or in response to the video information complexity of the target sub-video being less than the first threshold.

[0021] In a third aspect of the embodiments of the present application, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of any one of the above video processing methods are implemented.

[0022] In the fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the video processing method as described in any one of the above are implemented.

[0023] In the fifth aspect of the embodiments of the present application, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps of the video processing method as described in any one of the above are implemented.

[0024] The technical solution of the present application has the following beneficial effects:

[0025] Through this video processing method, by obtaining the video to be transcoded, the video to be transcoded is divided into multiple sub-videos; for each target sub-video among the multiple sub-videos, eigenvalue of the target sub-video is extracted from multiple video dimensions respectively to obtain multiple video eigenvalues; based on the multiple video eigenvalues, the scene category of the target sub-video is determined to obtain the initial scene category of each sub-video, and the target scene category corresponding to the video to be transcoded is determined according to the initial scene category of each sub-video; according to the target scene category, the scene encoding parameter corresponding to the video to be transcoded is determined, so as to transcode the video to be transcoded based on the scene encoding parameter to obtain the transcoded target video.

[0026] The above method divides the video to be transcoded into multiple sub-videos to identify the scene category for each sub-video, and finally determines the target scene category of the video to be transcoded from the overall video. This process realizes a two-layer mechanism from the local to the global of the video to achieve video content classification and recognition. And when performing scene recognition on each video, a multi-dimensional feature fusion strategy is adopted to comprehensively consider the multi-dimensional features of the video content, and then the scene category of each sub-video is determined by fusion. On the one hand, the above process only needs to rely on the general computing power of the computer to complete, without relying on specific hardware of the NVIDIA architecture, effectively avoiding the high purchase, maintenance and upgrade costs caused by using such hardware, and reducing the resource cost input in the transcoding process to a considerable extent. On the other hand, the above method ensures that similar content can obtain a consistent picture transcoding output effect, fundamentally eliminating the inherent uncertainty in the existing AI inference, and greatly improving the stability of the scene classification result. On the other hand, the above method does not need to perform complicated content classification annotation and model training on a large amount of video data, and its business logic can be directly integrated into the transcoding software, significantly simplifying the system construction process, greatly shortening the cycle from project start to final delivery and use, and providing an efficient and fast solution for practical applications. Description of the Drawings

[0027] The accompanying drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0028] Figure 1 Schematic diagram of a scenario recognition method for one of the related technical solutions provided in an embodiment of the present application;

[0029] Figure 2 Architecture diagram of a video processing system provided in an embodiment of the present application;

[0030] Figure 3 Flowchart of a video processing method provided in an embodiment of the present application;

[0031] Figure 4 Flowchart of a method for determining video feature values when the video dimension is the video information complexity provided in an embodiment of the present application;

[0032] Figure 5 Flowchart of a method for determining video feature values when the video dimension is the color complexity of the video content provided in an embodiment of the present application;

[0033] Figure 6 Flowchart of a method for determining video feature values when the video dimension is the inter-frame information difference of the video provided in an embodiment of the present application;

[0034] Figure 7 Flowchart of a complete video processing method provided in an embodiment of the present application;

[0035] Figure 8 Schematic diagram of the structure of a video processing device provided in an embodiment of the present application;

[0036] Figure 9 Schematic diagram of the structure of a computer device provided in an embodiment of the present application. Detailed implementation manners

[0037] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure may be practiced without one or more of the specific details, or may be implemented using other methods, components, devices, steps, etc. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring the various aspects of the present disclosure.

[0038] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0039] The flowcharts shown in the accompanying drawings are only exemplary illustrations and do not necessarily include all the steps. For example, some steps may be further decomposed, while some steps may be combined or partially combined, so the actual execution order may be changed according to the actual situation.

[0040] In the current context of the continuous evolution of new media, watching video content through terminals has become the mainstream video viewing form for the public, achieving coverage of video content in a relatively economical way. Among them, scene coding technology can optimize video coding parameters according to the scene category to which the video content belongs, so as to meet the requirement of providing a better video picture experience without increasing the bandwidth, or reducing the bit rate while ensuring the same picture quality. And an accurate and efficient scene recognition technology is a key step in scene coding.

[0041] With the development of artificial intelligence, AI technology has been introduced into the scene recognition link for video content. Based on this, scene coding can select transcoding templates and optimize parameters according to the recognized scene category, and then output a higher-quality subjective visual effect.

[0042] Figure 1 Schematic diagram of a scene recognition method for one of the related technical solutions provided in an embodiment of the present application; refer to Figure 1As shown in the figure, in this AI-based scene recognition technology, first, the video scene parameters in a large number of video materials are sent to the AI scene recognition engine, enabling the AI to have the recognition ability for different scene categories. For example, the AI is trained to recognize sports-related content through a large number of video materials marked with the "sports" scene, and the AI is trained to recognize variety show-related content through video materials marked with "variety shows". Then, according to the recognition results of the AI scene recognition, the corresponding scene coding parameters (i.e., transcoding parameters) are called respectively to achieve transcoding, and finally the visual effect of the transcoded output is obtained.

[0043] For the above AI-based scene recognition technology, there are mainly the following technical problems:

[0044] 1) Long delivery cycle and low efficiency.

[0045] Since the AI-based scene recognition method depends on the training degree of the AI model, it needs to feed a large number of training materials with video scene labels to the AI in the early stage. The period of material content annotation and AI training is long, so this method requires a certain time investment.

[0046] 2) Insufficient stability and low accuracy.

[0047] The AI-based scene recognition method has inherent uncertainties, which stem from the possibility of the AI outputting different results for the same problem during the reasoning process. Numerous studies have shown that current mainstream deep learning models, such as ResNet and ViT, can achieve a Top-1 accuracy of 85%-90% on the standard ImageNet dataset. However, when facing specific scenes, especially for video content containing a large amount of complex information, due to the differences in data distribution characteristics, the error of the AI inference output will increase significantly.

[0048] 3) High cost.

[0049] In the AI technology system, both the model training link and the reasoning process have high requirements for graphics processing capabilities, and mostly rely on NVIDIA's CUDA-architecture-based graphics cards to achieve efficient computing acceleration. For video content processing tasks, there are also strict requirements for the number of computing cores and video memory of the graphics cards. Therefore, the AI-based scene recognition method will increase the hardware cost of transcoding.

[0050] In view of the above problems, an exemplary embodiment of the present disclosure provides a video processing method, which can be applied to any application scenario for video content scene recognition. This method realizes video content classification through feature fusion from multiple video dimensions and a two-layer decision-making mechanism that combines local video scene recognition with global scene recognition. On the one hand, this process can be completed relying on the general computing power of a computer without relying on specific hardware with NVIDIA architecture, effectively avoiding the high acquisition, maintenance, and upgrade costs associated with using such hardware, and significantly reducing the resource cost input during the transcoding process to a considerable extent. On the other hand, this method ensures that similar content can obtain a consistent picture transcoding output effect, fundamentally eliminating the inherent uncertainty in the AI inference process in the prior art and greatly improving the stability of the classification results. Furthermore, different from the prior art's AI-based scene recognition transcoding system, the embodiment of the present application does not require complex content classification annotation and model training for a large amount of video data. Its business logic can be directly integrated into the transcoding software, significantly simplifying the system setup process and greatly shortening the cycle from project initiation to final delivery and use, providing an efficient and fast solution for practical applications.

[0051] For further understanding, the present disclosure provides a video processing method and apparatus, which can be applied to Figure 2 the system architecture of the exemplary application environment shown.

[0052] As Figure 2 shown, the system architecture 200 may include a terminal device 201, a server 202, and a network 203. The network 203 is used to provide a medium for the communication link between the terminal device 201 and the server 202. The network 203 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc. The terminal device 201 may be, for example, a smart phone, a personal digital assistant (PDA), a laptop computer, a server, a desktop computer, or any other computing device with networking capabilities, but is not limited thereto.

[0053] It should be understood that Figure 2 the numbers of terminal devices, networks, and servers in are merely illustrative. According to the implementation requirements, there may be any number of terminal devices, networks, and servers. For example, the server 106 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms.

[0054] The video processing method provided by the embodiments of the present application can be executed on the server 202. Correspondingly, the video processing device is generally set in the server 202. The video processing method provided by the embodiments of the present application can also be executed on the terminal device 201. Correspondingly, the video processing device can also be set in the terminal device 201. The video processing method provided by the embodiments of the present application can also be partially executed on the server 202 and partially executed on the terminal device. Correspondingly, some modules of the video processing device can be set in the server 202, and some modules can be set in the terminal device 201.

[0055] For example, in an exemplary embodiment, the server 202 can obtain the video to be transcoded, divide the video to be transcoded to obtain multiple sub-videos; then, for each target sub-video among the multiple sub-videos, extract the feature values of the target sub-video from multiple video dimensions to obtain multiple video feature values; determine the scene category of the target sub-video based on the multiple video feature values to obtain the initial scene categories of each sub-video, and determine the target scene category corresponding to the video to be transcoded according to the initial scene categories of each sub-video; finally, determine the scene encoding parameters corresponding to the video to be transcoded according to the target scene category, so as to transcode the video to be transcoded based on the scene encoding parameters to obtain the transcoded target video. The server 202 sends the transcoded target video to the terminal device 201 to play the target video with high quality through the graphical user interface provided by the terminal device 201.

[0056] However, those skilled in the art can easily understand that the above application scenarios are only for illustration, and the present exemplary embodiment is not limited thereto.

[0057] Next, taking the above server 202 as the execution subject and applying the video processing method to the above server 202 as an example for illustration. Figure 3 Schematically show the flowchart of an interface adaptation method in this exemplary embodiment. Please refer to Figure 3 , the interface adaptation method provided by the embodiments of the present application includes the following steps S301-S304:

[0058] S301. Obtain the video to be transcoded, and divide the video to be transcoded to obtain multiple sub-videos.

[0059] S302. For each target sub-video among the multiple sub-videos, extract the feature values of the target sub-video from multiple video dimensions to obtain multiple video feature values.

[0060] S303. Determine the scene category of the target sub-video based on the multiple video feature values to obtain the initial scene categories of each sub-video, and determine the target scene category corresponding to the video to be transcoded according to the initial scene categories of each sub-video.

[0061] S304. Determine the scene encoding parameters corresponding to the video to be transcoded according to the target scene category, and transcoding the video to be transcoded based on the scene encoding parameters to obtain the transcoded target video.

[0062] Based on the above Figure 3 In the provided technical solution, by dividing the video to be transcoded into multiple sub-videos, identifying the scene categories for each sub-video, and finally determining the target scene category of the video to be transcoded from the overall video, this process realizes a two-layer mechanism from local to global of the video to achieve video content classification and recognition. And when performing scene recognition on each video, a multi-dimensional feature fusion strategy is adopted to comprehensively consider the multi-dimensional features of the video content, and then the scene categories of each sub-video are fused and determined. On the one hand, the above process only needs to rely on the general computing power of the computer to complete, without relying on specific hardware of the NVIDIA architecture, effectively avoiding the high purchase, maintenance and upgrade costs caused by using such hardware, and reducing the resource cost input in the transcoding process to a considerable extent. On the other hand, the above method ensures that similar content can obtain a consistent picture transcoding output effect, fundamentally eliminating the inherent uncertainty in the existing technical solutions using AI inference, and greatly improving the stability of the scene classification results. On the other hand, the above method does not need to perform complicated content classification annotation and model training on a large amount of video data, and its business logic can be directly integrated into the transcoding software, significantly simplifying the system construction process and greatly shortening the cycle from project start to final delivery and use, providing an efficient and fast solution for practical applications.

[0063] The following will combine specific embodiments to Figure 3 elaborate in detail on the specific implementation manners of each step in the

[0064] In step S301, obtain the video to be transcoded, and divide the video to be transcoded into multiple sub-videos.

[0065] Among them, the multiple sub-videos are local video segments in the video to be transcoded.

[0066] Exemplarily, after obtaining the video to be transcoded, the video to be transcoded can be segmented to obtain multiple sub-videos. When performing the step of dividing the video to be transcoded into multiple sub-videos, it can be implemented based on the following embodiments:

[0067] In an optional embodiment of the present application, obtain the Group of Pictures (GOP) interval of the video to be transcoded; divide the video to be transcoded according to the GOP interval to obtain multiple sub-videos.

[0068] Among them, a Group of Pictures (GOP) is an important concept in video coding. In digital video processing, a GOP is used to describe a set of consecutive frames. The above video frames can be consecutive or intermittent, but they have a common point logically, that is, certain video characteristics are the same or similar during the encoding and decoding processes.

[0069] In the field of video processing, using GOP segmentation is a common method. Each GOP usually contains an I-frame (key frame) and subsequent P-frames (predicted frames) and B-frames (bi-directionally predicted frames), so that through GOP segmentation, a continuous video to be transcoded can be segmented into multiple independent picture sequences.

[0070] Exemplarily, the GOP interval can be a pre-configured fixed-duration interval to segment the video to be transcoded into multiple sub-video segments of a fixed duration. It can also be determined according to the video format of the video to be transcoded (such as H.264, H.265, etc., because different coding standards may have different GOP structures and parameters), and the embodiments of the present application do not impose any special restrictions on this.

[0071] In the above embodiments, the segmentation of the video to be transcoded is achieved through the GOP interval. From the perspective of efficiency: GOP segmentation can reduce the complexity of video encoding and decoding because the decoder only needs to store and decode the nearest few GOPs instead of the entire video sequence. From the perspective of error recovery: during transmission or storage, if some video frames are damaged, the decoder can resynchronize starting from the nearest I-frame instead of decoding from the beginning of the video. From the perspective of compression efficiency: by using P-frames and B-frames, video data can be effectively compressed.

[0072] In addition to the above embodiments, the number of groups can also be set in advance, and the target video frames can be divided according to the number of groups, or other grouping methods can also be used, and the embodiments of the present application do not impose any special restrictions on this.

[0073] In step S302, for each target sub-video among the multiple sub-videos, eigenvalue of the target sub-video is extracted from multiple video dimensions respectively to obtain multiple video eigenvalues.

[0074] Exemplarily, after the video to be transcoded is segmented into multiple sub-videos, for each target sub-video, feature analysis in multiple video dimensions can be performed to obtain multiple video eigenvalues.

[0075] In an optional embodiment of the present application, the multiple video dimensions include two or more of video information complexity, color complexity of the video picture, and difference in inter-frame information.

[0076] Exemplarily, for each target sub-video, multi-dimensional analysis can be performed from two or more of the video information complexity, the color complexity of the video picture, and the inter-frame information difference, so as to obtain the video feature value under the corresponding video dimension.

[0077] The following will, in combination with specific embodiments, respectively obtain multiple video feature values of the target sub-video from the perspectives of the above three video dimensions in step 302.

[0078] Embodiment 1: At least the video information complexity is included in multiple video dimensions.

[0079] Based on the above Embodiment 1, the steps of extracting the feature values of the target sub-video from multiple video dimensions respectively in step 302 to obtain multiple video feature values can refer to Figure 4 As shown, the process of determining multiple video feature values includes the following steps S401 to S403:

[0080] S401. Obtain the first number of video frames included in the target sub-video.

[0081] S402. For each first sub-video frame in the target sub-video, determine the variance value of the discrete cosine transform (DCT) coefficients of each first sub-video frame.

[0082] S403. Based on the DCT coefficient variance values of each first sub-video frame and the first number of video frames, determine the average value of the DCT coefficient variance values corresponding to the target sub-video, and obtain the video feature value of the target sub-video in the video information complexity dimension.

[0083] Among them, the discrete cosine transform (DCT) coefficients are the representations of an image or a signal in the frequency domain. By converting the data from the spatial domain to the frequency domain, its internal structure can be revealed. And DCT is a transformation method that decomposes a signal into different frequency components. For an image, DCT converts pixel values into a frequency coefficient matrix, which includes a direct current (DC) component and an alternating current (AC) component.

[0084] Exemplarily, for each first sub-video frame in the target sub-video, calculate the variance value of the DCT coefficients of each first sub-video frame, so that it can be calculated based on the DCT coefficient variance values of each first sub-video frame and the first number of video frames of each first sub-video frame, so as to obtain the average value of the DCT coefficient variance values corresponding to the target sub-video of the entire target sub-video, and then obtain the video feature value of the target sub-video in the video information complexity dimension.

[0085] The above calculation process can be based on the following formula (1):

[0086]

[0087] In formula (1), VIC is the video feature value of the target sub-video in the video information complexity dimension, which quantitatively represents the average value of the variance of the DCT coefficients corresponding to the target sub-video of the entire target sub-video; N is the first number of video frames of the target sub-video, that is, the total number of frames of the target sub-video; represents the variance value of the DCT coefficients of the i-th frame, which is used to reflect the spatial complexity, and its value range is 1 to N.

[0088] Through the above method, the video information complexity within the target sub-video segment can be obtained, and the higher the video information complexity, the more information the video carries.

[0089] Embodiment 2: At least one of the multiple video dimensions includes the color complexity of the video picture.

[0090] Based on the above Embodiment 1, the step 302 of extracting the feature values of the target sub-video from multiple video dimensions to obtain multiple video feature values refers to the following Figure 5 As shown, this process may include the following steps S501 to S502:

[0091] S501. Obtain multiple color channels corresponding to the target sub-video, and determine the histogram variance value of each color channel.

[0092] S502. Determine the average value of the histogram variance values of the multiple color channels to obtain the color complexity of the target sub-video.

[0093] Among them, the multiple color channels corresponding to the target sub-video may be the red channel (R), the green channel (G), and the blue channel (B). Of course, they may also be other types of color channels, and the embodiments of the present application do not impose any restrictions on this.

[0094] Exemplarily, taking the multiple color channels as the R, G, and B channels as an example, the histogram variance values of the target sub-video on the R, G, and B channels can be obtained, and the average value of the histogram variance values of the multiple color channels can be solved to obtain the color complexity of the target sub-video.

[0095] The above calculation process can be based on the following formula (2):

[0096]

[0097] In formula (2), CCC is the color complexity of the video picture of the target sub-video, which quantitatively represents the average value of the histogram variance values of the multiple color channels; σ Hist-R is the variance of the red channel color histogram; σ Hist-G is the variance of the green channel color histogram; σ Hist-B is the variance of the blue channel color histogram.

[0098] Through the above method, the color complexity of the entire target sub-video segment can be obtained, and the higher the color complexity of the picture, the more color information the video picture represents.

[0099] Embodiment 3: At least the inter-frame information difference of the video is included in multiple video dimensions.

[0100] In an alternative embodiment of the present application, the steps of separately extracting the feature values of the target sub-video from multiple video dimensions to obtain multiple video feature values are as follows Figure 6 shown, including the following steps S601 - S603:

[0101] S601. Obtain the second number of video frames included in the target sub-video.

[0102] S602. For each second sub-video frame in the target sub-video, determine the Manhattan norm of the motion vector in each second sub-video frame;

[0103] S603. Based on the Manhattan norm of the motion vector in each second sub-video frame and the second number of video frames, determine the average value of the Manhattan norm of the motion vector in the target sub-video, and obtain the video feature value of the target video in the dimension of inter-frame information difference.

[0104] Among them, the motion vector is the vector of the target object in motion in the video frame; the Manhattan norm, also known as the L1 norm, is used to characterize the sum of the absolute values of the vector elements. The second number of video frames is the total number of frames included in the target sub-video, and the second sub-video frame is one of the video sub-frames in the target sub-video.

[0105] Exemplarily, for each second sub-video frame in the target sub-video, the Manhattan norm (L1 norm) of the motion vector in each second sub-video frame can be determined, and the average value of the Manhattan norm of the motion vector in the target sub-video is calculated relative to the total number of video frames (i.e., the second number of video frames) of the target sub-video. Specifically, it can be referred to as shown in the following formula (3):

[0106]

[0107] In formula (3), IFD is the inter-frame information difference of the video picture of the target sub-video, which quantitatively represents the average value of the Manhattan norm of the motion vector in the target sub-video; N is the second number of video frames of the target sub-video, that is, the total number of frames of the target sub-video; ||MV i ||1 is the L1 norm of the motion vector of the i-th frame.

[0108] Exemplarily, the video features of the target sub-video can be collected based on two video dimensions in the above-mentioned Embodiment 1 and Embodiment 2, or the video features of the target sub-video can be collected based on two video dimensions in the above-mentioned Embodiment 1 and Embodiment 3, or the video features of the target sub-video can be collected based on two video dimensions in the above-mentioned Embodiment 2 and Embodiment 3, or the video features of the target sub-video can be collected based on three video dimensions in the above-mentioned Embodiment 1, Embodiment 2 and Embodiment 3.

[0109] Through the above embodiments, the initial scene categories of each sub-video are determined from multiple video dimensions. Especially for video content containing a large amount of complex information, the initial scene categories can be determined more accurately, and it is ensured that similar content can obtain a consistent picture transcoding output effect, fundamentally eliminating the inherent uncertainty in the AI inference process and greatly improving the stability of the classification results.

[0110] In step S303, based on multiple video feature values, the scene category of the target sub-video is determined to obtain the initial scene categories of each sub-video, and the target scene category corresponding to the video to be transcoded is determined according to the initial scene categories of each sub-video.

[0111] Exemplarily, after determining the multiple video feature values corresponding to each target sub-video based on the above embodiments, the scene category of the target sub-video can be determined based on the multiple video feature values, and then the initial scene categories of each divided sub-video can be determined.

[0112] The following will describe the above embodiments of determining the scene category of the target sub-video based on multiple video feature values in conjunction with specific embodiments:

[0113] In an optional embodiment of the present application, in response to the video information complexity of the target sub-video being greater than or equal to the first threshold and the color complexity of the video picture being greater than or equal to the second threshold, it is determined that the scene category of the target sub-video is the first scene category.

[0114] In response to the video information complexity of the target sub-video being greater than or equal to the first threshold and the inter-frame information difference being greater than or equal to the third threshold, it is determined that the scene category of the target sub-video is the second scene category.

[0115] In response to the video information complexity of the target sub-video being greater than or equal to the first threshold and the inter-frame information difference being less than the third threshold, it is determined that the scene category of the target sub-video is the third scene category;

[0116] In response to the video information complexity of the target sub-video being greater than or equal to the first threshold and the color complexity of the video picture being less than the second threshold, or in response to the video information complexity of the target sub-video being less than the first threshold, it is determined that the scene category of the target sub-video is the fourth scene category.

[0117] Exemplarily, the determination of the scene category of each target sub - video can be based on different video dimensions. For example, taking the TV drama scene as an example, as shown in the following formula (4):

[0118]

[0119] In formula (4), other cases refer to cases other than the above three cases. For example, the video information complexity (VIC) of the target sub - video is greater than or equal to the first threshold (T1) and the color complexity (CCC) of the video frame is less than the second threshold (T2), that is, VIC≥T1∧CCC<T2; or, in response to the video information complexity (VIC) of the target sub - video being less than the first threshold (T1), that is, VIC<T1.

[0120] It should be noted that T1, T2, and T3 can be calibrated through experiments and experience (reference thresholds: T1 = 0.7, T2 = 0.5, T3 = 15 pixels / frame).

[0121] Taking the TV drama scene as an example, based on the conditions of formula (4), the scene category of the target sub - video can be determined based on multiple video feature values.

[0122] Furthermore, after determining the initial scene categories of each sub - video, the target scene category corresponding to the video to be transcoded can be determined according to the initial scene categories of each sub - video.

[0123] Exemplarily, the target scene category corresponding to the video to be transcoded can be determined according to the following embodiments:

[0124] In an optional embodiment of the present application, the proportion of each scene category in multiple sub - videos is statistically calculated, and the scene category corresponding to the highest proportion is determined as the target scene category corresponding to the video to be transcoded.

[0125] Exemplarily, assuming that the video to be transcoded is divided into 10 sub - videos, then for the 10 sub - videos, their corresponding initial scene categories can be determined respectively. Then, various initial scene categories in the initial scene categories are statistically calculated, and the scene category corresponding to the highest proportion is determined as the target scene category corresponding to the video to be transcoded. Specifically, it can be referred to the following formula (5):

[0126]

[0127] In formula (5), GlobalType is the target scene category corresponding to the video to be transcoded; argmax Type () function is the maximum - value - taking function; Count type is the proportion of a certain scene category; TotalSegments is the total number of sub - videos into which the video to be transcoded is divided.

[0128] For example, continuing with the TV drama scenario, assume that among 10 sub - videos, the variety show scenarios are 2, the sports scenarios are 3, the TV drama scenarios are 1, and the news interview scenarios are 4. Then, it can be determined that the target scenario category corresponding to the video to be transcoded is the news interview scenario.

[0129] In addition to the above - mentioned embodiments, the importance level of the video content can also be determined, so as to add weight values to each sub - video. For example, the greater the importance level of the video content of a sub - video, the greater the corresponding weight value; conversely, the smaller the importance level of the video content of the sub - video, the smaller the corresponding weight value. Then, based on the scenario category of each sub - video and the corresponding weight values, the target scenario category corresponding to the video to be transcoded is determined.

[0130] For example, assume there are 10 sub - videos, and the weight value of one sub - video is greater than the weight threshold. The scenario category corresponding to this sub - video is determined to be 2. Therefore, among the 10 sub - videos, the variety show scenarios are 2, the sports scenarios are 3, the TV drama scenarios are 1, and the news interview scenarios are 4. Since the weight values of the two sub - videos determined to be sports scenarios are greater than the weight threshold, the number of sports scenarios can be made 6. At this time, it can be determined that the target scenario category corresponding to the video to be transcoded is the sports scenario.

[0131] It can be understood that other methods can also be used to determine the target scenario category corresponding to the video to be transcoded, and the embodiments of this application do not impose any restrictions or exhaustive enumeration on this.

[0132] In step S304, the scene coding parameters corresponding to the video to be transcoded are determined according to the target scene category, so as to transcode the video to be transcoded based on the scene coding parameters to obtain the transcoded target video.

[0133] Exemplarily, after determining the target scene category corresponding to the video to be transcoded based on the above - mentioned embodiments, the scene coding parameters corresponding to the video to be transcoded can be determined according to the target scene category, and then the video to be transcoded is transcoded based on the scene coding parameters to obtain the transcoded target video for playing the target video.

[0134] When performing the step of determining the scene coding parameters corresponding to the video to be transcoded according to the target scene category, the scene coding parameters corresponding to the target scene category can be determined based on the parameter mapping relationship between the scene category and the scene coding parameters, and finally the conversion output is performed according to the scene coding parameters. Or, the transcoding template can also be selected and the parameters optimized according to the identified target scene category to implement transcoding of the video to be transcoded to obtain the transcoded target video.

[0135] The following will refer to Figure 7A detailed description is given of the entire video processing process of the video processing method according to the exemplary embodiments of the present disclosure.

[0136] First step: Input the video to be transcoded. Specifically, the video to be transcoded is pre-loaded, and then the video to be transcoded is fragmented / split to obtain a plurality of sub-videos. In an alternative embodiment, the video to be transcoded can be split according to the Group of Pictures (GOP).

[0137] Second step: Feature extraction: That is, video feature values of each sub-video are extracted from multiple video dimensions; for example, video feature values such as Video Information Complexity (VIC), Color Complexity of Video Content (CCC), and Inter-Frame Difference (IFD) can be extracted.

[0138] Third step: Local classification, that is, the classification determination of the scene category of each sub-video is completed through a classifier to obtain the initial scene category of each sub-video.

[0139] Fourth step: Global determination surface, that is, the classification information of all sub-videos is integrated to determine the target scene category of the entire video to be transcoded.

[0140] Fifth step: Parameter mapping, transcoding parameters are mapped based on the material type.

[0141] Sixth step: Encoding output, that is, finally, conversion output is performed according to the transcoding parameters to obtain the target video.

[0142] Based on Figure 7 the embodiments shown, first, a quantitative analysis of the video content of the video to be transcoded is carried out from three video dimensions, namely information complexity, color complexity, and inter-frame difference. Subsequently, according to the established quantitative data classification rules, the category attribution of local content segments is determined. Finally, by counting the proportion of each local category in the overall content, the scene classification of the overall content is determined. The two-layer decision-making method from local to global shown in the above method fully considers the local detail features and overall distribution characteristics of the video content. Compared with the AI-based scene recognition method, it has higher stability and accuracy. Moreover, the above method does not require complex content classification annotation and model training work on a large amount of video data, and its business logic can be directly integrated into the transcoding software, significantly simplifying the system construction process and greatly shortening the cycle from project start to final delivery and use, providing an efficient and fast solution for practical applications. In addition, since the above method realizes video content classification by means of multi-dimensional feature fusion and a two-layer decision-making mechanism, it can be completed relying on the general computing power of a computer without relying on specific hardware of the NVIDIA architecture, effectively avoiding the high purchase, maintenance, and upgrade costs caused by using such hardware, and significantly reducing the resource cost input in the transcoding process.

[0143] It should be understood that although the steps in the flowchart are shown sequentially in the direction of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the figure may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0144] To implement the above video processing method, please refer to Figure 8 , an embodiment of the present application provides a video processing device. The video processing device 800 may include: a video segmentation module 801, a feature extraction module 802, a scene category determination module 803, and a video transcoding module 804.

[0145] Among them, the video segmentation module 801 is configured to obtain the video to be transcoded and divide the video to be transcoded into multiple sub-videos; the feature extraction module 802 is configured to extract the feature values of each target sub-video from multiple video dimensions for each target sub-video among the multiple sub-videos to obtain multiple video feature values; the scene category determination module 803 is configured to determine the scene category of the target sub-video based on the multiple video feature values to obtain the initial scene category of each sub-video, and determine the target scene category corresponding to the video to be transcoded according to the initial scene category of each sub-video; the video transcoding module 804 is configured to determine the scene coding parameters corresponding to the video to be transcoded according to the target scene category, and transcoding the video to be transcoded based on the scene coding parameters to obtain the transcoded target video.

[0146] In an optional embodiment of the present application, the video segmentation module 801 is specifically configured to obtain the Group of Pictures (GOP) interval of the video to be transcoded; and segment the video to be transcoded according to the GOP interval to obtain multiple sub-videos.

[0147] In an optional embodiment of the present application, the multiple video dimensions include two or more of video information complexity, color complexity of the video picture, and inter-frame information difference.

[0148] In an alternative embodiment of the present application, among multiple video dimensions, at least the video information complexity is included. The feature extraction module 802 is configured to obtain the number of first video frames included in the target sub-video; for each first sub-video frame in the target sub-video, determine the variance value of the discrete cosine transform (DCT) coefficients of each first sub-video frame; based on the DCT coefficient variance values of each first sub-video frame and the number of first video frames, determine the average value of the DCT coefficient variance values corresponding to the target sub-video, and obtain the video feature value of the target sub-video in the video information complexity dimension.

[0149] In an alternative embodiment of the present application, among multiple video dimensions, at least the color complexity of the video picture is included. The feature extraction module 802 is configured to obtain multiple color channels corresponding to the target sub-video, determine the variance value of the histogram of each color channel; determine the average value of the histogram variance values of multiple color channels, and obtain the color complexity of the target sub-video.

[0150] In an alternative embodiment of the present application, among multiple video dimensions, at least the inter-frame information difference is included. The feature extraction module 802 is configured to obtain the number of second video frames included in the target sub-video; for each second sub-video frame in the target sub-video, determine the Manhattan norm of the motion vector of each second sub-video frame; based on the Manhattan norm of the motion vector of each second sub-video frame and the number of second video frames, determine the average value of the Manhattan norm of the motion vectors in the target sub-video, and obtain the video feature value of the target video in the inter-frame information difference dimension.

[0151] In an alternative embodiment of the present application, the scene category determination module 803 is configured to determine that the scene category of the target sub-video is the first scene category in response to the video information complexity of the target sub-video being greater than or equal to the first threshold and the color complexity of the video picture being greater than or equal to the second threshold; the scene category determination module 803 is configured to determine that the scene category of the target sub-video is the second scene category in response to the video information complexity of the target sub-video being greater than or equal to the first threshold and the inter-frame information difference being greater than or equal to the third threshold; the scene category determination module 803 is configured to determine that the scene category of the target sub-video is the third scene category in response to the video information complexity of the target sub-video being greater than or equal to the first threshold and the inter-frame information difference being less than the third threshold; the scene category determination module 803 is configured to determine that the scene category of the target sub-video is the fourth scene category in response to the video information complexity of the target sub-video being greater than or equal to the first threshold and the color complexity of the video picture being less than the second threshold, or in response to the video information complexity of the target sub-video being less than the first threshold.

[0152] For the specific limitations of the above video processing device, reference can be made to the limitations of the video processing method in the foregoing text, which will not be elaborated here. Each module in the above video processing device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.

[0153] In one embodiment, a computer device is provided, and the internal structure diagram of the computer device can be as Figure 9 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a video processing method as described above. It includes: including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements any step in the above video processing method.

[0154] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by the processor, it can implement any step in the above video processing method.

[0155] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0156] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0157] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0158] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0159] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concept. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0160] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.

Claims

1. A video processing method, characterized in that, Including: Obtain the video to be transcoded, and divide the video to be transcoded to obtain multiple sub-videos; For each target sub-video among the multiple sub-videos, extract eigenvalue of the target sub-video from multiple video dimensions respectively to obtain multiple video eigenvalues; Determine the scene category of the target sub-video based on the multiple video eigenvalues to obtain the initial scene categories of each sub-video, and determine the target scene category corresponding to the video to be transcoded according to the initial scene categories of each sub-video; Determine the scene encoding parameter corresponding to the video to be transcoded according to the target scene category, so as to transcode the video to be transcoded based on the scene encoding parameter to obtain the transcoded target video.

2. The method according to claim 1, wherein The step of dividing the video to be transcoded to obtain multiple sub-videos includes: Obtain the group of pictures (GOP) interval of the video to be transcoded; Segment the video to be transcoded according to the GOP interval to obtain multiple sub-videos.

3. The method according to claim 1, wherein The multiple video dimensions include two or more of video information complexity, color complexity of video pictures, and inter-frame information difference.

4. The method according to claim 3, characterized in that, At least the video information complexity is included in the multiple video dimensions. The step of extracting eigenvalue of the target sub-video from multiple video dimensions respectively to obtain multiple video eigenvalues includes: Obtain the first number of video frames included in the target sub-video; For each first sub-video frame in the target sub-video, determine the variance value of the discrete cosine transform (DCT) coefficients of each first sub-video frame; Based on the variance value of the DCT coefficients of each first sub-video frame and the first number of video frames, determine the average value of the DCT coefficient variance value corresponding to the target sub-video to obtain the video eigenvalue of the target sub-video in the dimension of video information complexity.

5. The method according to claim 3, wherein At least the color complexity of video pictures is included in the multiple video dimensions. The step of extracting eigenvalue of the target sub-video from multiple video dimensions respectively to obtain multiple video eigenvalues includes: Obtain multiple color channels corresponding to the target sub-video, and determine the histogram variance value of each color channel; Determine the average value of the histogram variance values of the multiple color channels to obtain the color complexity of the target sub-video.

6. The method according to claim 3, characterized in that, At least the inter-frame information difference is included in the multiple video dimensions. The step of extracting eigenvalue of the target sub-video from multiple video dimensions respectively to obtain multiple video eigenvalues includes: Obtain the second number of video frames included in the target sub-video; For each second sub-video frame in the target sub-video, determine the Manhattan norm of the motion vector in each second sub-video frame; Based on the Manhattan norm of the motion vector in each second sub-video frame and the second number of video frames, determine the average value of the Manhattan norm of the motion vector in the target sub-video to obtain the video eigenvalue of the target video in the dimension of inter-frame information difference.

7. The method according to claim 3, characterized in that The step of determining the scene category of the target sub-video based on the multiple video eigenvalues includes: In response to the video information complexity of the target sub-video being greater than or equal to a first threshold and the color complexity of the video picture being greater than or equal to a second threshold, determine that the scene category of the target sub-video is the first scene category; In response to the video information complexity of the target sub-video being greater than or equal to the first threshold and the inter-frame information difference being greater than or equal to the third threshold, determine that the scene category of the target sub-video is the second scene category; In response to the video information complexity of the target sub-video being greater than or equal to the first threshold and the inter-frame information difference being less than the third threshold, determine that the scene category of the target sub-video is the third scene category; In response to the video information complexity of the target sub-video being greater than or equal to the first threshold and the color complexity of the video picture being less than the second threshold, or in response to the video information complexity of the target sub-video being less than the first threshold, determine that the scene category of the target sub-video is the fourth scene category.

8. A video processing device, characterized in that, Comprising: A video segmentation module configured to execute obtaining a video to be transcoded and dividing the video to be transcoded to obtain a plurality of sub-videos; A feature extraction module configured to execute extracting eigenvalue of each target sub-video from a plurality of video dimensions for each target sub-video among the plurality of sub-videos to obtain a plurality of video eigenvalues; A scene category determination module configured to execute determining the scene category of the target sub-video based on the plurality of video eigenvalues to obtain an initial scene category of each sub-video, and determining a target scene category corresponding to the video to be transcoded according to the initial scene category of each sub-video; A video transcoding module configured to execute determining scene encoding parameters corresponding to the video to be transcoded according to the target scene category, and transcoding the video to be transcoded based on the scene encoding parameters to obtain a transcoded target video.

9. A computer device, comprising: Comprising a memory and a processor, the memory stores a computer program, characterized in that when the processor executes the computer program, the steps of the video processing method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the video processing method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Classified scene video snapshot quality test system and method based on visual large model

    CN120808236A