Video processing method and related equipment
By identifying scene switching points in online videos and configuring personalized picture quality and sound quality parameters, the problem of the inability to adjust the video effect in the existing technology is solved, and the video presentation effect and user experience are improved.
Patent Information
- Application Number
- CN202510459220.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art cannot adjust the video effect from a more fine-grained basis based on video content, resulting in poor video presentation effect.
By identifying the scene switching points in online videos, obtaining video frames of different scene types, and configuring personalized picture quality and sound quality parameters according to the scene type, and inserting scene switching point indication information using custom supplementary enhancement information (SEI) in the H.264 code stream to realize personalized configuration of video frames.
It realizes the audio and video presentation effect of videos on a more fine-grained basis, improving the user's video viewing experience.
Smart Images

Figure CN120390119A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of display devices, and in particular, to a method for video processing and related devices. Background Art
[0002] With the rapid development of technology, the display industry has flourished. Display devices in various forms, such as smart TVs, tablets (pads), virtual reality (VR) / augmented reality (AR), etc., have become increasingly popular, and people's requirements for their sound quality and picture quality have also been continuously improving.
[0003] Currently, related technologies can adjust the picture quality according to the service type or network status parameters of the display device. However, this adjustment method is based on the entire service or network status as the granularity, resulting in poor video presentation effects. Summary of the Invention
[0004] This application provides a method for video processing to solve the problem that the video effect cannot be adjusted from a finer granularity based on the video content.
[0005] In a first aspect, a method for video processing is provided, which is applied to a display device for playing online videos, and includes: obtaining a scene switching point of the online video, where the scene switching point is the moment when video frames of different scene types in the online video are switched; configuring corresponding playback parameters for video frames of different scene types in the online video according to the scene switching point; and playing the online video based on the configured playback parameters.
[0006] Among them, different scene types in the online video may refer to scene contents with relatively large differences in the online video. For example, the content differences between the host's broadcast scene and the scene closely related to the news content in a news video are relatively large, so they belong to two different scene types. In addition, various scenes closely related to the news content can also be divided into different scene types according to their different contents.
[0007] The playback parameters may include picture quality parameters and / or sound quality parameters. Exemplarily, the picture quality parameters include, for example, color depth, color gamut, contrast, brightness, etc.; the sound quality parameters include, for example, volume, surround sound effect, voice clarity enhancement, etc.
[0008] In a possible implementation manner, the picture quality parameters and / or sound quality parameters here are parameters matched with the scene types of the video frames. Different scene types may correspond to different picture quality parameters and / or sound quality parameters. Compared with fixed and single parameters, personalized configuration of corresponding picture quality parameters and / or sound quality parameters for video frames can make the video picture effect and / or sound quality effect more conform to the picture content.
[0009] In a possible implementation, after configuring corresponding picture quality parameters and / or audio quality sampling parameters for video frames of different scene types, the display device can play the video pictures corresponding to each video frame in the online video, and the playback effect of the online video includes that the video pictures of different scene types have picture quality effects and / or audio quality effects that fit the content.
[0010] According to the video processing method provided in this implementation, by identifying the scene switching points in the video and obtaining different types of scenes in the video, and then performing personalized playback parameter configuration on different types of scenes, each scene has a playback effect that fits the scene content better. The method provided in the embodiments of the present application can achieve more fine-grained adjustment of the audio-visual presentation effect of the video, thereby improving the user's video viewing experience.
[0011] In combination with the first aspect, in some implementations of the first aspect, obtaining the scene switching points of the online video includes: receiving the H.264 bitstream corresponding to the online video sent by the cloud server, where the H.264 bitstream includes scene switching point indication information; wherein, the scene switching point indication information is located before the target I-frame or the target P-frame of the H.264 bitstream, and the target I-frame or the target P-frame is a video frame aligned with the scene switching point in the time domain, or the target I-frame or the target P-frame is respectively the I-frame or P-frame closest to the scene switching point in the time domain; obtaining the scene switching points of the online video according to the scene switching point indication information.
[0012] According to the video processing method provided in this implementation, since there is a clear time and sequence relationship between I-frames and P-frames in the video stream, inserting scene switching point indication information before an I-frame or a P-frame can associate this custom data with the time sequence of the video frames. When playing the video, the scene switching point indication information can be synchronized with the corresponding video frame in time. In addition, I-frames and P-frames have their own corresponding formats and coding methods respectively. Inserting custom scene switching point indication information before an I-frame or a P-frame will not damage the structure and coding information of the video frame itself, and when parsing the bitstream, the parsing of the I-frame or P-frame will not be interfered by the custom data.
[0013] In combination with the first aspect, in some implementations of the first aspect, the scene switching point indication information is a custom supplementary enhancement information SEI.
[0014] According to the video processing method provided by this implementation manner, by encoding the video data of an online video into an H.264 bitstream, the data volume can be reduced and the transmission efficiency can be improved; and as supplementary enhancement information, SEI can not only insert indication information of scene switching points additionally in the H.264 bitstream, but also, because the SEI information itself is usually small in data volume, it can effectively avoid significantly increasing the data volume of the video bitstream.
[0015] Combined with the first aspect, in some implementation manners of the first aspect, obtaining the scene switching point of the online video where it is located includes: determining whether the online video is a premiere program; if the online video is a premiere program, performing content recognition on the online video to obtain the scene switching point of the online video; storing the scene switching point of the online video in the video scene database; if the online video is not a premiere program, querying the scene switching point of the stored online video from the video scene database.
[0016] Among them, content recognition may include audio analysis and / or image analysis.
[0017] According to the video processing method provided by this implementation manner, by directly querying the corresponding scene switching point from the corresponding video scene database for a non-premiere video, rather than identifying the scene switching point by means of audio analysis and / or image analysis, the video processing efficiency can be improved and the video processing effect can be optimized.
[0018] Combined with the first aspect, in some implementation manners of the first aspect, performing content recognition on the online video to obtain the scene switching point of the online video specifically includes: obtaining the audio features of the video frames of the online video; calculating the difference value between adjacent video frames according to the audio features; when the difference value is equal to or greater than the first threshold, determining the switching moment of the adjacent video frames as the scene switching point.
[0019] Combined with the first aspect, in some implementation manners of the first aspect, performing content recognition on the online video to obtain the scene switching point of the online video specifically includes: obtaining the image features of the video frames of the online video; calculating the frame similarity between adjacent video frames according to the image features; when the frame similarity is less than the second threshold, determining the switching moment of the adjacent video frames as the scene switching point.
[0020] It can be understood that since video scenes usually match the image frames and audio features with each other, and the image information and audio information under different video scenes are quite different, such as image color, image content, audio energy, audio frequency, etc., therefore, by performing audio analysis and / or image analysis on the online video, the switching points between different scenes can be accurately identified, thereby improving the accuracy of video effect adjustment.
[0021] In combination with the first aspect, in certain implementations of the first aspect, the determination of whether the online video is a premiere program specifically includes: determining whether the online video is digital intelligent television (DTV) data. If the online video is the DTV data, it is determined that the online video is a premiere program; or, determining whether the online video is bound to a set-top box. If the online video is bound to the set-top box, it is determined that the online video is a premiere program.
[0022] In a second aspect, a video processing method is provided, which is applied to a cloud server and includes: performing content recognition on an online video to obtain a scene switching point of the online video, where the scene switching point is the moment when video frames of different scene types in the online video are switched; sending an H.264 bitstream corresponding to the online video to a display device, where the H.264 bitstream includes scene switching point indication information for indicating the scene switching point in the online video, so that the display device configures corresponding playback parameters for video frames of different scene types in the online video according to the scene switching point.
[0023] According to the video processing method provided by this implementation, by the cloud server identifying the scene switching points in the video and obtaining different types of scenes in the video, the video can be divided into finer granularity, which is convenient for the display device to configure personalized playback parameters for different types of scenes, optimize the video effect at a finer granularity, make each scene have a playback effect more fitting the scene content, and improve the user's viewing experience.
[0024] In combination with the second aspect, in certain implementations of the second aspect, the scene switching point indication information is a custom supplementary enhancement information (SEI); the method further includes: inserting the SEI before a target I-frame or a target P-frame in the H.264 bitstream, where the target I-frame or the target P-frame is a video frame aligned with the scene switching point in the time domain, or the target I-frame or the target P-frame is respectively the I-frame or P-frame closest to the scene switching point in the time domain.
[0025] In combination with the second aspect, in certain implementations of the second aspect, the performing audio analysis on the online video to identify the scene switching point of the online video specifically includes: obtaining audio features of video frames of the online video; calculating a difference value between adjacent video frames according to the audio features; when the difference value is equal to or greater than a first threshold, determining the switching moment of the adjacent video frames as the scene switching point.
[0026] In combination with the second aspect, in some implementation manners of the second aspect, the image analysis of the online video to identify the scene switching points of the online video specifically includes: obtaining the image features of the video frames of the online video; calculating the picture similarity between adjacent video frames according to the image features; when the picture similarity is less than a second threshold, determining the switching moment of the adjacent video frames as the scene switching point.
[0027] In combination with the second aspect, in some implementation manners of the second aspect, the method further includes: establishing a video scene database; storing the scene switching points of the online video into the video scene database.
[0028] In combination with the second aspect, in some implementation manners of the second aspect, the method further includes: classifying the online video according to the scene switching points to obtain different scene types.
[0029] In a third aspect, there is provided a device, including: a communication device; a processor; a memory configured to store a computer program, the computer program including instructions which, when executed by the processor, cause the device to execute the method according to any one of the implementation manners in the first aspect or the second aspect above.
[0030] In a fourth aspect, there is provided a communication system, including a cloud server and a display device, the display device being configured to execute the method according to any one of the implementation manners in the first aspect above, and the cloud server being configured to execute the method according to any one of the implementation manners in the second aspect above.
[0031] In a fifth aspect, there is provided a computer-readable storage medium storing computer-executable program instructions, which, when running on a computer, cause the computer or the processor to execute the method according to any one of the implementation manners in the first aspect or the second aspect above.
[0032] It can be seen from the above technical solutions that some embodiments of the present application provide a method for video processing to solve the problem that video parameters cannot be adjusted from a finer granularity based on the scene type. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a schematic diagram of a communication system applicable to the method for video processing provided by an embodiment of the present application.
[0034] Figure 2 It is a schematic diagram of the hardware configuration of the display device provided by an embodiment of the present application.
[0035] Figure 3 It is a schematic diagram of the software configuration of the display device provided by an embodiment of the present application.
[0036] Figure 4 Schematic diagram of a basic data block NALU of an H.264 bitstream provided by an embodiment of the present application.
[0037] Figure 5 Example of combination of H.264 bitstream data blocks provided by an embodiment of the present application.
[0038] Figure 6 Example of a GOP provided by an embodiment of the present application.
[0039] Figure 7A and Figure 7B Schematic diagrams of an I-frame / IDR-frame and a P-frame respectively provided by an embodiment of the present application.
[0040] Figure 8 Schematic flowchart of a video processing method provided by an embodiment of the present application.
[0041] Figure 9A and Figure 9B Schematic diagrams of SEI field formats of custom data respectively provided by an embodiment of the present application.
[0042] Figure 10A and Figure 10B Schematic flowcharts of video processing methods executed on the cloud server side and the display device side respectively provided by an embodiment of the present application.
[0043] Figure 11 Schematic flowchart of another video processing method provided by an embodiment of the present application.
[0044] Figure 12 Schematic flowchart of yet another video processing method provided by an embodiment of the present application. Detailed implementation manners
[0045] Embodiments will be described in detail below, and examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following examples do not represent all embodiments consistent with the present application. They are merely examples of systems and methods consistent with some aspects of the present application detailed in the claims.
[0046] It should be noted that the brief description of terms in the present application is only for facilitating the understanding of the following described embodiments, rather than intending to limit the embodiments of the present application. Unless otherwise specified, these terms should be understood in their ordinary and common meanings.
[0047] In this application, terms such as "first", "second", "third", etc. in the specification, claims, and the above-mentioned drawings are used to distinguish similar or like objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that such terms can be interchanged under appropriate circumstances.
[0048] The terms "comprising" and "having" and any variations thereof are intended to cover but not exclusively include. For example, a product or device comprising a series of components does not necessarily have to be limited to all the components clearly listed, but may include other components not clearly listed or inherent to these products or devices.
[0049] The term "module" refers to any known or later-developed hardware, software, firmware, artificial intelligence, fuzzy logic, or a combination of hardware and / or software code that can perform functions related to that element.
[0050] Figure 1 Schematic diagram of a system applicable to a video processing method provided for the embodiments of this application. The system includes a cloud processor 100 and a display device 200.
[0051] In some embodiments, the cloud server 100 can be a device with multiple functions. For example, taking the online video playback scenario as an example, the cloud server 100 can be a device for storing video content, distributing video data, encoding and format-converting video data, sending video stream data in a specific encoding format to the display device 100, and performing user authentication and authorization, etc. The cloud server 100 can also be used for audio analysis and / or image analysis of videos, identifying the scene types included in the videos and the switching points between scenes. Exemplarily, an artificial intelligence (AI) model can be deployed on the cloud server 100 side, which can use this AI model to extract and calculate the audio features and / or image features of video frames to identify the corresponding scene switching points. In addition, in the embodiments of this application, the cloud server 100 can also classify video scenes and establish a corresponding video scene database to record information such as the scene switching points and scene types of the videos.
[0052] In some embodiments, the display device 200 may refer to a device with screen display and data processing capabilities. For example, the display device 200 may configure relevant parameters (such as picture quality parameters and / or sound quality parameters) for video personalization in different scenarios. The display device 200 may be various types of electronic devices, such as smart TVs, mobile terminals, computers, monitors, advertising screens, wearable devices, VR / AR, etc., and the embodiments of the present application do not limit this. In addition, the display device 200 may provide a broadcast receiving smart TV function, and may also additionally provide a smart network smart TV function with computer support functions, such as network smart TVs, smart TVs, Internet Protocol Smart TVs (IPTV), etc.
[0053] In some embodiments, the display device 200 may communicate with the cloud server 100 through various communication methods. For example, the display device 200 may communicate and connect with the cloud server 100 through a local area network (LAN), a wireless local area network (WLAN), or other types of networks. The embodiments of the present application do not limit this.
[0054] Figure 2 For some embodiments of the present application Figure 1 The hardware configuration block diagram of the display device 200 in
[0055] In some embodiments, the display device 200 may include at least one of a tuner demodulator 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface.
[0056] In some embodiments, the detector 230 is used to collect signals from the external environment or for external interaction. The display 260 includes a display function component for presenting a picture and a driving component for driving image display. The communication device 220 is a component for communicating with external devices or the server 400 according to various communication protocol types.
[0057] In some embodiments, the user may input a user command on the graphical user interface (GUI) displayed on the display 260, and then the user input interface receives the user input command through the graphical user interface (GUI). The audio output device 270 may be the native speaker of the display device 200 or an external audio output device connected to the display device 200. The user input interface 280 can be used to receive instructions from the user input.
[0058] To perform user interaction, in some embodiments, the display device 200 may run an operating system. As Figure 3 shown, in some embodiments, the operating system is divided into four layers, from top to bottom are the Applications layer (referred to as the "application layer" for short), the Application Framework layer (referred to as the "framework layer" for short), the system library layer, and the kernel layer.
[0059] The application layer is used to provide services and interfaces for applications, so that the display device 200 can run applications and interact with users based on the applications.
[0060] The framework layer provides application programming interfaces (APIs) and programming frameworks for applications. Through the API interfaces, applications can access resources in the system and obtain system services during execution.
[0061] The application framework layer includes a view system, managers, content providers, etc.
[0062] The kernel layer is a functional layer between the hardware and software of the display device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, and memory management. For example, as Figure 3 shown, hardware drivers can be configured in the kernel layer, and the drivers included in the kernel layer can be at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor drivers (such as fingerprint sensors, temperature sensors, pressure sensors, etc.), and power drivers, etc.
[0063] It should be noted that the above examples are only simple divisions of the operating system functions and do not limit the specific form of the operating system of the display device 200 in the embodiments of the present application.
[0064] In current video processing technologies, in order to improve the video presentation effect, a method for adjusting the picture quality in a cloud game video scenario is: obtaining the game type of the cloud game logged in by the display device and the network status parameters corresponding to the display device, and then determining the target picture quality parameters matching the display device and the cloud game according to the game type and network status parameters, and then controlling the display device to process and display the game screen of the cloud game using the target picture quality parameters.
[0065] The method of configuring video quality parameters based on the service type and network status corresponding to the video has low resource consumption, but a large adjustment granularity, and the configured video quality parameters are fixed and single. However, in current various video works, the same video often covers rich and diverse scene contents. For videos in different scenes, if they want to achieve the optimal effect, the video quality requirements and / or audio quality requirements may not be the same. For example, news videos generally include the scene of the host's broadcast and various scenes closely related to the news content. The host's broadcast scene usually focuses on clearly showing the image details, so accurate color restoration and uniform and soft light are required; for the scenes related to the news content, real scene restoration is emphasized, and specific details need to be highlighted, and dynamic elements need to be captured smoothly, etc. Another example is that in movie videos, grand scenes require a wide color gamut and high contrast to highlight the gorgeous colors and strong light and dark contrasts, so as to create a shocking visual impact; while for the delicate emotional communication scenes of characters, more attention is paid to the accuracy of color restoration, making the nuances such as skin color and expression more real, and reducing the sharpness to avoid the picture being too rigid. And the single and coarse-grained parameter configuration method cannot make the videos in different scenes present better effects.
[0066] In view of this, the embodiments of the present application provide a video processing method. This method identifies the scene switching points in the video, obtains different types of scenes in the video, and then configures personalized video quality parameters and / or audio quality parameters (or collectively referred to as playback parameters) for different types of scenes, so that each scene has a video quality effect and / or audio quality effect that better fits the scene content. The method provided by the embodiments of the present application can achieve more fine-grained adjustment of the audio-visual presentation effect of the video, thereby improving the audio-visual experience of users when watching videos.
[0067] To better understand the video processing method provided by the embodiments of the present application, the following introduces the terms or concepts that may be involved in the embodiments of this article.
[0068] 1. H.264 bitstream
[0069] H.264 is a standard widely used in video coding, and the H.264 bitstream is the data stream generated after encoding the video according to this standard. H.264 consists of a series of independent network abstraction layer units (NALUs). The NALU data blocks are interrelated and their order cannot be reversed, but data in a specified format can be inserted between the NALU data blocks, and the length of the inserted data is theoretically unlimited. Exemplarily, as Figure 4 shown, it is a schematic diagram of the basic data block NALU of an H.264 bitstream. Using the start code as the delimiter, each NALU contains one byte of header information and payload data. Among them, the header information is, for example, Figure 4The NALU header shown is used to identify the type of NALU. The NALU header may include, for example, the following: supplemental enhancement information (SEI) (identified by 06), video frame type (e.g., 25 / 65 represents an I-frame, 21 / 61 represents a P-frame), sequence parameter set (SPS), picture parameter set (PPS), etc. The NALU payload data is, for example, Figure 4 the NALU body shown.
[0070] Custom data can be inserted into the H.264 bitstream. Specifically, the custom data can be encapsulated into an array in a specific format, such as SEI, and then inserted into the H.264 bitstream. Exemplarily, as Figure 5 shown, it is an example of the combination of H.264 bitstream data blocks. When decoding the H.264 bitstream, the decoder can first traverse the start code (00 00 00 01) in the bitstream, and immediately after finding the start code, parse the NALU unit. During parsing, traverse from front to back until the next start code is found, and then start parsing the NALU unit of the next frame.
[0071] 2. Group of pictures (GOP)
[0072] Generally, an H.264 bitstream contains multiple GOPs, and each GOP contains multiple video coding frames (or video frames), such as P-frames or B-frames. Figure 6 An example of a GOP is shown. Combining Figure 6 shown, the division of the GOP in the H.264 bitstream is that the images between two adjacent key frames (IDR frames) form a GOP, which includes the previous IDR frame, does not include the subsequent IDR frame, and includes all P-frames and B-frames after the first IDR frame. Exemplarily, Figure 6 the GOP image shown contains 5 image coding frames, one IDR frame, two P-frames, and two B-frames. GOPs are further divided into open GOPs and close GOPs. An open GOP means that the P-frames and B-frames in the current GOP can use the images of the previous GOP as reference frames. There is no IDR frame in an open GOP, but there will be I-frames. A close GOP means that the P-frames and B-frames in the current GOP only refer to the I-frames or P-frames within this GOP, and do not refer to the frames in other GOPs, like an independent unit that is "closed" externally.
[0073] Assume Figure 6What is shown is an open GOP. Then, the P-frame numbered 6 can reference the IDR-frame numbered 0 (which should be called an I-frame at this time), or the P-frame numbered 6 can reference the P-frames numbered 1 and 4. Suppose Figure 5 is a closed GOP. Then, the P-frame numbered 6 cannot reference the IDR-frame numbered 0, nor can it reference the P-frames numbered 1 and 4, but it can reference the IDR-frame numbered 5.
[0074] 3. H.264 Encoded Frame Types: IDR-frame, I-frame, P-frame, B-frame
[0075] The H.264 bitstream divides the video sequence into individual video frames (or image frames). The video frames can be I-frames (key frames), P-frames (forward frames), or B-frames (bi-directionally predicted frames).
[0076] An IDR-frame refers to an instantaneous decoding refresh picture, or a key frame. An IDR-frame is essentially also an I-frame. When encoding an IDR-frame, intra-frame encoding technology is used, that is, the encoding and decoding of an IDR-frame do not require reference to other video frames, and only spatial redundancy compression encoding is performed on the image. During the decoding process, if an IDR-frame is encountered, the decoding parameters will be re-parsed and calculated, and the previous decoding information will be cleared. Therefore, IDR-frames can prevent errors in the previous GOP from continuing into the current GOP.
[0077] An I-frame refers to an intra-coded picture. When encoding an I-frame, intra-frame compression encoding technology is used. Intra-frame compression is an encoding method that solves spatial redundancy based on the similarity of adjacent pixels in space (spatial redundancy means that there are similarities or identities between the current coding block / pixel and the surrounding blocks / pixels). For example, common jpeg images are compressed images through intra-frame encoding. An I-frame is not necessarily an IDR-frame (i.e., a key frame), but an IDR-frame is an I-frame. Decoding an I-frame does not require dependence on video frame images in front of or behind it like a P-frame or a B-frame. An I-frame can be decoded independently as a single frame. The biggest difference between an I-frame and an IDR-frame lies in whether the decoding process clears the previous decoding information. Among them, an IDR-frame clears the previous decoding information, while an I-frame does not.
[0078] A P-frame refers to a predicted picture, and usually uses an inter-frame and intra-frame hybrid coding method. Usually, there are some similar contents and some different contents between the current video frame image and the previous video frame image. The corresponding coding method for a P-frame encodes by removing the similar contents and retaining the difference values, which can eliminate the temporal redundancy of the image. Decoding a P-frame requires dependence on a reference frame, and the P-frame can only be decoded after the decoding of its reference frame is completed.
[0079] A B-frame refers to a bidirectionally predicted picture. Similar to a P-frame, a B-frame usually also adopts a hybrid coding method of inter-frame and intra-frame coding. The difference is that a P-frame is encoded based on the difference from the previous video frame image, while a B-frame needs to be encoded based on the difference not only from the previous video frame image but also from the next video frame image. Decoding a B-frame requires referring to the previous frame image and also the next frame image, and the B-frame can only be decoded after the decoding of both the previous and next reference frames is completed.
[0080] Both P-frames and B-frames can also adopt the intra-frame coding mode, but B-frames and P-frames mainly rely on inter-frame coding to improve the compression ratio.
[0081] Figure 7A This is a schematic diagram of an I-frame / IDR-frame. Among them, the shaded squares are the data blocks encoded using intra-frame coding in the I-frame / IDR-frame. Figure 7B This is a schematic diagram of a P-frame. Among them, the shaded squares are the data blocks encoded using intra-frame coding, and the blank squares are the data blocks encoded using inter-frame coding. It can be seen that the I-frame / IDR-frame and the P-frame have different characteristics in the coding method. The I-frame / IDR-frame uses all intra-frame coding, does not rely on reference frames, and is encoded only based on its own picture information; while the P-frame mainly uses inter-frame coding but also contains some data blocks encoded using intra-frame coding.
[0082] The video processing method provided in the embodiments of this application can be applied to the scenario of online video playback. Online video playback in this application can refer to that users watch video content stored on a cloud server through the Internet with the help of display devices such as smart TVs and tablet computers. In the embodiments of this application, video live broadcast can be regarded as a special form of online video playback. The main difference is that the content of online video playback is mostly pre-produced and stored on the cloud server, and the display device needs to obtain the video resources from the cloud server (such as transmitted in the form of an H.264 bitstream) and decode them for playback; while video live broadcast emphasizes real-time performance, and it transmits the events that are happening to the audience in real time through the network, such as live sports events and live e-commerce product promotions. That is to say, video live broadcast can bypass the cloud server. The following will be combined with Figure 8 Embodiment Figure 11 Embodiment will respectively introduce the video processing methods in the scenarios of online video playback and video live broadcast.
[0083] Please refer to Figure 1Taking an example of an intelligent TV playing online videos, for the system shown, video content can be stored in a cloud server in a specific format. To facilitate data transmission and meet the playback requirements of the intelligent TV, the cloud server can perform encoding processing on the video, converting the original video into a more efficient encoding format (H.264 bitstream) to reduce the data volume and improve the transmission efficiency. Then, through the wireless communication link between the cloud and the intelligent TV, the video data is transmitted to the intelligent TV in the form of an H.264 bitstream. The intelligent TV receives the video data, decodes it, and then presents the video to the user.
[0084] In some embodiments, the process of the intelligent TV decoding the H.264 bitstream, for example, includes: (1) Creating a decoder and initializing the decoder. Specifically, a lookup function (such as avcodec_find_decoder) can be used to find the corresponding decoder according to the encoder ID, obtaining the decoder pointer and recording decoder information, such as decoder name, type, and supported features. (2) Creating a decoder context and initializing it. Specifically, by calling the avcodec_alloc_context3 function and passing in the decoder pointer, a decoder context can be created, obtaining the decoder context pointer, which records the stream information during the encoding process. (3) Creating decoder input information. Specifically, a data packet AVpacket can be created through the av_packet_alloc function to store the compressed data. (4) Creating a parser. A parser context can be created through the av_parser_init function to parse the bitstream. (5) Creating decoder output information. For example, an AVFrame can be created by calling the av_frame_alloc function to store the original audio or video data after decoding. (6) Opening the decoder by calling the avcodec_open2 function to initialize the decoder thread and configuration.
[0085] After the decoder is turned on, it is in a working state and can be used to decode the received video data. Exemplarily, the decoder decoding process may include: First, read data from the input stream data and store it in a buffer; Then, use the av_parser_parse2 function to parse the input data in the buffer. This function can extract a frame of data from the input bitstream, store it in an AVPacket, and at the same time return the index number corresponding to the end of a frame of data; After that, input the parsed AVPacket into the decoder for decoding. Use the avcodec_send_packet function to input the data packet into the decoder, and obtain the decoded data through the avcodec_receive_frame function. These data use AVFrame as the output carrier. After decoding is completed, perform subsequent processing on the obtained AVFrame data, such as storing the data as a file, displaying it, etc. Continuously repeat the steps from reading data to processing the decoded data until all data is decoded.
[0086] Exemplarily, such as Figure 8 As shown, it is a schematic flowchart of a video processing method provided by an embodiment of the present application. For ease of understanding, here, the scenario of a display device playing an online video is used as an example for introduction. The display device may be, for example, a smart TV.
[0087] S801, the cloud server performs audio analysis and / or image analysis on the video frame to identify scene transition points.
[0088] It can be understood that video scenes usually match the image content and audio features. The image content and audio features under different scene types usually vary greatly, such as image color, image content, audio energy, audio frequency, etc. Taking audio as an example, in a video scene showing a bustling market, the audio may include the noise of the crowd, the cries of vendors, etc. The audio features corresponding to these sounds have high audio energy, rich frequency components, and audio texture; while in a video scene of a quiet forest, the audio is mainly the sounds of birdsong and the breeze, and the audio features corresponding to these sounds show low audio energy and simple frequency distribution. Therefore, scene transition points can be identified through audio analysis and / or image analysis.
[0089] In some embodiments, the cloud server can perform audio analysis and / or image analysis on the video through an AI model to identify scene switching points. Exemplarily, the process of the cloud server identifying scene switching points through audio analysis includes, for example: the cloud server can first extract the audio features corresponding to the video frame; then, compare the audio features corresponding to the two adjacent frames of video, and determine the video scene switching point based on the comparison results; wherein, if the comparison result indicates that the audio features of the two adjacent frames of video have changed significantly, it is determined that the two adjacent frames of video belong to different scene types, and then the scene switching point can be determined based on the moment when the audio features change. Optionally, if the scene switching point determined by the aforementioned method is not in the same dimension as the video time, the scene switching point can also be mapped to a time point in the video.
[0090] The audio features corresponding to the video frame may include, for example, one or more of audio energy (or volume), zero-crossing rate, audio frequency, Mel-frequency cepstral coefficients (MFCC), harmonic features, rhythm features, timbre, etc., and the embodiments of the present application do not specifically limit this.
[0091] Taking the example of an audio feature being the audio frequency corresponding to a video frame, a method for identifying a video scene switching point is as follows: first, extract the audio frequency feature corresponding to the video frame, for example, by converting the audio signal corresponding to the video frame from the time domain to the frequency domain through a Fourier transform or other method, obtaining the spectrum information of the audio, and then obtaining its frequency feature; then, calculate the difference value between the audio frequency features corresponding to two adjacent frames of video, for example, by calculating the distance value between the audio frequency features of the two adjacent frames of video through a distance function; then, determine whether the two adjacent frames of video belong to different scene types based on the calculated difference value, and if so, determine the video scene switching point based on the moment when the audio feature changes. If the difference value between the audio frequency features corresponding to the two adjacent frames of video is equal to or greater than a first threshold, then it is determined that the two adjacent frames of video belong to different scene types; otherwise, then it is determined that the two adjacent frames of video belong to the same scene type.
[0092] Exemplarily, the process of the cloud server identifying scene switching points through image analysis includes, for example: the cloud server can calculate the image similarity between two adjacent frames of video based on the image features of the video screen; then, determine whether the two adjacent frames of video belong to different scene types based on the image similarity result, and if so, determine the scene switching point based on the moment when the video screen changes. Among them, if the image similarity between two adjacent frames of video is less than a second threshold (such as 50%), it is determined that the two adjacent frames of video belong to different scene types; otherwise, it is determined that the two adjacent frames of video belong to the same scene type. Optionally, if the scene switching point determined in this way is not in the same dimension as the video time, the scene switching point can also be mapped to a time point in the video.
[0093] It should be noted that there are various ways to calculate the similarity between adjacent video frames based on the image features of the video frames. For example, calculating the mean square error, peak signal-to-noise ratio, etc. of the corresponding pixel values of two frames of images; for another example, calculating the similarity between frames by using the color histograms of adjacent video frames; and for yet another example, using a feature point detection algorithm to extract the image feature points of adjacent video frames, and then matching the image feature points, and using the number or ratio of the matched feature points as an index for measuring the similarity between frames. The embodiments of the present application do not limit the specific way of calculating the similarity between video frames.
[0094] S802, classify video scenes based on scene switching points, and establish a video scene database.
[0095] In some embodiments, the cloud server can also identify the scene content. The cloud server can perform scene classification based on scene switching points and scene content, obtain the scene type corresponding to each video frame, and add the corresponding scene type label to the video frame.
[0096] Taking a 30-minute competition video as an example, it includes two types of scenes: host broadcast and competition site. The time information corresponding to each scene is shown in Table 1 below. Through audio analysis and / or image analysis, the scene switching points in this competition video can be identified as: 05:00, 18:50, 20:10, and 29:00. Therefore, the preliminary classification result of the video frames according to the scene switching points can be that the video frames within each time period of 00:00-05:00, 05:00-18:50, 18:50-20:10, 20:10-29:00, and 29:00-30:00 belong to the same scene type. Then, the cloud server can determine the specific scene content corresponding to each type of video frame (such as host broadcast or competition site), and then add the corresponding scene type label to the video frame. Optionally, the video frames with the same scene type label can ultimately be classified into the same category. Among them, the ways for the cloud server to determine the scene type corresponding to the video frame include, for example: image recognition, audio analysis, scene classification model, target detection algorithm, etc., and the embodiments of the present application do not limit this.
[0097] Table 1
[0098] Scene type Host broadcast Competition site Host broadcast Competition site Host broadcast Time 00:00-05:00 05:00-18:50 18:50-20:10 20:10-29:00 29:00-30:00
[0099] In some embodiments, the cloud server can establish a video scene database and store each type of video frame and its corresponding scene type in the database. Exemplarily, the information of a type of video frame and its corresponding scene type stored in the database can be shown in Table 2:
[0100] Table 2
[0101]
[0102] Alternatively, the storage of a video frame and its corresponding scene type information in the database can be as shown in Table 3:
[0103] Table 3
[0104]
[0105] It should be noted that the information types shown in Table 2 and Table 3 are only examples. In practical applications, the video scene database can also store more video frame information or the information corresponding to its scene type. For example, the video scene database can also store the video name, video duration, scene description, video frame image data (such as thumbnail data), etc. The embodiments of the present application do not limit this.
[0106] It should also be noted that the recognition of scene content or the performance of scene classification by the cloud server is an optional operation. In practical applications, the cloud server can also only recognize the scene switching points, and the display device is responsible for performing the tasks of scene content recognition and scene classification.
[0107] S803, generate video stream data, which includes indication information of scene switching points.
[0108] In some embodiments, the video stream can be encoded using the H.264 video coding standard, and the format of the video stream data can be an H.264 bitstream.
[0109] In some embodiments, the indication information of scene switching points can be inserted into the H.264 bitstream. The indication information of scene switching points can be custom data. For example, it can be the information encapsulated in the general format of the SEI field according to the custom data. Among them, the SEI field is an optional H.264 bitstream information, which can be used to carry some auxiliary information, such as scene information, timestamp, etc. The custom data can be encapsulated according to the format of the SEI field, including adding header information, data length, etc.
[0110] In some embodiments, if the cloud server performs scene classification, the indication information of scene switching points can also be used to indicate the scene type. Exemplarily, Figure 9AShows a format of possible scene transition point indication information. The scene transition point indication information can be an SEI field of custom data, and its format includes, for example, a start code, a network abstraction layer (NAL) unit reference indication (NRI), a payload type, a universally unique identifier (UUID), a scene category length, a scene category value, and an end alignment code. Among them, the scene category length is used to indicate the duration of the video scene, and the scene category value can be used to indicate the type of the video scene. According to these two key fields, the display device can configure the picture quality parameters and / or the sound quality parameters for the corresponding scene.
[0111] In some embodiments, the cloud server can also only identify the scene transition point without performing scene classification. In this case, the scene transition point indication information may not indicate the scene type. Exemplarily, Figure 9B Shows another possible SEI field format of custom data. Different from the Figure 9A format, this format is more concise and does not include the scene category length and the scene category value. This information can be used to indicate the scene transition point.
[0112] In some embodiments, the scene transition point indication information can be inserted into the H.264 bitstream at a target position that is temporally aligned with or closest to the scene transition point, so that the indication information matches the scene transition point on the video timeline. The target position can be, for example, before the target I-frame or before the target P-frame in the video stream data, where the target I-frame and the target P-frame refer to the video frames that are temporally aligned with the scene transition point, or the target I-frame or the target P-frame is respectively the I-frame or P-frame that is closest to the scene transition point in terms of time domain.
[0113] Optionally, in some other embodiments, the indication information can also be inserted before each I-frame or P-frame. Among them, if the moment corresponding to the I-frame or P-frame is the scene transition point, the first indication information can be inserted before the I-frame or P-frame, and the first indication information is used to indicate that this moment is the scene transition point of the video; if the moment corresponding to the I-frame or P-frame is not the scene transition point, the second indication information can be inserted, and the second indication information is used to indicate that this moment is not the scene transition point of the video.
[0114] It is understandable that when inserting scene transition point indication information into the H.264 bitstream, inserting the scene transition point indication information at a specific position, such as the gap position between NALUs, or before the I-frame position, or before the P-frame position, can maintain the integrity and correctness of the bitstream and ensure that the inserted data does not damage the structure of the bitstream. Since there is a clear temporal and sequential relationship between I-frames and P-frames in the video stream, inserting scene transition point indication information before an I-frame or a P-frame can associate this custom data with the time sequence of the video frames. When playing the video, the scene transition point indication information can be synchronized with the corresponding video frames in time. In addition, I-frames and P-frames have their respective corresponding formats and encoding methods. Inserting custom scene transition point indication information before an I-frame or a P-frame will not damage the structure and encoding information of the video frame itself. When parsing the bitstream, the parsing of the I-frame or P-frame will not be interfered by the custom data.
[0115] In some embodiments, after inserting scene transition point indication information into the H.264 bitstream, the relevant information of the bitstream, such as data length, number of frames, etc., can be updated. The cloud server can store the updated relevant information of the bitstream locally, or send it to the display device.
[0116] S804. The cloud server sends video stream data to the display device.
[0117] In some embodiments, before the cloud server sends video stream data to the display device, the display device can first send a video play request to the cloud server.
[0118] An exemplary scenario where the display device sends a video play request to the cloud server can be: when the user opens a video application (App) on the display device and selects a video to watch, the display device sends a video play request to the cloud server, which can include a video identifier, and can also include user account information, device information, etc. In response to the video play request sent by the display device, the cloud server sends video stream data (such as an H.264 bitstream) to the display device.
[0119] Optionally, in response to the video play request sent by the display device, the cloud server can also first send a response message to the display device, which can include the basic information of the video. After receiving the response message, the display device can first initialize the local video player according to the basic information of the video.
[0120] S805. The display device configures corresponding picture quality parameters and / or sound quality parameters for different video scenes according to the video stream data.
[0121] In some embodiments, after receiving the video stream data sent by the cloud server, the display device may first decode the video stream data. After decoding the video stream data, the display device may obtain the scene switching points in the video according to the scene switching point indication information in the video stream data. Specifically, the display device may obtain the scene switching points in the video by identifying the SEI data in the video stream data; or, the display device may identify the scene switching points in the video according to the first indication information inserted in the video stream data. The process of the display device decoding the video stream data may refer to the introduction in the above text and will not be elaborated here.
[0122] In some embodiments, the display device may obtain the scene type in the video according to the scene switching point indication information in the video stream data, and then configure corresponding picture quality parameters and / or sound quality parameters for different types of video scenes. For example, when the cloud server can identify the scene type, it may indicate the corresponding scene information, such as the scene type value, to the display device through the video stream data. At this time, the scene switching point indication information may correspond to Figure 9A the information shown; next, the decoder of the display device may send the parsed scene information to the parameter configuration module, so that it sets the corresponding picture quality parameters and / or sound quality parameters according to the scene information.
[0123] Or, the display device may also directly configure the corresponding picture quality parameters according to the mapping relationship between the preset scene type and the picture quality parameters; and / or, the display device may also directly configure the corresponding sound quality parameters according to the mapping relationship between the preset scene type and the sound quality parameters. For example, when the cloud server does not identify the scene type or the identification of the scene type is inaccurate, it may indicate the scene switching point to the display device without indicating the scene type. At this time, the scene switching point indication information may correspond to Figure 9B the information shown; the display device itself may perform the scene type identification task, obtain the scene type in the video, and then configure the corresponding parameters according to the mapping relationship between the scene type and the parameters (picture quality parameters and / or sound quality parameters).
[0124] In some embodiments, the display device may pre-configure corresponding picture quality parameters and / or sound quality parameters for different scene types and store the mapping relationship between the scene type and the picture quality parameters and / or sound quality parameters. Or, the cloud server may pre-configure corresponding picture quality parameters and / or sound quality parameters for different scene types and send the picture quality parameters and / or sound quality parameters corresponding to different scene types to the display device, and the display device receives and stores the mapping relationship between the scene type and the picture quality parameters and / or sound quality parameters.
[0125] According to the video processing method provided by the embodiments of the present application, by identifying scene transition points in a video based on audio analysis and / or image analysis, and inserting custom SEI data before the I-frame or P-frame of the H.264 bitstream to indicate the scene transition points, a display device can personalized configure audio-visual parameters for the pictures in different scenes according to the scene transition points, and can improve the audio quality and picture effect of the video from a finer granularity.
[0126] Continuing with the example of the online video playback scenario, in the video processing method provided by the embodiments of the present application, the main steps executed on the cloud server side are as Figure 10A shown, including: performing audio analysis and / or image analysis (or called shot analysis) on video frames to identify scene transition points; then, inserting custom SEI data at a specific position in the H.264 bitstream corresponding to the video to indicate the scene transition points; and, scene classification can also be performed, and a video scene database can be established. The main steps executed on the display device side are as Figure 10B shown, including: the display device pulls the stream from the cloud server, that is, sends a video playback request to the cloud server, and receives the H.264 bitstream (that is, video stream data) sent by the cloud server; then decodes the video to identify the custom SEI data (that is, scene transition point indication information) in the H.264 bitstream; and then, based on the specific scene obtained according to the scene transition point indication information, configure picture quality parameters and / or audio quality parameters matching the scene type for different video frames.
[0127] It should be noted that the above embodiments take the cloud server performing the operations of identifying scene transition points and scene types as an example. In actual applications, the display device can also perform the operations of identifying scene transition points and / or scene types. For example, in the video live broadcast scenario, the video data may not pass through the cloud server, but is directly sent to the display device in a point-to-point (P2P) manner. At this time, the display device can perform the operations of identifying scene transition points and scene types.
[0128] When the display device performs the operation of identifying scene transition points, the display device can also adjust the specific process of configuring picture quality parameters and / or audio quality parameters for different scenes according to whether the video is a premiere video. Exemplarily, as Figure 11 shown, it is a schematic flowchart of another video processing method provided by the embodiments of the present application. Taking watching a live video through a smart TV as an example, this process can specifically include the following steps:
[0129] S1101, the display device obtains video stream data.
[0130] S1102, determine whether the video is a premiere program.
[0131] In some embodiments, the ways for the display device to determine whether a video is a premiere program may include: determining whether the video data is bound to a set-top box, whether it is data of the digital television (DTV) type, etc.
[0132] If the video is a premiere program (i.e., the judgment result of this step is "yes"), usually there is no corresponding video scene database established for this video yet. Then, step S1103A can be executed next, that is, the display device performs audio analysis and / or image analysis on the video to identify scene transition points. After identifying the scene transition points, a video scene database corresponding to this video can also be established to record information such as the scene transition points corresponding to this video.
[0133] If the video is a non-premiere program (i.e., the judgment result of this step is "no"), usually there is already a video scene database corresponding to this video. At this time, the display device can execute step S1103B, that is, the display device queries the scene transition points corresponding to this video from the established video scene database.
[0134] In the non-premiere scenario, the display device can identify the video name or video identification (ID), and query the matching scene transition points from the video scene database according to the video name or video ID.
[0135] S1104, the display device configures corresponding picture quality parameters and / or sound quality parameters for different video scenes according to the scene transition points.
[0136] It should be noted that step S1104 can be an optional step. In a possible situation, for the video of a non-premiere program, if the video scene database records the scene type of this video and its corresponding picture quality parameters and / or sound quality parameters, then the display device can directly process the corresponding video frames according to the picture quality parameters and / or sound quality parameters recorded in this database.
[0137] Among them, the specific ways of step S1101 and step S1104 can refer to the relevant introductions above and will not be elaborated here.
[0138] According to the video processing method provided by the embodiments of the present application, by directly querying the corresponding scene transition points from the corresponding video scene database for non-premiere videos, instead of identifying the scene transition points by means of audio analysis and / or image analysis, the efficiency of video processing can be improved and the video processing effect can be optimized.
[0139] Exemplarily, Figure 12 is a schematic flowchart of a video processing method provided by the embodiments of the present application. Specifically, it may include the following steps:
[0140] S1201, obtain the scene switching points of the online video, where the scene switching points are the moments when video frames of different scene types in the online video are switched.
[0141] S1202, configure corresponding playback parameters for video frames of different scene types in the online video according to the scene switching points.
[0142] S1203, play the online video based on the configured playback parameters.
[0143] In a possible implementation manner, the playback parameters may include picture quality parameters and / or sound quality parameters. The picture quality parameters include, for example, color depth, color gamut, contrast, brightness, etc.; the sound quality parameters include, for example, volume, surround sound effect, voice clarity enhancement, etc.
[0144] In some embodiments, obtaining the scene switching points of the online video includes: receiving the H.264 bitstream corresponding to the online video sent by the cloud server, where the H.264 bitstream includes scene switching point indication information; wherein, the scene switching point indication information is located before the target I-frame or the target P-frame of the H.264 bitstream, and the target I-frame or the target P-frame is a video frame aligned with the scene switching point in the time domain, or the target I-frame or the target P-frame is respectively the I-frame or P-frame closest to the scene switching point in the time domain; obtaining the scene switching points of the online video according to the scene switching point indication information.
[0145] Since there is a clear time and sequence relationship between I-frames and P-frames in the video stream, inserting scene switching point indication information before an I-frame or a P-frame can associate this custom data with the time sequence of video frames. When playing the video, the scene switching point indication information can be synchronized with the corresponding video frames in time. In addition, I-frames and P-frames have their own corresponding formats and encoding methods. Inserting custom scene switching point indication information before an I-frame or a P-frame will not damage the structure and encoding information of the video frame itself. When parsing the bitstream, the parsing of the I-frame or P-frame will not be interfered by the custom data.
[0146] In some embodiments, the scene switching point indication information is a custom supplementary enhancement information SEI.
[0147] In some embodiments, obtaining the scene switching points of the online video includes: determining whether the online video is a premiere program; if the online video is a premiere program, perform content recognition on the online video to obtain the scene switching points of the online video; store the scene switching points of the online video in the video scene database; if the online video is not a premiere program, query the stored scene switching points of the online video from the video scene database.
[0148] Among them, content recognition may include audio analysis and / or image analysis.
[0149] In some embodiments, the content recognition of the online video to obtain the scene switching point of the online video specifically includes: obtaining the audio features of the video frames of the online video; calculating the difference value between adjacent video frames according to the audio features; when the difference value is equal to or greater than the first threshold, determining the switching moment of the adjacent video frames as the scene switching point.
[0150] In some embodiments, the content recognition of the online video to obtain the scene switching point of the online video specifically includes: obtaining the image features of the video frames of the online video; calculating the similarity between adjacent video frames according to the image features; when the similarity is less than the second threshold, determining the switching moment of the adjacent video frames as the scene switching point.
[0151] In some embodiments, the judgment of whether the online video is a premiere program specifically includes: judging whether the online video is DTV data, wherein if the online video is the DTV data, it is determined that the online video is a premiere program; or, judging whether the online video is bound to the set-top box, wherein if the online video is bound to the set-top box, it is determined that the online video is a premiere program.
[0152] According to the video processing method provided by the embodiments of the present application, by identifying the scene switching points in the video, different types of scenes in the video are obtained, and then personalized picture quality parameters and / or sound quality parameters are configured for different types of scenes, so that each scene has a picture quality effect and / or sound quality effect that is more suitable for the scene content. The method provided by the embodiments of the present application can achieve more fine-grained adjustment of the audio-visual presentation effect of the video, thereby improving the user's video viewing experience.
[0153] The embodiments of the present application also provide a video processing method, which is applied to a cloud server and includes: performing content recognition on an online video to obtain the scene switching point of the online video, where the scene switching point is the moment when video frames of different scene types in the online video are switched; sending the H.264 bitstream corresponding to the online video to a display device, where the H.264 bitstream includes scene switching point indication information, and the scene switching point indication information is used to indicate the scene switching point in the online video, so that the display device configures corresponding playback parameters for video frames of different scene types in the online video according to the scene switching point.
[0154] In some embodiments, the scene transition point indication information is custom supplementary enhancement information (SEI); the method further includes: inserting the SEI before the target I-frame or target P-frame of the H.264 bitstream, where the target I-frame or target P-frame is a video frame that is temporally aligned with the scene transition point, or the target I-frame or target P-frame is respectively the I-frame or P-frame that is closest to the scene transition point in the time domain.
[0155] In some embodiments, the audio analysis of the online video to identify the scene transition point of the online video specifically includes: obtaining the audio features of the video frames of the online video; calculating the difference value between adjacent video frames according to the audio features; when the difference value is equal to or greater than a first threshold, determining the switching moment of the adjacent video frames as the scene transition point.
[0156] In some embodiments, the image analysis of the online video to identify the scene transition point of the online video specifically includes: obtaining the image features of the video frames of the online video; calculating the similarity between adjacent video frames according to the image features; when the similarity is less than a second threshold, determining the switching moment of the adjacent video frames as the scene transition point.
[0157] In some embodiments, the method further includes: establishing a video scene database; storing the scene transition points of the online video in the video scene database.
[0158] In some embodiments, the method further includes: classifying the online video according to the scene transition point to obtain different scene types.
[0159] Based on the same technical concept, an embodiment of the present application further provides a device, including: a communication device; a processor; a memory configured to store a computer program, the computer program including instructions that, when executed by the processor, cause the device to execute the method described in any of the above embodiments.
[0160] Based on the same technical concept, an embodiment of the present application further provides a communication system, including a cloud server and a display device, where the cloud server and the display device are used to execute one or more steps of the functions respectively belonging to the cloud server and the display device in any of the above methods.
[0161] Based on the same technical concept, an embodiment of the present application further provides a chip system, the chip system includes: a processing circuit, a receiving pin, and a transmitting pin; wherein, the receiving pin, the transmitting pin, and the processing circuit communicate with each other through an internal connection path, and the processing circuit executes one or more steps in any of the above methods to control the receiving pin to receive a signal and control the transmitting pin to transmit a signal.
[0162] Based on the same technical concept, an embodiment of the present application further provides a computer-readable storage medium, in which computer-executable program instructions are stored. When the computer-executable program instructions run on a computer, the computer or the processor is caused to execute one or more steps in any of the above methods.
Claims
1. A method for video processing, characterized in that, Applied to a display device, the display device is used to play online videos, including: Obtain the scene switching points of the online video, where the scene switching points are the moments when video frames of different scene types in the online video are switched; Configure corresponding playback parameters for video frames of different scene types in the online video according to the scene switching points; Play the online video based on the configured playback parameters.
2. The method according to claim 1, characterized in that, The obtaining of the scene switching points of the online video includes: Receive the H.264 bitstream corresponding to the online video sent by the cloud server, where the H.264 bitstream includes scene switching point indication information; wherein, The scene switching point indication information is located before the target I-frame or the target P-frame of the H.264 bitstream, and the target I-frame or the target P-frame is a video frame aligned with the scene switching point in the time domain, or the target I-frame or the target P-frame is respectively the I-frame or P-frame closest to the scene switching point in the time domain; Obtain the scene switching points of the online video according to the scene switching point indication information.
3. The method according to claim 2, wherein The scene switching point indication information is a custom supplementary enhancement information SEI.
4. The method according to claim 1, characterized in that, The obtaining of the scene switching points of the online video includes: Judge whether the online video is a premiere program; If the online video is a premiere program, perform content recognition on the online video to obtain the scene switching points of the online video; Store the scene switching points of the online video in the video scene database; If the online video is not a premiere program, query the stored scene switching points of the online video from the video scene database.
5. The method according to claim 4, wherein The performing of content recognition on the online video to obtain the scene switching points of the online video specifically includes: Obtain the audio features of the video frames of the online video; Calculate the difference value between adjacent video frames according to the audio features; When the difference value is equal to or greater than the first threshold, determine the switching moment of the adjacent video frames as the scene switching point.
6. The method according to claim 4 or 5, characterized in that, The performing of content recognition on the online video to obtain the scene switching points of the online video specifically includes: Obtain the image features of the video frames of the online video; Calculate the picture similarity between adjacent video frames according to the image features; When the picture similarity is less than the second threshold, determine the switching moment of the adjacent video frames as the scene switching point.
7. The method according to claim 4, wherein The judging of whether the online video is a premiere program specifically includes: Judge whether the online video is digital intelligent television DTV data, where, If the online video is the DTV data, determine that the online video is a premiere program; or, Judge whether the online video is bound to the set-top box, where, If the online video is bound to the set-top box, determine that the online video is a premiere program.
8. A method for video processing, characterized in that Applied to a cloud server, including: Perform content recognition on an online video to obtain the scene switching points of the online video, where the scene switching points are the moments when video frames of different scene types in the online video are switched; Send the H.264 bitstream corresponding to the online video to the display device, where the H.264 bitstream includes scene change point indication information for indicating the scene change points in the online video, so that the display device configures corresponding playback parameters for video frames of different scene types in the online video according to the scene change points.
9. The method according to claim 8, wherein The scene change point indication information is a custom supplementary enhancement information SEI; the method further includes: Insert the SEI before the target I-frame or the target P-frame of the H.264 bitstream, where the target I-frame or the target P-frame is a video frame aligned with the scene change point in the time domain, or the target I-frame or the target P-frame is respectively the I-frame or P-frame closest to the scene change point in the time domain.
10. A device, characterized in that, Comprising: A communication device; A processor; A memory configured to store a computer program, the computer program including instructions that, when executed by the processor, cause the device to execute the method according to any one of claims 1 to 7 or 8 to 9.