Video processing method and apparatus, and storage medium and electronic device
By performing layered frame loss processing on video data and optimizing video transmission according to the client's network status and decoding capabilities, the problems of video playback freezes and decoding delays are solved, achieving more efficient video data transmission and playback effects.
Patent Information
- Application Number
- PCT/CN2025/072846
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-07
- Filing Date
- 2025-01-16
- Publication Date
- 2025-10-16
AI Technical Summary
When network bandwidth is limited and client device hardware performance varies, playback freezes and decoding delays may occur during video data transmission, affecting video playback quality.
SVC encoding technology is used to perform layered frame loss processing on the original video data. The frame loss strategy is determined according to the client's network status and decoding capabilities, and video transmission is optimized to meet playback requirements.
It improves the compatibility between network bandwidth and client decoding capabilities during video transmission, reduces playback pauses, and improves the smoothness and quality of video playback.
Smart Images

Figure CN2025072846_16102025_PF_FP_ABST
Abstract
Description
Video processing method and device, storage medium and electronic device
[0001] This application claims priority to Chinese Patent Application No. 202410413136.9, filed on April 7, 2024, the disclosure of which is incorporated herein in its entirety as part of the present application. TECHNICAL FIELD
[0002] Embodiments of the present disclosure relate to a video processing method and device, a storage medium and an electronic device. BACKGROUND
[0003] With the continuous development of Internet technology and mobile devices, the consumption of network multimedia content is growing rapidly. Among them, the demand for real-time audio and video data, such as video conferencing, live streaming and other services, is increasing, and is used more and more frequently in work and life.
[0004] Most network environments have limited bandwidth and large fluctuations. At the same time, due to differences in hardware performance, client devices also have differences in video stream decoding capabilities. The above problems lead to problems such as playback lag and decoding delay caused by network congestion during video data transmission, affecting the video playback effect of the client. SUMMARY
[0005] The present disclosure provides a video processing method and device, a storage medium and an electronic device to implement frame dropping processing on video data, taking into account the playback quality and smoothness requirements of the client.
[0006] In a first aspect, embodiments of the present disclosure provide a video processing method, comprising:
[0007] receiving a video request of a client and determining a request parameter corresponding to the video request;
[0008] obtaining original video data corresponding to the video request and encoding information of the original video data;
[0009] determining a frame dropping strategy based on the encoding information of the original video data and the request parameter, performing frame dropping processing on the original video data based on the frame dropping strategy to obtain target video data, and pushing the target video data to the client.
[0010] In a second aspect, embodiments of the present disclosure also provide a video processing device, comprising:
[0011] a request receiving module configured to receive a video request of a client and determine a request parameter corresponding to the video request;
[0012] an information obtaining module, configured to obtain original video data corresponding to the video request and encoding information of the original video data;
[0013] a frame dropping processing module, configured to determine a frame dropping strategy based on the encoding information of the original video data and the request parameter, perform frame dropping processing on the original video data based on the frame dropping strategy to obtain target video data, and push the target video data to the client.
[0014] In a third aspect, the embodiments of the present disclosure further provide an electronic device, which comprises:
[0015] one or more processors;
[0016] a storage device configured to store one or more programs,
[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement a video processing method provided by the embodiments of the present disclosure.
[0018] In a fourth aspect, the embodiments of the present disclosure further provide a storage medium containing computer executable instructions, which, when executed by a computer processor, are used to perform a video processing method provided by the embodiments of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0019] The above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description when taken in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals are used to represent the same or similar elements. It should be understood that the drawings are schematic, and the original and elements are not necessarily drawn according to the scale.
[0020] FIG. 1 is a schematic diagram of video data obtained by SVC encoding according to an embodiment of the present disclosure;
[0021] FIG. 2 is a schematic diagram of a video processing method according to an embodiment of the present disclosure;
[0022] FIG. 3 is a schematic diagram of a video processing device according to an embodiment of the present disclosure; and
[0023] FIG. 4 is a schematic diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0024] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein; rather, these embodiments are provided so as to more completely and thoroughly understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for illustrative purposes and are not intended to limit the scope of protection of the present disclosure.
[0025] It should be understood that each of the steps recited in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0026] The term "comprising" and variations thereof as used herein are open-ended, that is "including but not limited to". The term "based on" is "based, at least in part, on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related terms are defined in the following description.
[0027] It should be noted that the terms "first", "second", and the like in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.
[0028] It should be noted that the terms "one", "multiple" in the present disclosure are illustrative and not limiting, and those skilled in the art should understand that unless otherwise explicitly stated in the context, it should be understood as "one or more".
[0029] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0030] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type, scope of use, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained in a proper manner according to relevant laws and regulations.
[0031] For example, when responding to the active request of the user, the user is sent a prompt message to explicitly prompt the user that the operation requested to be performed will require the acquisition and use of the personal information of the user. Thus, the user can voluntarily choose whether to provide personal information to the software or hardware, such as electronic devices, application programs, servers or storage media, etc. that perform the operation of the technical solutions of the present disclosure according to the prompt information.
[0032] As an optional but non-limiting implementation, in response to receiving the active request of the user, the manner of sending the prompt information to the user may be, for example, a pop-up window manner in which the prompt information may be presented in a text manner. In addition, the pop-up window may also carry a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0033] It can be understood that the above notification and user authorization obtaining process is only illustrative and does not limit the implementation of the present disclosure. Other manners that meet the relevant laws and regulations can also be applied to the implementation of the present disclosure.
[0034] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the relevant laws and regulations and the relevant provisions.
[0035] The SVC (Scalable Video Coding, scalable video coding) coding technology is used for hierarchical coding of video in the time domain, and outputs a multi-layer code stream including a base layer and an enhancement layer, wherein the enhancement layer can be at least one. For example, refer to FIG. 1, which is a schematic diagram of video data obtained by SVC coding according to an embodiment of the present disclosure. In FIG. 1, arrows are used to represent the dependency relationship between video frames, wherein L0 is a base layer, L1, L2 and L3 are enhancement layers, and a frame of an upper layer depends on a frame of a lower layer or a coded frame of the same layer.
[0036] In view of the problem of weak network environment or poor client decoding capability, an embodiment of the present disclosure provides a video processing method for performing hierarchical frame dropping processing on original video data to be transmitted, and performing frame dropping on partial enhancement frames in the original video data, thereby reducing the consumption of network bandwidth and the consumption of client decoding capability after video transmission while ensuring the image quality.
[0037] FIG. 2 is a flowchart of a video processing method according to an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to a case of performing frame dropping processing on original video data to be pushed according to the processing requirement of a client for the video data. The method can be executed by a video processing device, which can be implemented in the form of software and / or hardware. Optionally, the video processing device can be implemented by an electronic device, which can be a computer or a server.
[0038] As shown in FIG. 2, the method comprises the following steps.
[0039] S110, receiving a video request of a client and determining the request parameter corresponding to the video request.
[0040] S120, obtaining original video data corresponding to the video request and coding information of the original video data.
[0041] In S130, a frame dropping strategy is determined based on the encoding information of the original video data and the request parameter, frame dropping processing is performed on the original video data based on the frame dropping strategy, target video data is obtained, and the target video data is pushed to the client.
[0042] The client can be an electronic device with a video playing function, for example, a mobile phone, a tablet computer, or a PC, etc. The client generates a video request in response to a video selection operation of a user, and sends the video request to the server. The video request can include video information, which can be a video name or a video identifier. The video data requested by the video request can include online video stream data, offline video stream data, etc., wherein the online video stream data includes but is not limited to live video stream data, conference video data, etc. Correspondingly, the server matches the video information in the video data pushed by the streaming end to determine the original video data corresponding to the video request. In some embodiments, the streaming end can push multiple candidate video data for the same video information, and the priority of different candidate video data is different. The server can match the corresponding multiple candidate video data based on the video information, and determine the original video data corresponding to the video request based on the priority of the candidate video data.
[0043] In some embodiments, the video request can include network state information and / or client type, wherein the network state information can be network speed data or network type, etc., and the client type can be used to represent the decoding capability of the client. By carrying the network state information and / or the client type in the video request, the request parameter corresponding to the video request of the client can be determined based on the network state information and / or the client type. The request parameter can be used to represent the frame dropping strength of the layered frame dropping processing on the original video data. Different clients can correspond to different request parameters, or the same client can correspond to different request parameters at different time (the network state information at different time can be different).
[0044] In some embodiments, the video request can include a request parameter, which can be determined by the client based on the network state information and / or the client type. Correspondingly, the server receives the video request and parses the request parameter in the video request.
[0045] By adaptively determining a request parameter corresponding to each video request, the request parameter feeds back the video transmission demand and the playing demand of the client. Based on the request parameter corresponding to the video request, the original video data is subjected to adaptive hierarchical frame dropping processing, so as to obtain target video data satisfying the video transmission demand and the playing demand of the client. Different network environments of different clients and different network environments of the same client at different time or different locations determine different request parameters for the video request of the client, improve the adaptability and accuracy of the frame dropping processing, and achieve the reduction of frame dropping to enhance the picture quality in a strong network environment and the increase of frame dropping to avoid lag in a weak network environment, thereby improving the compatibility to the network environment and the client type.
[0046] Optionally, the request parameter comprises offset information and frame dropping interval information; the offset information is an offset layer number of the frame dropping layer relative to the maximum encoding layer number; and the frame dropping interval information is an interval number between the frame dropping layers. In some embodiments, the request parameter can be determined based on a pre-set parameter prediction model, can be determined by the server or by the client, for example, the network state information and / or the client type is input into the parameter prediction model, and the request parameter output by the parameter prediction model is obtained. The parameter prediction model can be a pre-trained neural network model, for example, can be a classification model, etc., by classifying the network state information and / or the client type, determining the request parameter corresponding to the classification result, and pre-establishing the corresponding relationship between the classification result and the request parameter. In some embodiments, a mapping relationship between the network state information and / or the client type and the request parameter is pre-set, the network state information and / or the client type is matched in the mapping relationship, and the corresponding request parameter is determined.
[0047] In the embodiment, the original video data can be encoded based on the SVC encoding technology, and the encoding information of the original video data is obtained, the encoding information of the original video data comprising a maximum encoding layer number of the original video data. Taking FIG. 1 as an example, the maximum encoding layer number of the original video data can be 3. FIG. 1 is only an example, and the maximum encoding layer number of different video data can be different.
[0048] The encoding information of the original video data is carried in the original video data. The encoding information of the original video data is obtained by information analysis on the original video data. Optionally, the method for obtaining the encoding information of the original video data comprises: analyzing a first maximum coding layer number in a video header of the original video data. The video header of the original video data is obtained, and the first maximum coding layer number in the VPS (Video Parameter Set) in the video header of the original video data is obtained. It can be understood that the video parameters in the VPS corresponding to different encoding formats (for example, H264 / H265 / H266) are different. The parameter used to represent the maximum coding layer number can be nuh_temporal_id or nuh_temporal_id_plus1. The maximum value in the parameter corresponding to nuh_temporal_id can be determined as the first maximum coding layer number, or the maximum value in nuh_temporal_id_plus1-1 can be determined as the first maximum coding layer number.
[0049] Optionally, the method for obtaining the encoding information of the original video data comprises: analyzing a second maximum coding layer number in the supplemental enhancement information in the original video data. The second maximum coding layer number can be obtained by analyzing the type 100 supplemental enhancement information (SEI, Supplemental Enhancement Information) corresponding to the video frame in the original video data. The video frame in the original video data can be a key frame or a non-key frame. The key frame can be an I frame. Correspondingly, the type 100 SEI in front of the I frame in the original video data is analyzed to obtain the second maximum coding layer number. Specifically, the maximum value in nuh_temporal_id_plus1-1 obtained from the SEI of each video frame can be determined as the second maximum coding layer number.
[0050] The priority of the second maximum coding layer number is higher than the priority of the first maximum coding layer number. When the first maximum coding layer number and the second maximum coding layer number are obtained, and the first maximum coding layer number is different from the second maximum coding layer number, the second maximum coding layer number is determined as the encoding information of the original video data.
[0051] Optionally, the encoding information of the original video data is stored, for example, in a preset data structure, which can be an avPkt. Optionally, after parsing the encoding layers of each video frame in the original video frames, a preset identifier of the original video data is set, which indicates that the original video data supports the layered frame dropping processing, for example, the preset identifier can be avpkt.svc_parse_flag=true. The preset identifier of the original video data is stored in the preset data structure. For example, the video identifier of the original video data, the encoding information of the original video data and the preset identifier are stored in the preset data structure in association, and the encoding information and the preset identifier corresponding to the original video data are read from the preset data structure based on the video information in the video request.
[0052] It can be explained that the encoding information of the original video data can be determined and stored by the server after receiving the original video data uploaded by the pushing end, and the encoding information and the preset identifier corresponding to the original video data are read from the preset data structure when receiving any video request, so as to reduce the repeated parsing process of the encoding information and the preset identifier.
[0053] On the basis of the above embodiment, before determining the frame dropping strategy based on the encoding information of the original video data and the request parameter, it further includes: determining whether the original video data meets the frame dropping processing condition. The layered frame dropping processing is cancelled for the original video data which does not meet the frame dropping processing condition, which can be that the original video data is taken as the target video data, or other ways of frame dropping processing are performed on the original video data, for example, the frame dropping processing is performed according to the type (such as I frame / P frame / B frame) of the video frame in the original video data to obtain the target video data, which is not limited here.
[0054] Among them, the determination method of the frame dropping processing condition includes one or more of the following: the original video data is provided with a preset identifier; the encoding information of the original video data and / or the encoding layers of the video frame in the original video data is greater than the layered processing layer threshold.
[0055] The original video data is not provided with a preset identifier, which indicates that the original video data does not support the layered frame dropping processing, for example, the preset identifier position is empty, or avpkt.svc_parse_flag=false indicates that the preset identifier is not set. The 0 layer in the original video data is the base layer, and the base layer is not subjected to frame dropping processing. Correspondingly, the original video data is 0, and the encoding information (i.e. the maximum encoding layer) of the original video data is equal to the layered processing layer threshold, or the encoding layers of the video frame in the original video data is equal to the layered processing layer threshold, which indicates that the video frame in the original video data cannot be subjected to frame dropping processing.
[0056] In some embodiments, a preset identifier in the original video data is read, and if the preset identifier is empty or avpkt.svc_parse_flag = false, the layered frame dropping processing is cancelled. If the preset identifier is set in the original video data, the encoding information of the original video data and / or the encoding layer number of the video frame in the original video data are compared based on the layered processing layer number threshold, and if the encoding information of the original video data and / or the encoding layer number of the video frame in the original video data is equal to the layered processing layer number threshold, the layered frame dropping processing is cancelled. If the encoding information of the original video data and / or the encoding layer number of the video frame in the original video data is equal to or greater than the layered processing layer number threshold, the original video data is subjected to the layered frame dropping processing.
[0057] In the embodiment, the frame dropping strategy suitable for the frame dropping processing of the original video data is determined based on the request parameter corresponding to the video request and the encoding information of the original video data, and the target video data corresponding to the video request is obtained.
[0058] Optionally, the frame dropping strategy is determined based on the encoding information of the original video data and the request parameter, including: the frame dropping strategy is determined based on the offset information, the frame dropping interval information and the encoding information of the original video data. In the embodiment, the offset information can include at least one preset offset value offset, for example, a first offset value and a second offset value, wherein the first offset value can be 0, and the second offset value can be 1. The frame dropping interval information can be empty (i.e. non-existent), or the frame dropping interval information can be an integer greater than or equal to 0. It should be noted that if the offset information is empty or the offset information is other than the first offset information and the second offset information, it indicates that the parameter is abnormal, and the layered frame dropping processing is cancelled. If the frame dropping interval information is other than less than 0, it indicates that the parameter is abnormal, and the layered frame dropping processing is cancelled.
[0059] In some embodiments, a plurality of frame dropping strategies are preset, each frame dropping strategy can correspond to a combination of different offset information and frame dropping interval information, and the corresponding frame dropping strategy is determined based on the parameter combination formed by the offset information and the frame dropping interval information corresponding to the video request. The frame dropping strategy includes the target video frame layer number subjected to the frame dropping processing and the frame dropping manner corresponding to the target video frame layer number. The target video frame layer numbers in different frame dropping strategies are different and / or the frame dropping manners corresponding to the target video frame layer numbers are different.
[0060] The target video frame layer number is the layer number of the frame dropping layer. Taking an example of the target video frame layer number being empty 3, the corresponding frame dropping layer is T3 layer, and at least part of the T3 layer in the original video data is subjected to frame dropping.
[0061] The target video frame layer number in the frame loss strategy includes a maximum coding layer number of the original video data and / or a frame loss offset layer number, which is determined based on the maximum coding layer number of the original video data and the offset information. For example, the frame loss offset layer number can be determined as a difference between the maximum coding layer number of the original video data and the offset information. For example, when the maximum coding layer number of the original video data is 3 and the offset information is 1, the frame loss offset layer number can be 3-1=2. Correspondingly, the target video frame layer number can be one or more of the maximum coding layer number 2 and the frame loss offset layer number 2. It can be understood that when the frame loss offset layer number is 0, it indicates that the coding layer corresponding to the frame loss offset layer number is the base layer, and the base layer is not subjected to frame loss processing.
[0062] Optionally, the frame loss manner corresponding to the target video frame layer number includes one or more of no frame loss, full frame loss, and interval frame loss. When the frame loss manner corresponding to the target video frame layer number is no frame loss, all video frames corresponding to the target video frame layer number in the original video data are retained. When the frame loss manner corresponding to the target video frame layer number is full frame loss, all video frames corresponding to the target video frame layer number in the original video data are subjected to frame loss. When the frame loss manner corresponding to the target video frame layer number is interval frame loss, partial video frames corresponding to the target video frame layer number in the original video data are subjected to frame loss, and the partial video frames are determined based on frame loss interval information. When the frame loss interval information exists, the frame loss interval information is greater than or equal to 0. When the frame loss interval information is 0, it indicates that the interval number of the frame loss layer is 0, that is, all video frames corresponding to the target video frame layer number in the original video data are subjected to frame loss.
[0063] In some embodiments, the frame loss strategy is determined based on the offset information, the frame loss interval information, and the coding information of the original video data, including: matching the offset information and the frame loss interval information in a frame loss strategy table to determine a frame loss strategy for which the matching is successful, wherein the frame loss strategy table includes a mapping relationship between a plurality of frame loss strategies and the offset information and the frame loss interval information. For example, refer to Table 1, which is a schematic diagram of a frame loss strategy table provided by an embodiment of the present disclosure.
[0064] Table 1
[0065] In some embodiments, the frame loss strategy is determined based on the offset information, the frame loss interval information, and the coding information of the original video data, including:
[0066] In a case where the offset information is the first offset value (e.g., 0), if the frame dropping interval information is empty, the target video frame layer number in the frame dropping strategy is zero, and no layered frame dropping processing is performed on the original video data. If the frame dropping interval information is greater than or equal to zero, the target video frame layer number in the frame dropping strategy includes the maximum coding layer number and the frame dropping offset layer number, and the frame dropping manner is interval frame dropping. In a case where the offset information is the second offset value (e.g., 1), if the frame dropping interval information is empty, the target video frame layer number in the frame dropping strategy is the maximum coding layer number, and the frame dropping manner is full frame dropping. If the frame dropping interval information is greater than or equal to zero, the target video frame layer number in the frame dropping strategy includes the maximum coding layer number and the frame dropping offset layer number, and the frame dropping manner corresponding to the maximum coding layer number is full frame dropping, and the frame dropping manner corresponding to the frame dropping offset layer number is interval frame dropping.
[0067] Based on the frame dropping strategy, the original video data is processed to obtain target video data. Specifically, the video frame layer number of each video frame in the original video data is obtained, and the video frame corresponding to the target video frame layer number is processed based on the frame dropping manner corresponding to the target video frame layer number. If the video frame layer number of a video frame is the target video frame layer number, it is determined whether to process the video frame based on the frame dropping manner of the target video frame layer number. If yes, the video frame is discarded, and if no, the video frame is retained.
[0068] Taking a case where the target video frame layer number in the frame dropping strategy includes the maximum coding layer number and the frame dropping offset layer number, the frame dropping manner corresponding to the maximum coding layer number is full frame dropping, and the frame dropping manner corresponding to the frame dropping offset layer number is interval frame dropping as an example, for example, the maximum coding layer number of the original video data is 3, the frame dropping offset layer number is 2, the frame dropping interval information is 2, and the offset information is 1. In the process of processing the video frames in the original video data, the video frame layer number of a video frame and the sequence number of the video frame in the video frame layer number are obtained. Taking FIG. 1 as an example, the video frame layer number of the first video frame is 0, which does not belong to the target video frame layer number, and the first video frame is retained. The video frame layer number of the second video frame is 3, which belongs to the target video frame layer number (i.e., the maximum coding layer number), and the frame dropping manner of the maximum coding layer number 3 is full frame dropping, so the second video frame is processed. The video frame layer number of the third video frame is 2, which belongs to the target video frame layer number (i.e., the frame dropping offset layer number), and the frame dropping manner of the frame dropping offset layer number is interval frame dropping. The sequence number of the third video frame is 0, which satisfies the interval frame dropping condition (cnt%3==0), so the third video frame is processed. The interval frame dropping condition is a remainder processing on the sequence number, if the remainder is zero, the interval frame dropping condition is satisfied, cnt is the sequence number of the third video frame, and 3 is the frame dropping interval information+1. By analogy, the video frames with the video frame layer number 3 in FIG. 1 are all processed, and the video frames with the video frame layer number 2 discard the video frames with the sequence numbers 0, 3, 6, and the like, respectively.
[0069] The technical scheme provided in the embodiments of the present disclosure reduces the requirements for network bandwidth and client decoding capability in the transmission process of video data, and improves the compatibility of network bandwidth and client types.
[0070] FIG. 3 is a schematic diagram of a video processing device structure provided by an embodiment of the present disclosure. As shown in FIG. 3, the device includes a request receiving module 210, an information obtaining module 220, and a frame dropping processing module 230.
[0071] The request receiving module 210 is configured to receive a video request of a client, and determine a request parameter corresponding to the video request.
[0072] The information obtaining module 220 is configured to obtain original video data corresponding to the video request and encoding information of the original video data.
[0073] The frame dropping processing module 230 is configured to determine a frame dropping strategy based on the encoding information of the original video data and the request parameter, perform frame dropping processing on the original video data based on the frame dropping strategy, obtain target video data, and push the target video data to the client.
[0074] The technical scheme provided in the embodiments of the present disclosure reduces the requirements for network bandwidth and client decoding capability in the transmission process of video data, and improves the compatibility of network bandwidth and client types.
[0075] Optionally, the encoding information of the original video data includes a maximum encoding layer number of the original video data.
[0076] The information obtaining module 220 is configured to parse a first maximum encoding layer number in a video header of the original video data, and / or parse a second maximum encoding layer number in supplemental enhancement information in the original video data.
[0077] Optionally, the request parameter includes offset information and frame dropping interval information.
[0078] The frame loss processing module 230 is configured to determine the frame loss strategy based on the offset information, the frame loss interval information and the encoding information of the original video data, wherein the frame loss strategy comprises a target video frame layer number and a frame loss mode corresponding to the target video frame layer number.
[0079] Optionally, the target video frame layer number in the frame loss strategy comprises a maximum encoding layer number of the original video data and / or a frame loss offset layer number, the frame loss offset layer number being determined based on the maximum encoding layer number of the original video data and the offset information; and the frame loss mode comprises one or more of no frame loss, full frame loss and interval frame loss.
[0080] Optionally, the frame loss processing module 230 is further configured to match the offset information and the frame loss interval information in a frame loss strategy table to determine a frame loss strategy with which the matching is successful, wherein the frame loss strategy table comprises a mapping relationship between a plurality of frame loss strategies and the offset information and the frame loss interval information.
[0081] On the basis of the above-mentioned embodiments, the frame loss processing module 230 is further configured to obtain a video frame layer number of each video frame in the original video data; and perform frame loss processing on a video frame corresponding to the target video frame layer number based on the frame loss mode corresponding to the target video frame layer number.
[0082] On the basis of the above-mentioned embodiments, the apparatus further comprises:
[0083] The determination module is configured to determine whether the original video data satisfies a frame loss processing condition before determining the frame loss strategy based on the encoding information of the original video data and the request parameter, wherein the determination manner of the frame loss processing condition comprises one or more of the following: the original video data is provided with a preset identifier, the preset identifier indicating that the original video data supports layered frame loss processing; and the encoding information of the original video data and / or the encoding layer number of a video frame in the original video data is greater than a layered processing layer number threshold.
[0084] The video processing apparatus provided in the embodiments of the present disclosure can perform the video processing method provided in any of the embodiments of the present disclosure, and has the corresponding function modules and beneficial effects of the execution method.
[0085] It should be noted that each unit and module included in the above apparatus is only divided according to the function logic, but is not limited to the above division, as long as the corresponding function can be implemented; in addition, the specific names of each functional unit are only for convenient distinction, and do not limit the protection scope of the embodiments of the present disclosure.
[0086] FIG. 4 is a structural diagram of an electronic device according to an embodiment of the disclosure. Below, referring to FIG. 4, a structural diagram of an electronic device (e.g., a terminal device or a server in FIG. 4) 500 suitable for implementing an embodiment of the disclosure is illustrated. The terminal device in an embodiment of the disclosure can include, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), a car terminal (e.g., a car navigation terminal), and the like, and a stationary terminal such as a digital TV, a desktop computer, and the like. The electronic device illustrated in FIG. 4 is merely an example, and should not impose any limitation on the functions and use range of an embodiment of the disclosure.
[0087] As illustrated in FIG. 4, the electronic device 500 can include a processing device (e.g., a central processor, a graphic processor, etc.) 501 that can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0088] Generally, the following devices can be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; a storage device 508 including, for example, a magnetic tape, a hard disk, and the like; and a communication device 509. The communication device 509 can allow the electronic device 500 to communicate with other devices wirelessly or via a wire to exchange data. Although FIG. 4 illustrates the electronic device 500 having various devices, it should be understood that all of the illustrated devices are not required to be implemented or possessed. More or fewer devices can be alternatively implemented or possessed.
[0089] In particular, according to an embodiment of the disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, an embodiment of the disclosure includes a computer program product including a computer program carried on a non-transitory computer readable medium, the computer program containing program code for executing the methods illustrated in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-described functions defined in the methods of an embodiment of the disclosure are performed.
[0090] Names of messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes, and are not used to limit the scope of the messages or information.
[0091] The electronic device provided by the embodiments of the present disclosure and the video processing method provided by the above embodiments belong to the same inventive concept, and the technical details not described in detail in the present embodiment can be referred to the above embodiments, and the present embodiment has the same beneficial effects as the above embodiments.
[0092] The embodiments of the present disclosure provide a computer storage medium, which stores a computer program, and the program is executed by a processor to implement the video processing method provided by the above embodiments.
[0093] It should be noted that the computer readable medium of the present disclosure can be a computer readable signal medium or a computer readable storage medium or any combination of the two. The computer readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus. In the present disclosure, the computer readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer readable program code. Such a propagated data signal can take many forms, including but not limited to an electromagnetic signal, an optical signal or any suitable combination of the above. The computer readable signal medium can also be any computer readable medium other than the computer readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or apparatus. The program code contained in the computer readable medium can be transmitted by any suitable medium, including but not limited to a wire, a cable, an RF (radio frequency) or the like, or any suitable combination of the above.
[0094] In some embodiments, the client, server, can communicate using any known or later developed network protocols, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (for example, a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), the Internet, and peer-to-peer networks (for example, ad hoc peer-to-peer networks), as well as any then-existing or later-developed networks.
[0095] The computer-readable medium described above can be included in the electronic device described above; alternatively, it can exist separately from the electronic device and be not incorporated into the electronic device.
[0096] The computer-readable medium described above carries one or more programs, which, when executed by the electronic device, cause the electronic device to:
[0097] The computer-readable medium described above carries one or more programs, which, when executed by the electronic device, cause the electronic device to: receive a video request of a client, determine a request parameter corresponding to the video request; acquire original video data corresponding to the video request and encoding information of the original video data; determine a frame dropping strategy based on the encoding information of the original video data and the request parameter, perform frame dropping processing on the original video data based on the frame dropping strategy to obtain target video data, and push the target video data to the client.
[0098] Computer program code for carrying out operations of the present disclosure can be written in any one or combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network ("LAN") or a wide area network ("WAN"), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0099] The flow and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow and block diagrams can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may
[0100] The units described in the embodiments of the present disclosure can be implemented by software, or by hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself. For example, the first obtaining unit can also be described as a unit for obtaining at least two Internet protocol addresses.
[0101] The functions described above in the specification can be performed by one or more hardware logic components. For example, and without limitation, example types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0102] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0103] According to one or more embodiments of the present disclosure, example one provides a video processing method, comprising:
[0104] receiving a video request of a client, determining a request parameter corresponding to the video request;
[0105] obtaining original video data corresponding to the video request and encoding information of the original video data;
[0106] determining a frame dropping strategy based on the encoding information of the original video data and the request parameter, performing frame dropping processing on the original video data based on the frame dropping strategy to obtain target video data, and pushing the target video data to the client.
[0107] According to one or more embodiments of the present disclosure, example two provides the video processing method of example one, further comprising:
[0108] The encoding information of the original video data comprises a maximum encoding layer number of the original video data.
[0109] The encoding information of the original video data is obtained in a manner comprising: parsing a first maximum encoding layer number in a video header of the original video data; and / or, parsing a second maximum encoding layer number in supplemental enhancement information in the original video data.
[0110] According to one or more embodiments of the present disclosure, example three provides the video processing method of example one, further comprising:
[0111] The request parameter comprises offset information and frame dropping interval information.
[0112] The frame dropping strategy is determined based on the offset information, the frame dropping interval information and the encoding information of the original video data, wherein the frame dropping strategy comprises a target video frame layer number to be processed by frame dropping and a frame dropping manner corresponding to the target video frame layer number.
[0113] According to one or more embodiments of the present disclosure, example four provides the video processing method of example one, further comprising:
[0114] The target video frame layer number in the frame dropping strategy comprises a maximum encoding layer number of the original video data and / or a frame dropping offset layer number, the frame dropping offset layer number being determined based on the maximum encoding layer number of the original video data and the offset information.
[0115] The frame dropping manner comprises one or more of no frame dropping, all frame dropping and interval frame dropping.
[0116] According to one or more embodiments of the present disclosure, example five provides the video processing method of example one, further comprising:
[0117] The determining the frame dropping strategy based on the offset information, the frame dropping interval information and the encoding information of the original video data comprises: matching the offset information and the frame dropping interval information in a frame dropping strategy table to determine a frame dropping strategy for which matching is successful, wherein the frame dropping strategy table comprises a mapping relationship between a plurality of frame dropping strategies and the offset information and the frame dropping interval information.
[0118] According to one or more embodiments of the present disclosure, example six provides the video processing method of example one, further comprising:
[0119] The frame dropping processing based on the frame dropping strategy comprises: obtaining a video frame layer number of each video frame in the original video data; and performing frame dropping processing on a video frame corresponding to the target video frame layer number based on a frame dropping manner corresponding to the target video frame layer number.
[0120] According to one or more embodiments of the present disclosure, example seven provides the video processing method of example one, further comprising:
[0121] Before determining the frame dropping strategy based on the encoding information of the original video data and the request parameter, further comprising: determining whether the original video data satisfies a frame dropping processing condition.
[0122] The determination manner of the frame dropping processing condition comprises one or more of the following: the original video data is provided with a preset identifier, the preset identifier indicating that the original video data supports layered frame dropping processing; the encoding information of the original video data and / or the encoding layer number of a video frame in the original video data is greater than a layered processing layer number threshold.
[0123] According to one or more embodiments of the present disclosure, example eight provides a video processing device, comprising:
[0124] A request receiving module is configured to receive a video request of a client and determine a request parameter corresponding to the video request.
[0125] An information obtaining module is configured to obtain original video data corresponding to the video request and encoding information of the original video data.
[0126] A frame dropping processing module is configured to determine a frame dropping strategy based on the encoding information of the original video data and the request parameter, perform frame dropping processing on the original video data based on the frame dropping strategy to obtain target video data, and push the target video data to the client.
[0127] The above description merely illustrates the preferred embodiments of the disclosure and a principle for applying the technologies. It is understood by those skilled in the art that the disclosed scope of the disclosure is not limited to the technical solutions formed by the specific combinations of the technical features described above, and should also cover other technical solutions formed by the combinations of the technical features described above or their equivalent features without departing from the disclosed concept. For example, the technical solutions formed by the mutual replacement of the above-described features and the technical features with similar functions disclosed in the disclosure (but not limited to) can be formed.
[0128] Further, although operations are depicted in a particular, sequential order, this should not be understood as requiring or implying that the operations are performed in the order illustrated or sequentially. In certain circumstances, multitasking and parallel processing can be advantageous. Likewise, although specific implementation details are included for the purpose of providing a thorough disclosure, these should not be construed as limitations on the scope of the disclosure. Certain features that are described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.
[0129] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. A video processing method, comprising: Receive a video request from a client, and determine request parameters corresponding to the video request; Obtaining original video data corresponding to the video request and encoding information of the original video data; A frame loss strategy is determined based on the encoding information of the original video data and the request parameters, frame loss processing is performed on the original video data based on the frame loss strategy to obtain target video data, and the target video data is pushed to the client.
2. The method according to claim 1, wherein The encoding information of the original video data includes the maximum number of encoding layers of the original video data; The method for obtaining the encoding information of the original video data includes: Parsing a first maximum number of coding layers in a video header of the original video data; and / or, Parse the second maximum number of coding layers in the supplementary enhancement information in the original video data.
3. The method according to claim 1, wherein The request parameters include offset information and frame loss interval information; The determining of the frame dropping strategy based on the encoding information of the original video data and the request parameter includes: The frame loss strategy is determined based on the offset information, the frame loss interval information and the encoding information of the original video data, wherein the frame loss strategy includes a target number of video frame layers for frame loss processing and a frame loss method corresponding to the target number of video frame layers.
4. The method according to claim 3, wherein: The target number of video frame layers in the frame drop strategy includes the maximum number of coding layers of the original video data and / or the number of frame drop offset layers, where the number of frame drop offset layers is determined based on the maximum number of coding layers of the original video data and the offset information; The frame dropping mode includes one or more of no frame dropping, all frame dropping and interval frame dropping.
5. The method according to claim 3, wherein The determining the frame dropping strategy based on the offset information, the frame dropping interval information, and the encoding information of the original video data includes: The offset information and the frame loss interval information are matched in a frame loss strategy table to determine a frame loss strategy that matches successfully, wherein the frame loss strategy table includes mapping relationships between multiple frame loss strategies and the offset information and the frame loss interval information respectively.
6. The method according to any one of claims 3 to 5, wherein: The performing frame loss processing on the original video data based on the frame loss strategy includes: Obtaining the number of video frame layers of each video frame in the original video data; Frame dropping processing is performed on the video frames corresponding to the target video frame layer number based on the frame dropping method corresponding to the target video frame layer number.
7. The method according to any one of claims 1 to 6, wherein: Before determining the frame drop strategy based on the encoding information of the original video data and the request parameter, the method further includes: Determine whether the original video data meets a frame loss processing condition, where the frame loss processing condition is determined by one or more of the following methods: The original video data is provided with a preset flag, wherein the preset flag indicates that the original video data supports layered frame loss processing; The coding information of the original video data and / or the number of coding layers of the video frames in the original video data is greater than a layered processing layer number threshold.
8. A video processing device, comprising: A request receiving module is configured to receive a video request from a client and determine request parameters corresponding to the video request; An information acquisition module is configured to acquire original video data corresponding to the video request and encoding information of the original video data; The frame loss processing module is configured to determine a frame loss strategy based on the encoding information of the original video data and the request parameters, perform frame loss processing on the original video data based on the frame loss strategy, obtain target video data, and push the target video data to the client.
9. An electronic device comprising: one or more processors; A storage device configured to store one or more programs, wherein When the one or more programs are executed by the one or more processors, the one or more processors implement the video processing method according to any one of claims 1 to 7.
10. A storage medium containing computer-executable instructions, wherein: When the computer executable instructions are executed by a computer processor, the computer executable instructions are used to perform the video processing method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Video transmission method and device, electronic equipment and storage medium
CN112312229A
Video transmission method and electronic equipment
CN115623288A
Layered coding method and device, equipment and storage medium
CN116567256A
Video processing method and device, equipment and storage medium
CN116708938A
Video processing method and device, storage medium and electronic equipment
CN118138848A