Video Coding Optimization Method, Device, Equipment and Storage Medium for Adaptive Network Environment
By matching and adjusting video encoding parameters in real time, video playback problems caused by changes in the network environment are solved, more stable and efficient video transmission is achieved, and user experience and product stability are improved.
Patent Information
- Application Number
- CN202111637639.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-29
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-12-29
AI Technical Summary
The existing technology cannot adjust the video transmission rate and effect in real time when the network environment changes, resulting in problems such as black screen and lag in video. Users need to have professional knowledge to manually adjust the parameters, which is inconvenient to use.
By obtaining the decoding parameters supported by the decoding device and the encoding parameters supported by the camera, matching and filtering out the candidate encoding parameters, confirming the best encoding parameters based on the network evaluation information, and adjusting the video encoding parameters in real time to adapt to changes in the network environment.
It realizes that when the network environment changes, automatically adjusts the video encoding parameters to reduce problems such as stopping video playback, stuttering, intermittent black screen and video loss, improves user experience and product stability, and reduces customer complaints and corporate costs.
Smart Images

Figure CN114302145B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio and video coding and decoding. Specifically, it relates to a video coding optimization method adaptable to network environments, as well as a video coding optimization device, equipment, and storage medium. Background Art
[0002] With the popularization of broadband networks, WiFi networks, and 4G / 5G networks, as well as the progress of chip technology, artificial intelligence technology, Internet of Things technology, and storage technology, high-definition network cameras have been vigorously developed, used more and more widely, and the usage scenarios have become more and more complex. Especially, the network environment often changes dynamically. After the network environment changes, when users view videos through mobile phone APPs or network video recorders, phenomena such as frequent video black screens and freezes often occur, making it impossible to use normally; even causing important recordings to be lost, resulting in relatively serious consequences. Existing technologies require manual modification of various professional parameters. Users need to have professional knowledge, which has high requirements and poor real-time performance. Troubleshooting and handling can only be carried out after problems occur, so installation and troubleshooting are not convenient enough. Due to the large differences among user groups, some users are very difficult to troubleshoot, often causing customer complaints.
[0003] In the process of using network cameras, scenarios with unstable network environments often occur. For example, when connecting to a network video recorder and adding a new network camera, it may cause a reduction in the network bandwidth of the originally connected cameras, resulting in video data transmission failures; when network cameras are connected to the network through wifi or 4G / 5G, the network itself has large jitters, and at the same time, it is affected by the network environment of the decoding device, which will cause problems such as image freezes and loading failures. Existing technologies are to manually adjust the resolution, encoding parameters, etc. of the camera when the decoding device observes phenomena such as video playback freezes, black screens, or recording losses. Existing technologies require users to set manually, which brings inconvenience to users; existing technologies have hysteresis, and when users discover, consequences such as recording losses may have occurred.
[0004] Therefore, it is necessary to provide a video coding optimization method that reduces the impact of network environment changes on the use of cameras, optimizes video coding, and adjusts the video transmission rate and effect in real time, which can facilitate user use and improve the video transmission and viewing effects, as well as a camera applying this method. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a video coding optimization method adaptable to network environments, a video coding optimization device applying this method, a video coding optimization equipment executing this method, and a storage medium storing instructions for executing the method to solve the above problems.
[0006] To achieve the above purpose, the present invention provides the following technical solutions:
[0007] An embodiment of the present invention provides a video coding optimization method for adapting to a network environment. The method includes:
[0008] Obtain decoding parameters supported by a decoding device, and output a decoding capability list, where the decoding capability list includes at least one set of the decoding parameters;
[0009] Query encoding parameters supported by a camera to obtain an initial encoding list, match the decoding capability list with the initial encoding list, and screen out a candidate encoding list from the initial encoding list. The candidate encoding list includes at least one set of candidate encoding parameters;
[0010] Obtain network evaluation information, and based on the network evaluation information, confirm the best encoding parameters from the at least one set of candidate encoding parameters, and perform encoding according to the best encoding parameters.
[0011] In some embodiments, after confirming the best encoding parameters from the at least one set of candidate encoding parameters, it further includes: after obtaining encoded frame data by encoding according to the best encoding parameters, sending the frame data to the decoding device through a network protocol.
[0012] In some embodiments, after sending the frame data to the decoding device, it further includes: evaluating the process of sending the frame data to obtain the network evaluation information, and using the network evaluation information for confirming the best encoding parameters.
[0013] Optionally, the obtaining of the decoding parameters supported by the decoding device includes obtaining the encoding methods that can support decoding and the resolutions supported by each encoding method. The encoding parameters include encoding method, resolution, target bit rate, GOP, and frame rate.
[0014] Optionally, the confirmation of the best encoding parameters includes:
[0015] Select an encoding method with the lowest encoding bit rate;
[0016] Set the expected target bit rate to αVs, where α is a reference coefficient;
[0017] Statistically calculate the data sizes of I frames and P frames in a group of GOPs by group, calculate the average bit rate according to the frame rate, and use the maximum GOP when the GOP has reached the maximum value within the optional range;
[0018] Calculate the target bitrate according to the maximum GOP. If the target bitrate is lower than the expected target bitrate αVs, use the original resolution. If the target bitrate is higher than the expected target bitrate αVs, lower the frame rate until the minimum value. If the result calculated according to the maximum GOP is still higher than αVs after lowering the frame rate, then reduce the resolution. After reducing the resolution, the GOP and frame rate resume using the default values, and the target bitrate takes αVs as the output.
[0019] Optionally, evaluate the frame data sending process by statistically calculating the average sending rate of data packets within an evaluation period T.
[0020] Optionally, to statistically calculate the average sending rate of data packets, within an evaluation period T, count the number of successfully sent data packets N, and the payload sizes of the data packets are P1, P2,..., Pn. The average sending rate of the data packets is:
[0021] The embodiment of the present application also provides a video coding optimization device for adapting to the network environment. The video coding optimization device includes:
[0022] A video decoding capability acquisition module that acquires the decoding parameters supported by the decoding device and outputs a decoding capability list;
[0023] A video coding parameter optimization module that queries the coding parameters supported by the camera, filters out a candidate coding list, and confirms the optimal coding parameters according to the network evaluation information;
[0024] A video coding module that encodes according to the optimal coding parameters to obtain the encoded frame data.
[0025] Optionally, the video coding optimization device for adapting to the network environment of the present application further includes a video data sending module that sends the frame data to the decoding device through a network protocol.
[0026] Optionally, the video coding optimization device for adapting to the network environment of the present application further includes a network environment evaluation module that evaluates the frame data sending process, obtains the network evaluation information, and uses the network evaluation information for confirming the optimal coding parameters.
[0027] Optionally, after the video coding parameter optimization module queries the coding parameters supported by the camera to obtain an initial coding list, it matches the decoding capability list with the initial coding list, filters out the candidate coding list from the initial coding list. The candidate coding list includes at least one set of candidate coding parameters, and the optimal coding parameters are confirmed from the candidate coding list according to the network evaluation information.
[0028] An embodiment of the present application also provides a storage medium with a storage function. Instructions are stored on the storage medium, and when the instructions are executed by a processor, the steps of the video encoding optimization method for the adaptive network environment of the present application are implemented.
[0029] An embodiment of the present application also provides a video encoding optimization device, which includes a memory, a processor, and instructions stored in the memory and executable on the processor. The processor executes the instructions to implement the video encoding optimization method for the adaptive network environment of the present application.
[0030] The present application can be applied to electronic products such as network cameras.
[0031] The present application can dynamically evaluate the network environment. When the network environment changes, the device intelligently adjusts the video encoding parameters to achieve the best encoding scheme, provides the optimal video data in the current network environment to the decoding device, and reduces problems such as video playback stop, stuttering, intermittent black screen, and video loss. When there is network fluctuation, the present application does not require the user to adjust, which is more convenient to use. At the same time, it also improves the product stability, reduces customer complaints, returns, etc., reduces the enterprise cost, and improves the competitiveness of the product. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0033] Figure 1 It is a schematic diagram of the video encoding optimization method for the adaptive network environment of the present application;
[0034] Figure 2 It is a schematic diagram of an embodiment for determining the best encoding parameters in the video encoding optimization method of the present application;
[0035] Figure 3 It is a schematic diagram of the video encoding optimization device for the adaptive network environment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] All the defects existing in the above prior art solutions are the results obtained by the inventors through practice and careful research. Therefore, the discovery process of the above problems and the solutions proposed by the embodiments of the present invention below for the above problems should be the contributions made by the inventors to the present invention during the process of the present invention.
[0037] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Usually, the components of the embodiments of the present invention described and illustrated in the accompanying drawings herein can be arranged and designed in various different configurations.
[0038] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0039] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In the description of the present invention, the terms "first", "second", "third", "fourth", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0040] It should be noted that the reference numerals in the following description are only used for distinguishing descriptions and cannot be understood as the sequence numbers in order. Each operation of the video coding optimization method for an adaptive network environment has no fixed sequence, and there is an organic connection between each operation.
[0041] Please refer to Figure 1 , a video coding optimization method for an adaptive network environment provided by a preferred embodiment of the present invention includes:
[0042] S1 Obtain the decoding parameters supported by the decoding device and output a decoding capability list, where the decoding capability list includes at least one set of the decoding parameters;
[0043] S2 Query the encoding parameters supported by the camera to obtain an initial encoding list, where the initial encoding list includes at least one set of encoding parameters. Match the decoding capability list of the decoding device with the initial encoding list supported by the camera, and screen out a candidate encoding list from the initial encoding list. The candidate encoding list includes at least one set of candidate encoding parameters, and the candidate encoding parameters are supported by the camera and match the decoding parameters supported by the decoding device;
[0044] S3 Obtain network evaluation information, and confirm the best encoding parameters from at least one set of candidate encoding parameters in the candidate encoding list according to the network evaluation information;
[0045] S4 Encodes according to the optimal encoding parameters to obtain encoded frame data, and saves the size of each frame of data within a period of time. The frame data includes data of I-frames and P-frames;
[0046] S5 Sends the frame data to a decoding device for decoding via a network protocol. The sending uses reliable transmission to ensure successful data sending.
[0047] S6 Evaluates the process of sending the frame data to obtain the network evaluation information, and uses the network evaluation information for confirming the optimal encoding parameters.
[0048] As an example, the decoding parameters supported by the decoding device include obtaining the encoding methods that can support decoding and the resolutions supported by each encoding method. For example, obtain encoding methods ENC1, ENC2,..., ENCn, and the resolutions R1, R2, Rn supported by each encoding method. The decoding device can be a mobile phone APP, a network video recorder, a web, etc. It is executed when the decoding device is connected to a camera, and outputs the decoding capability list of the decoding device. The encoding parameters include the encoding method and the resolution.
[0049] As an example, the encoding parameters include the encoding method, the resolution, the target bit rate, GOP (Group of picture, a picture group composed of one I-frame and multiple P-frames), and the frame rate.
[0050] Among them, the resolution is the frame size, and each frame is an image.
[0051] The bitstream / data rate refers to the data traffic used by a video file per unit time. It is also called the bit rate or bitstream rate. A more popular understanding is the sampling rate, which is the most important part of picture quality control in video encoding. The commonly used units are kb / s or Mb / s. Generally speaking, at the same resolution, the larger the bitstream of a video file, the smaller the compression ratio and the higher the picture quality. The larger the bitstream, the larger the sampling rate per unit time, the higher the data stream and precision, and the closer the processed file is to the original file. The better the image quality and the clearer the picture quality, the higher the decoding ability required for the playback device. Of course, the larger the bitstream, the larger the file size. Its calculation formula is file size = time × bit rate / 8. For example, a common 720P RMVB file with a 1Mbps bitstream and a duration of 90 minutes on the Internet has a size of 5400 seconds × 1Mb / 8 = 675MB. Usually, a video file includes pictures and sound. For example, an RMVB video file contains video information and audio information. Both audio and video have their own different sampling methods and bit rates. That is to say, the bit rates of audio and video in the same video file are not the same. And the bitstream rate of a video file we mentioned generally refers to the sum of the bitstream rates of audio and video information in the video file.
[0052] GOP (Group of picture) means a group of pictures. A GOP is a set of consecutive pictures. A GOP is a collection of pictures in a sequence used to assist random access. The first image of a GOP must be an I-frame, so that it can be ensured that the GOP does not need to refer to other images and can be decoded independently. The period of key frames, that is, the distance between two IDR frames, and the maximum number of frames in a frame group. Generally speaking, at least 1 key frame is required per second of video. Increasing the number of key frames can improve the quality, but at the same time increase the bandwidth and network load. It should be noted that improving the image quality by increasing the GOP value has limitations. When a scene change occurs, the H.264 encoder will automatically force the insertion of an I-frame, and at this time the actual GOP value is shortened. On the other hand, in a GOP, P and B frames are predicted from the I-frame. When the image quality of the I-frame is relatively poor, it will affect the image quality of subsequent P and B frames in a GOP until the next GOP starts before it may be restored. Therefore, the GOP value should not be set too large. At the same time, since the complexity of P and B frames is greater than that of I-frames, too many P and B frames will affect the encoding efficiency and reduce the encoding efficiency. In addition, too long a GOP will also affect the response speed of the Seek operation. Since P and B frames are predicted from the previous I or P frame, the Seek operation needs to be directly located. When decoding a certain P or B frame, it is necessary to first decode the I-frame and the previous N predicted frames within this GOP. The longer the GOP value, the more predicted frames need to be decoded, and the longer the seek response time.
[0053] The frame rate is the number of frames of pictures transmitted in 1 second. It can also be understood as how many times the graphics processor can refresh per second, usually expressed in fps (Frames Per Second). Each frame is a static image. Displaying frames quickly and continuously creates the illusion of motion. A high frame rate can result in smoother and more realistic animations. The more frames per second (fps), the smoother the displayed action. The reason we can use a camera to see continuous images is that the image sensor continuously captures pictures and transmits them to the screen. When the transmission speed reaches a certain level, the human eye cannot distinguish the time gap between the pictures, so people can see continuous dynamic pictures. A high frame rate can result in smoother and more realistic animations. Generally, 30 fps is acceptable, but increasing the performance to 60 fps can significantly enhance the sense of interaction and realism. However, generally speaking, when the frame rate exceeds 75 fps, it is not easy to notice a significant improvement in smoothness. If the frame rate exceeds the screen refresh rate, it will only waste the graphics processing ability because the monitor cannot update at such a fast speed, and the frame rate exceeding the refresh rate is wasted.
[0054] Therefore, choosing appropriate encoding parameters is important for adapting to the network environment. In view of the problems of the prior art, through research efforts, the inventor has created a video encoding optimization method and device for adapting to the network environment, as well as a camera applying the method and device.
[0055] As Figure 2 shown, as an embodiment, the confirmation of the optimal encoding parameters includes:
[0056] a. Select the encoding method with the lowest encoding bit rate;
[0057] b. Set the expected target bit rate to αVs, where α is a reference coefficient;
[0058] c. Statistically calculate the data sizes of I-frames and P-frames in a group of GOPs by group, calculate the average bit rate according to the frame rate, and when the GOP has reached the maximum value within the optional range, use the maximum GOP;
[0059] d. Calculate the target bit rate according to the maximum GOP. If the target bit rate is lower than the expected target bit rate αVs, use the original resolution; if the target bit rate is higher than the expected target bit rate αVs, lower the frame rate until the minimum value. If the result calculated according to the maximum GOP is still higher than αVs after lowering the frame rate, then reduce the resolution. After reducing the resolution, the GOP and frame rate resume using the default values, and the target bit rate takes αVs as the output.
[0060] Where α ranges from 0.8 to 1.0, and the maximum does not exceed 1.0, that is, the maximum value of αVs is Vs.
[0061] Among them, evaluating the frame data sending process is to count the average sending rate of data packets within an evaluation period T.
[0062] Among them, counting the average sending rate of data packets is to count the number of successfully sent data packets N within an evaluation period T. The effective payload sizes of the data packets are P1, P2,..., Pn, and the average sending rate of the data packets is:
[0063] As Figure 3 shown, an embodiment of the present invention provides a video coding optimization device for an adaptive network environment, including:
[0064] A video decoding capability acquisition module, which acquires the decoding parameters supported by the decoding device and outputs a decoding capability list;
[0065] A video coding parameter optimization module, which queries the coding parameters supported by the camera, filters out a candidate coding list, and confirms the optimal coding parameters according to the network evaluation information;
[0066] A video coding module, which encodes according to the optimal coding parameters to obtain encoded frame data;
[0067] A video data sending module, which sends the frame data to the decoding device through a network protocol;
[0068] A network environment evaluation module, which evaluates the frame data sending process, obtains the network evaluation information, and uses the network evaluation information for confirming the optimal coding parameters.
[0069] In one embodiment, the video coding parameter optimization module obtains an initial coding list after querying the coding parameters supported by the camera, matches the decoding capability list with the initial coding list, filters out the candidate coding list in the initial coding list. The candidate coding list includes at least one set of candidate coding parameters, and the optimal coding parameters are confirmed from the candidate coding list according to the network evaluation information.
[0070] In one embodiment, the coding parameters include coding mode, resolution, target bit rate, GOP, and frame rate.
[0071] The video coding parameter optimization module's confirmation of the optimal coding parameters includes:
[0072] Selecting the coding mode with the lowest coding bit rate;
[0073] Setting the expected target bit rate to αVs, where α is a reference coefficient;
[0074] Statistically calculate the data sizes of I-frames and P-frames in a group of GOPs by group, calculate the average bitrate according to the frame rate, and when the GOP has reached the maximum value within the optional range, use the maximum GOP;
[0075] Calculate the target bitrate according to the maximum GOP. If the target bitrate is lower than the expected target bitrate αVs, use the original resolution; if the target bitrate is higher than the expected target bitrate αVs, lower the frame rate until the minimum value. If the result calculated according to the maximum GOP is still higher than αVs after lowering the frame rate, then reduce the resolution. After reducing the resolution, the GOP and frame rate resume using the default values, and the target bitrate takes αVs as the output.
[0076] An embodiment of the present application also provides a storage medium with a storage function. Instructions are stored on the storage medium, and when the instructions are executed by a processor, the steps of the video encoding optimization method for an adaptive network environment described in the present application are implemented.
[0077] An embodiment of the present application also provides a video encoding optimization device, which includes a memory, a processor, and instructions stored in the memory and executable on the processor. The processor executes the instructions to implement the video encoding optimization method for an adaptive network environment of the present application.
[0078] An embodiment of the present application provides a camera that adopts a video encoding optimization method for an adaptive network environment. The camera is connected to a video encoding optimization device, and the video encoding optimization device includes:
[0079] A video decoding capability acquisition module that acquires the decoding parameters supported by the decoding device and outputs a decoding capability list;
[0080] A video encoding parameter optimization module that queries the encoding parameters supported by the camera, filters out a candidate encoding list, and confirms the optimal encoding parameters according to the network evaluation information;
[0081] A video encoding module that encodes according to the optimal encoding parameters to obtain encoded frame data;
[0082] A video data sending module that sends the frame data to the decoding device through a network protocol;
[0083] A network environment evaluation module that evaluates the process of sending the frame data, obtains the network evaluation information, and uses the network evaluation information for confirming the optimal encoding parameters.
[0084] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An adaptive video coding optimization method for network environment, characterized in that The method includes: Obtaining decoding parameters supported by a decoding device and outputting a decoding capability list, where the decoding capability list includes at least one set of the decoding parameters; Querying encoding parameters supported by a camera to obtain an initial encoding list, matching the decoding capability list with the initial encoding list, and screening out a candidate encoding list from the initial encoding list, where the candidate encoding list includes at least one set of candidate encoding parameters; Obtaining network evaluation information, and based on the network evaluation information, confirming optimal encoding parameters from the at least one set of candidate encoding parameters, and performing encoding according to the optimal encoding parameters; After confirming the optimal encoding parameters from the at least one set of candidate encoding parameters, it further includes: After performing encoding according to the optimal encoding parameters to obtain encoded frame data, sending the frame data to the decoding device through a network protocol; The confirmation of the optimal encoding parameters includes: Selecting an encoding method with the lowest encoding bitrate; Setting the expected target bitrate to αVs, where α is a reference coefficient; Statistically calculating the data sizes of I-frames and P-frames in a group of GOPs by group, calculating the average bitrate according to the frame rate, and when the GOP has reached the maximum value within the optional range, using the maximum GOP; Calculating the target bitrate according to the maximum GOP. If the target bitrate is lower than the expected target bitrate αVs, using the original resolution; if the target bitrate is higher than the expected target bitrate αVs, reducing the frame rate until the minimum value. If the result calculated according to the maximum GOP is still higher than αVs after reducing the frame rate, then reducing the resolution. After reducing the resolution, the GOP and the frame rate resume using the default values, and the target bitrate takes αVs as the output.
2. The video coding optimization method for an adaptive network environment according to claim 1, wherein After sending the frame data to the decoding device, it further includes: Evaluating the process of sending the frame data to obtain the network evaluation information, and using the network evaluation information for the confirmation of the optimal encoding parameters.
3. The video coding optimization method for an adaptive network environment according to any one of claims 1 to 2, characterized in that, The obtaining of the decoding parameters supported by the decoding device includes obtaining the encoding methods that can support decoding and the resolutions supported by each encoding method. The encoding parameters include encoding method, resolution, target bitrate, GOP, and frame rate.
4. The video coding optimization method for an adaptive network environment according to claim 2, wherein Evaluating the process of sending the frame data is to statistically calculate the average packet sending rate within an evaluation period T.
5. The video coding optimization method for an adaptive network environment according to claim 4, characterized in that To calculate the average transmission rate of data packets, within an evaluation period T, count the number of successfully transmitted data packets N. The payload sizes of the data packets are P1, P2, ..., Pn. The average transmission rate of the data packets is: 。 6. A video coding optimization device for an adaptive network environment, which is used based on the video coding optimization method for an adaptive network environment described in any one of claims 1-5, and is characterized in that, The video encoding optimization device includes: A video decoding capability obtaining module, which obtains decoding parameters supported by a decoding device and outputs a decoding capability list; A video encoding parameter optimization module, which queries encoding parameters supported by a camera, screens out a candidate encoding list, and confirms optimal encoding parameters according to network evaluation information; A video encoding module, which performs encoding according to the optimal encoding parameters to obtain encoded frame data.
7. A storage medium with a storage function, characterized in that, Instructions are stored on the storage medium, and when the instructions are executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
8. A video encoding optimization device, characterized in that It includes a memory, a processor, and instructions stored in the memory and executable on the processor. The processor executes the instructions to implement the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Video stream transmission method and device
CN107529069A