Video encoding method, apparatus, computer readable medium, and electronic device

CN115701709BActive Publication Date: 2026-05-22TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2021-08-02
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

Existing video encoding methods based on 5G private networks suffer from long transmission latency. Existing adaptive encoding technologies cannot adjust in a timely manner in 5G air interface scenarios, resulting in long video transmission latency.

Method used

By acquiring wireless network signal strength information within a historical time period, the signal strength of the next key frame is predicted, and the target data volume is determined based on the prediction results. Intra-frame coding is then performed to ensure that the key frame is transmitted within one uplink time slot.

Benefits of technology

It reduces video transmission latency, increases video stream transmission rate, and reduces transmission latency while ensuring video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115701709B_ABST
    Figure CN115701709B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of communication, and particularly relates to a video coding method, device, computer readable medium and electronic equipment. The video coding method comprises the following steps: obtaining historical intensity information in a historical time period, wherein the historical intensity information is used for representing network signal intensity of wireless video transmission corresponding to each historical moment in the historical time period; predicting intensity information of a next key frame according to the historical intensity information, wherein the intensity information of the next key frame is used for representing wireless network signal intensity of transmitting the next key frame; determining a target data amount of the next key frame according to the intensity information of the next key frame, and performing intra-frame coding on the next key frame according to the target data amount. The target data amount is determined in advance before the frame is coded, and then the next key frame is coded according to the target data amount, so that the transmission of the key frame can be completed in one uplink time slot, and the transmission is completed without waiting for a frame period, thereby reducing the transmission delay.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of communication technology, and specifically relates to a video encoding method, a video encoding device, a computer-readable medium, and an electronic device. Background Technology

[0002] Video coding aims to eliminate redundant information between video signals. With the continuous development of multimedia digital video applications, the amount of raw video data has become unsustainable for existing transmission network bandwidth and storage resources. Therefore, only encoded and compressed video is suitable for transmission over networks. Video coding technology has become one of the hot topics in academic research and industrial applications both domestically and internationally. Furthermore, with the development of artificial intelligence and the arrival of the 5G era, the even larger volume of video data places higher demands on video coding standards.

[0003] Currently, video encoding methods based on 5G private networks suffer from long transmission latency, and reducing this latency is an urgent problem to be solved.

[0004] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] The purpose of this application is to provide a video encoding method, apparatus, computer-readable medium, and electronic device that at least to some extent overcomes technical problems in the related art, such as reducing transmission latency.

[0006] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.

[0007] According to one aspect of the embodiments of this application, a video encoding method is provided, comprising:

[0008] Obtain historical intensity information within a historical time period, wherein the historical intensity information is used to represent the wireless network signal strength of video transmission at each historical moment within the historical time period;

[0009] The intensity information of the next key frame is predicted based on the historical intensity information, and the intensity information of the next key frame is used to represent the wireless network signal strength for transmitting the next key frame.

[0010] Based on the intensity information of the next keyframe, the target data volume of the next keyframe is determined, and intra-frame coding is performed on the next keyframe according to the target data volume.

[0011] According to one aspect of the embodiments of this application, a video encoding apparatus is provided, comprising:

[0012] The acquisition module is used to acquire historical intensity information within a historical time period, wherein the historical intensity information is used to represent the wireless network signal strength of video transmission at each historical moment within the historical time period.

[0013] The prediction module is used to predict the intensity information of the next key frame based on the historical intensity information, wherein the intensity information of the next key frame is used to represent the wireless network signal strength for transmitting the next key frame.

[0014] The determination module is used to determine the target data volume of the next keyframe based on the intensity information of the next keyframe, and to perform intra-frame coding on the next keyframe according to the target data volume.

[0015] In some embodiments of this application, based on the above technical solutions, the determining module includes:

[0016] The first determining unit is used to determine the modulation and coding strategy of the next key frame based on the intensity information of the next key frame.

[0017] The calculation unit is used to calculate the data volume of the next key frame according to the modulation and coding strategy of the next key frame;

[0018] The second determining unit is used to determine the target data volume of the next key frame by comparing the calculated data volume with the actual data volume.

[0019] In some embodiments of this application, based on the above technical solutions, the first determining unit is used to obtain a mapping table between modulation and coding strategies and intensity information; and to determine the modulation and coding strategy of the next key frame by searching the mapping table between the modulation and coding strategies and intensity information according to the intensity information of the next key frame.

[0020] In some embodiments of this application, based on the above technical solutions, the first determining unit is used to acquire in real time the historical modulation and coding strategy and the historical intensity information allocated by the base station, and use the coding information corresponding to the historical modulation and coding strategy as a classification label, wherein the historical modulation and coding strategy is used to represent the modulation and coding strategy corresponding to the historical moment; the historical intensity information is clustered according to the coding information classification label to obtain a cluster classifier; each intensity value in the distribution range of the historical intensity information value is input into the cluster classifier to obtain the classification information corresponding to the coding information classification label, so as to obtain the mapping relationship table between the modulation and coding strategy and the intensity information.

[0021] In some embodiments of this application, based on the above technical solutions, the second determining unit is used to adjust the calculated data amount; compare the adjusted data amount with the actual data amount, and select the data amount with the smaller value as the target data amount for the next key frame.

[0022] In some embodiments of this application, based on the above technical solutions, the second determining unit is used to obtain the current network load redundancy; and to obtain the adjusted data volume by multiplying the calculated data volume, the load redundancy, the preset protection ratio, and the preset adjustment coefficient.

[0023] In some embodiments of this application, based on the above technical solutions, the second determining unit is configured to, if the next key frame adopts forward error correction coding, then the preset adjustment coefficient is 1 / (1+R), where R represents the redundancy rate of forward error correction coding; if the next key frame does not adopt forward error correction coding, then the preset adjustment coefficient is 1.

[0024] In some embodiments of this application, based on the above technical solutions, the second determining unit is used to obtain the latency information of historical key frames; and to determine the load redundancy of the current network based on the latency information.

[0025] In some embodiments of this application, based on the above technical solutions, the second determining unit is used to use the latency information fed back by the receiving end as the latency information of the historical key frame, or to use the latency information of clearing the transmission buffer obtained by monitoring as the latency information of the historical key frame.

[0026] In some embodiments of this application, based on the above technical solutions, the second determining unit is used to set an initial load redundancy; if the time window of the delay information is greater than the time window of the preset delay, the load redundancy is reduced according to the preset first step size, and the adjusted final load redundancy is used as the load redundancy of the current network; if the time window of the delay information is less than the time window of the preset delay, the load redundancy is increased according to the second preset step size, and the adjusted final load redundancy is used as the load redundancy of the current network; wherein, the first preset step size is greater than the second preset step size.

[0027] In some embodiments of this application, based on the above technical solutions, in the second determining unit, the time window of the preset delay is determined according to the delay of the forward reference frame transmission, or according to the test value of the no-load network environment.

[0028] In some embodiments of this application, based on the above technical solutions, in the acquisition module, the historical time period and the encoding time of the next key frame have a preset time interval sliding time window.

[0029] According to one aspect of the embodiments of this application, a computer-readable medium is provided, on which a computer program is stored, which, when executed by a processor, implements the video encoding method as described in the above technical solutions.

[0030] According to one aspect of the embodiments of this application, an electronic device is provided, the electronic device comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform a video encoding method as described above by executing the executable instructions.

[0031] According to one aspect of the embodiments of this application, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the video encoding method as described above.

[0032] In the technical solution provided in this application embodiment, the intensity information of the next key frame is predicted by historical intensity information, and the target data volume of the next key frame is determined based on the predicted intensity information. The target data volume is predetermined before the frame is encoded, and then encoded in the next key frame according to the target data volume. In this way, the transmission of the key frame can be completed within one uplink time slot, without waiting for a frame period to complete the transmission, thereby reducing transmission latency.

[0033] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0034] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0035] Figure 1 The diagram illustrates an exemplary 5G network data transmission process.

[0036] Figure 2 An exemplary system architecture block diagram illustrating the application of the technical solution of this application is shown schematically.

[0037] Figure 3 The flowchart of the video encoding method provided in the embodiments of this application is illustrated schematically.

[0038] Figure 4 The illustration schematically shows the steps of determining the target data volume of the next keyframe based on the intensity information of the next keyframe in one embodiment of this application.

[0039] Figure 5 The illustration schematically shows the steps of determining the modulation and coding strategy of the next keyframe based on the intensity information of the next keyframe in one embodiment of this application.

[0040] Figure 6 The flowchart illustrates the steps of obtaining the modulation coding strategy mapping table between modulation coding strategy and intensity information in one embodiment of this application.

[0041] Figure 7 The illustration schematically shows the steps in one embodiment of this application to determine the target data volume of the next keyframe by comparing the calculated data volume with the actual data volume.

[0042] Figure 8 The steps for adjusting the calculated data volume are illustrated schematically in one embodiment of this application.

[0043] Figure 9 The flowchart illustrating the steps for obtaining the current network load redundancy in one embodiment of this application is shown in the illustration.

[0044] Figure 10 The flowchart illustrating the steps of determining the current network load redundancy based on latency information in one embodiment of this application is shown in the illustration.

[0045] Figure 11 A schematic block diagram of the video encoding apparatus provided in an embodiment of this application is shown.

[0046] Figure 12 A schematic diagram of a computer system architecture suitable for implementing the embodiments of this application is shown. Detailed Implementation

[0047] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art.

[0048] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.

[0049] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0050] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0051] With the development of artificial intelligence and the arrival of the 5G era, the massive amount of video data has placed higher demands on video coding standards. For the same video, a higher compression ratio results in higher compression distortion and a worse video quality for the user experience; conversely, a lower compression ratio increases the cost of video storage and transmission. Finding a balance between these two factors is a challenge in video coding technology. Currently, network-aware adaptive coding technology has become a key technology in real-time audio and video communication.

[0052] In 5G networks, uplink and downlink wireless transmission resources are configured using different frames. Uplink frames only allow the transmission of uplink data (from the terminal to the base station), while downlink frames only allow the transmission of downlink data (from the base station to the terminal).

[0053] See Figure 1 , Figure 1 This diagram schematically illustrates an exemplary 5G network data transmission process. Currently, the common 5G network frame configuration is 3D1U, meaning one time period is 5ms, with 3 downlink frames, 1 uplink frame, and 1 special subframe, each frame lasting 1ms. The amount of uplink data that one U-frame can carry depends on the signal strength between the terminal and the base station, as well as the base station's scheduling. When the I-frame data of the terminal's uplink video stream cannot be transmitted completely on the current U-frame, it will be scheduled to the next U-frame or the U-frame after that for uplink transmission, until all data is transmitted. See also... Figure 1If only Part 1 of the I-frame data is uploaded in the current U-frame, the remaining Part 2 can only be scheduled for uplink transmission in the next U-frame, requiring a further 5ms frame period wait. Therefore, existing 5G networks suffer from transmission latency issues during video transmission.

[0054] To address the transmission latency issue, the commonly used approach is to employ adaptive coding techniques based on network awareness.

[0055] Currently, existing adaptive coding mainly includes two methods. The first is direct feedback adjustment, which directly adjusts coding parameters based on the latency and packet loss reported by the sending end. The second is prediction adjustment based on a routing congestion model, which uses Wiener filtering or time-series prediction to predict network latency and packet loss rate based on the latency and jitter reported by the sending end, and then performs adaptive coding adjustments. The prediction model is mainly for public network scenarios, considering that network latency is mainly caused by router forwarding. When routers are congested, there will be significant latency jitter, and even packet loss may occur. However, both of the above-mentioned adaptive coding methods have problems. Specifically, the method of adjusting directly based on the receiving end feedback has a certain lag in encoder adjustment, while the method of prediction adjustment based on a router congestion model cannot adapt to 5G private network scenarios where 5G air interface latency is the main concern.

[0056] In response to the above situation, this application proposes a coding method designed for the characteristics of 5G air interface to improve the latency problem in 5G private network scenarios.

[0057] Specifically, see Figure 2 , Figure 2 An exemplary system architecture block diagram illustrating the application of the technical solution of this application is shown schematically.

[0058] See Figure 2 As shown, the system architecture includes a video transmitting terminal 210 and a network 220, which are communicatively connected. The video transmitting terminal 210 includes an encoding end 211 and a 5G module 212. The video stream is encoded by the encoding end 211, and the encoded video stream is then sent to the 5G module 212. To reduce transmission latency, the target data size of the key frame to be transmitted needs to be determined before encoding at the encoding end 211. The target data size of the next key frame is determined using the following video encoding method.

[0059] It should be noted that video keyframes are usually I-frames, which are frames encoded using intra-frame coding techniques in video encoding. They can be decoded independently without relying on other frames during decoding. Therefore, they are usually larger in size and are different from frames encoded using inter-frame coding techniques. Therefore, in this application, I-frames are used to refer to video keyframes.

[0060] The video encoding method provided in this application will be described in detail below with reference to specific implementation methods.

[0061] See Figure 3 , Figure 3 The flowchart illustrating the steps of the video encoding method provided in this application is shown in the embodiment. This application discloses a video encoding method that mainly includes the following steps S301 to S303.

[0062] Step S301: The terminal obtains historical intensity information within a historical time period. The historical intensity information is used to represent the wireless network signal strength of video transmission at the corresponding historical moment within the historical time period.

[0063] When the terminal transmits an uplink video stream, it begins executing the video encoding method of this application at time T1 before the arrival of the next intra-frame keyframe (I-frame). The specific video encoding method is as follows: the terminal obtains historical strength information within a historical time period, where RSRP (Reference Signal Receiving Power) represents the network signal strength. First, it obtains the RSRP received within the most recent time window T2 time series.

[0064] In one embodiment of this application, the historical time period and the encoding time of the next keyframe have a preset time interval sliding time window, which slides forward as time progresses. This facilitates obtaining the network signal strength of the video transmission at the corresponding historical moment.

[0065] Step S302: Predict the strength information of the next key frame based on the historical strength information. The strength information of the next key frame is used to represent the wireless network signal strength for transmitting the next key frame.

[0066] The intensity information of the next keyframe is predicted based on the historical intensity information within a preset time window. The RSRP at the keyframe transmission time is predicted based on the RSRP time series values ​​received within time window T2. By setting multiple RSRP values ​​corresponding to a time period for prediction, the prediction results are more stable. If only a single value is used for prediction, jitter will occur. Specifically, linear regression, zero-order hold, XGBoost, or neural network models can be used to predict the RSRP of the next keyframe.

[0067] In one embodiment, when predicting the RSRP of the next keyframe using linear regression, specifically, the following steps are taken: first, the RSRP dataset corresponding to the historical frames is obtained; then, clusters are generated based on the RSRP dataset corresponding to the historical frames; simultaneously, linear regression coefficients are calculated based on the generated clusters; then, the linear regression function of the clusters is obtained through the generated clusters and the linear regression coefficients; then, the linear regression of the Kth clusters is calculated based on the obtained linear regression function of the clusters; finally, the RSRP of the next keyframe is obtained based on the obtained linear regression results of the Kth clusters.

[0068] In one embodiment, when predicting the RSRP of the next keyframe using a zero-order hold approach, specifically, a continuous model is constructed based on the RSRP corresponding to historical frames. The continuous model is then discretized using the step response invariance method to obtain a zero-order hold discretized model. A computation time delay is then introduced to optimize the discretized model, thereby establishing the zero-order hold discretized model. Based on the zero-order hold discretized model, computation time delay and disturbance terms are introduced to obtain delay and disturbance models, respectively. Based on the disturbance model, an extended state observer is designed using a state-space approach. A zero-order hold discretized state equation considering the disturbance is established, and a disturbance state observer is designed based on this zero-order hold discretized state equation to estimate the disturbance. Considering both system robustness and system dynamic performance, a direct pole-zero placement method is used to select appropriate RSRP parameters, thereby obtaining the RSRP of the next keyframe.

[0069] In one embodiment, when using the XGBOOST algorithm to predict the RSRP of the next keyframe, specifically, the RSRP corresponding to the historical frame is obtained, the RSRP corresponding to the historical frame is used as historical features, the features affecting the RSRP data and the historical features are used as the training dataset of XGBOOST, and the training dataset is trained using XGBOOST to obtain the prediction model; the prediction model is used to predict the RSRP of the prediction time to obtain the predicted RSRP value of the next keyframe.

[0070] In one embodiment, when using a neural network model to predict the RSRP of the next keyframe, specifically, consecutive video frames are selected as training samples, and the inter-frame difference of the training samples is extracted. The inter-frame difference is used as the input of the encoder in the generator model. The neural network weights of the encoder and decoder are obtained based on the loss function training. The prediction frame that minimizes the loss function value is solved, thereby obtaining the RSRP prediction value of the next keyframe.

[0071] Step S303: Determine the target data amount of the next key frame based on the intensity information of the next key frame, and perform intra-frame coding on the next key frame according to the target data amount.

[0072] Since the amount of uplink data that can be carried in one uplink frame is specifically related to the signal strength from the terminal to the base station, the target data amount of the next key frame can be obtained based on the predicted strength information of the next key frame. When performing intra-frame coding in the next key frame, the coding is performed using the determined target data amount.

[0073] Since a video stream consists of many frames, it is transmitted frame by frame. Before encoding the next keyframe, the size of the keyframe encoding is pre-calculated. This pre-calculated keyframe encoding size refers to the target data size of the next keyframe. After determining the target data size of the next keyframe, the encoder's encoding rate is set so that the encoded keyframe size is the pre-set size. This target data size allows the keyframe to be transmitted within a single uplink time slot, thereby reducing transmission latency.

[0074] In the technical solution provided in this application embodiment, the intensity information of the next keyframe is predicted through historical intensity information, and the target data volume of the next keyframe is determined based on the predicted intensity information. By pre-determining the target data volume before encoding the next keyframe, and then performing intra-frame encoding on the next keyframe according to the target data volume, the keyframe transmission can be completed within one uplink time slot, without waiting for a full frame period, thereby reducing transmission latency. This application sacrifices some image quality for shorter latency, significantly improving the transmission rate of the video stream.

[0075] It should be noted that by adaptively adjusting the keyframe size, keyframe transmission can be completed within a single 5G uplink time slot, eliminating the need to wait for the 5ms frame period. Currently, the encoding range of a 1080P keyframe is approximately 50-200KB, and the amount of data that can be transmitted in an idle 5G uplink time slot is around 125KB. Adaptive keyframe transmission is feasible depending on network conditions.

[0076] See Figure 4 , Figure 4 The illustration schematically depicts the steps of determining the target data volume of the next keyframe based on the intensity information of the next keyframe in one embodiment of this application. In one embodiment of this application, step S303, determining the target data volume of the next keyframe based on the intensity information of the next keyframe, mainly includes the following steps S401 to S403.

[0077] Step S401: Determine the modulation and coding strategy of the next key frame based on the intensity information of the next key frame.

[0078] In wireless communication, the Modulation and Coding Scheme (MCS) is generally used to describe the selection of channel coding and constellation modulation during physical layer transmission. Transmitted data is typically processed through the MCS, and different MCSs have different coding rates, affecting the amount of data that can actually be transmitted on the same radio resource block. To predetermine the target data volume before encoding the next keyframe, the MCS corresponding to the next keyframe needs to be obtained. This MCS is determined by the intensity information of the next keyframe.

[0079] Step S402: Calculate the data volume of the next key frame according to the modulation and coding strategy of the next key frame.

[0080] Since different modulation and coding strategies have different coding rates, different coding rates affect the amount of data that can actually be transmitted on the same radio resource block. By obtaining the modulation and coding strategy of the next key frame, the corresponding coding rate is obtained, and the amount of data in the next key frame is calculated accordingly.

[0081] Step S403: Based on the comparison between the calculated data volume and the actual data volume, determine the target data volume for the next key frame.

[0082] To ensure normal video streaming while reducing transmission latency, the calculated data amount is not directly used as the target data amount for the next keyframe. Instead, a comparison is made between the calculated data amount and the actual data amount of the video stream, and an appropriate data amount is selected as the target data amount. This ensures normal video streaming while reducing transmission latency, without significantly degrading video quality.

[0083] In this way, the target data volume of the next keyframe is estimated by using the intensity information of the next keyframe. By pre-determining the target data volume before encoding the next keyframe, and comparing the calculated data volume with the actual data volume when determining the target data volume, a more appropriate data volume is selected as the target data volume, thereby reducing transmission latency while ensuring that the video stream can be transmitted normally.

[0084] See Figure 5 , Figure 5 The illustration schematically depicts the steps of determining the modulation and coding strategy of the next key frame based on the intensity information of the next key frame in one embodiment of this application. In one embodiment of this application, step S401, determining the modulation and coding strategy of the next key frame based on the intensity information of the next key frame, mainly includes the following steps S501 to S502.

[0085] Step S501: The terminal obtains the mapping table between modulation and coding strategies and intensity information.

[0086] The mapping table between modulation and coding strategies and strength information, namely the mapping table between MCS and RSRP, describes the MCS level corresponding to different RSRP segments.

[0087] Optionally, rules for determining the MCS table can be pre-configured, so that during the access process of the video transmitting terminal, the MCS table used by each channel can be determined according to the pre-configured rules. For example, the pre-configured rules could be based on the current RSRP measurement value to determine the MCS table used by the video transmitting terminal for all or some channels during the access process.

[0088] Since the current RSRP measurement value of the video transmitting terminal reflects the current user capability and actual transmission performance of the video transmitting terminal, the MCS table used by each channel during random access can be determined based on the current RSRP measurement value to better meet the real-time performance requirements of the video transmitting terminal. Therefore, the video transmitting terminal can measure RSRP during access to determine the MCS table used by each channel based on the range of the current RSRP measurement value. Furthermore, after measuring the current RSRP value, the video transmitting terminal can send the current RSRP measurement value to the base station so that the base station can determine the MCS table used by each channel that matches the current RSRP measurement value, thereby obtaining the mapping relationship table between MCS and RSRP.

[0089] Step S502: Based on the intensity information of the next key frame, look up the mapping table between modulation and coding strategies and intensity information to determine the modulation and coding strategy of the next key frame.

[0090] By looking up the mapping table between modulation and coding strategies and intensity information, the modulation and coding strategy corresponding to the intensity information of the next keyframe can be obtained. Specifically, the mapping table between modulation and coding strategies and intensity information is the MCS and RSRP mapping table. Since MCS is discrete and RSRP is continuous, generally, an interval of RSRP corresponds to one MCS, thus establishing an interval corresponding to one MCS.

[0091] After obtaining the new RSRP, i.e., the intensity information of the next keyframe, a comparison is performed. Since all historical RSRPs are categorized into their corresponding MCS classes, the new RSRP is identified by finding the MCS with the smallest average distance to all RSRPs within that MCS. This determines which MCS the new RSRP belongs to; that is, the MCS with the smallest distance to the historical RSRPs. If the MCS is 0 or 1, there are many RSRPs. Historical RSRPs belong to either 0 or 1. The new RSRP is then classified into MCS by whether its average distance to all historical RSRPs in 0 or 1 is minimized. If it's minimized, the new RSRP is classified as class 0, thus determining the modulation and coding strategy for the next keyframe. This achieves the goal of finding the mapping table between the modulation and coding strategy and the intensity information based on the intensity information of the next keyframe to determine the modulation and coding strategy for the next keyframe.

[0092] In this way, using a lookup table facilitates determining the modulation and coding strategy for the next keyframe based on the obtained intensity information. Furthermore, it should be noted that adjusting the coding strategy based on network awareness at the video transmitting terminal does not require waiting for feedback from the receiving end, allowing for more timely adjustments.

[0093] See Figure 6 , Figure 6 The illustration schematically shows the steps of obtaining the mapping relationship table between modulation and coding strategies and intensity information in one embodiment of this application. In one embodiment of this application, the method for obtaining the mapping relationship table between modulation and coding strategies and intensity information specifically includes step S501, which mainly includes the following steps S601 to S603.

[0094] Step S601: The terminal obtains the historical modulation and coding strategy and the historical intensity information allocated by the base station in real time, and uses the coding information corresponding to the historical modulation and coding strategy as a classification label. The historical modulation and coding strategy is used to represent the modulation and coding strategy corresponding to the historical moment.

[0095] Step S602: Cluster the historical intensity information according to the coded information classification labels to obtain a cluster classifier.

[0096] Step S603: Input the intensity values ​​in the historical intensity information value distribution range into the cluster classifier to obtain the classification information corresponding to the coding information classification label, so as to obtain the mapping relationship table between modulation coding strategy and intensity information.

[0097] The terminal records the MCS and RSRP information allocated by the base station in recent uplink transmissions obtained from the 5G module, and establishes an MCS-RSRP mapping table. The specific table construction method can employ clustering methods to minimize the average distance between the historical real MCS and the MCS obtained by looking up the mapping table using RSRP.

[0098] In this way, by obtaining the modulation and coding strategy and strength information allocated by the base station, it is beneficial to establish a mapping table between the modulation and coding strategy and the strength information.

[0099] If the terminal cannot obtain the MCS information allocated by the base station from the 5G module in real time, it can use simulation or testing methods to measure offline, establish an MCS and RSRP mapping table in advance, and store it in the terminal.

[0100] See Figure 7 , Figure 7 The illustration schematically depicts the steps in one embodiment of this application to determine the target data volume of the next keyframe by comparing the calculated data volume with the actual data volume. In one embodiment of this application, step S403, which involves comparing the calculated data volume with the actual data volume to determine the target data volume of the next keyframe, mainly includes the following steps S701 to S702.

[0101] Step S701: The terminal adjusts the amount of data calculated.

[0102] After the terminal calculates the amount of data, it adjusts the calculated amount of data, specifically by adjusting the coefficients, to ensure that the data can be transmitted normally.

[0103] Step S702: Compare the adjusted data volume with the actual data volume, and select the data volume with the smaller value as the target data volume for the next keyframe.

[0104] The adjusted data size is compared with the actual data size of the video stream. If the adjusted data size is smaller than the actual data size, the adjusted data size is selected as the target data size for the next keyframe, and intra-frame encoding is performed on the next keyframe according to the target data size. If the actual data size is smaller than the adjusted data size, the actual data size is selected as the target data size for the next keyframe, and intra-frame encoding is performed on the next keyframe according to the target data size.

[0105] In this way, by comparing the adjusted data volume with the actual data volume, the smaller data volume is selected as the target data volume for the next keyframe. Selecting the smaller data volume as the target data volume can greatly reduce transmission latency. In addition, while reducing transmission latency, the video stream can be transmitted normally.

[0106] In one embodiment of this application, after determining the target data volume of the next keyframe, further adjustments can be made to the quantization parameters to reduce the bitrate while ensuring video quality. Specifically, the current quantization parameter value is obtained; it is determined whether the output bitrate corresponding to the current quantization parameter value meets the requirements of a preset threshold. If so, no adjustment is needed; otherwise, the current quantization parameter value is adjusted to the target quantization parameter value.

[0107] The Quality Parameter (QP) is one of the key parameters in video encoding. A minimum QP value of 0 indicates the finest quantization, while a maximum QP value indicates the coarsest quantization. Typically, video content providers, such as video websites, transcode raw videos to ensure they meet the requirements for transmission and playback over the internet. Video transcoding is fundamental to almost all internet video services, including live streaming and video-on-demand. The goal of video transcoding is simple: to obtain smooth and clear video data. However, smoothness and clarity are contradictory requirements. Smoothness requires a lower bitrate, while clarity requires a higher bitrate. Video transcoding must prioritize smooth playback; on this basis, it should strive to improve the transcoded image quality and compression ratio as much as possible. Bitrate refers to the number of bits of data transmitted per unit of time, usually measured in kbps (kilobits per second). Generally, for a video, a bitrate that is too low will result in a blurry image; while a bitrate that is too high will prevent smooth playback over the internet.

[0108] When encoding video using the current quantization parameter values, if the average or instantaneous bitrate of the output video fails to meet the preset requirements, the current quantization parameter values ​​can be adjusted to obtain the target quantization parameter values. Generally, the bitrate of the output video is inversely proportional to the quantization parameter values ​​used by the video encoder; a larger quantization parameter value results in a lower bitrate for the transcoded output video, and vice versa. Therefore, the quantization parameter values ​​of the video encoder can be adjusted based on a comparison between the bitrate of the transcoded output video and the preset requirements. For example, if the bitrate of the transcoded output video is greater than the preset requirements, it indicates that the currently used quantization parameter value is too low, and it can be increased accordingly to obtain the target quantization parameter value.

[0109] See Figure 8 , Figure 8 The illustration schematically shows the steps for adjusting the calculated data volume in one embodiment of this application. In one embodiment of this application, step S701, adjusting the calculated data volume, mainly includes the following steps S801 to S802.

[0110] Step S801: The terminal obtains the current network load redundancy.

[0111] The current network load redundancy represents the current state of the network. By obtaining the current network state, it is easier to adjust the data volume in the future.

[0112] Optionally, the current network load redundancy can be dynamically adjusted based on packet loss and network bandwidth estimation results.

[0113] Step S802: The adjusted data volume is obtained by multiplying the calculated data volume, load redundancy, preset protection ratio, and preset adjustment coefficient.

[0114] Regarding the setting of the protection ratio P, for example, if the maximum target data volume is 100k, since there will be deviations in the prediction process, in order to prevent some unexpected situations from causing the data to fail to be transmitted normally, a protection ratio is set, for example, the protection ratio is set to 90%. By setting the protection ratio, the target data volume is reduced to 100*90% for transmission. Even if there are deviations in the prediction, the target data volume can still be transmitted normally. Therefore, setting a protection ratio is to reduce the impact of deviations on the final result, thereby eliminating the impact of prediction deviations.

[0115] In this way, considering the current network load redundancy, it is beneficial to obtain the adjusted data volume.

[0116] In one embodiment of this application, if the next keyframe adopts forward error correction coding, the preset adjustment coefficient is 1 / (1+R), where R represents the redundancy rate of forward error correction coding.

[0117] If the next keyframe does not use forward error correction coding, the preset adjustment factor is 1.

[0118] Since forward error correction coding (FEC) is used, the product size will increase when encoding the next keyframe. Therefore, a preset adjustment factor is used to adjust the size to obtain a target data size that can be transmitted.

[0119] In this way, by using different encoding methods and setting different preset adjustment coefficients, different application scenarios can be adapted.

[0120] See Figure 9 , Figure 9 The flowchart illustrating the steps for obtaining the current network load redundancy in one embodiment of this application is shown. In one embodiment of this application, step S801, obtaining the current network load redundancy, mainly includes the following steps S901 to S902.

[0121] Step S901: Obtain the latency information of historical keyframes.

[0122] The current network state is determined by using the latency information of historical keyframes as feedback.

[0123] Step S902: Determine the current network load redundancy based on the latency information.

[0124] By obtaining the latency information of historical keyframes, the current network load redundancy can be obtained, which is beneficial for obtaining a more accurate current network load redundancy and thus a more accurate target data volume.

[0125] In one embodiment of this application, the load redundancy of the current network is determined based on latency information. Specifically, the video transmitting terminal performs real-time statistics on the network transmission status and dynamically adjusts the load redundancy in real time. This includes: statistically analyzing the round-trip latency of N consecutive data packets within a certain period to obtain the average value and standard deviation of the round-trip latency of data packets within that initial period; statistically determining a latency threshold based on the round-trip latency of consecutive data packets within that period; obtaining the round-trip latency of a single data packet transmitted by the transmitting end; comparing the current round-trip latency of the data packet with the latency threshold obtained by statistically analyzing the transmission latency of consecutive data packets preceding that packet; if the current round-trip latency value is greater than or equal to the latency threshold, it indicates that the round-trip latency of the data packet is too long, and it is determined that the current data packet was lost during network transmission, recording a packet loss; if the current data... If the round-trip time (RTT) value is less than the time threshold, it indicates that the RTT is normal. Using a sliding window approach with a window size of M, the oldest RTT in the window is removed and the new result is added after each new data packet acknowledgment. This allows for real-time monitoring of RTT and data packet loss, resulting in a packet loss rate. Based on the statistical analysis of packet loss rates from N data packet transmissions, the average and standard deviation of the packet loss rate over that period are obtained. A reference ratio for adjusting the current network load redundancy is calculated. Based on this reference ratio, the redundancy of the video sending terminal is adjusted, updating the number of data packets and redundant packets. Again, using a sliding window approach with a window size of M, the oldest packet loss rate data in the window is removed and the new result is added after each new packet loss rate is obtained, enabling real-time monitoring of network conditions.

[0126] In this way, real-time continuous statistics based on video sending terminals are used to accurately assess network conditions and predict short-term network conditions, which serves as the basis for dynamically adjusting load redundancy. Ultimately, this decouples the system from network packet loss while reducing the bandwidth consumption of redundant packets, thereby improving network utilization and transmission efficiency.

[0127] Furthermore, a system was implemented that uses statistical calculations on packet loss at the video transmitting terminal to estimate network conditions in real time. This on-site statistical analysis of network packet loss at the video transmitting terminal reduces the delay caused by feedback from the receiving end. By calculating packet loss rate at the video transmitting terminal, the transmission results of data packets can be obtained within one round-trip time, and the packet loss statistics are available at the end of the statistical period, whereas the receiving end requires additional feedback and processing. Additionally, the mean and standard deviation used for online statistical testing are updated using a sliding window, making the statistical process more timely. Moreover, compared to transmission methods that use video frames as units and have varying numbers of data packets per encoded group, the constant-rate transmission method achieves stable and efficient data packet transmission. This ensures a constant amount of data entering the network per unit time and reduces delays caused by varying data packet transmission intervals.

[0128] In one embodiment of this application, step S901, obtaining the latency information of historical keyframes, includes:

[0129] The latency information fed back by the receiver can be used as the latency information of historical key frames, or the latency information of clearing the transmission buffer obtained by monitoring can be used as the latency information of historical key frames.

[0130] In this way, latency information can be obtained based on the latency feedback from the receiving end or the latency information of clearing the transmission buffer observed from the 5G module. This makes it easier to obtain latency information and makes the obtained latency information data more accurate.

[0131] See Figure 10 , Figure 10 The illustration schematically depicts the steps of determining the current network load redundancy based on latency information in one embodiment of this application. In one embodiment of this application, step S902, determining the current network load redundancy based on latency information, mainly includes the following steps S1001 to S1003.

[0132] Step S1001: Set the initial load redundancy.

[0133] Set the initial load redundancy to the default value, for example, 80%.

[0134] Step S1002: If the time window of the delay information is greater than the preset delay time window, then reduce the load redundancy according to the first preset step size, and use the adjusted final load redundancy as the load redundancy of the current network.

[0135] For example, if the time window of the delay information is greater than the preset delay time window, such as by 5ms, the load redundancy is reduced according to a certain step size until the time window of the delay information is equal to the preset delay time window. The final adjusted load redundancy is then used as the current network load redundancy.

[0136] Step S1003: If the time window of the delay information is less than the time window of the preset delay, the load redundancy is increased according to the second preset step size, and the adjusted final load redundancy is used as the load redundancy of the current network, wherein the first preset step size is greater than the second preset step size.

[0137] Until the time window of the delay information is less than the preset delay time window, for example, by 5ms, the load redundancy is increased according to the second preset step size until the time window of the delay information is equal to the preset delay time window. The final adjusted load redundancy is then used as the current network load redundancy.

[0138] In this way, by setting different step sizes for different latency information, the load redundancy can be adjusted to achieve rapid adjustment of the load redundancy and dynamic adjustment of the load redundancy, so as to obtain a more accurate current network load redundancy, which is conducive to obtaining a more accurate target data volume.

[0139] In one embodiment of this application, the time window for the preset delay is determined based on the delay of the forward reference frame transmission, or based on test values ​​from an idle network environment.

[0140] A Group of Pictures (GOP) refers to a group of consecutive frames in a video, serving as a frame group in video encoding. In low-latency video transmission, typically the first encoded frame within a GOP is an I-frame, and subsequent frames are forward reference frames (P-frames). The preset latency window can be determined based on the latency of the P-frames. Furthermore, the preset latency window can also be set according to the corresponding network environment under no-load conditions.

[0141] In this way, the preset delay time window is combined with the actual delay situation, making the setting more reasonable.

[0142] In one embodiment of this application, the method further includes: the terminal acquiring a media frame and placing the acquired media frame into a buffer queue; determining the total number of media frames in the buffer queue and the total length of media frames sent between the last calculated bandwidth and the current time; and calculating the current bandwidth and the current network congestion level based on the total number and length of media frames; determining the bitstream adjustment type based on the current bandwidth and network congestion level, and calculating encoding adjustment parameters, which requires calculating not only the adjusted bitstream value but also the adjusted frame rate value. This way, when the bitstream is reduced, the frame rate is also reduced accordingly, because when the bitstream is low, an excessively high frame rate is not very meaningful, and proportionally reducing the frame rate can effectively reduce the image quality degradation caused by bitstream reduction. Furthermore, when the bitstream adjustment type is reduction, the bitstream value to be reduced is calculated based on the current bandwidth, thus ensuring smooth operation while maximizing bandwidth utilization. Finally, the terminal adjusts the encoding configuration based on the calculated encoding adjustment parameters. This adaptive bandwidth-adjusted encoding configuration reduces the transmission of invalid media frames and improves smoothness.

[0143] It should be noted that although the steps of the method in this application are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0144] The following describes an apparatus embodiment of this application, which can be used to execute the video encoding method described in the above embodiments of this application. Figure 11 A schematic block diagram of a video encoding apparatus provided in an embodiment of this application is shown. Figure 11 As shown, the video encoding device 1100 includes:

[0145] The acquisition module 1110 is used to acquire historical intensity information within a historical time period. The historical intensity information is used to represent the wireless network signal strength of video transmission at each historical moment within the historical time period.

[0146] The prediction module 1120 is used to predict the strength information of the next key frame based on historical strength information. The strength information of the next key frame is used to represent the strength of the wireless network signal transmitting the next key frame.

[0147] The determination module 1130 is used to determine the target data volume of the next key frame based on the intensity information of the next key frame, and to perform intra-frame coding on the next key frame according to the target data volume.

[0148] In some embodiments of this application, based on the above technical solutions, the determining module 1130 includes:

[0149] The first determining unit is used to determine the modulation and coding strategy of the next key frame based on the intensity information of the next key frame.

[0150] The calculation unit is used to calculate the amount of data for the next key frame based on the modulation and coding strategy of the next key frame.

[0151] The second determining unit is used to determine the target data volume of the next key frame by comparing the calculated data volume with the actual data volume.

[0152] In some embodiments of this application, based on the above technical solutions, the first determining unit is used to obtain a mapping table between modulation and coding strategies and intensity information; and to determine the modulation and coding strategy of the next key frame by searching the mapping table between modulation and coding strategies and intensity information according to the intensity information of the next key frame.

[0153] In some embodiments of this application, based on the above technical solutions, the first determining unit is used to acquire historical modulation and coding strategies and historical intensity information allocated by the base station in real time, and use the coding information corresponding to the historical modulation and coding strategies as classification labels. The historical modulation and coding strategies are used to represent the modulation and coding strategies corresponding to historical moments. The historical intensity information is clustered according to the coding information classification labels to obtain a cluster classifier. The intensity values ​​in the distribution range of historical intensity information values ​​are input into the cluster classifier to obtain the classification information corresponding to the coding information classification labels, so as to obtain a mapping relationship table between modulation and coding strategies and intensity information.

[0154] In some embodiments of this application, based on the above technical solutions, the second determining unit is used to adjust the calculated data amount; compare the adjusted data amount with the actual data amount, and select the data amount with the smaller value as the target data amount of the next key frame.

[0155] In some embodiments of this application, based on the above technical solutions, the second determining unit is used to obtain the current network load redundancy; and to obtain the adjusted data volume by multiplying the calculated data volume, load redundancy, preset protection ratio and preset adjustment coefficient.

[0156] In some embodiments of this application, based on the above technical solutions, the second determining unit is used to preset the adjustment coefficient as 1 / (1+R) if the next key frame adopts forward error correction coding, where R represents the redundancy rate of forward error correction coding; and to preset the adjustment coefficient as 1 if the next key frame does not adopt forward error correction coding.

[0157] In some embodiments of this application, based on the above technical solutions, the second determining unit is used to obtain the latency information of historical key frames and determine the load redundancy of the current network based on the latency information.

[0158] In some embodiments of this application, based on the above technical solutions, the second determining unit is used to use the latency information fed back by the receiving end as the latency information of the historical key frame, or to use the latency information of clearing the transmission buffer obtained by monitoring as the latency information of the historical key frame.

[0159] In some embodiments of this application, based on the above technical solutions, the second determining unit is used to set an initial load redundancy; if the time window of the delay information is greater than the time window of the preset delay, the load redundancy is reduced according to the first preset step size, and the adjusted final load redundancy is used as the load redundancy of the current network; if the time window of the delay information is less than the time window of the preset delay, the load redundancy is increased according to the second preset step size, and the adjusted final load redundancy is used as the load redundancy of the current network; wherein, the first preset step size is greater than the second preset step size.

[0160] In some embodiments of this application, based on the above technical solutions, in the second determining unit, the time window for the preset delay is determined according to the delay of the forward reference frame transmission, or according to the test value of the no-load network environment.

[0161] In some embodiments of this application, based on the above technical solutions, the acquisition module has a sliding time window with a preset time interval between the historical time period and the encoding time of the next key frame.

[0162] The specific details of the video encoding apparatus provided in the various embodiments of this application have been described in detail in the corresponding method embodiments, and will not be repeated here.

[0163] Figure 12 A schematic block diagram of a computer system architecture for implementing an electronic device according to embodiments of the present application is shown.

[0164] It should be noted that, Figure 12 The computer system 1200 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0165] like Figure 12As shown, the computer system 1200 includes a central processing unit (CPU) 1201, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 1202 or programs loaded from storage section 1208 into random access memory (RAM). The RAM 1203 also stores various programs and data required for system operation. The CPU 1201, ROM 1202, and RAM 1203 are interconnected via a bus 1204. An input / output interface 1205 (I / O interface) is also connected to the bus 1204.

[0166] The following components are connected to the input / output interface 1205: an input section 1206 including a keyboard, mouse, etc.; an output section 1207 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1208 including a hard disk, etc.; and a communication section 1209 including a network interface card such as a local area network card, modem, etc. The communication section 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to the input / output interface 1205 as needed. A removable medium 1211, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 1210 as needed so that computer programs read from it can be installed into the storage section 1208 as needed.

[0167] Specifically, according to embodiments of this application, the processes described in the various method flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1209, and / or installed from removable medium 1211. When the computer program is executed by central processing unit 1201, it performs various functions defined in the system of this application.

[0168] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such transmitted data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0169] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0170] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0171] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the method according to the embodiments of this application.

[0172] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.

[0173] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A video encoding method, characterized in that, include: Obtain historical intensity information within a historical time period, wherein the historical intensity information is used to represent the wireless network signal strength of video transmission at each historical moment within the historical time period; The intensity information of the next key frame is predicted based on the historical intensity information, and the intensity information of the next key frame is used to represent the wireless network signal strength for transmitting the next key frame. Based on the intensity information of the next keyframe, the target data volume of the next keyframe is determined, and the encoder encoding rate is set according to the target data volume. The encoder is then used to perform intra-frame encoding on the next keyframe. The target data volume is less than or equal to the data volume carried by the uplink time slot of one frame period of the wireless network.

2. The video encoding method according to claim 1, characterized in that, Determining the target data volume of the next keyframe based on the intensity information of the next keyframe includes: Based on the intensity information of the next key frame, determine the modulation and coding strategy of the next key frame; The data volume of the next key frame is calculated based on the modulation and coding strategy of the next key frame. The target data volume for the next keyframe is determined by comparing the calculated data volume with the actual data volume.

3. The video encoding method according to claim 2, characterized in that, The step of determining the modulation and coding strategy of the next key frame based on the intensity information of the next key frame includes: Obtain the mapping table between modulation and coding strategies and intensity information; Based on the intensity information of the next key frame, the mapping table between the modulation and coding strategy and the intensity information is looked up to determine the modulation and coding strategy of the next key frame.

4. The video encoding method according to claim 3, characterized in that, The mapping table between the modulation and coding strategy and the intensity information includes: The historical modulation and coding strategies and the historical intensity information allocated by the base station are acquired in real time, and the coding information corresponding to the historical modulation and coding strategies is used as a classification label. The historical modulation and coding strategies are used to represent the modulation and coding strategies corresponding to historical moments. The historical intensity information is clustered according to the coded information classification labels to obtain a cluster classifier; The intensity values ​​in the distribution range of the historical intensity information values ​​are input into the cluster classifier to obtain the classification information corresponding to the classification label of the coding information, so as to obtain the mapping relationship table between the modulation coding strategy and the intensity information.

5. The video encoding method according to claim 2, characterized in that, The step of comparing the calculated data volume with the actual data volume to determine the target data volume for the next keyframe includes: Adjust the calculated data volume; The adjusted data volume is compared with the actual data volume, and the smaller data volume is selected as the target data volume for the next keyframe.

6. The video encoding method according to claim 5, characterized in that, The adjustment of the calculated data volume includes: Obtain the current network load redundancy; The adjusted data volume is obtained by multiplying the calculated data volume, the load redundancy, the preset protection ratio, and the preset adjustment coefficient.

7. The video encoding method according to claim 6, characterized in that, If the next keyframe uses forward error correction coding, the preset adjustment coefficient is 1 / (1+R), where R represents the redundancy rate of forward error correction coding. If the next keyframe does not use forward error correction coding, then the preset adjustment coefficient is 1.

8. The video encoding method according to claim 6, characterized in that, The process of obtaining the current network load redundancy includes: Obtain latency information for historical keyframes; The load redundancy of the current network is determined based on the latency information.

9. The video encoding method according to claim 8, characterized in that, The process of obtaining the latency information of historical keyframes includes: The latency information fed back by the receiving end is used as the latency information of the historical key frame, or the latency information of clearing the transmission buffer obtained by monitoring is used as the latency information of the historical key frame.

10. The video encoding method according to claim 8, characterized in that, Determining the load redundancy of the current network based on the latency information includes: Set initial load redundancy; If the time window of the delay information is greater than the time window of the preset delay, the load redundancy is reduced according to the first preset step size, and the adjusted final load redundancy is used as the load redundancy of the current network. If the time window of the delay information is less than the time window of the preset delay, the load redundancy is increased according to the second preset step size, and the adjusted final load redundancy is used as the load redundancy of the current network. Wherein, the first preset step size is greater than the second preset step size.

11. The video encoding method according to claim 10, characterized in that, The preset delay time window is determined based on the delay of the forward reference frame transmission, or based on the test value of the no-load network environment.

12. The video encoding method according to any one of claims 1-11, characterized in that, The historical time period and the encoding time of the next keyframe have a preset time interval sliding time window.

13. A video encoding device, characterized in that, include: The acquisition module is used to acquire historical intensity information within a historical time period, wherein the historical intensity information is used to represent the wireless network signal strength of video transmission at each historical moment within the historical time period. The prediction module is used to predict the intensity information of the next key frame based on the historical intensity information, wherein the intensity information of the next key frame is used to represent the wireless network signal strength for transmitting the next key frame. The determination module is used to determine the target data volume of the next key frame based on the intensity information of the next key frame, and set the encoder encoding rate value according to the target data volume, and perform intra-frame encoding on the next key frame through the encoder; wherein the target data volume is less than or equal to the data volume carried by the uplink time slot of one frame period of the wireless network.

14. A computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the video coding method according to any one of claims 1 to 12.

15. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the video encoding method of any one of claims 1 to 12 by executing the executable instructions.

16. A computer program product, characterized in that, The computer instructions, when executed by a processor of a computer device, cause the computer device to perform the video encoding method as described in any one of claims 1 to 12.