Media data processing method and device, electronic equipment and storage medium

By predicting the network packet loss rate and dynamically adjusting the encoding bit rate, the problem of low media data transmission efficiency is solved, and efficient media data transmission under different network quality conditions is achieved.

CN120021252APending Publication Date: 2025-05-20TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311548788.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-17
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

The real-time transmission of media data is affected by network quality. The prior art fights against high network packet loss by reducing the redundancy ratio. However, when the network surplus bandwidth is limited, the media data transmission effect is poor and it is difficult for the receiver to obtain complete media data.

Method used

By obtaining the historical packet loss rate of the historical time window, predicting the packet loss rate of the current time window, and obtaining the encoded bit rate based on the predicted packet loss rate mapping, encoding the transmission media data to generate media data that adapts to network quality.

Benefits of technology

By dynamically adjusting the encoded bit rate, the packet loss rate of media data packets during transmission is alleviated, the transmission efficiency of media data packets is improved, and the receiver can receive media data more completely, improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120021252A_ABST
    Figure CN120021252A_ABST
Patent Text Reader

Abstract

The invention provides a media data processing method and device, electronic equipment and a storage medium. The method comprises the steps that the historical packet loss rate of at least one historical time window is acquired, and the historical packet loss rate is the ratio of the number of lost media data packets in the historical time window to the number of transmitted media data packets; determining a predicted packet loss rate in the current time window based on the historical packet loss rate; mapping the predicted packet loss rate into a coding bit rate of the current time window; obtaining to-be-transmitted media data in the current time window; and coding the to-be-transmitted media data based on the coding bit rate to obtain a to-be-transmitted media data packet of the current time window. According to the invention, the transmission efficiency of the media data can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to computer technology, and in particular to a method, apparatus, electronic device, and storage medium for processing media data. Background Art

[0002] The real-time transmission of media data is affected by network quality. In related technologies, high network packet loss is countered by reducing the redundancy ratio. However, when the surplus network bandwidth is limited, by adjusting the redundancy ratio of media data packets, the expected egress bandwidth of media data will be higher than the surplus bandwidth, resulting in the deterioration of packet loss, poor transmission effect of media data, and it is difficult for the receiving party to obtain complete media data.

[0003] In related technologies, there is no good way to improve the transmission efficiency of media data. Summary of the Invention

[0004] Embodiments of this application provide a method, apparatus, electronic device, computer-readable storage medium, and computer program product for processing media data, which can improve the transmission efficiency of media data.

[0005] The technical solution of the embodiments of this application is implemented as follows:

[0006] Embodiments of this application provide a method for processing media data, the method includes:

[0007] Obtain the historical packet loss rate of at least one historical time window, where the historical packet loss rate is the ratio of the number of media data packets lost within the historical time window to the number of media data packets transmitted;

[0008] Determine the predicted packet loss rate within the current time window based on the historical packet loss rate;

[0009] Map the predicted packet loss rate to the encoding bit rate of the current time window;

[0010] Obtain the media data to be transmitted within the current time window;

[0011] Encode the media data to be transmitted based on the encoding bit rate to obtain the media data packets to be transmitted within the current time window.

[0012] Embodiments of this application provide a media data processing apparatus, including:

[0013] A data acquisition module, configured to obtain the historical packet loss rate of at least one historical time window, where the historical packet loss rate is the ratio of the number of media data packets lost within the historical time window to the number of media data packets transmitted;

[0014] A prediction module, configured to determine a predicted packet loss rate within a current time window based on the historical packet loss rate;

[0015] The prediction module is further configured to map the predicted packet loss rate to an encoded bit rate of the current time window;

[0016] A data acquisition module, configured to acquire media data to be transmitted within the current time window;

[0017] An encoding module, configured to perform encoding processing on the media data to be transmitted based on the encoded bit rate to obtain media data packets to be transmitted in the current time window.

[0018] An embodiment of the present application provides an electronic device, where the electronic device includes:

[0019] A memory, configured to store computer-executable instructions;

[0020] A processor, configured to implement the media data processing method provided by the embodiment of the present application when executing the computer-executable instructions stored in the memory.

[0021] An embodiment of the present application provides a computer-readable storage medium, storing computer-executable instructions, and configured to implement the media data processing method provided by the embodiment of the present application when being executed by a processor.

[0022] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions, and the computer program or computer-executable instructions implement the media data processing method provided by the embodiment of the present application when being executed by a processor.

[0023] The embodiment of the present application has the following beneficial effects:

[0024] By predicting the packet loss rate of the current time window through the packet loss rate within the historical time window, pre-judging the packet loss rate of the time window, and mapping to obtain the encoded bit rate based on the packet loss rate, the size of the media data packets to be transmitted is adjusted according to the network quality, thereby alleviating the packet loss rate of the media data packets during the transmission process, and further improving the transmission efficiency of the media data packets. Description of the Drawings

[0025] Figure 1 is a schematic diagram of an application mode of the media data processing method provided by the embodiment of the present application;

[0026] Figure 2 is a schematic diagram of the structure of the electronic device provided by the embodiment of the present application;

[0027] Figures 3A to 3C is a schematic flowchart of the media data processing method provided by the embodiment of the present application;

[0028] Figure 4 It is the first schematic diagram of the media data processing method provided by the embodiments of the present application;

[0029] Figure 5 It is the second schematic diagram of the media data processing method provided by the embodiments of the present application;

[0030] Figure 6 It is the flowchart of the media data processing method provided by the embodiments of the present application;

[0031] Figure 7 It is the structural schematic diagram of the media data processing system provided by the embodiments of the present application. Detailed implementation manners

[0032] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limitations on the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0033] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0034] In the following description, the terms "first / second / third" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0035] It should be noted that the collection and processing of relevant data in the present application (for example: video data uploaded by users, audio data of voices issued by users) should strictly comply with the requirements of relevant national laws and regulations during actual application, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing behaviors within the scope authorized by laws and regulations and the personal information subject.

[0036] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0037] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.

[0038] 1) Round-Trip Time (RTT): This is a concept in computer networks, which is the time it takes for a data packet to travel from the source to the destination and then back to the source.

[0039] 2) Adaptive adjustment of coding bitrate: It refers to a system or method that automatically adjusts the coding bitrate according to the current environment or conditions. Bitrate refers to the number of bits transmitted per second. Generally speaking, the higher the bitrate, the more information the digital file contains, and the better the audio quality or video quality. However, correspondingly, the required storage space and transmission bandwidth are also larger.

[0040] 3) Constant Bitrate (CBR): A term used to describe the Quality of Service (QoS) of a communication service. In the constant bitrate mode, the output bitrate of the encoder (or the input bitrate of the decoder) should be a fixed value (constant).

[0041] 4) Long Short-Term Memory (LSTM) neural network model: A special type of Recurrent Neural Network (RNN), specifically designed to handle and predict long-term dependencies in time series data. Traditional recurrent neural networks perform poorly in handling long-term dependency information due to the vanishing gradient and exploding gradient problems. The long short-term memory neural network model solves this problem by introducing a structure called "gates" (including input gates, forget gates, and output gates). These gate structures allow the model to learn when to remember, forget, or output certain information, enabling it to effectively learn and process long-term dependencies in time series data.

[0042] 5) Opus: A format for lossy audio coding, aiming to include both audio and speech in a single format, replacing Speex and Vorbis, and is suitable for low-latency real-time audio transmission over the network. The standard format is defined in RFC6716.

[0043] 6) Opus encoder: An open-source, free, and highly flexible audio coding standard suitable for the entire range of Internet audio, from as low as 2 kbps to high-fidelity audio transmission. Its excellent low-latency performance makes Opus perform well in real-time and interactive applications, such as VoIP, video conferencing, and online games.

[0044] 7) Floor Function: Also known as the floor function, it is a real function in mathematics. Its symbol is a number between two vertical bars with a horizontal line at the bottom, for example: The definition of the floor function is: for any real number x, is the largest integer not greater than x. Written as: For example: for the value 7.6, its floor function value will be 7, and 7 is the largest integer not exceeding 7.6.

[0045] 8) Self-Refreshing Transmit Terminal (SRTT): In the field of mobile communications, the main function of self-refreshing transmission is to continuously send empty data packets to inform the receiving party system to maintain the connection state for a period of time, and take appropriate measures to adjust the transmission rate when encountering network congestion or reduced link quality, thereby improving the communication effect.

[0046] 9) Loss Tolerance or Packet Loss Rate: It refers to the ratio of the number of lost data packets in data transmission to the number of data groups sent.

[0047] The real-time transmission of media data is affected by network quality. Related technologies combat high network packet loss by reducing the redundancy ratio. However, only improving the redundancy ratio will still face the problem that the surplus bandwidth is not enough to transmit data packets, resulting in packet loss and affecting the transmission efficiency.

[0048] The embodiments of the present application provide a method for processing media data, a device for processing media data, an electronic device, a computer-readable storage medium, and a computer program product, which can improve the transmission efficiency of media data.

[0049] The following describes the exemplary applications of the electronic device provided by the embodiments of the present application. The electronic device provided by the embodiments of the present application can be implemented as various types of user terminals such as terminal devices, such as laptop computers, tablet computers, desktop computers, set-top boxes, smart TVs, mobile devices (for example, mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable game devices), vehicle-mounted terminals, Virtual Reality (VR) devices, Augmented Reality (AR) devices, etc., or can also be implemented as a server. Below, the exemplary applications when the electronic device is implemented as a terminal device or a server will be described.

[0050] Reference Figure 1 , Figure 1 is a schematic diagram of the application mode of the method for processing media data provided by the embodiments of the present application; for example, Figure 1It involves a server 200, a network 300, a first terminal device 401, a second terminal device 402, and a database 500. The first terminal device 401 and the second terminal device 402 are connected to the server 200 through the network 300. The network 300 can be a wide area network, a local area network, or a combination of the two.

[0051] In some embodiments, the database 500 can be a game database, the server 200 is a server of a game platform, the first terminal device 401 and the second terminal device 402 are installed with corresponding game application programs, the user is a player, the game application program supports the user to conduct voice chat during the game, and the media data can be audio data, video data, and image data. The following will take audio data as an example for illustration.

[0052] Exemplarily, during the game, the user using the first terminal device 401 emits voice, the first terminal device 401 converts the voice into corresponding audio data, the server 200 sends the packet loss rate of the historical time window to the first terminal device 401, and the first terminal device 401 encodes the audio data by invoking the media data processing method of the embodiments of the present application based on the packet loss rate to obtain audio data packets. The first terminal device 401 sends the encoded audio data packets to the second terminal device 402 through the network 300, and the second terminal device 402 decodes the audio data packets so that the user using the second terminal device 402 can hear the voice information emitted by the user using the first terminal device 401.

[0053] In some embodiments, the media data processing method of the embodiments of the present application can also be applied to the following application scenarios: 1. Voice or video conferencing, where users access an online conference through corresponding terminal devices, and the terminal device of any speaking user encodes and processes the voice and shared video by invoking the media data processing method provided by the embodiments of the present application to obtain transmission data packets, and sends the transmission data packets to the terminal devices of other participating users. 2. Online live broadcast, where the first terminal device is used by the anchor and the second terminal device is used by the audience. The anchor conducts content live broadcast through the first terminal device. The first terminal device acquires the video and voice emitted by the anchor, encodes and processes the voice and video by invoking the media data processing method provided by the embodiments of the present application to obtain transmission data packets, and sends the transmission data packets to the second terminal device of the audience.

[0054] The embodiments of the present application can be implemented through database technology. A database, in short, can be regarded as a place for storing electronic files in an electronic filing cabinet, and users can perform operations such as adding, querying, updating, and deleting data in the files. The so-called "database" is a data set stored together in a certain way, shared by multiple users, having as little redundancy as possible, and independent of application programs.

[0055] A database management system (DBMS) is a computer software system designed to manage databases and generally has basic functions such as storage, retrieval, security, backup, etc. Database management systems can be classified according to the database models they support, such as relational, XML (Extensible Markup Language); or according to the types of computers they support, such as server clusters, mobile phones; or according to the query languages they use, such as Structured Query Language (SQL), XQuery; or according to the key performance metrics, such as maximum scale, highest running speed; or other classification methods. Regardless of the classification method used, some DBMSs can span categories. For example, they can support multiple query languages simultaneously.

[0056] In some embodiments of the present application, it can also be implemented through cloud technology. Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model, which can form a resource pool, be used on demand, and is flexible and convenient. Cloud computing technology will become an important support. The back-end services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the high development and application of the Internet industry, and the promotion of demands such as search services, social networks, mobile commerce, and open collaboration, in the future, each item may have its own hash code identification mark and needs to be transmitted to the back-end system for logical processing. Data at different levels will be processed separately, and various industry data requires a powerful system back-end support, which can only be achieved through cloud computing.

[0057] In some embodiments, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The electronic device can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal device and the server can be directly or indirectly connected through wired or wireless communication methods, and there is no limitation in the embodiments of the present application.

[0058] In some embodiments, in the game application scenario, for the solution jointly implemented by the terminal device and the server, it mainly involves two game modes, namely the local game mode and the cloud game mode. Among them, the local game mode means that the terminal device and the server jointly run the game processing logic. For the operation instructions input by the player in the terminal device, part of them are processed by the terminal device running the game logic, and the other part are processed by the server running the game logic. Moreover, the game logic processing run by the server is often more complex and requires more computing power. The cloud game mode means that the game logic processing is completely run by the server, and the cloud server renders the game scene data into an audio-video stream and transmits it to the terminal device for display through the network. The terminal device only needs to have the basic ability to play streaming media and the ability to obtain the player's operation instructions and send them to the server.

[0059] In another implementation scenario, it is applied to the terminal device (for example: the first terminal device 401 and the second terminal device 402) and the server 200, and is suitable for the application mode that depends on the computing power of the server 200 to complete virtual scene calculation and output the virtual scene on the terminal device.

[0060] Taking the formation of the visual perception of the virtual scene as an example, the server 200 calculates the virtual scene-related display data (such as scene data) and sends it to the terminal device through the network 300. The terminal device depends on the graphics computing hardware to complete the loading, parsing, and rendering of the calculation display data, and depends on the graphics output hardware to output the virtual scene to form visual perception. For example, it can present two-dimensional video frames on the display screen of a smart phone, or project video frames that achieve three-dimensional display effects on the lenses of augmented reality / virtual reality glasses; for the perception of the form of the virtual scene, it can be understood that it can be output by means of the corresponding hardware of the terminal device. For example, a microphone is used to form auditory perception, and a vibrator is used to form tactile perception, etc.

[0061] As an example, a client (such as an online game application) runs on the terminal device. During the operation of the client, a virtual scene including role-playing is output. The virtual scene can be an environment for game characters to interact. For example, it can be a plain, a street, a valley, etc. for game characters to fight. The first virtual object can be a game character controlled by the user, that is, the first virtual object is controlled by the real user and will move in the virtual scene in response to the operation of the real user on the controller (such as a touch screen, a voice control switch, a keyboard, a mouse, and a joystick, etc.). For example, when the real user moves the joystick to the right, the first virtual object will move to the right in the virtual scene, and can also stay still, jump, and control the first virtual object to perform shooting operations, etc.

[0062] In some embodiments, the terminal device can implement the media data processing method provided in the embodiments of the present application by running a computer program. For example, the computer program can be a native program or software module in the operating system; it can be a local (Native) application (APP, APPlication), that is, a program that needs to be installed in the operating system to run, such as a game APP; it can also be a small program that can be embedded in any APP, that is, a program that only needs to be downloaded to the browser environment to run. In short, the above computer program can be any form of application program, module or plug-in.

[0063] Taking the computer program as an application program as an example, in actual implementation, the terminal device installs and runs an application program that supports virtual scenarios. The application program can be any one of a first-person shooting game (FPS, First-Person Shootinggame), a third-person shooting game, a virtual reality application program, a three-dimensional map program, or a multi-player survival game. The user uses the terminal device to operate virtual objects located in the virtual scenario for activities, and the activities include, but are not limited to: adjusting the body posture, crawling, walking, running, cycling, jumping, driving, picking up, shooting, attacking, throwing, and building at least one of virtual buildings. Schematically, the virtual object can be a virtual character, such as a simulated character or an anime character.

[0064] See Figure 2 , Figure 2 is a schematic structural diagram of an electronic device provided in the embodiments of the present application. The electronic device can be a terminal device 400 (for example: Figure 1 the first terminal device 401 and the second terminal device 402 in Figure 2 The terminal device 400 shown in includes: at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. Each component in the terminal device 400 is coupled together through a bus system 440. It can be understood that the bus system 440 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in Figure 2 all kinds of buses are labeled as the bus system 440.

[0065] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0066] The user interface 430 includes one or more output devices 431 that enable the presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons, and controls.

[0067] The memory 450 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid state memory, hard disk drives, optical disk drives, etc. The memory 450 optionally includes one or more storage devices that are physically remote from the processor 410.

[0068] The memory 450 includes volatile memory or non-volatile memory, and may also include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), and the volatile memory can be random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0069] In some embodiments, the memory 450 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are described below by way of example.

[0070] The operating system 451 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic services and handling hardware-based tasks;

[0071] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include: Bluetooth, Wi-Fi (Wireless Fidelity), and USB (Universal Serial Bus), etc.;

[0072] The presentation module 453 is used to enable the presentation of information (such as a user interface for operating peripheral devices and displaying content and information) via one or more output devices 431 associated with the user interface 430 (such as a display screen, speaker, etc.).

[0073] The input processing module 454 is used to detect and translate one or more user inputs or interactions from one of the one or more input devices 432.

[0074] In some embodiments, the device provided by the embodiments of the present application may be implemented in software. Figure 2 FIG. Figure 2 shows a media data processing device 455 stored in a memory 450, which may be software in the form of a program and a plug-in, etc., including the following software modules: a data acquisition module 4551, a prediction module 4552, a data acquisition module 4553, and an encoding module 4554. These modules are logical, and thus can be arbitrarily combined or further split according to the functions to be implemented. The functions of each module will be described below.

[0075] The media data processing method provided by the embodiments of the present application will be described in combination with the exemplary applications and implementations of the terminal device provided by the embodiments of the present application.

[0076] Next, the media data processing method provided by the embodiments of the present application will be described. As before, the electronic device implementing the media data processing method of the embodiments of the present application may be a terminal device or a server, or a combination of both. Therefore, the execution subject of each step will not be repeated hereinafter.

[0077] It should be noted that in the example of the encoding process of media data hereinafter, the media data is taken as audio data for illustration. Those skilled in the art can apply the media data processing method provided by the embodiments of the present application to the processing of media data including other types of objects according to the understanding of the following text, such as: image data, video data, etc.

[0078] See Figure 3A Figure 3A FIG. Figure 3A is a schematic flowchart of the media data processing method provided by the embodiments of the present application, which will be described in combination with the steps shown in Figure 3A FIG. Figure 3A .

[0079] In step 301, at least one historical packet loss rate of a historical time window is obtained.

[0080] Here, the historical packet loss rate is the ratio of the number of media data packets lost within the historical time window to the number of media data packets transmitted.

[0081] Exemplarily, the historical time window may be a fixed duration, or a dynamic duration for sending a preset number of data packets. For example: each historical time window is 1 minute (fixed duration); for another example: each historical time window is the duration consumed for sending 10 data packets (preset number). In the embodiments of the present application, the duration for sending a preset number of data packets is taken as the time window for illustration. Therefore, the specific length of each historical time window may be the same or different, the number of media data packets sent in each historical time window is the same, and the number of media data packets lost may be different or the same.

[0082] ​Exemplarily, at least one historical time window may be the at least one historical time window closest to the current moment, and the historical packet loss rate may be the average of the packet loss rates of multiple historical time windows.

[0083] In step 302, the predicted packet loss rate within the current time window is determined based on the historical packet loss rate.

[0084] Exemplarily, if the difference between multiple historical packet loss rates is less than the difference threshold, it indicates that the network state does not change frequently, and the average of the multiple historical packet loss rates can be used as the predicted packet loss rate. If the difference between multiple historical packet loss rates is greater than or equal to the difference threshold, it indicates that the network state changes frequently, and the change trend of the packet loss rate can be determined based on at least one historical packet loss rate, and the predicted packet loss rate within the current time window can be predicted based on the change trend. The prediction process can be implemented through a neural network model, and the neural network model can be a classification model or a long short-term memory neural network model.

[0085] In some embodiments, refer to Figure 3B , Figure 3B is a schematic flowchart of the media data processing method provided by the embodiments of the present application. Figure 3A Step 302 of Figure 3B can be implemented through steps 3021 to 3024 in

[0086] In step 3021, the historical time window corresponding to each historical packet loss rate is determined, and the first quantity and the second quantity within each historical time window are obtained.

[0087] Here, the first quantity is the number of media data packets transmitted within the historical time window, and the second quantity is the number of media data packets lost within the historical time window.

[0088] Exemplarily, the principle of the neural network model predicting the probability of whether each media data packet will be lost is to perform binary classification processing on the media data packets, and the types include: lost and not lost. However, in the scenario of data transmission, whether each data packet is lost has great randomness. Predicting for each data packet requires high computing resources on the one hand and affects the accuracy on the other hand. Therefore, in the embodiments of the present application, the data packets are grouped by time window, each data packet within each time window is regarded as a data block, and the number of lost packets is predicted based on the actual number of lost packets within each time window, thereby improving the prediction accuracy.

[0089] Exemplarily, the media data packets transmitted in each historical time window can be regarded as a data block, a data block includes a first quantity of data packets, and each first quantity is the same. The number of data packets transmitted in the current time window can be determined based on the first quantity, and the number of lost packets within the current time window can be predicted based on the second quantity.

[0090] In step 3022, based on each first quantity, determine a third quantity of media data packets transmitted within the current time window.

[0091] Exemplarily, when the duration of transmitting a preset quantity is used as the time window, the third quantity is the same as the first quantity, and the quantity of media data packets transmitted in each time window is the same.

[0092] When the preset duration is used as the time window, the third quantity that may be transmitted within the current time window can be obtained by converting multiple first quantities. For example: taking the average value of the first quantities as the third quantity. Or, obtain the relationship curve between each first quantity and the number of the time window, and determine the third quantity based on the number of the current time window and the relationship curve.

[0093] In step 3023, based on the second quantity, determine a fourth quantity of data packets lost within the current time window.

[0094] Exemplarily, the third quantity is a predicted value. The third quantity can be determined by a neural network model. The neural network model is used to predict the quantity (the fourth quantity) of data packets lost in the current time window when the third quantity of data packets is transmitted and the packet loss quantities transmitted in the previous several time windows are the corresponding second quantities respectively. The neural network model can be a long short-term memory neural network model.

[0095] Before step 3023, the neural network model can be trained in the following manner: obtain a sample training set, where the sample training set includes the sample transmission quantities and sample loss quantities of data packets within multiple time windows; based on each sample transmission quantity, call the initialized neural network model for prediction processing to obtain the predicted loss quantity; based on the difference between each predicted loss quantity and the corresponding sample loss quantity, determine the loss function of the neural network model; based on the loss function, update the parameters of the initialized neural network model to obtain the trained neural network model.

[0096] Exemplarily, the loss function can be a cross-entropy loss function, and the cross-entropy loss is used to measure the difference between the actual value and the predicted value; the way to update the parameters of the initialized neural network model can be backpropagation processing, and the update process can be iterative until the loss function is reduced to a preset range.

[0097] Continue to refer to Figure 3B , in step 3024, use the ratio between the fourth quantity and the third quantity as the predicted packet loss rate within the current time window.

[0098] Exemplarily, the packet loss rate is the ratio between the transmission quantity and the loss quantity, the third quantity is a predicted value, and the ratio between the third quantity and the fourth quantity is used as the predicted packet loss rate.

[0099] In the embodiments of the present application, the loss quantity in the case of transmitting multiple data packets is predicted through a neural network model to obtain the packet loss rate corresponding to the current time window, which is more in line with the environmental factors in the data transmission scenario, avoids the interference caused by individual randomness to the prediction result, and improves the accuracy of obtaining the packet loss rate.

[0100] Continue to refer to Figure 3A , in step 303, the predicted packet loss rate is mapped to the encoding bit rate of the current time window.

[0101] Exemplarily, the mapping method may be to call a mapping function or query a preset mapping relation table.

[0102] In some embodiments, refer to Figure 3C , Figure 3C is a schematic flowchart of the media data processing method provided by the embodiments of the present application, Figure 3A step 303 of Figure 3C can be implemented through steps 3031 to 3032 in

[0103] In step 3031, the mapping relation between the candidate packet loss rate and the candidate encoding bit rate is obtained.

[0104] Exemplarily, the mapping relation represents the conversion process between the candidate packet loss rate and the candidate encoding bit rate. The mapping relation may be a mapping function or a mapping relation table.

[0105] In step 3032, based on the mapping relation, the predicted packet loss rate is mapped to obtain the encoding bit rate of the current time window.

[0106] In some embodiments, there is a negative correlation between the candidate packet loss rate and the candidate encoding bit rate. Taking the method of determining the encoding bit rate of the current time window by querying the mapping relation table as an example, in the mapping relation table, the higher the candidate packet loss rate, the lower the corresponding candidate encoding bit rate, and vice versa, the lower the candidate packet loss rate, the higher the corresponding candidate encoding bit rate. The mapping relation table may be determined by empirical values obtained through experiments.

[0107] In some embodiments, mapping processing can also be performed by calling a mapping function. Step 3032 can be implemented in the following manner: Obtain the maximum coding bit rate and the minimum coding bit rate of the encoder. Obtain the first difference between the preset value and the predicted packet loss rate. Obtain the second difference between the maximum coding bit rate and the minimum coding bit rate. Obtain the first product between the first difference and the second difference. Round up the sum of the first product and the minimum coding bit rate to obtain the coding bit rate of the current time window.

[0108] Exemplarily, taking audio coding as an example, the encoder can be an Opus encoder, which is suitable for video conferencing and online games. The rounding processing can be rounding down or rounding up. In this embodiment of the present application, rounding down is taken as an example for illustration. The preset value can be 1, and the above process can be represented by the following formula (1):

[0109] preferred_ratio = floor((1 - loss_rate) * (ratio_max - ratio_min) + ratio_min) (1)

[0110] Wherein, preferred_ratio is the predicted coding bit rate of the current time window, loss_rate is the predicted packet loss rate, (1 - loss_rate) is the first difference, ratio_max is the maximum coding bit rate, ratio_min is the minimum coding bit rate, (ratio_max - ratio_min) is the second difference, (1 - loss_rate) * (ratio_max - ratio_min) is the first product, and floor() is the rounding-down function.

[0111] In the embodiments of the present application, the coding bit rate is determined based on the mapping function, so that the coding bit rate can be proportionally superimposed on the basis of the minimum coding bit rate according to the predicted packet loss rate. On the one hand, it can make the coding bit rate of the current time window be dynamically adjusted according to the network quality, thereby improving the transmission efficiency. On the other hand, it can avoid the coding bit rate exceeding the range of the maximum coding bit rate, save the computing resources required for coding, and relieve the pressure on bandwidth and processor memory occupancy.

[0112] Continue to refer to Figure 3A In step 304, obtain the media data to be transmitted within the current time window.

[0113] Exemplarily, different types of media data to be transmitted can be obtained in different ways. For example: The audio data can be the voice input by the user or the uploaded audio file.

[0114] In some embodiments, the types of media data to be transmitted include at least one of the following: audio, video, and image. Step 304 can be implemented in the following manner: Obtain the media data to be transmitted within the current time window in at least one of the following ways:

[0115] Method 1: In response to a recording trigger operation, obtain a sound signal, convert the sound signal into a digital signal, and generate the media audio data to be transmitted within the current time window based on the digital signal.

[0116] Exemplarily, the trigger operation can be an operation such as a click, a slide, or a long press on a recording control. For example: In a game application, the user clicks on the recording control to set the recording function to be always on. The microphone of the terminal device collects the sound signal of the voice input by the user, converts the sound signal into an electrical signal, the electrical signal into a digital signal, and generates a corresponding data file from the digital signal to obtain the media audio data to be transmitted.

[0117] Method 2: In response to a sending operation for video data, use the video data as the media video data to be transmitted within the current time window.

[0118] Exemplarily, the video data can be currently recorded by the terminal device. For example: In a live broadcast or video conferencing scenario, the camera of the terminal device records the current environment in real time to generate video frames, and a sequence of video frames forms video data. The video data can also be uploaded by the user from the local storage of the terminal device.

[0119] Method 3: In response to a sending operation for image data, use the video data as the media image data to be transmitted within the current time window.

[0120] Exemplarily, the acquisition of the image data is the same as that of the video data, and will not be elaborated here.

[0121] In the embodiments of the present application, it is applicable to the transmission processing of different media data and has strong generalization ability. Therefore, the media data processing method provided by the embodiments of the present application can be applied to different application scenarios such as game voice, video conferencing, and live broadcast.

[0122] In step 305, perform encoding processing on the media data to be transmitted based on the encoding bit rate to obtain the media data packet to be transmitted in the current time window.

[0123] Exemplarily, the encoding process of the media data is performed in units of frames. Each media data packet to be transmitted can also add repeated data according to the redundancy ratio to improve the anti-interference ability of the data code.

[0124] In some embodiments, when the type of media data to be transmitted is audio, step 305 may be implemented in the following manner: Each audio frame of the media data to be transmitted is encoded based on the encoding bit rate and the preconfigured redundancy ratio to obtain an encoded audio bitstream, where the encoded audio bitstream includes encoded audio frames; The encoded audio bitstream is divided into at least one media data packet to be transmitted in the current time window.

[0125] Exemplarily, the encoding process may be implemented by an Opus encoder, and the encoding format output by the Opus encoder is Opus, which is suitable for low-latency real-time voice transmission over the network.

[0126] In some embodiments, when the type of media data to be transmitted is video, step 305 may be implemented in the following manner: Each video frame of the media data to be transmitted is encoded based on the encoding bit rate to obtain an encoded video bitstream, where the encoded video bitstream includes encoded video frames; The encoded video bitstream is divided into at least one media data packet to be transmitted in the current time window.

[0127] Exemplarily, the encoded video bitstream occupies less storage space compared to the media data to be transmitted. The media data of the uncompressed video may include a series of pictures, and each picture has spatial dimensions such as 1920×1080 luminance sampling and associated chrominance sampling. The video encoding process is used to compress and reduce the redundancy in the input video signal, and the compression helps to reduce the bandwidth or storage space requirements.

[0128] Exemplarily, for ease of understanding, the effects of the media data processing method provided in the embodiments of the present application are explained below with specific examples. For example: Multiple application programs are running in the terminal device, such as: instant messaging software, video conferencing software, office text software (such as: PPT presented in the video conference). The surplus bandwidth for the video conferencing software is 20 kbps. Assuming the bit depth of the data sample is 16 bits and the sampling rate is 16 KHz, then 16,000 samples are generated per second. Therefore, the number of bits required per second (i.e., the uplink bandwidth) is 16,000 (samples) * 16 (bit depth) = 256,000 bits / second, that is, 256 kbps. The encoding bit rate in the current time window is mapped according to the predicted network packet loss rate, and the encoding compression ratio corresponding to this encoding bit rate is 16, 256 / 16 = 16 kbps. 16 kbps is less than the surplus bandwidth, and thus the video conferencing software can run stably, reducing the possibility of packet loss after the media data is sent.

[0129] In some embodiments, after step 305, the terminal device sends the media data packets to be transmitted to the media server through the network. The media server forwards the media data packets to be transmitted to other terminal devices. The other terminal devices perform decoding processing on the media data packets to be transmitted and call a player to play the decoding result.

[0130] Exemplarily, the encoding processing and decoding processing of media data are inverse to each other. Similarly, during the decoding process, the decoding bit rate can be determined by calling the media data processing method provided in the embodiments of the present application, improving the decoding efficiency, and further improving the smoothness of the media data played by the decoding-side terminal device.

[0131] In some embodiments, after step 305, the neural network model can also be optimized in the following manner: taking the ended current time window as the historical time window; constructing new samples based on the number of lost media data packets and the number of transmitted media data packets within each historical time window, and adding the new samples to the sample training set; training the neural network model based on the sample training set added with the new samples.

[0132] In the embodiments of the present application, after the media data transmission is completed, new samples are generated using the actual packet loss rate to update the neural network model and optimize the neural network model, which can improve the accuracy of predicting the packet loss rate and further improve the transmission efficiency.

[0133] In the embodiments of the present application, the packet loss rate of the current time window is predicted through the packet loss rate within the historical time window, and the packet loss rate of the time window is pre-judged. The encoding bit rate is mapped based on the packet loss rate, so that the size of the transmitted media data packets is adjusted according to the network quality, thereby alleviating the packet loss rate of the media data packets during the transmission process, and further improving the transmission efficiency of the media data packets. It has stronger smoothness in the case of a weak network, and at the same time reduces the bandwidth required to transmit media data, reducing bandwidth occupancy. In the game application scenario, it can improve the smoothness of the game program. Reducing the packet loss rate enables the receiving party to receive the media data more completely, reducing the information difference between the receiving party and the sending party, and improving the user experience.

[0134] Next, the exemplary application of the media data processing method in the embodiments of the present application in an actual application scenario will be described.

[0135] During the real-time transmission of media data between terminal devices, the data received by the receiving terminal device is affected by network state fluctuations. In related technologies, media data is encoded at a fixed coding bit rate, and the redundancy ratio is reduced to combat high network packet loss. However, in extreme environments, simply reducing the redundancy ratio is difficult to resist network state fluctuations, which in turn leads to an increase rather than a decrease in the packet loss rate and poor media data transmission effects. An embodiment of this application proposes a media data processing method that combines learning the mutation law of network bandwidth to obtain the optimal coding bit rate in the network situation after self-refresh transmission. Taking the transmission of audio data as an example, in a network environment with a specific packet loss rate, the optimal coding bit rate helps to improve the transmission success rate, and thus greatly improves the fluency of speech.

[0136] The media data processing method provided by the embodiments of this application will be explained below in conjunction with the accompanying drawings. Refer to Figure 4 , Figure 4 which is the first schematic diagram of the principle of the media data processing method provided by the embodiments of this application; the terminal device of the data sender collects the user's voice or the uploaded video to generate audio or video source data. The adaptive coding manager 404 of the terminal device includes a mapping module 4041 and a media encoder 4042. The mapping module 4041 predicts the predicted packet loss rate of the current time window based on the packet loss feedback information (including the packet loss rate within the historical time window) sent by the terminal device of the data receiver, and maps the predicted packet loss rate to the coding bit rate. The media encoder 4042 encodes and processes the video / audio source data based on the coding bit rate to obtain the encoded data packets, and sends the encoded data packets to the terminal device of the data receiver through the network 300.

[0137] Refer to Figure 6 , Figure 6 which is the flowchart of the media data processing method provided by the embodiments of this application. The media data processing method provided by the embodiments of this application is executed by the terminal device. Taking the terminal device as the execution subject and the scenario of audio data transmission as an example, an explanation will be given.

[0138] In step 601, the source audio data is obtained.

[0139] For example, in a game application scenario, the terminal device can collect the voice emitted by the user through the microphone and convert the sound signal into the source audio data.

[0140] In some embodiments, a media data processing system can be integrated in the electronic device that executes the media data processing method provided by the embodiments of this application. Refer to Figure 7 , Figure 7It is a schematic structural diagram of a media data processing system provided by an embodiment of the present application. The media data processing system 700 includes: an audio acquisition module 701, a bit rate prediction module 702, an audio encoding module 703, a forward error correction module 705, a packet loss statistics module 704, and an audio transmission module 706.

[0141] The audio acquisition module 701 is used to acquire the audio input by the user to the terminal device (for example: collected or uploaded by the microphone), and transmit the original audio data (pcm) to the audio encoding module 703. The packet loss statistics module 704 is used to count the packet loss rate of the audio in the current network environment, that is, to count the historical packet loss rate of the already sent data packets. The bit rate prediction module 702 predicts the encoding bit rate of the best encoding process based on the historical packet loss rate. The prediction process can be implemented in the following way: use the trained neural network model to infer the most suitable encoding bit rate for the current network environment. The audio encoding module 703 encodes according to the predicted encoding bit rate.

[0142] The forward error correction module 705 calculates the audio packet with redundancy through the set redundancy ratio for the audio transmission module 706 to send. Forward error correction is a method to increase the reliability of data communication. By sending additional information together with the data for error recovery, the bit error rate is reduced. Redundancy, generally speaking, is the degree of repetition of data. The repeated data in a data set is called data redundancy, and the ratio between the repeated data and the total data is the redundancy. Redundancy is in data transmission. Due to attenuation or interference, the data code will mutate. At this time, it is necessary to improve the anti-interference ability of the data code so that the corresponding data has a certain redundancy.

[0143] The audio transmission module 706 sends the audio packet with redundancy processed by the forward error correction module 705 to the media server, and the media server forwards it to the terminal device of the receiving party.

[0144] In step 602, forward error correction processing is performed on the source audio data to obtain a redundancy ratio.

[0145] Exemplarily, the meaning of forward error correction processing has been explained above and will not be elaborated here.

[0146] In step 603, the packet loss rate of the audio data in the current network environment is counted.

[0147] Exemplarily, the packet loss rate in the current network environment can characterize the network quality in the current network environment. The packet loss rate is statistically calculated in units of time windows, and the time window can be a preset time period or the time consumed to transmit a preset number of data packets. In the embodiments of the present application, the time consumed to transmit a preset number of data packets is taken as an example for illustration. For example: The data is chunked (package window) with 10 packets as a data block, and the transmission process of each data block corresponds to a time window. Also, the number of lost voice packets in each block (10 voice packets as a data block) is used as training data.

[0148] 3. In the field of audio transmission, it is of little value to specifically analyze the loss status of a certain packet. To combat unpredictable loss patterns, in the embodiments of the present application, 10 packets are taken as a time window (i.e., the number of historical data packets observed in the past), and the number of lost packets in each time window is used as the training and prediction target, thus solving the problem that it is difficult for the long short-term memory neural network model to predict the loss status of a certain packet with extremely strong randomness. For example: Multiple voice packets are divided into multiple chunks according to the time sequence, and the number of lost packets in each chunk is used as the input data of the long short-term memory neural network model to predict the number of lost packets in the upcoming next chunk (corresponding to the next time window).

[0149] Exemplarily, the execution of step 603 is before step 604, and step 602 can also be executed after step 603 or step 604.

[0150] In step 604, the optimal coding bit rate in the current network environment is predicted based on the packet loss rate.

[0151] Exemplarily, after predicting the number of lost packets, the predicted packet loss rate can be determined based on the predicted number of lost packets, and mapping processing is performed according to the defined mapping relationship in the manner of low coding bit rate for high packet loss rate and high coding bit rate for low packet loss rate to obtain the optimal coding bit rate.

[0152] For ease of understanding, refer to Figure 5 , Figure 5It is the second schematic diagram of the media data processing method provided by the embodiments of this application; one data packet chunk corresponds to one time window. The data packet chunk 501 includes 10 data packets. The loss quantity (packet loss quantity) corresponding to the data packet chunk 501 is counted. The long short-term memory neural network model 502 predicts the loss quantity of the current time window based on the loss quantity of the historical time window. The loss quantity of the current time window is subjected to mapping processing to obtain the encoded bit rate. Based on the packet loss quantity of the voice packet window sequence in the historical network environment as data, a long short-term memory neural network model is trained. The packet loss data of the voice window in a period of time before the current moment online is used as the input data for prediction, and the packet loss data of the next voice window is output. According to the mapping relationship, the packet loss data is mapped to the corresponding encoded bit rate to be adjusted.

[0153] Exemplarily, the mapping processing can be implemented in the following manner: obtain the encoded bit rate range of the encoder, obtain the bit rate difference between the minimum bit rate and the maximum bit rate, determine the predicted non-packet-loss rate based on the predicted packet loss rate. The predicted non-packet-loss rate is the packet loss rate difference between 1 and the predicted packet loss rate. Multiply the bit rate difference by the packet loss rate difference, and perform a floor operation on the sum of the product and the minimum bit rate to obtain the encoded bit rate. That is, starting from the minimum encoded bit rate, the desired encoded bit rate is increased proportionally.

[0154] The above process can be represented by the following formula (1):

[0155] preferred_ratio = floor((1 - loss_rate) * (ratio_max - ratio_min) + ratio_min) (1)

[0156] Among them, preferred_ratio is the predicted encoded bit rate, loss_rate is the predicted packet loss rate, (1 - loss_rate) is the predicted non-packet-loss rate, and floor() is the floor function.

[0157] In the embodiments of this application, 12 kbps <= Encoding_Ratio <= 96 kbps is taken as an example for illustration, where Encoding_Ratio is the encoded bit rate. Among them, 12 kbps is the minimum bit rate (ratio_min), and 96 kbps is the maximum bit rate (ratio_max). In the embodiments of this application, the predicted packet loss quantity is converted into the packet loss rate of the current window, and the packet loss rate of the current window is converted into the predicted encoded bit rate. The encoded bit rate is proportionally adjusted to a certain value between 12 kbps and 96 kbps, and the decimal part is directly removed using the floor function (downward rounding function).

[0158] In step 605, encoding is performed on the source data based on the optimal encoding bit rate and redundancy ratio to obtain an encoded data packet.

[0159] For example, the principle of encoding processing is referred to the introduction of the forward error correction module above, and will not be repeated here.

[0160] In step 606, the encoded data packet is sent to the receiving end.

[0161] For example, the terminal device sends the encoded data packet to the receiving terminal device through the network, realizing the real-time transmission of media data. Compared with the strategy of counteracting weak network based on the redundancy ratio of the forward error correction module in the related art, the media data processing method provided by the embodiment of the present application has a long short-term memory neural network model, which can adaptively adjust the encoding bit rate and can more effectively cope with various different network environments.

[0162] For example: the surplus bandwidth for voice in the current game scenario is 20kbps. Assume that the bit depth (the number of bits occupied by a data sample) is 16 bits, that is, each sample occupies 16 bits. Then, since the sampling rate is 16KHz, that is, 16,000 samples are generated per second, the number of bits required per second (i.e., the uplink bandwidth) is 16,000 (samples) * 16 (bit depth) = 256,000 bits / second, i.e., 256Kbps. Assume that the encoding compression ratio corresponding to the encoding bit rate of the opus encoder at this time is 11, 256 / 11 = 23Kbps. Then, even if the forward error correction module is completely turned off at this time, the surplus bandwidth cannot send out the audio after the current encoding because it exceeds 3kbps. At this time, the packet loss rate will surge. The relevant technology simply increases the redundancy ratio of the forward error correction module, and the expected export bandwidth of voice is higher than 23Kbps, then the packet loss situation will be worse.

[0163] For the above scenario, the media data processing method provided in the embodiment of the present application timely reduces the encoding bit rate according to the packet loss situation. Under the same forward error correction module redundancy ratio, the volume of the encoded output voice packet will be much smaller and will not exceed the surplus bandwidth limit, thus ensuring the flow of voice, preventing the deterioration of packet loss, and avoiding voice freeze and loss caused by packet loss.

[0164] In some embodiments, the media data processing method provided in the embodiments of the present application can be applied to the voice engine (GVoice) game voice. When providing game voice services, in order to better adapt to network environments with different packet loss rates, the adaptive encoding bit rate adjustment system of the embodiments of the present application can better adapt to the network environment. In high packet loss rate scenarios, an optimal low encoding bit rate configuration is selected to encode audio, and in low packet loss rate scenarios, a higher encoding bit rate configuration is selected.

[0165] ​In the embodiments of the present application, the packet loss rate in the current time window is predicted based on the packet loss rate in the historical time window, and the coding bit rate is determined based on the predicted packet loss rate, so that the coding bit rate changes dynamically following the network quality, making the media data transmission more fluent in the case of a weak network, and reducing the bandwidth required for media data transmission to a certain extent. During real-time transmission such as voice and video, the stuttering during the playback of media data at the receiving end is reduced, improving the user experience.

[0166] The following continues to describe the exemplary structure of the media data processing device 455 provided in the embodiments of the present application as a software module. In some embodiments, as Figure 2 shown, the software module in the media data processing device 455 stored in the memory 450 may include: a data acquisition module 4551, configured to obtain the historical packet loss rate of at least one historical time window, where the historical packet loss rate is the ratio of the number of media data packets lost in the historical time window to the number of media data packets transmitted; a prediction module 4552, configured to determine the predicted packet loss rate in the current time window based on the historical packet loss rate; the prediction module 4552 is further configured to map the predicted packet loss rate to the coding bit rate of the current time window; a data acquisition module 4553, configured to obtain the media data to be transmitted in the current time window; and a coding module 4554, configured to perform coding processing on the media data to be transmitted based on the coding bit rate to obtain the media data packets to be transmitted in the current time window.

[0167] In some embodiments, the prediction module 4552 is configured to determine the historical time window corresponding to each historical packet loss rate, obtain the first quantity and the second quantity in each historical time window, where the first quantity is the number of media data packets transmitted in the historical time window, and the second quantity is the number of media data packets lost in the historical time window; determine the third quantity of the media data packets transmitted in the current time window based on each first quantity; determine the fourth quantity of the data packets lost in the current time window based on the second quantity; and use the ratio between the fourth quantity and the third quantity as the predicted packet loss rate in the current time window.

[0168] In some embodiments, the third quantity is determined by a neural network model; a prediction module 4552, configured to obtain a sample training set before determining a fourth quantity of data packets lost in the current time window based on the second quantity, where the sample training set includes sample transmission quantities and sample loss quantities of data packets within a plurality of time windows; perform a prediction process on the initialized neural network model based on each of the sample transmission quantities to obtain predicted loss quantities; determine a loss function of the neural network model based on the difference between each predicted loss quantity and the corresponding sample loss quantity; and update parameters of the initialized neural network model based on the loss function to obtain a trained neural network model.

[0169] In some embodiments, the prediction module 4552 is configured to, after encoding the media data to be transmitted based on the encoding bit rate to obtain media data packets to be transmitted in the current time window, use the ended current time window as a historical time window; construct new samples based on the number of lost media data packets and the number of transmitted media data packets within each historical time window, add the new samples to the sample training set; and train the neural network model based on the sample training set with the new samples added thereto.

[0170] In some embodiments, the prediction module 4552 is configured to obtain a mapping relationship between a candidate packet loss rate and a candidate encoding bit rate, where the mapping relationship characterizes a conversion process between the candidate packet loss rate and the candidate encoding bit rate; and perform a mapping process on the predicted packet loss rate based on the mapping relationship to obtain the encoding bit rate of the current time window.

[0171] In some embodiments, there is a negative correlation between the candidate packet loss rate and the candidate encoding bit rate.

[0172] In some embodiments, the prediction module 4552 is configured to obtain a maximum encoding bit rate and a minimum encoding bit rate of an encoder; obtain a first difference between a preset value and the predicted packet loss rate; obtain a second difference between the maximum encoding bit rate and the minimum encoding bit rate; obtain a first product between the first difference and the second difference; and perform a rounding process on the sum of the first product and the minimum encoding bit rate to obtain the encoding bit rate of the current time window.

[0173] In some embodiments, the type of the media data to be transmitted includes at least one of the following: audio, video, and image. A data acquisition module 4553 is configured to acquire media data to be transmitted in the current time window by at least one of the following methods:

[0174] In response to a recording trigger operation, a sound signal is acquired, the sound signal is converted into a digital signal, and audio data to be transmitted within the current time window is generated based on the digital signal; in response to a sending operation for video data, the video data is used as the video data to be transmitted within the current time window; in response to a sending operation for image data, the video data is used as the image data to be transmitted within the current time window.

[0175] In some embodiments, when the type of the media data to be transmitted is audio, an encoding module 4554 is configured to perform encoding processing on each audio frame of the media data to be transmitted based on the encoding bit rate and preconfigured redundancy, so as to obtain an encoded audio bitstream, where the encoded audio bitstream includes encoded audio frames; and divide the encoded audio bitstream into at least one media data packet to be transmitted within the current time window.

[0176] In some embodiments, when the type of the media data to be transmitted is video, an encoding module 4554 is configured to perform encoding processing on each video frame of the media data to be transmitted based on the encoding bit rate, so as to obtain an encoded video bitstream, where the encoded video bitstream includes encoded video frames; and divide the encoded video bitstream into at least one media data packet to be transmitted within the current time window.

[0177] An embodiment of the present application provides a computer program product, which includes a computer program or computer executable instructions, and the computer program or computer executable instructions are stored in a computer-readable storage medium. A processor of an electronic device reads the computer program or computer executable instructions from the computer-readable storage medium, and the processor executes the computer program or computer executable instructions, so that the electronic device executes the media data processing method described above in the embodiments of the present application.

[0178] An embodiment of the present application provides a computer-readable storage medium storing computer executable instructions, where computer executable instructions or a computer program are stored, and when the computer executable instructions or the computer program are executed by a processor, the processor will be caused to execute the media data processing method provided by the embodiments of the present application. For example, Figure 3A the media data processing method shown.

[0179] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or may be various devices including one or any combination of the above memories.

[0180] In some embodiments, the computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as a stand-alone program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0181] As an example, the computer-executable instructions may or may not correspond to a file in the file system, and may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a HyperText Markup Language (HTML) document, stored in a single file dedicated to the program in question, or, stored in multiple cooperating files (for example, files that store one or more modules, subroutines, or portions of code).

[0182] As an example, the executable instructions may be deployed to execute on one electronic device, or on multiple electronic devices located at one location, or, on multiple electronic devices distributed across multiple locations and interconnected by a communication network.

[0183] In summary, in the embodiments of the present application, the packet loss rate of the current time window is predicted through the packet loss rate within the historical time window, the packet loss rate of the time window is pre-judged, and the coding bit rate is obtained based on the packet loss rate mapping, so that the size of the transmitted media data packet is adjusted according to the network quality, thereby alleviating the packet loss rate of the media data packet during the transmission process, and further improving the transmission efficiency of the media data packet. It has stronger fluency in the face of weak network conditions, and at the same time reduces the bandwidth required to transmit media data, reducing bandwidth occupancy, and can improve the fluency of game programs in game application scenarios. Reducing the packet loss rate enables the receiving party to receive the media data more completely, reduces the information difference between the receiving party and the sending party, and improves the user experience.

[0184] The above is only the embodiments of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.

Claims

1. A media data processing method, characterized in that: The method comprises: Obtaining a historical packet loss rate of at least one historical time window, wherein the historical packet loss rate is a ratio of the number of media data packets lost to the number of media data packets transmitted within the historical time window; Determining a predicted packet loss rate within a current time window based on the historical packet loss rate; Mapping the predicted packet loss rate to the encoding bit rate of the current time window; Acquire the media data to be transmitted within the current time window; The media data to be transmitted is encoded based on the encoding bit rate to obtain a media data packet to be transmitted in the current time window.

2. The method according to claim 1, characterized in that The determining the predicted packet loss rate in the current time window based on the historical packet loss rate includes: Determine the historical time windows corresponding to each of the historical packet loss rates, and obtain a first quantity and a second quantity in each of the historical time windows, wherein the first quantity is the number of media data packets transmitted in the historical time window, and the second quantity is the number of media data packets lost in the historical time window; determining, based on each of the first quantities, a third quantity of media data packets transmitted in the current time window; determining a fourth number of packets lost in the current time window based on the second number; The ratio between the fourth number and the third number is used as the predicted packet loss rate in the current time window.

3. The method according to claim 2, characterized in that The third quantity is determined by a neural network model; Before determining a fourth number of data packets lost in the current time window based on the second number, the method further includes: Acquire a sample training set, where the sample training set includes a sample transmission quantity and a sample loss quantity of data packets within a plurality of time windows; Based on each of the sample transmission quantities, the initialized neural network model is called to perform prediction processing to obtain a predicted loss quantity; Determining a loss function of the neural network model based on a difference between each of the predicted loss quantities and the corresponding sample loss quantity; The parameters of the initialized neural network model are updated based on the loss function to obtain a trained neural network model.

4. The method according to claim 3, characterized in that After encoding the media data to be transmitted based on the encoding bit rate to obtain the media data packet to be transmitted in the current time window, the method further includes: Taking the ended current time window as the historical time window; Based on the number of media data packets lost and the number of media data packets transmitted in each of the historical time windows, construct a new sample, and add the new sample to the sample training set; The neural network model is trained based on the sample training set to which the new sample is added.

5. The method according to claim 1, characterized in that Mapping the predicted packet loss rate to the encoding bit rate of the current time window includes: Acquire a mapping relationship between a candidate packet loss rate and a candidate encoding bit rate, wherein the mapping relationship represents a conversion process between the candidate packet loss rate and the candidate encoding bit rate; The predicted packet loss rate is mapped based on the mapping relationship to obtain the encoding bit rate of the current time window.

6. The method according to claim 5, characterized in that There is a negative correlation between the candidate packet loss rate and the candidate encoding bit rate.

7. The method according to claim 5, characterized in that The mapping process is performed on the predicted packet loss rate based on the mapping relationship to obtain the encoding bit rate of the current time window, including: Get the maximum encoding bit rate and the minimum encoding bit rate of the encoder; Obtaining a first difference between a preset value and the predicted packet loss rate; Obtaining a second difference between the maximum encoding bit rate and the minimum encoding bit rate; Obtaining a first product between the first difference and the second difference; The sum of the first product and the minimum coding bit rate is rounded to obtain the coding bit rate of the current time window.

8. The method according to any one of claims 1 to 6, characterized in that: The type of the media data to be transmitted includes at least one of the following: audio, video, and image, and the acquiring the media data to be transmitted in the current time window includes: Acquire the media data to be transmitted in the current time window by at least one of the following methods: In response to a recording trigger operation, acquiring a sound signal, converting the sound signal into a digital signal, and generating the audio data to be transmitted within the current time window based on the digital signal; In response to a sending operation on video data, taking the video data as video data to be transmitted within the current time window; In response to a sending operation for image data, the video data is used as image data to be transmitted within the current time window.

9. The method according to any one of claims 1 to 6, characterized in that: When the type of the media data to be transmitted is audio, encoding the media data to be transmitted based on the encoding bit rate to obtain a media data packet to be transmitted in the current time window includes: Encoding each audio frame of the media data to be transmitted based on the encoding bit rate and the preconfigured redundancy ratio to obtain an encoded audio bit stream, wherein the encoded audio bit stream includes the encoded audio frame; The encoded audio bit stream is divided into at least one media data packet to be transmitted in the current time window.

10. The method according to any one of claims 1 to 6, characterized in that: When the type of the media data to be transmitted is video, encoding the media data to be transmitted based on the encoding bit rate to obtain a media data packet to be transmitted in the current time window includes: Encoding each video frame of the media data to be transmitted based on the encoding bit rate to obtain an encoded video bit stream, wherein the encoded video bit stream includes the encoded video frames; The encoded video bitstream is divided into at least one media data packet to be transmitted in the current time window.

11. A media data processing device, characterized in that: The device comprises: A data collection module, used to obtain a historical packet loss rate of at least one historical time window, wherein the historical packet loss rate is a ratio of the number of media data packets lost to the number of media data packets transmitted within the historical time window; A prediction module, configured to determine a predicted packet loss rate within a current time window based on the historical packet loss rate; The prediction module is further used to map the predicted packet loss rate to the encoding bit rate of the current time window; A data acquisition module, used to acquire the media data to be transmitted within the current time window; The encoding module is used to encode the media data to be transmitted based on the encoding bit rate to obtain the media data packet to be transmitted in the current time window.

12. An electronic device, characterized in that: The electronic device comprises: A memory for storing computer executable instructions; The processor is used to implement the media data processing method according to any one of claims 1 to 10 when executing the computer executable instructions or computer programs stored in the memory.

13. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the media data processing method according to any one of claims 1 to 10 is implemented.

14. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the media data processing method according to any one of claims 1 to 10 is implemented.