Audio and video transmission method, device, electronic device and storage medium based on network bandwidth

By obtaining the historical bandwidth change curve of the target user and using the user matching model to match the encoding bit rate and transmission mode of the sample user, the problem of audio and video quality fluctuation caused by inaccurate network bandwidth prediction in the existing technology is solved, and more efficient and reliable audio and video data transmission is achieved.

CN120547369BActive Publication Date: 2025-09-26GUANGDONG CHANGE INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511037989.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-09-26
Estimated Expiration
2045-07-28

AI Technical Summary

Technical Problem

In the existing technology, the prediction method based on network bandwidth is not accurate enough in audio and video data transmission, resulting in frequent adjustment of encoding bit rate, causing fluctuations in audio and video quality and low encoding efficiency.

Method used

By obtaining the historical bandwidth change curve of the target user and using the user matching model to match the encoding bit rate and transmission mode of the sample user, the audio and video data are transmitted to the target user, avoiding inaccurate prediction of future bandwidth and adopting the target encoding bit rate and transmission mode for transmission.

Benefits of technology

It improves coding efficiency and audio and video quality, reduces quality fluctuations, and improves transmission reliability, especially when bandwidth changes frequently.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120547369B_ABST
    Figure CN120547369B_ABST
Patent Text Reader

Abstract

The present invention discloses a network bandwidth-based audio and video transmission method, device, electronic device and storage medium. The method inputs a historical bandwidth change curve of a target user into a user pairing model to obtain a target user pair, wherein the target user pair includes a target user and a sample user. The method matches the sample user through the target user's historical bandwidth change curve, and transmits audio and video data with reference to the target encoding rate and target transmission mode of the sample user. There is no need to dynamically adjust the encoding rate based on the current network bandwidth to predict the future bandwidth, thereby avoiding the problem of frequent fluctuations and inaccuracies in the predicted network bandwidth, which leads to frequent adjustment of the encoding rate and causes drastic fluctuations in audio and video quality and low encoding efficiency. The method can improve the audio and video encoding efficiency and quality, and transmits the target audio and video data to the target user with reference to the encoding rate and transmission mode of the sample user who has played the audio and video data, thereby improving the reliability of audio and video transmission.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio and video transmission, and in particular to a method, device, electronic device and storage medium for audio and video transmission based on network bandwidth. Background Art

[0002] In the fields of audio and video transmission technologies such as remote monitoring, remote operation, and multi-party video, cameras and microphones collect audio and video data and upload them to a server, which then distributes them to user terminals that request to play the audio and video data.

[0003] Currently, after audio and video data is uploaded to the server, the server stores it in a database (non-real-time audio and video services) or immediately forwards it to each user terminal (real-time audio and video services). When decomposing audio and video data, the server often considers the network bandwidth of the user terminal and dynamically adjusts the encoding bit rate of the audio and video data to adapt to the network bandwidth of the user terminal.

[0004] The method of transmitting audio and video data by dynamically adjusting the encoding bit rate based on network bandwidth usually predicts the user's network bandwidth based on the user's current network status. For example, the current packet loss rate, delay and other parameters are used to predict the user's network bandwidth. This method of network bandwidth prediction has no factual reference, and the accuracy of the predicted network bandwidth is difficult to guarantee. In addition, the predicted network bandwidth fluctuates frequently, resulting in drastic fluctuations in audio and video quality and the encoder needs to frequently adjust the encoding bit rate, which reduces encoding efficiency. Summary of the Invention

[0005] The present invention provides a network bandwidth-based audio and video transmission method, device, electronic device and storage medium, which can transmit audio and video data to the target user based on the historical bandwidth changes of the target user and with reference to the target encoding rate and transmission mode of sample users who have played the audio and video data.

[0006] In a first aspect, the present invention provides an audio and video transmission method based on network bandwidth, which is applied to a server and includes:

[0007] Upon receiving a target user's audio or video playback request, determining the target audio or video data and the target playback time point;

[0008] Acquire a historical bandwidth change curve of the target user, where the historical bandwidth change curve is a bandwidth change curve when the target user plays audio and video data on the network at the target playback time point within a first historical time period;

[0009] Inputting the historical bandwidth change curve into a user pairing model to obtain a target user pair, wherein the target user pair includes the target user and sample users, wherein the sample users are users who received audio and video data transmitted by the server at the target encoding rate and target transmission mode during the second historical time period;

[0010] The target audio and video data are transmitted to the target user using the target encoding bit rate and the target transmission mode.

[0011] Optionally, obtaining a historical bandwidth change curve of the target user includes:

[0012] Determining at least one audio or video data on the network played by the target user at a target playback time point within a first historical time period;

[0013] The bandwidth of the target user when playing the audio and video data is sampled, and a historical bandwidth change curve of the target user when playing the audio and video data is generated with the sampling time as the horizontal axis and the bandwidth obtained by sampling as the vertical axis.

[0014] Optionally, the user pairing model is trained in the following manner:

[0015] Determine a sample user, and obtain a bandwidth change curve when the sample user plays audio and video data on the network as a bandwidth change curve sample, and obtain a target encoding bit rate and a target transmission mode of the server when playing the audio and video data, wherein the sample user is a user who scores the clarity and smoothness of the played audio and video data higher than a preset score;

[0016] Generating user pair samples using the sample users, each of the user pair samples including a first sample user and a second sample user, wherein the first sample user and the second sample user are sample users having a bandwidth change curve sample similarity greater than a first threshold, a target encoding rate difference less than a second threshold, and the same target transmission mode;

[0017] Randomly extracting N bandwidth change curve samples of the first user pairs from the user pair samples and inputting them into the user pairing model to generate N second user pairs;

[0018] Calculating loss values ​​using the N first user pairs and second user pairs;

[0019] Determine whether the preset training conditions are met;

[0020] If so, determining that the user pairing model has completed training, and storing the bandwidth feature of the bandwidth change curve sample of the sample user in a feature library;

[0021] If not, adjust the model parameters of the user pairing model according to the loss value, and return to the step of randomly extracting bandwidth change curve samples of N user pair samples and inputting them into the user pairing model to generate N user pairs.

[0022] Optionally, randomly extracting N bandwidth variation curve samples of first user pairs from the user pair samples and inputting them into a user pairing model to generate N second user pairs includes:

[0023] Randomly extract N bandwidth change curve samples of the first user pair from the user pair samples and input them into the user pairing model;

[0024] extracting bandwidth features from bandwidth change curve samples in the user pairing model;

[0025] Calculating the feature similarity of bandwidth features between each sample user and other sample users in the N first user pairs, and calculating the absolute value of the difference between the target encoding bit rate of each sample user and other sample users;

[0026] Calculating a weighted sum using the feature similarity, the absolute value of the difference, a preset curve weight, and an encoding rate weight as the similarity between each sample user and other sample users;

[0027] The two sample users with the greatest similarity are used to generate the second user pair.

[0028] Optionally, the user pairing model is provided with a feature library including bandwidth features of each sample user, and further includes an input layer, a feature extraction layer, a feature similarity calculation layer, and an output layer connected in sequence. Inputting the historical bandwidth change curve into the user pairing model to obtain a target user pair includes:

[0029] Inputting the historical bandwidth change curve into the input layer of the user pairing model;

[0030] Extracting bandwidth features from the historical bandwidth change curve in the feature extraction layer;

[0031] Calculating feature similarity between the extracted bandwidth feature and each feature in the feature library in the feature similarity calculation layer;

[0032] The target sample user with the greatest feature similarity is determined in the output layer, and a user pair including the target user and the target sample user is output.

[0033] Optionally, transmitting the target audio and video data to the target user using the target encoding bit rate and the target transmission mode includes:

[0034] If the target transmission mode is a combined audio and video transmission mode, encoding the target audio and video data using the target encoding bit rate to obtain an audio and video stream;

[0035] The audio and video stream is transmitted to the target user through a first transmission path.

[0036] Optionally, the target transmission mode includes an audio and video separation transmission mode, the target encoding bit rate includes an audio encoding bit rate and a video encoding bit rate, and transmitting the target audio and video data to the target user using the target encoding bit rate and the target transmission mode includes:

[0037] If the target transmission mode is an audio and video separation transmission mode, separating audio data and video data from the target audio and video data;

[0038] Encoding the audio data at the audio encoding rate to obtain an audio stream, and encoding the video data at the video encoding rate to obtain a video stream;

[0039] The audio stream is transmitted to the target user through a first transmission path, and the video stream is transmitted to the target user through a second transmission path, wherein the first transmission path and the second transmission path are different routing paths from a server to the target user.

[0040] In a second aspect, the present invention provides an audio and video transmission device based on network bandwidth, comprising:

[0041] The play request response module is used to determine the target audio and video data and the target play time point when receiving the audio and video play request of the target user;

[0042] a historical bandwidth change curve acquisition module, configured to acquire a historical bandwidth change curve of the target user, wherein the historical bandwidth change curve is a bandwidth change curve when the target user plays audio and video data on the network at the target playback time point within a first historical time period;

[0043] a user pair generation module, configured to input the historical bandwidth change curve into a user pairing model to obtain a target user pair, wherein the target user pair includes the target user and a sample user, wherein the sample user is a user who receives audio and video data transmitted by the server at a target encoding rate and a target transmission mode during a second historical time period;

[0044] The audio and video data transmission module is used to transmit the target audio and video data to the target user by adopting the target encoding bit rate and the target transmission mode.

[0045] In a third aspect, the present invention provides an electronic device, comprising:

[0046] at least one processor; and

[0047] a memory communicatively connected to the at least one processor; wherein,

[0048] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the network bandwidth-based audio and video transmission method described in any one of the first aspects of the present invention.

[0049] In a fourth aspect, the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the network bandwidth-based audio and video transmission method described in any one of the first aspects of the present invention when executed.

[0050] The present invention inputs a historical bandwidth change curve of a target user into a user pairing model to obtain a target user pair, wherein the target user pair includes a target user and a sample user, wherein the sample user is a user who receives audio and video data transmitted by a server at a target encoding rate and a target transmission mode. The target audio and video data is transmitted to the target user using the target encoding rate and the target transmission mode, thereby matching the sample user with the target user's historical bandwidth change curve, and transmitting the audio and video data with reference to the target encoding rate and the target transmission mode of the sample user. This eliminates the need to dynamically adjust the encoding rate based on the current network bandwidth to predict the future bandwidth, thereby avoiding the problem of frequent fluctuations and inaccuracies in the predicted network bandwidth, which results in drastic fluctuations in audio and video quality and low encoding efficiency caused by frequent encoding rate adjustments. This improves encoding efficiency and audio and video quality, and transmits the target audio and video data to the target user with reference to the encoding rate and transmission mode of the sample user who has played the audio and video data, thereby improving the reliability of audio and video transmission.

[0051] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0053] Figure 1This is a flow chart of a method for audio and video transmission based on network bandwidth provided in Example 1 of the present invention;

[0054] Figure 2 This is a flow chart of a method for audio and video transmission based on network bandwidth provided by the second embodiment of the present invention;

[0055] Figure 3 is a schematic diagram of the bandwidth variation curve;

[0056] Figure 4 It is a schematic diagram of the user pairing model;

[0057] Figure 5 It is a schematic diagram of the transmission path;

[0058] Figure 6 This is a schematic structural diagram of a network bandwidth-based audio and video transmission device provided in a third embodiment of the present invention;

[0059] Figure 7 It is a structural diagram of an electronic device provided in Embodiment 4 of the present invention. DETAILED DESCRIPTION

[0060] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0061] Example 1

[0062] Figure 1 A flowchart of a method for transmitting audio and video based on network bandwidth is provided in the first embodiment of the present invention. Figure 1 As shown, the audio and video transmission method based on network bandwidth includes:

[0063] S101. Upon receiving an audio or video playback request from a target user, determine target audio or video data and a target playback time point.

[0064] This embodiment can be applied to a server, and an audio and video encoder can be set on the server. The audio and video encoder can be a hardware encoder or a software encoder. After the camera in scenarios such as remote monitoring and video conferencing collects audio and video data, it can be stored in the server first. When the audio and video playback request of the target user is received, the audio and video identifier in the audio and video playback request can be used to search the target audio and video data in the database. The audio and video data can refer to media data including audio and / or video, wherein the target playback time point can refer to the time point when the audio and video playback request is received. It should be noted that the target playback time point refers to any time of the day and does not include the date. For example, the time when the audio and video playback request is received is 14:30 on July 14, 2025, and the playback time point refers to 14:30.

[0065] S102: Obtain a historical bandwidth change curve of the target user, where the historical bandwidth change curve is a bandwidth change curve when the target user plays audio and video data on the network at a target playback time point within a first historical time period.

[0066] The historical bandwidth change curve may refer to a curve showing bandwidth changes when the target user plays audio and video data on the network in a first historical time period (e.g., the past month, week, or day). Alternatively, it may be a curve showing bandwidth changes when the audio and video data on the network is played at the target playback time point. In this embodiment, when the server transmits audio and video data to each user, it may record bandwidth changes when each user plays the audio and video data to obtain a historical bandwidth change curve.

[0067] S103: Input the historical bandwidth change curve into the user pairing model to obtain a target user pair, where the target user pair includes a target user and sample users, where the sample users are users who receive audio and video data transmitted by the server at the target encoding rate and target transmission mode during the second historical time period.

[0068] The user matching model of this embodiment includes a feature library of bandwidth variation curves, which includes bandwidth features of bandwidth variation curves of sample users. The sample users may be users who played audio and video data over the network during a second historical time period (a month, a week, or a day) and rated the clarity and smoothness of the playback above a preset threshold. In other words, the sample users are users who are satisfied with the audio and video transmission. In another embodiment, the sample users may also be users whose packet loss rate and latency when playing audio and video data over the network are below a preset threshold. It should be noted that when a sample user plays multiple audio and video data over the network during the second historical time period, the sample user has multiple bandwidth features in the feature library, each of which is associated with a corresponding encoding bit rate and transmission mode. The user matching model can be a model that, after training, can generate user pairs with similar bandwidth variation curve features. Furthermore, the second historical time period can be the same as the first historical time period, and the audio and video data played by the sample user can be the target audio and video data.

[0069] After obtaining the historical bandwidth change curve of the target user, the historical bandwidth change curve can be input into the user matching model. The bandwidth features of the target user in the user matching model are extracted, and the extracted bandwidth features are matched with the bandwidth features in the feature library to obtain a sample user with the greatest similarity to the bandwidth features of the target user. Each bandwidth feature of the sample user is associated with a target encoding bit rate and a target transmission mode, indicating that under the network bandwidth represented by each bandwidth feature, when the audio and video data is encoded and transmitted to the sample user at the target encoding bit rate and target transmission mode, the audio and video transmission can be smooth and the clarity can meet the needs of the sample user. The encoding bit rate refers to the bit rate used for encoding audio and video, and the transmission mode can include a joint audio and video transmission mode and an audio and video separate transmission mode.

[0070] S104: Transmit the target audio and video data to the target user using the target encoding rate and target transmission mode.

[0071] Specifically, if the target transmission mode is a combined audio and video transmission mode, the target audio and video data can be directly encoded using the target encoding bit rate to generate an audio and video stream, which is then transmitted to the target user via the same transmission path. If the target transmission mode is a separate audio and video transmission mode, the audio and video data is separated into audio data and video data, which are then encoded using the target encoding bit rate to generate an audio stream and a video stream, respectively. The audio stream is transmitted to the target user via a first transmission path, and the video stream is transmitted to the target user via a second transmission path. After receiving the audio and video streams, the target user's terminal decodes and combines them to generate the audio and video streams, and then plays the audio and video streams.

[0072] The present invention inputs a historical bandwidth change curve of a target user into a user pairing model to obtain a target user pair, wherein the target user pair includes a target user and a sample user, wherein the sample user is a user who receives audio and video data transmitted by a server at a target encoding rate and a target transmission mode. The target audio and video data is transmitted to the target user using the target encoding rate and the target transmission mode, thereby matching the sample user with the target user's historical bandwidth change curve, and transmitting the audio and video data with reference to the target encoding rate and the target transmission mode of the sample user. This eliminates the need to dynamically adjust the encoding rate based on the current network bandwidth to predict the future bandwidth, thereby avoiding the problem of frequent fluctuations and inaccuracies in the predicted network bandwidth, which results in drastic fluctuations in audio and video quality and low encoding efficiency caused by frequent encoding rate adjustments. This improves encoding efficiency and audio and video quality, and transmits the target audio and video data to the target user with reference to the encoding rate and transmission mode of the sample user who has played the audio and video data, thereby improving the reliability of audio and video transmission.

[0073] Example 2

[0074] Figure 2 A flowchart of a method for transmitting audio and video based on network bandwidth is provided in the second embodiment of the present invention, such as Figure 2 As shown, the audio and video transmission method based on network bandwidth includes:

[0075] S201. Upon receiving an audio or video playback request from a target user, determine target audio or video data and a target playback time point.

[0076] In this embodiment, the server can store audio and video data collected by cameras in scenarios such as remote monitoring and video conferencing. Each audio and video data is set with an identifier, such as an audio and video ID. When an audio and video playback request is received, the target audio data can be searched in the database through the audio and video ID in the request, and the time point when the audio and video playback request is received can be determined as the target playback time point.

[0077] S202: Determine at least one piece of audio or video data on the network played by the target user at a target playback time point within a first historical time period.

[0078] For example, all audio and video data played by the target user in a first historical time period (such as the past month, week or day) can be searched through the user ID of the target user, and all audio and video data can be filtered to obtain at least one audio and video data played at the target playback time point. For example, audio and video data with a playback time greater than a preset time length and played at the target playback time point can be filtered out, such as audio and video data with a viewing time greater than 10 minutes played by the target user at the target playback time point can be filtered out, so as to avoid the situation where the playback time is too short and has no reference significance for generating a bandwidth change curve.

[0079] It should be noted that the audio and video data played at the target playback time point may mean that the time period for playing the audio data includes the target playback time point, or the time period for playing the audio data does not include the target time point, but the duration between the start time point or the end time point of the time period for playing the audio data and the target playback time point is less than the preset duration, for example, the preset duration is 1 hour.

[0080] If the number of audio and video data finally determined is greater than 1, the audio and video data played most recently from the current moment (the moment when the audio and video playback request is received) can be selected as the final audio and video data, or the audio and video data with the shortest duration between the start time point of playback and the target playback time point can be selected.

[0081] S203: Sampling the bandwidth of the target user when playing the audio and video data, and generating a historical bandwidth change curve of the target user when playing the audio and video data with the sampling time as the horizontal axis and the bandwidth obtained by sampling as the vertical axis.

[0082] For example, the bandwidth of the user playing audio and video data can be sampled by embedding points. For example, when the server sends audio and video data to the user, it can sample the user's bandwidth every 0.5 minutes, 1 minute, etc., and then generate a bandwidth change curve with the sampling time as the horizontal axis and the bandwidth as the vertical axis, such as Figure 3 The figure shows a schematic diagram of a bandwidth variation curve, where the horizontal axis includes the broadcast date and the bandwidth value at each sampling time point under the broadcast date.

[0083] S204: Input the historical bandwidth change curve into the input layer of the user pairing model.

[0084] The user pairing model of this embodiment may be a neural network model for generating user pairs with similar bandwidth variation curves. That is, after inputting the historical bandwidth variation curve of a target user, the user pairing model may match the target user with sample users to generate user pairs. The user pairing model is trained through the following steps:

[0085] S1. Determine a sample user, and obtain a bandwidth change curve when the sample user plays audio and video data on the network as a bandwidth change curve sample, and obtain a target encoding bit rate and a target transmission mode of the server when playing the audio and video data, wherein the sample user is a user who scores the clarity and smoothness of the played audio and video data higher than a preset score.

[0086] In this embodiment, the sample user may be a user who has received and played audio and video data from the server, and the sample user may be a user who scores the clarity and smoothness of the played audio and video data higher than the preset score, that is, the sample user is a user who is satisfied with the encoding bit rate and transmission mode used in the audio and video transmission process. Of course, the sample user may also be a user whose packet loss rate during audio and video transmission is less than the preset packet loss rate and whose delay is less than the preset delay threshold.

[0087] After determining the sample users, the bandwidth change curve when the sample users play audio and video data on the network can be obtained as a bandwidth change curve sample, as well as the target encoding bit rate and target transmission mode of the server for the audio and video data when playing the audio and video data.

[0088] S2. Generate user pair samples using sample users. Each user pair sample includes a first sample user and a second sample user. The first sample user and the second sample user are sample users for which the similarity of bandwidth change curve samples is greater than a first threshold, the difference in target encoding bit rates is less than a second threshold, and the first sample user and the second sample user have the same target transmission mode.

[0089] In one embodiment, based on a manual labeling operation, a user pair sample including a first sample user and a second sample user can be generated in response to a manual selection operation. For example, the bandwidth change curve samples of the two sample users are manually compared (curve similarity, playback time period, etc.), the encoding bit rate and transmission mode are compared, and then two sample users are selected to generate a user pair sample.

[0090] Of course, the bandwidth change curve, encoding bit rate and transmission mode of the sample user can also be input into the feature extraction network to extract features and then calculate the feature similarity, and determine two sample users with feature similarity greater than a threshold as user pair samples.

[0091] S3. Randomly extract N bandwidth variation curve samples of first user pairs from the user pair samples and input them into a user pairing model to generate N second user pairs.

[0092] After initializing the user pairing model, bandwidth change curves of multiple sample users can be randomly extracted and input into the user pairing model to generate a second user pair. Specifically, bandwidth change curve samples of N first user pairs can be randomly extracted from the user pair samples and input into the user pairing model. Bandwidth features are extracted from the bandwidth change curve samples in the user pairing model. The feature similarity of the bandwidth features of each sample user and other sample users in the N first user pairs is calculated, and the absolute value of the difference between the target encoding bit rate of each sample user and other sample users is calculated. A weighted sum is calculated using the feature similarity, the absolute value of the difference, a preset curve weight, and the encoding bit rate weight as the similarity between each sample user and other sample users. The two sample users with the greatest similarity are used to generate the second user pair. For example, among multiple sample users, the two sample users with the greatest similarity are first selected to generate the second user pair, and then the two sample users with the greatest similarity are again selected from the remaining sample users to generate the second user pair. This process is repeated until all sample users are paired to generate the second user pair.

[0093] Exemplarily, bandwidth change curves of three first user pairs (A1, A2), (B1, B2), and (C1, C2), a total of six sample users A1, A2, B1, B2, C1, and C2, are extracted and input into a user pairing model. Bandwidth features are extracted from the bandwidth change curves of A1, A2, B1, B2, C1, and C2 in the user pairing model, and similarities are calculated to generate second user pairs, resulting in three second user pairs (A1, A2), (B1, C2), and (C1, B2).

[0094] S4. Calculate loss values ​​using N first user pairs and N second user pairs.

[0095] The loss value represents the degree of difference between the generated second user pair and the first user pair. See the above example. The sample users B1, B2, C1, and C2 are incorrectly paired. When calculating the loss value, the ratio of the total number of incorrectly paired sample users to the total number of extracted sample users can be calculated as the loss value. As in the above example, the loss value can be 4 / 6, which is approximately equal to 0.66. Of course, the ratio of incorrect user pairs to total user pairs can also be used as the loss value. As in the above example, the loss value can be 2 / 3, which is approximately equal to 0.66.

[0096] S5. Determine whether the preset training conditions are met.

[0097] The preset training condition may be that the number of iterative training reaches a preset number, or the loss value is less than a preset threshold. When the preset training condition is met, S6 may be executed, and when the preset training condition is not met, S7 may be executed.

[0098] S6. Determine that the user pairing model has completed training, and store the bandwidth features of the bandwidth change curve samples of the sample users in a feature library.

[0099] After determining that the user pairing model has completed training, a feature library can be established to store the bandwidth features of the bandwidth change curve samples of the sample users in the feature library, and the encoding bit rate and transmission mode of the audio and video data of the sample users are associated and stored, such as the bandwidth features are associated and stored with the encoding bit rate and transmission mode.

[0100] S7. Adjust the model parameters of the user pairing model according to the loss value and return to S3.

[0101] For example, the loss value can be used to adjust the model parameters of the model through various gradient descent algorithms, and return to S3. The method of adjusting the model parameters can refer to the existing supervised training method and will not be described in detail here.

[0102] After the user pairing model is trained, the historical bandwidth change curve of the target user can be input into the input layer of the user pairing model. Figure 4 As shown, the user pairing model of this embodiment includes an input layer, a feature extraction layer, a feature similarity calculation layer, and an output layer connected in sequence. The feature extraction layer is used to extract bandwidth features from the bandwidth change curve, the feature similarity calculation layer is used to calculate the similarity between the extracted bandwidth features and the features in the feature library, and the output layer is used to output the user IDs corresponding to the two bandwidth features with the greatest similarity to generate a user pair.

[0103] S205 . Extract bandwidth features from the historical bandwidth change curve in a feature extraction layer.

[0104] In this embodiment, the bandwidth change curve is an image, and the feature extraction layer may include a multi-layer convolutional neural network. In the feature extraction layer, a convolution operation can be performed on the bandwidth change curve through the multi-layer convolutional neural network, such as a downsampling convolution operation to obtain a feature map as a bandwidth feature.

[0105] S206 : Calculate feature similarity between the extracted bandwidth feature and each feature in the feature library in the feature similarity calculation layer.

[0106] Specifically, in the feature similarity calculation layer, the distance (Euclidean distance, Manhattan distance, etc.) between the extracted bandwidth feature and each feature in the feature library can be calculated as the feature similarity.

[0107] S207. Determine the target sample user with the greatest feature similarity in the output layer, and output a user pair including the target user and the target sample user. The sample user is a user who receives audio and video data transmitted by the server at the target encoding rate and target transmission mode during the second historical time period.

[0108] After calculating the feature similarity between the extracted bandwidth feature and each feature in the feature library, the target sample user with the greatest feature similarity can be determined, and a user pair including the target user ID and the target sample user ID can be output. The target encoding bit rate and target transmission mode associated with the bandwidth feature of the target sample user can be found through the target sample user ID.

[0109] S208: If the target transmission mode is the audio and video co-transmission mode, encode the target audio and video data using the target encoding bit rate to obtain an audio and video stream.

[0110] If the target transmission mode is the audio and video co-transmission mode, it means that the network bandwidth changes of the target user and the target sample user are similar and relatively abundant, and audio and video can be transmitted simultaneously through the same transmission path. The target audio and video data can be encoded at the target encoding bit rate to obtain an audio and video stream.

[0111] S209: Transmit the audio and video stream to the target user through the first transmission path.

[0112] like Figure 5 Figure a is a schematic diagram of the transmission path. The server to the target user may include multiple network nodes, and the audio and video streams can be transmitted to the target user through the first transmission path L1 (which can be the default network communication path between the server and the target user).

[0113] S210: If the target transmission mode is the audio and video separation transmission mode, separate the audio data and the video data from the target audio and video data.

[0114] If the target transmission mode is the audio and video separation transmission mode, it means that the network bandwidth changes of the target user and the target sample user are similar and the bandwidth is insufficient. Using the same transmission path to transmit too much audio and video data may cause audio and video playback to be stuck. Audio and video need to be transmitted separately through two transmission paths. Therefore, audio data and video data need to be separated from the target audio and video data. The separation of audio data and video data from audio and video data can refer to the existing technology and will not be described in detail here.

[0115] S211 . Encode the audio data using an audio encoding rate to obtain an audio stream, and encode the video data using a video encoding rate to obtain a video stream.

[0116] The target encoding bit rate may include an audio encoding bit rate and a video encoding bit rate. The audio encoding bit rate may be used to encode audio data to obtain an audio stream, and the video encoding bit rate may be used to encode video data to obtain a video stream.

[0117] S212: Transmit the audio stream to the target user through a first transmission path, and transmit the video stream to the target user through a second transmission path, where the first transmission path and the second transmission path are different routing paths from the server to the target user.

[0118] like Figure 5 Figure b in the figure is a schematic diagram of the transmission path. The server to the target user may include multiple network nodes. The audio stream can be transmitted to the target user through the first transmission path L1, and the video stream can be transmitted to the target user through the second transmission path L2. After receiving the audio stream and video stream, the target user's terminal decodes and synthesizes them to obtain the audio and video streams, and plays the audio and video streams. The first transmission path L1 and the second transmission path L2 can be the two paths with the least amount of data currently transmitted among the multiple network paths from the server to the target user.

[0119] In this embodiment, the historical bandwidth change curve of the target user is input into the user matching model to obtain a target user pair. The target user pair includes the target user and a sample user. The sample user is a user who receives audio and video data transmitted by the server at a target encoding rate and a target transmission mode. The target audio and video data is transmitted to the target user using the target encoding rate and the target transmission mode. This achieves matching the sample user with the target user's historical bandwidth change curve, and transmits the audio and video data with reference to the target encoding rate and the target transmission mode of the sample user. This eliminates the need to dynamically adjust the encoding rate based on the current network bandwidth to predict the future bandwidth. This avoids frequent and inaccurate fluctuations in the predicted network bandwidth, which in turn leads to drastic fluctuations in audio and video quality and low encoding efficiency caused by frequent encoding rate adjustments. This improves encoding efficiency and audio and video quality. In addition, the target audio and video data is transmitted to the target user using the encoding rate and transmission mode of the sample user, with reference to the sample user who has played audio and video data. This improves the reliability of audio and video transmission.

[0120] Furthermore, the transmission mode can include an audio and video separation transmission mode. When the bandwidth of a transmission path from the server to the target user is too small, audio and video can be transmitted separately through two transmission paths, avoiding the problem of reducing the encoding bit rate due to the bandwidth of one transmission path being too small, resulting in a decrease in audio and video quality, and can simultaneously ensure the quality and speed of audio and video transmission.

[0121] Example 3

[0122] Figure 6 This is a structural diagram of an audio and video transmission device based on network bandwidth provided by the third embodiment of the present invention. Figure 6 As shown, the audio and video transmission device based on network bandwidth includes:

[0123] The play request response module 601 is used to determine the target audio and video data and the target play time point when receiving the audio and video play request from the target user;

[0124] A historical bandwidth change curve acquisition module 602 is configured to acquire a historical bandwidth change curve of the target user, wherein the historical bandwidth change curve is a bandwidth change curve when the target user plays audio and video data on the network at the target playback time point within the first historical time period;

[0125] A user pair generation module 603 is configured to input the historical bandwidth change curve into a user pairing model to obtain a target user pair, where the target user pair includes the target user and sample users, where the sample users are users who received audio and video data transmitted by the server at the target encoding rate and target transmission mode during the second historical time period;

[0126] The audio and video data transmission module 604 is configured to transmit the target audio and video data to the target user by using the target encoding bit rate and the target transmission mode.

[0127] Optionally, the historical bandwidth change curve acquisition module 602 includes:

[0128] An audio and video data acquisition unit, configured to determine at least one audio and video data on the network played by the target user at a target playback time point within a first historical time period;

[0129] The bandwidth sampling unit is used to sample the bandwidth of the target user when playing audio and video data, and generate a historical bandwidth change curve of the target user when playing the audio and video data with the sampling time as the horizontal axis and the sampled bandwidth as the vertical axis.

[0130] Optionally, a model training module is also included for:

[0131] Determine a sample user, and obtain a bandwidth change curve when the sample user plays audio and video data on the network as a bandwidth change curve sample, and obtain a target encoding bit rate and a target transmission mode of the server when playing the audio and video data, wherein the sample user is a user who scores the clarity and smoothness of the played audio and video data higher than a preset score;

[0132] Generating user pair samples using the sample users, each of the user pair samples including a first sample user and a second sample user, wherein the first sample user and the second sample user are sample users having a bandwidth change curve sample similarity greater than a first threshold, a target encoding rate difference less than a second threshold, and the same target transmission mode;

[0133] Randomly extracting N bandwidth change curve samples of the first user pairs from the user pair samples and inputting them into the user pairing model to generate N second user pairs;

[0134] Calculating loss values ​​using the N first user pairs and second user pairs;

[0135] Determine whether the preset training conditions are met;

[0136] If so, determining that the user pairing model has completed training, and storing the bandwidth feature of the bandwidth change curve sample of the sample user in a feature library;

[0137] If not, adjust the model parameters of the user pairing model according to the loss value, and return to the step of randomly extracting bandwidth change curve samples of N user pair samples and inputting them into the user pairing model to generate N user pairs.

[0138] Optionally, the model training module is further configured to include:

[0139] Randomly extract N bandwidth change curve samples of the first user pair from the user pair samples and input them into the user pairing model;

[0140] extracting bandwidth features from bandwidth change curve samples in the user pairing model;

[0141] Calculating the feature similarity of bandwidth features between each sample user and other sample users in the N first user pairs, and calculating the absolute value of the difference between the target encoding bit rate of each sample user and other sample users;

[0142] Calculating a weighted sum using the feature similarity, the absolute value of the difference, a preset curve weight, and an encoding rate weight as the similarity between each sample user and other sample users;

[0143] The two sample users with the greatest similarity are used to generate the second user pair.

[0144] Optionally, the user pairing model is provided with a feature library including bandwidth features of each sample user, and further includes an input layer, a feature extraction layer, a feature similarity calculation layer, and an output layer connected in sequence. The user pair generation module 603 includes:

[0145] a bandwidth change curve input unit, configured to input the historical bandwidth change curve into an input layer of a user pairing model;

[0146] a bandwidth feature extraction unit, configured to extract bandwidth features from the historical bandwidth change curve in the feature extraction layer;

[0147] A feature similarity calculation unit, configured to calculate feature similarity between the extracted bandwidth feature and each feature in the feature library in the feature similarity calculation layer;

[0148] The user pair generating unit is configured to determine a target sample user with the greatest feature similarity in the output layer, and output a user pair including the target user and the target sample user.

[0149] Optionally, the audio and video data transmission module 604 includes:

[0150] A first encoding unit is configured to encode the target audio and video data using the target encoding bit rate to obtain an audio and video stream if the target transmission mode is an audio and video co-transmission mode;

[0151] The first transmission unit is used to transmit the audio and video stream to the target user through a first transmission path.

[0152] Optionally, the target transmission mode includes an audio and video separation transmission mode, the target encoding bit rate includes an audio encoding bit rate and a video encoding bit rate, and the audio and video data transmission module 604 includes:

[0153] a separation unit, configured to separate audio data and video data from the target audio and video data if the target transmission mode is an audio and video separation transmission mode;

[0154] a second encoding unit, configured to encode the audio data using the audio encoding rate to obtain an audio stream, and to encode the video data using the video encoding rate to obtain a video stream;

[0155] The second transmission unit is used to transmit the audio stream to the target user through a first transmission path, and to transmit the video stream to the target user through a second transmission path, wherein the first transmission path and the second transmission path are different routing paths from the server to the target user.

[0156] The network bandwidth-based audio and video transmission device provided in the embodiment of the present invention can execute the network bandwidth-based audio and video transmission method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0157] Example 4

[0158] Figure 7This is a schematic diagram of the structure of an electronic device 70 provided in Example 4 of the present application. The electronic device 70 may specifically include: at least one processor 71, at least one memory 72, a power supply 73, a communication interface 74, an input / output interface 75, and a communication bus 76. The memory 72 is used to store a computer program, which is loaded and executed by the processor 71 to implement the relevant steps of the network bandwidth-based audio and video transmission method disclosed in any of the aforementioned embodiments. In addition, the electronic device 70 in this embodiment may specifically be an electronic computer.

[0159] In this embodiment, the power supply 73 is used to provide operating voltage for each hardware device on the electronic device 70; the communication interface 74 can create a data transmission channel between the electronic device 70 and external devices. The communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not specifically limited here; the input and output interface 75 is used to obtain external input data or output data to the outside world. Its specific interface type can be selected according to specific application needs and is not specifically limited here.

[0160] Furthermore, memory 72, as a resource storage medium, may be read-only memory, random access memory, a magnetic disk, or an optical disk. The resources stored therein may include an operating system 721, a computer program 722, and the like. The storage may be either transient or permanent. Operating system 721 is used to manage and control the hardware devices on electronic device 70, as well as computer program 722. Operating system 721 may be Windows Server, NetWare, Unix, Linux, or the like. Computer program 722 may include computer programs capable of implementing the network bandwidth-based audio and video transmission method disclosed in any of the aforementioned embodiments and executed by electronic device 70, and may further include computer programs capable of performing other specific tasks.

[0161] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when executed by a processor, the computer program implements the aforementioned method for audio and video transmission based on network bandwidth. The specific steps of this method can be referred to the corresponding contents disclosed in the aforementioned embodiments and will not be repeated here.

[0162] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. Reference can be made to the descriptions of the identical or similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions of the methods.

[0163] Those skilled in the art may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0164] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0165] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0166] The above is a detailed introduction to the technical solution provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the ideas of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A method for transmitting audio and video based on network bandwidth, characterized in that: Applicable to servers, including: Upon receiving a target user's audio or video playback request, determining the target audio or video data and the target playback time point; Acquire a historical bandwidth change curve of the target user, where the historical bandwidth change curve is a bandwidth change curve when the target user plays audio and video data on the network at the target playback time point within a first historical time period; Inputting the historical bandwidth change curve into a user pairing model to obtain a target user pair, wherein the target user pair includes the target user and sample users, wherein the sample users are users who received audio and video data transmitted by the server at the target encoding rate and target transmission mode during the second historical time period; The target audio and video data are transmitted to the target user using the target encoding bit rate and the target transmission mode.

2. The method according to claim 1, characterized in that Obtaining a historical bandwidth change curve of the target user includes: Determining at least one audio or video data on the network played by the target user at a target playback time point within a first historical time period; The bandwidth of the target user when playing the audio and video data is sampled, and a historical bandwidth change curve of the target user when playing the audio and video data is generated with the sampling time as the horizontal axis and the bandwidth obtained by sampling as the vertical axis.

3. The method according to claim 1, characterized in that The user pairing model is trained in the following way: Determine a sample user, and obtain a bandwidth change curve when the sample user plays audio and video data on the network as a bandwidth change curve sample, and obtain a target encoding bit rate and a target transmission mode of the server when playing the audio and video data, wherein the sample user is a user who scores the clarity and smoothness of the played audio and video data higher than a preset score; Generating user pair samples using the sample users, each of the user pair samples including a first sample user and a second sample user, wherein the first sample user and the second sample user are sample users having a bandwidth change curve sample similarity greater than a first threshold, a target encoding rate difference less than a second threshold, and the same target transmission mode; Randomly extracting N bandwidth change curve samples of the first user pairs from the user pair samples and inputting them into the user pairing model to generate N second user pairs; Calculating loss values ​​using the N first user pairs and second user pairs; Determine whether the preset training conditions are met; If so, determining that the user pairing model has completed training, and storing the bandwidth feature of the bandwidth change curve sample of the sample user in a feature library; If not, adjust the model parameters of the user pairing model according to the loss value, and return to the step of randomly extracting bandwidth change curve samples of N user pair samples and inputting them into the user pairing model to generate N user pairs.

4. The method according to claim 3, characterized in that Randomly extracting N bandwidth change curve samples of first user pairs from the user pair samples and inputting them into the user pairing model to generate N second user pairs includes: Randomly extract N bandwidth change curve samples of the first user pair from the user pair samples and input them into the user pairing model; extracting bandwidth features from bandwidth change curve samples in the user pairing model; Calculating the feature similarity of bandwidth features between each sample user and other sample users in the N first user pairs, and calculating the absolute value of the difference between the target encoding bit rate of each sample user and other sample users; Calculating a weighted sum using the feature similarity, the absolute value of the difference, a preset curve weight, and an encoding rate weight as the similarity between each sample user and other sample users; The two sample users with the greatest similarity are used to generate the second user pair.

5. The method according to any one of claims 1 to 4, characterized in that The user pairing model is provided with a feature library including bandwidth features of each sample user, and further includes an input layer, a feature extraction layer, a feature similarity calculation layer, and an output layer connected in sequence. The historical bandwidth change curve is input into the user pairing model to obtain a target user pair, including: Inputting the historical bandwidth change curve into the input layer of the user pairing model; Extracting bandwidth features from the historical bandwidth change curve in the feature extraction layer; Calculating feature similarity between the extracted bandwidth feature and each feature in the feature library in the feature similarity calculation layer; The target sample user with the greatest feature similarity is determined in the output layer, and a user pair including the target user and the target sample user is output.

6. The method according to any one of claims 1 to 4, characterized in that Transmitting the target audio and video data to the target user using the target encoding bit rate and the target transmission mode includes: If the target transmission mode is a combined audio and video transmission mode, encoding the target audio and video data using the target encoding bit rate to obtain an audio and video stream; The audio and video stream is transmitted to the target user through a first transmission path.

7. The method according to any one of claims 1 to 4, characterized in that The target transmission mode includes an audio and video separation transmission mode, the target encoding bit rate includes an audio encoding bit rate and a video encoding bit rate, and the target audio and video data are transmitted to the target user using the target encoding bit rate and the target transmission mode, including: If the target transmission mode is an audio and video separation transmission mode, separating audio data and video data from the target audio and video data; Encoding the audio data at the audio encoding rate to obtain an audio stream, and encoding the video data at the video encoding rate to obtain a video stream; The audio stream is transmitted to the target user through a first transmission path, and the video stream is transmitted to the target user through a second transmission path, wherein the first transmission path and the second transmission path are different routing paths from a server to the target user.

8. An audio and video transmission device based on network bandwidth, characterized in that: Applicable to servers, including: The play request response module is used to determine the target audio and video data and the target play time point when receiving the audio and video play request of the target user; a historical bandwidth change curve acquisition module, configured to acquire a historical bandwidth change curve of the target user, wherein the historical bandwidth change curve is a bandwidth change curve when the target user plays audio and video data on the network at the target playback time point within a first historical time period; a user pair generation module, configured to input the historical bandwidth change curve into a user pairing model to obtain a target user pair, wherein the target user pair includes the target user and a sample user, wherein the sample user is a user who receives audio and video data transmitted by the server at a target encoding rate and a target transmission mode during a second historical time period; The audio and video data transmission module is used to transmit the target audio and video data to the target user by adopting the target encoding bit rate and the target transmission mode.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the network bandwidth-based audio and video transmission method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the network bandwidth-based audio and video transmission method according to any one of claims 1 to 7 when executed.

Citation Information

Patent Citations

  • Method and system for adjusting video encoding rates

    CN106658072A

  • Video transmission method and device, equipment and storage medium

    CN118175356A