Audio data transmission method and device
By carrying attribute information when the client requests audio data, the server sends data packets, which solves the problem of delay and resource consumption caused by the client parsing MPD files, and achieves more efficient audio data transmission.
Patent Information
- Application Number
- CN202510476988.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, the client needs to download and parse the MPD file when playing streaming audio data, resulting in a large startup delay and resource consumption, especially the impact on low-memory devices is more obvious.
The client directly carries the attribute information of the audio data fragment in the request message, such as audio format, sampling rate and bit depth, instead of downloading the MPD file first, the server determines the appropriate attribute information and sends data packets based on the network delay.
It reduces the cache usage and startup delay of the client, improves playback performance, adapts to poor network environments, and reduces transmission overhead.
Smart Images

Figure CN120343011A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data transmission, and particularly to an audio data transmission method and apparatus. Background Art
[0002] The Dynamic Adaptive Streaming over HTTP (DASH) protocol is commonly used for online audio and video playback, especially for streaming services that need to support multiple resolutions to adapt to different network speeds.
[0003] In practical applications, content providers, based on the DASH protocol, cut the content to be played into multiple segments and send them to the client one by one through requests. Among them, DASH uses a Media Presentation Description (MPD) file to carry attribute information such as the Uniform Resource Locator (URL) address, playback order, and sampling rate of the segments. The client can obtain the URL address of the segment by parsing the MPD file and dynamically select the most suitable sampling rate for playback according to the current network condition and the performance of the playback device.
[0004] However, since the client needs to download and parse the MPD file when requesting to play the content, it causes a startup delay. In addition, for clients with low memory, since the MPD file is in XML format, the consumption of parsing the MPD file is relatively large, affecting the playback experience. Summary of the Invention
[0005] In view of this, embodiments of this application provide an audio data transmission method and apparatus, which reduce the startup delay of playing audio data and the overhead on the client side, and improve the playback performance of audio data.
[0006] To solve the above problems, the technical solutions provided by embodiments of this application are as follows:
[0007] In the first aspect of this application, an audio data transmission method is provided. The method is applied to a server providing a streaming service and includes:
[0008] Receiving a request message sent by a client, where the client is used to play streaming audio data;
[0009] Determining a target audio data segment based on the request message;
[0010] Sending a data packet to the client, where the data packet includes the target audio data segment and attribute information corresponding to the target audio data segment.
[0011] In a second aspect of the present application, an audio data transmission method is provided. The method is applied to a client for playing streaming media audio data and includes:
[0012] Sending a request message to a server for providing streaming media services;
[0013] Receiving a data packet sent by the server, where the data packet includes a target audio data segment determined by the server based on the request message and attribute information corresponding to the target audio data segment;
[0014] Playing the target audio data segment based on the attribute information.
[0015] In a third aspect of the present application, an audio data transmission device is provided. The device is applied to a server for providing streaming media services and includes:
[0016] A receiving unit for receiving a request message sent by a client;
[0017] A processing unit for determining a target audio data segment based on the request message;
[0018] A sending unit for sending a data packet to the client, where the data packet includes the target audio data segment and attribute information corresponding to the target audio data segment.
[0019] In a fourth aspect of the present application, an audio data transmission device is provided. The device is applied to a client for providing streaming media services and includes:
[0020] A sending unit for sending a request message to a server;
[0021] A receiving unit for receiving a data packet sent by the server, where the data packet includes a target audio data segment determined by the server based on the request message and attribute information corresponding to the target audio data segment;
[0022] A processing unit for playing the target audio data segment based on the attribute information.
[0023] In a fifth aspect of the embodiments of the present application, an electronic device is provided, including: a processor and a memory;
[0024] The memory is used to store computer-readable instructions or computer programs;
[0025] The processor is used to read the computer-readable instructions or the computer program so that the electronic device implements the audio data transmission method described in the first aspect or the second aspect.
[0026] In a sixth aspect of the present application, a computer-readable storage medium is provided. Instructions are stored in the computer-readable storage medium, and when the instructions run on a device, the device is caused to execute the audio data transmission method described in the first aspect or the second aspect.
[0027] In a seventh aspect of the present application, a computer program product is provided. When the computer program product runs on a computer, the computer is caused to execute the audio data transmission method described in the first aspect or the second aspect.
[0028] Thus, the embodiments of the present application have the following beneficial effects:
[0029] In the present application, when a client needs to play a certain streaming audio data, a request message is sent to the server. After receiving the request message, the server determines the target audio data segment and sends it to the client through a data packet. The data packet includes not only the target audio data segment to be played, but also the attribute information corresponding to the target audio data segment, such as information related to playing the target audio data, such as audio format, sampling rate, bit depth, etc., so that the client can directly parse and play the target audio data segment after receiving the data packet. It can be seen that through the solution provided by the present application, before the client requests to play a certain audio data, there is no need to first download and parse the MPD file, reducing the occupation of the client-side cache. Moreover, since there is no need to first download the MPD file, the time for buffering the MPD file is omitted, reducing the startup delay and improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 FIG. is a handshake interaction diagram provided by an embodiment of the present application;
[0031] Figure 2 FIG. is an interaction diagram of an audio data transmission method provided by an embodiment of the present application;
[0032] Figure 3 FIG. is a schematic diagram of a data packet structure provided by an embodiment of the present application;
[0033] Figure 4 FIG. is a technical framework diagram of an audio data transmission provided by an embodiment of the present application;
[0034] Figure 5 FIG. is a structural diagram of an audio data transmission device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] To make the above objects, features, and advantages of the present application more obvious and understandable, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0036] Currently, before the client requests to play streaming audio data from the server, it first requests to download the MPD file from the server. The MPD file does not contain the actual audio data to be played, but rather information such as the path, format, various supported sample rates, bit depths, etc. of the audio data. After receiving the MPD file, the client parses the MPD file and, based on an assessment of the network bandwidth, selects appropriate information such as the sample rate and bit depth from the MPD file according to the assessment result, and requests the server to download the audio data corresponding to that sample rate.
[0037] Since there is a need to first download and parse the MPD file and then buffer the audio data between the client's request to play the audio data and the start of audio data playback, there is a playback delay. Moreover, for a client built with a low-memory microcontroller unit (MCU), downloading and parsing the MPD file will consume a large amount of resources of the client, affecting the performance of the client.
[0038] Based on this, the present application provides a solution. In this solution, when the client requests to cache audio data from the server, it does not need to first download the MPD file, but directly carries the attribute information included in the MPD file, such as audio format, sample rate, bit depth, etc. in the data packet. In this way, after receiving the data packet, the client can directly parse and play it, not only reducing the occupation of client resources, but also reducing the playback delay. In addition, since there is no need to transmit the MPD file, the transmission overhead is reduced. Since the attribute information carried in the data packet is less than that of the MPD file, the network overhead is further reduced, thus being able to adapt to a poor network environment.
[0039] It should be noted that the technical solution provided by the present application can be applied to scenarios with relatively high requirements for real-time response, such as the separated AI dialogue scenario, the portable network audio data playback scenario, etc. Among them, the separated AI dialogue scenario means that the user uses the end-side device to ask questions to the AI server, and the AI server provides answers. For example, the user asks the AI server through the end-side device "How to get to XX University?".
[0040] To facilitate understanding of the specific implementation of the present application, the technologies involved in the present application will be described first.
[0041] Since the present application is applied to the network audio data transmission scenario, the client and the server are connected through the Hypertext Transfer Protocol (HTTP). In a scenario without data communication, a two-way handshake is performed once per second to know the current network latency status. The handshake packet only needs to be short, 1-2 bytes, to ensure that there will not be too much network consumption.
[0042] As Figure 1 shown in the handshake schematic diagram, after the client connects to the network, it initiates a handshake with the server and real-time detects the network latency. When the client needs to upload voice data, it packs and uploads the corresponding audio data packets according to the current network latency. Similarly, when the server sends down voice data, it also starts sending the data of the first packet according to the real-time network latency. In this way, the impact of the DASH startup delay can be avoided when starting audio transmission.
[0043] In this application, there are various delays involved, such as startup delay, buffering delay, and adaptive switching delay. Startup delay refers to the time between when a user requests to start playing and when the audio actually starts playing. DASH usually requires some time to buffer the initial data (i.e., the MPD file), resulting in a startup delay usually ranging from a few seconds to more than a dozen seconds.
[0044] Buffering delay. To ensure smooth playback, the DASH client usually maintains a certain buffer during playback. The size of the buffer affects the delay. A larger buffer can reduce the risk of the playback terminal but increase the delay.
[0045] Adaptive switching delay. During playback, DASH dynamically adjusts the audio quality according to the current network conditions. Since it needs to re-download and parse the MPD file after the adjustment, the switching will cause a short delay.
[0046] The sampling rate of audio data refers to the number of times the analog signal is sampled per unit time, usually expressed in Hertz (Hz). The sampling rate determines the number of samples collected per second, thus affecting the audio quality and frequency response range. The higher the sampling rate, the higher the signal restoration degree and the better the audio quality.
[0047] The change in the sampling rate leads to a change in the playback speed. The higher the sampling rate, the better the signal restoration degree and the slower the sound.
[0048] Bit depth describes the detail accuracy that the hardware or software for processing audio data can achieve. The higher the bit depth, the greater the amount of audio information and the more level values that can be stored.
[0049] To better understand the technical solution of this application, it will be described below in conjunction with specific embodiments.
[0050] Refer to Figure 2 , this figure shows a method for transmitting audio data provided by an embodiment of this application. The method includes:
[0051] S201: The client sends a request message to the server.
[0052] Among them, the request message is used to request streaming media audio data. The client can support playing streaming media audio data, and the server can provide streaming media audio data.
[0053] Specifically, in the scenario of a separate AI conversation, the request message carries the question sent by the user through the client, such as "How to get to XX for study?". In the scenario of playing portable network audio data, the request message carries the identifier of the audio data requested by the user through the client. For example, when the user requests to play a certain radio drama through an audio client, in response to the user's trigger operation on the client, a request message is sent to the server, and the identifier of the radio drama is carried in the request message so that the server can determine the audio data segment to be played through this identifier.
[0054] S202: The server determines the target audio data segment based on the request message.
[0055] In this embodiment, after receiving the request message, the server determines the audio data to be played based on the information carried in the request message. If the audio data is segmented into multiple audio data segments, the server also needs to further determine the target audio data segment to be played. For example, if the sequence number of the audio data segment recently received by the client is carried in the request message, the server determines the next audio data segment according to this sequence number.
[0056] To ensure that there are no breakpoints and lags during the transmission of audio data, the server also needs to determine the matching bit depth, sampling rate and other attribute information according to the current network environment. Specifically, determine the network delay between the client and the server; determine the bit depth and / or sampling rate corresponding to the target audio data segment based on this network delay. Among them, the server can determine the network delay based on the handshake mechanism, and then determine the bit depth and / or sampling rate matching this network delay according to the pre-configured corresponding relationship. As shown in Table 1, the server can pre-store the corresponding relationship between the network delay and (bit depth, sampling rate).
[0057] Table 1 Corresponding relationship between network delay, bit depth and sampling rate
[0058] Type Network Delay Bit Depth Sampling Rate 1 <20ms 32 bit 192 Khz 2 <50ms 24 bit 96 Khz 3 <80ms 24 bit 48 Khz 4 <100ms 16 bit 44.1 Khz 5 <120ms 16 bit 22.05 Khz 6 <150ms 16 bit 16 Khz 7 Others 8 bit 8 Khz
[0059] It can be seen from Table 1 that the greater the network delay, the smaller the bit depth and sampling rate, thus ensuring the smooth playback of audio data without affecting listening.
[0060] S203: The server sends a data packet to the client, and the data packet includes the target audio data segment and the attribute information.
[0061] Among them, the attribute information includes, but is not limited to, one or more of audio format, data packet serial number, sampling rate, and bit depth. Since the data packet only needs to carry the determined sampling rate and bit depth, and the attribute information such as audio format, data packet serial number, sampling rate, and bit depth occupies a small number of bytes, the overhead is very small compared to DASH.
[0062] Among them, the structure of the data packet is shown in Table 2, including fields such as data type, packet serial number, parameter length, parameter content, data length, and data content. Among them, the data type is used to indicate the audio format, occupying 1 Byte, and the audio format can be defined according to its own needs. For example, 0x01 indicates the WAV format, and 0x02 indicates the MP3 format; the packet serial number refers to the serial number of the data packet, which can occupy 1 Byte; the parameter length refers to the length occupied by attribute information such as sampling rate and bit rate, which can occupy 2 Bytes; the parameter content refers to the specific attribute information, and the specific length X Byte occupied is determined according to the actual situation; the data length refers to the length of the target audio data segment, which can occupy 4 Bytes; the data content refers to the specific target audio data segment, occupying N Bytes.
[0063] Table 2 Data Packet Format
[0064]
[0065] For example, the format of a WAV data packet is as Figure 3 shown. Among them, 0x01 indicates that the audio format is WAV; 0x00 indicates that the data packet serial number is 0; 0x0002 indicates that the parameter length is 2 Bytes; 0x20 indicates that the bit depth is 16 bit; 0xC0 indicates that the sampling rate is 96K; the audio data is 1 KByte.
[0066] Considering that the storage space of the client is limited, an audio data segment is usually not too large, for example, about 1-2 K bytes. Through Figure 3 it can be known that the data packet with attribute information added only has 10 more bytes compared to the original audio data segment, and the overhead is greatly reduced compared to separately sending the MPD file and the audio data segment.
[0067] S204: The client plays the target audio data segment based on the attribute information.
[0068] After receiving the data packet, the client obtains the attribute information by parsing the data packet, and then plays the target audio data segment according to the attribute information without downloading and parsing the MPD file in advance. In this way, for a client with limited storage space, the occupation of storage space is reduced, and the playback performance of the client is improved.
[0069] S205: The client sends a response message to the server.
[0070] In this embodiment, the client caches audio data segments using a buffer. Considering the limited capacity, after completing one transmission, the client needs to send a response message to the server so that the server can calculate the network latency and determine attribute information such as the sampling rate and bit depth corresponding to the next audio data segment based on the network latency.
[0071] Due to the limited cache of the client, after caching some audio data segments, when the cache occupancy rate reaches a certain level, to avoid affecting the playback performance, the waiting duration will be carried in the response message. The waiting duration is the time the server needs to wait to send the next data packet after receiving the response message. That is, to relieve the cache pressure, the server is instructed to adjust the time to send the next data packet by sending the waiting duration, so that the server delays sending the audio data and avoids the heavy cache load of the client from affecting the playback experience.
[0072] Specifically, the client can periodically determine whether the current cache occupancy rate is less than a preset threshold. If so, there is no need to carry the waiting duration in the sent response message; if not, the waiting duration for the server to send the next data packet is determined based on the playback speed of the cached data packets, and this waiting duration is carried in the response message.
[0073] In the scenario of streaming media file playback, after playing an audio data segment, the client releases the audio data segment, thus relieving the cache pressure. Based on this, the client will determine the playback duration according to the playback speed of the cached data packets and the amount of unplayed audio data; and determine the waiting duration according to this playback duration and a target ratio. The target ratio is usually less than 1, so the waiting duration is less than the playback duration.
[0074] It should be noted that the playback speed of audio data is related to its sampling rate, and the sampling rate is related to the current network latency. Since the network environment may be constantly changing, the sampling rates corresponding to different segments of the same audio data sent at different times are different, and thus the playback speeds of different segments are different. Based on this, it is necessary to determine the playback duration required to play each audio data segment according to the playback speed corresponding to the cached and unplayed audio data segment; obtain the total playback duration of all unplayed audio frequency bands; and determine the waiting duration based on this total playback duration and the target ratio. For example, the client determines that the playback duration of the cached audio segment 1 is 10 ms; the playback duration of audio segment 2 is 8 ms; and the playback duration of audio segment 3 is 15 ms. Suppose that after receiving audio segment 3, the client finds that it is currently playing 10% of audio segment 1, and the cache of the currently unplayed audio data reaches the warning threshold, then the server needs to wait for a certain proportion of the time of the unplayed audio. With the target ratio being 50%, the waiting duration is (10 * 90% + 8 + 15) * 50% = 16 ms.
[0075] Among them, the target ratio is dynamically adjusted based on the current cache occupancy rate, and this ratio can be debugged and optimized in segments according to different systems and actual hardware. Specifically, the larger the cache occupancy rate, the larger the target ratio; the smaller the cache occupancy rate, the smaller the target ratio, so as to avoid the server waiting for a long time. For example, when the cache occupancy rate reaches the warning threshold of 80%, the target ratio is 50%, and it is necessary to wait for 16 ms; after a period of time, when the cache occupancy rate drops to 50%, the audio segment can continue to be received. If the waiting duration is still determined according to the ratio of 50%, it may cause the server to wait for a long time. To reduce the waiting duration, the target ratio is lowered to 30%, thus reducing the waiting duration of the server.
[0076] Among them, the format of the response message is shown in Table 3, including a 1-Byte data type field, which indicates the audio format; a 1-Byte packet sequence number field, which indicates the sequence number of the most recently received data packet; and a 4-Byte waiting duration field, which indicates the duration that the server needs to wait before sending the next data packet.
[0077] Table 3 Format of the Response Message
[0078] Data Packet Data Type Packet Sequence Number Waiting Time ms Length 1 byte 1 byte 4 byte Description Audio Format Data Packet Sequence Number The next packet of data needs to wait for X ms
[0079] It can be seen that when the client requests cached audio data from the server, it does not need to download the MPD file first. Instead, it directly carries the attribute information included in the MPD file in the data packet, such as audio format, sampling rate, bit depth, etc. In this way, after receiving the data packet, the client can directly parse and play it, which not only reduces the occupation of client resources but also reduces the playback delay. In addition, since there is no need to transmit the MPD file, the transmission overhead is reduced. Since the attribute information carried in the data packet is less than that in the MPD file, the network overhead is further reduced, thus enabling adaptation to a poor network environment.
[0080] See Figure 4 the audio transmission technology framework diagram shown in Figure 4 As shown, in the absence of data communication, the client and the server detect the current network latency through periodic handshakes. After the client sends a request message for requesting audio data to the server, the server selects matching attribute information such as sampling rate and bit depth based on the latest determined network latency, and sends the first packet of audio data including the attribute information to the client.
[0081] After receiving the first packet of audio data, the client sends a response message ACK to the server. After receiving the ACK, the server calculates the network latency based on the ACK and packs the next audio data packet. If the waiting duration in the ACK is non-zero, the next audio data packet needs to be sent with a delay; if the waiting duration in the ACK is zero, it can be sent immediately.
[0082] After receiving the next audio data packet, the client still sends an ACK to the server. The above interaction process is repeated until all the required audio data is sent. If there is no data communication between the client and the server, the handshake mechanism is triggered.
[0083] Based on the above method content, an embodiment of the present application provides an audio data transmission device, which will be described below in combination with specific embodiments.
[0084] See Figure 5 , which is a structural diagram of an audio data transmission device provided by an embodiment of the present application. As Figure 5 shown, the device 500 may include a receiving unit 501, a processing unit 502, and a sending unit 503.
[0085] In some embodiments, the device 500 can implement the functions of the above server, and specifically may include:
[0086] The receiving unit 501 is configured to receive a request message sent by a client, where the client is used to play streaming media audio data;
[0087] A processing unit 502, configured to determine a target audio data segment based on the request message;
[0088] A sending unit 503, configured to send a data packet to the client, where the data packet includes the target audio data segment and attribute information corresponding to the target audio data segment.
[0089] In some embodiments, the attribute information at least includes one or more of audio format, data packet sequence number, sampling rate, and bit depth.
[0090] In some embodiments, the processing unit 502 is further configured to determine a network delay between the client and the server before sending a data packet to the client; and determine the bit depth and / or sampling rate corresponding to the target audio data segment based on the network delay.
[0091] In some embodiments, the receiving unit 501 is further configured to receive a response message sent by the client, where the response message includes a waiting duration for the server to send the next data packet, and the waiting duration is determined by the client based on the playback speed of the cached data packets.
[0092] In some embodiments, the apparatus 500 can also implement the functions of a client, specifically including:
[0093] A sending unit 503, configured to send a request message to a server, where the server is used to provide a streaming media service;
[0094] A receiving unit 501, configured to receive a data packet sent by the server, where the data packet includes a target audio data segment determined by the server based on the request message and attribute information corresponding to the target audio data segment;
[0095] A processing unit 502, configured to play the target audio data segment based on the attribute information.
[0096] In some embodiments, the attribute information at least includes one or more of audio format, data packet sequence number, sampling rate, and bit depth.
[0097] In some embodiments, the processing unit 502 is further configured to, if the cache occupancy rate is not less than a preset threshold, determine a waiting duration for the server to send the next data packet based on the playback speed of the cached data packets and the amount of unplayed audio data;
[0098] The sending unit 503 is further configured to send a response message to the server, where the response message includes the waiting duration.
[0099] In some embodiments, the processing unit 502 is specifically configured to determine the playback duration based on the playback speed of the cached data packets and the amount of unplayed audio data; determine the waiting duration for the server to send the next data packet based on the playback duration and a target ratio, where the target first ratio changes dynamically based on the cache occupancy rate.
[0100] It should be noted that for the information execution processes and the like of each unit in the above device, specifically, reference can be made to the descriptions in the method embodiments shown previously in this application, which will not be elaborated here.
[0101] In addition, an embodiment of the present application provides an electronic device, including: a processor, a memory;
[0102] The memory is used to store computer-readable instructions or computer programs;
[0103] The processor is used to read the computer-readable instructions or the computer programs so that the device implements the audio data transmission method described above.
[0104] An embodiment of the present application provides a computer-readable storage medium, including instructions or a computer program, which when running on a computer, causes the computer to execute the audio data transmission method described above.
[0105] It should be noted that the various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the systems or devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions in the method section.
[0106] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Here, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression means any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a, b, and c", where a, b, and c can be single or multiple.
[0107] It should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0108] The steps of the methods or algorithms described in connection with the embodiments disclosed herein can be implemented directly in hardware, in software modules executed by a processor, or in a combination of both. The software modules can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well known in the art.
[0109] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An audio data transmission method, characterized in that, The method is applied to a server providing streaming media services and includes: Receiving a request message sent by a client, where the client is used to play streaming media audio data; Determining a target audio data segment based on the request message; Sending a data packet to the client, where the data packet includes the target audio data segment and the attribute information corresponding to the target audio data segment.
2. The method according to claim 1, wherein The attribute information includes at least one or more of audio format, data packet sequence number, sampling rate, and bit depth.
3. The method according to claim 1 or 2, characterized in that, Before sending the data packet to the client, the method further includes: Determining the network latency between the client and the server; Determining the bit depth and / or sampling rate corresponding to the target audio data segment based on the network latency.
4. The method according to claim 1, characterized in that, The method further includes: Receiving a response message sent by the client, where the response message includes the waiting duration for the server to send the next data packet, and the waiting duration is determined by the client based on the playback speed of the cached data packets.
5. An audio data transmission method, characterized in that, The method is applied to a client playing streaming media audio data and includes: Sending a request message to a server, where the server is used to provide streaming media services; Receiving the data packet sent by the server, where the data packet includes the target audio data segment determined by the server based on the request message and the attribute information corresponding to the target audio data segment; Playing the target audio data segment based on the attribute information.
6. The method according to claim 5, characterized in that, The attribute information includes at least one or more of audio format, data packet sequence number, sampling rate, and bit depth.
7. The method according to claim 5 or 6, characterized in that, The method further includes: If the cache occupancy rate is not less than a preset threshold, determining the waiting duration for the server to send the next data packet based on the playback speed of the cached data packets and the amount of unplayed audio data; Sending a response message to the server, where the response message includes the waiting duration.
8. The method according to claim 7, wherein The determining the waiting duration for the server to send the next data packet based on the playback speed of the cached data packets and the amount of unplayed audio data includes: Determining the playback duration based on the playback speed of the cached data packets and the amount of unplayed audio data; Determining the waiting duration for the server to send the next data packet based on the playback duration and a target ratio, where the target ratio is dynamically changed based on the cache occupancy rate.
9. An audio data transmission device, characterized in that, The apparatus is applied to a server providing streaming media services and includes: A receiving unit, configured to receive a request message sent by a client; A processing unit, configured to determine a target audio data segment based on the request message; A sending unit, configured to send a data packet to the client, where the data packet includes the target audio data segment and the attribute information corresponding to the target audio data segment.
10. An audio data transmission device, characterized in that, The apparatus is applied to a client providing streaming media services and includes: A sending unit, configured to send a request message to a server; A receiving unit, configured to receive the data packet sent by the server, where the data packet includes the target audio data segment determined by the server based on the request message and the attribute information corresponding to the target audio data segment; A processing unit, configured to play the target audio data segment based on the attribute information.