Audio transmission method and related apparatuses

CN122614318BActive Publication Date: 2026-09-29LINKPLAY TECHNOLOGY INC NANJING
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611040878.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-14
Publication Date
2026-09-29
Estimated Expiration
2046-07-14

AI Technical Summary

Technical Problem

[0003]传统方案通常由应用程序直接调用厂商提供的多房间 SDK 或私有接口完成音频发送和播放控制,但是这种处理方式需要应用程序理解并处理设备发现、分组、时钟同步、音频缓存、网络传输及播放状态控制等步骤,使得应用程序处理效率慢且兼容性差,尤其是对于第三方播放器等未集成特定 SDK 的应用程序,无法直接支持多房间同步播放,导致音频输出无法覆盖多房间设备,限制了音频应用的音频输出效果,且进一步降低了兼容性

Benefits of technology

本申请的一种音频传输方法及其相关装置,该方法应用于电子设备,电子设备包括音频传输系统,电子设备与至少一个音频播放设备连接,方法包括:在音频传输系统的系统框架层中构建虚拟音频设备;其中,虚拟音频设备的音频传输接口的接口类型与音频传输系统的音频应用层的音频输出接口的接口类型一致,以使音频应用层通过标准音频输出接口无感知地接入虚拟音频设备;在系统框架层接收到音频应用层传输的第一音频数据的情况下,通过虚拟音频设备接收第一音频数据,并对第一音频数据进行解析,得到第一音频数据对应的原始音频参数;基于原始音频参数和音频播放设备的设备参数,对第一音频数据进行参数调整处理,得到与至少一个音频播放设备匹配的第二音频数据;通过虚拟音频设备将第二音频数据传输至音频传输系统的传输层,并通过传输层将第二音频数据传输至音频播放设备。该方式中,通过在系统框架层构建虚拟音频设备实现音频传输处理,将应用层与传输层解耦,避免了应用层直接与音频播放设备互联,提高了音频传输的稳定性和质量,且提高了多播放设备同时进行播放的一致性和稳定性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122614318B_ABST
    Figure CN122614318B_ABST
Patent Text Reader

Abstract

The application relates to an audio transmission method and a related device thereof, the method comprising: constructing a virtual audio device in a system framework layer of an audio transmission system; in the case that first audio data transmitted by an audio application layer is received in the system framework layer, receiving the first audio data through the virtual audio device, analyzing the first audio data to obtain original audio parameters corresponding to the first audio data, performing parameter adjustment processing on the first audio data based on the original audio parameters and device parameters of an audio playing device to obtain second audio data matched with at least one audio playing device, and transmitting the second audio data to a transmission layer through the virtual audio device and transmitting the second audio data to the audio playing device through the transmission layer. The scheme provided by the application can construct a virtual audio device to realize audio transmission processing, improve the stability and quality of audio transmission, and improve the consistency and stability of simultaneous playing of multiple playing devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technology, and in particular to an audio transmission method and related apparatus. Background Technology

[0002] In existing Linux or embedded audio devices, ALSA is a commonly used low-level audio framework. Audio applications typically output audio data to a sound card or virtual sound card via a standard PCM interface. However, in multi-room audio scenarios, audio data is not simply output to local hardware. Instead, it needs to be distributed to multiple playback nodes based on the user-selected room group, playback devices, master-slave relationships, and network status, ensuring low latency and high synchronization accuracy between nodes.

[0003] Traditional solutions typically involve applications directly calling the manufacturer's multi-room SDK or proprietary interface to control audio transmission and playback. However, this approach requires the application to understand and handle steps such as device discovery, grouping, clock synchronization, audio caching, network transmission, and playback status control. This results in slow processing efficiency and poor compatibility, especially for third-party players and other applications that do not integrate specific SDKs. These applications cannot directly support synchronized playback across multiple rooms, leading to audio output that cannot cover multiple room devices, limiting the audio output quality of audio applications, and further reducing compatibility. Summary of the Invention

[0004] To address or partially address the problems existing in related technologies, this application provides an audio transmission method and related apparatus. By constructing a virtual audio device at the system framework layer to achieve audio transmission processing, the application layer and the transmission layer are decoupled, avoiding direct interconnection between the application layer and the audio playback device. This improves the stability and quality of audio transmission, and also enhances the consistency and stability of simultaneous playback from multiple playback devices.

[0005] This application provides an audio transmission method for an electronic device, the electronic device including an audio transmission system, the electronic device being connected to at least one audio playback device, the method comprising: constructing a virtual audio device in the system framework layer of the audio transmission system; wherein the interface type of the audio transmission interface of the virtual audio device is consistent with the interface type of the audio output interface of the audio application layer of the audio transmission system, so that the audio application layer can seamlessly access the virtual audio device through a standard audio output interface; when the system framework layer receives first audio data transmitted by the audio application layer, receiving the first audio data through the virtual audio device and parsing the first audio data to obtain the original audio parameters corresponding to the first audio data; based on the original audio parameters and the device parameters of the audio playback device, performing parameter adjustment processing on the first audio data to obtain second audio data matching the at least one audio playback device; transmitting the second audio data to the transmission layer of the audio transmission system through the virtual audio device, and transmitting the second audio data to the audio playback device through the transmission layer.

[0006] In conjunction with the first aspect, in one possible implementation of the first aspect, the step of adjusting the parameters of the first audio data based on the original audio parameters and the device parameters of the audio playback device to obtain second audio data matching the at least one audio playback device includes: obtaining the device parameters corresponding to each audio playback device in the audio playback device group through the virtual audio device; determining a target transmission parameter based on the device parameters and the original audio parameters, wherein the value of the target transmission parameter is not higher than the lowest value of the corresponding parameter supported by each audio playback device in the audio playback device group; wherein the target transmission parameter is an audio parameter that matches each audio playback device in the audio playback device group; and performing format conversion processing on the first audio data according to the target transmission parameter to obtain the second audio data.

[0007] In conjunction with the first aspect, in one possible implementation of the first aspect, both the original audio parameters and the target transmission parameters include at least a sampling rate, bit depth, and number of channels. The step of performing format conversion processing on the first audio data according to the target transmission parameters to obtain the second audio data includes: when the original audio parameters and the target transmission parameters are inconsistent, performing resampling processing and / or bit depth conversion processing and / or channel number conversion processing on the first audio data through the virtual audio device to obtain the second audio data; wherein the audio parameters of the second audio data are consistent with the target transmission parameters.

[0008] In conjunction with the first aspect, in one possible implementation of the first aspect, before transmitting the second audio data to the transport layer of the audio transmission system via the virtual audio device, and transmitting the second audio data to the audio playback device via the transport layer, the method further includes: constructing playback text corresponding to the second audio data via the virtual audio device based on the audio parameters of the second audio data; wherein the playback text includes at least one of a session identifier, a playback group identifier, an audio format, and an audio payload; and dividing the second audio data into multiple audio data blocks with the same playback duration based on the playback text; wherein the block number of the audio data block is associated with the division order of the division process.

[0009] In conjunction with the first aspect, in one possible implementation of the first aspect, after dividing the second audio data into multiple audio blocks with the same playback duration based on the playback text, the method further includes: generating a target playback timestamp corresponding to each audio data block through the virtual audio device based on the playback duration and the original timestamp corresponding to the audio data block; and transmitting the audio data block to the audio playback device through the transport layer based on the target playback timestamp, so as to control the audio playback device to synchronously output the audio data block.

[0010] In conjunction with the first aspect, in one possible implementation of the first aspect, the method further includes: constructing a multi-level cache structure in the virtual audio device, the multi-level cache structure including at least an application receive cache, a format processing cache, a transmission send cache, and an exception recovery cache; receiving the first audio data through the application receive cache, and performing the parameter adjustment processing on the first audio data in the format processing cache to obtain the second audio data; transmitting the second audio data to the transmission send cache, and dynamically adjusting the cache load of the multi-level cache structure based on the network transmission rate of the transmission layer and the transmission speed of the audio application layer; transmitting the second audio data to the transmission layer through the transmission send cache; when the transmission layer detects a network connection abnormality in at least one of the audio playback devices, retaining the audio data corresponding to the current playback progress through the exception recovery cache, and after the audio playback device restores its network connection, sending pre-buffered data to the audio playback device based on the playback progress data retained in the exception recovery cache, so that the audio playback device rejoins synchronous playback; and maintaining the continuous state of the audio application layer writing audio data to the virtual audio device throughout the above process.

[0011] In conjunction with the first aspect, in one possible implementation of the first aspect, the method further includes: when the virtual audio device receives an audio control command sent by the audio application layer, identifying a playback state corresponding to the audio control command, wherein the playback state includes at least one of pause, resume, stop, volume adjustment, and mute; generating a playback control command corresponding to the playback state through the virtual audio device, and sending the playback control command to the transport layer; and sending the playback control command to the at least one audio playback device through the transport layer to control the at least one audio playback device to switch playback states according to the playback control command.

[0012] In conjunction with the first aspect, in one possible implementation of the first aspect, the method further includes: identifying the current playback mode of the audio transmission system through the virtual audio device; the current playback mode being a local playback mode or a multi-target playback mode; if the current playback mode is a local playback mode, then transmitting the second audio data to the current playback sound card of the electronic device through the transport layer, so as to output the first audio data through the current playback sound card; if the current playback mode is a multi-target playback mode, then distributing the second audio data to multiple audio playback devices in the audio playback device group through the transport layer.

[0013] In conjunction with the first aspect, in one possible implementation of the first aspect, the audio application layer includes multiple audio applications, and the method further includes: when the system framework layer receives third audio data transmitted by the multiple audio applications respectively, constructing a playback session corresponding to each of the third audio data through the virtual audio device; determining the playback priority of each playback session based on the audio source type corresponding to the playback session; and performing mixing processing or volume reduction processing on the third audio data based on the playback priority.

[0014] In conjunction with the first aspect, in one possible implementation of the first aspect, the method further includes: performing channel mapping processing on the second audio data based on the channel roles of each audio playback device in the audio playback device group; wherein the channel mapping processing includes at least one of the following: assigning the left and right channels of the second audio data to the corresponding audio playback devices respectively; mixing the multi-channel signal of the second audio data into a mono signal and assigning it to a mono audio playback device; performing low-pass filtering processing on the second audio data to generate a subwoofer channel signal and assigning it to the subwoofer device.

[0015] A second aspect of this application provides an audio transmission system applied to an electronic device, the electronic device being connected to at least one audio playback device. The audio transmission system includes: an audio application layer, a system framework layer, and a transmission layer. A virtual audio device is constructed in the system framework layer. The system framework layer is configured to receive first audio data transmitted by the audio application layer through the virtual audio device, and parse the first audio data to obtain the original audio parameters corresponding to the first audio data. The interface type of the audio transmission interface of the virtual audio device is consistent with the interface type of the audio output interface of the audio application layer of the audio transmission system, so that the audio application layer can seamlessly access the virtual audio device through a standard audio output interface. The virtual audio device is configured to perform parameter adjustment processing on the first audio data based on the original audio parameters and the device parameters of the audio playback device to obtain second audio data matching the at least one audio playback device, and transmit the second audio data to the transmission layer. The transmission layer is configured to receive the second audio data and transmit the second audio data to the audio playback device.

[0016] A third aspect of this application provides an audio transmission device, comprising: a construction module for constructing a virtual audio device in the system framework layer of the audio transmission system; a parsing module for receiving the first audio data transmitted by the audio application layer through the virtual audio device and parsing the first audio data to obtain the original audio parameters corresponding to the first audio data when the system framework layer receives the first audio data transmitted by the audio application layer; wherein the interface type of the audio transmission interface of the virtual audio device is consistent with the interface type of the audio output interface of the audio application layer of the audio transmission system, so that the audio application layer can seamlessly access the virtual audio device through a standard audio output interface; a processing module for performing parameter adjustment processing on the first audio data based on the original audio parameters and the device parameters of the audio playback device to obtain second audio data matching the at least one audio playback device; and a transmission module for transmitting the second audio data to the transmission layer of the audio transmission system through the virtual audio device, and transmitting the second audio data to the audio playback device through the transmission layer.

[0017] A fourth aspect of this application provides an electronic device, comprising: Processor; and A memory that stores executable code, which, when executed by the processor, causes the processor to perform the method described above.

[0018] A fifth aspect of this application provides a computer-readable storage medium having executable code stored thereon, which, when executed by a processor of an electronic device, causes the processor to perform the method described above.

[0019] A sixth aspect of this application provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the method described above.

[0020] The technical solution provided in this application may include the following beneficial effects: This application discloses an audio transmission method and related apparatus. The method is applied to an electronic device, which includes an audio transmission system and is connected to at least one audio playback device. The method includes: constructing a virtual audio device in the system framework layer of the audio transmission system; wherein the interface type of the audio transmission interface of the virtual audio device is consistent with the interface type of the audio output interface of the audio application layer of the audio transmission system, so that the audio application layer can seamlessly access the virtual audio device through a standard audio output interface; when the system framework layer receives first audio data transmitted by the audio application layer, receiving the first audio data through the virtual audio device and parsing the first audio data to obtain the original audio parameters corresponding to the first audio data; based on the original audio parameters and the device parameters of the audio playback device, performing parameter adjustment processing on the first audio data to obtain second audio data matching at least one audio playback device; transmitting the second audio data to the transmission layer of the audio transmission system through the virtual audio device, and transmitting the second audio data to the audio playback device through the transmission layer. In this approach, by constructing a virtual audio device in the system framework layer to realize audio transmission processing, the application layer and the transmission layer are decoupled, avoiding direct interconnection between the application layer and the audio playback device, improving the stability and quality of audio transmission, and improving the consistency and stability of simultaneous playback by multiple playback devices.

[0021] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0022] The above and other objects, features and advantages of this application will become more apparent from the more detailed description of exemplary embodiments thereof in conjunction with the accompanying drawings, wherein the same reference numerals generally represent the same components in the exemplary embodiments thereof.

[0023] Figure 1 This is a schematic flowchart illustrating the audio transmission method in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of the audio transmission device shown in the embodiments of this application; Figure 3This is a schematic diagram of the structure of an audio transmission system shown in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an electronic device shown in an embodiment of this application. Detailed Implementation

[0024] Embodiments of this application will now be described in more detail with reference to the accompanying drawings. While embodiments of this application are shown in the drawings, it should be understood that this application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to make this application more thorough and complete, and to fully convey the scope of this application to those skilled in the art.

[0025] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0026] It should be understood that although the terms "first," "second," "third," etc., may be used in this application to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0027] In existing Linux or embedded audio devices, ALSA is a commonly used low-level audio framework. Audio applications typically output audio data to the sound card or virtual sound card through the standard PCM interface. However, in multi-room audio scenarios, audio data is not simply output to the local hardware. Instead, it needs to be distributed to multiple playback nodes based on the user-selected room group, playback devices, master-slave relationships, and network status, while ensuring low latency and high synchronization accuracy between the nodes.

[0028] Traditional solutions typically involve the application directly calling the manufacturer's multi-room SDK or proprietary interface to control audio transmission and playback. However, this approach requires the application to understand and handle steps such as device discovery, grouping, clock synchronization, audio caching, network transmission, and playback status control. This results in slow processing efficiency and poor compatibility, especially for third-party players and other applications that do not integrate specific SDKs. These applications cannot directly support synchronized playback in multiple rooms, leading to audio output that cannot cover multiple room devices, limiting the audio output effect of audio applications, and further reducing compatibility.

[0029] To address the aforementioned issues, this application provides an audio transmission method that implements audio transmission processing by constructing a virtual audio device at the system framework layer. This decouples the application layer from the transmission layer, avoiding direct interconnection between the application layer and the audio playback device, thereby improving the stability and quality of audio transmission, and enhancing the consistency and stability of simultaneous playback from multiple devices.

[0030] The technical solutions of the embodiments of this application are described in detail below with reference to the accompanying drawings.

[0031] Figure 1 This is a schematic flowchart illustrating the audio transmission method in an embodiment of this application.

[0032] See Figure 1 An audio transmission method is applied to an electronic device, the electronic device including an audio transmission system, the electronic device being connected to at least one audio playback device, the method comprising: S110: Construct a virtual audio device in the system framework layer of the audio transmission system; wherein the interface type of the audio transmission interface of the virtual audio device is consistent with the interface type of the audio output interface of the audio application layer of the audio transmission system, so that the audio application layer can seamlessly access the virtual audio device through the standard audio output interface.

[0033] Specifically, electronic devices are hardware platforms with audio processing and transmission capabilities, such as smart speakers, streaming media players, and embedded Linux audio devices. These devices contain audio transmission systems. The system framework layer of the audio transmission system provides various audio processing services and interfaces, manages system resources, and can construct virtual audio devices. These virtual audio devices are software-simulated audio output interfaces within the system framework layer. They do not directly correspond to physical sound cards but act as intermediate processing nodes for audio data streams, providing standard audio interfaces to upper-layer applications and connecting to the actual audio transmission layer to lower layers. For example, a virtual audio device can be configured as a virtual PCM (Pulse Code Modulation) device within an ALSA (Advanced Linux Sound Architecture) system. This virtual audio device can receive audio data from the application layer, pass it to the next processing stage, and store the received audio data.

[0034] Specifically, when the audio application layer transmits the first audio data, due to the consistency of the interface type, the virtual audio device can directly receive the first audio data without converting the audio data format, which can improve the transmission efficiency and accuracy of the audio data.

[0035] S120: When the system framework layer receives the first audio data transmitted from the audio application layer, the system receives the first audio data through a virtual audio device, parses the first audio data, and obtains the original audio parameters corresponding to the first audio data.

[0036] Specifically, the audio application layer consists of various audio applications running on electronic devices. This audio application layer can be used to acquire raw audio data. When the audio application layer outputs audio data to the system, the virtual audio device intercepts and receives this data. Subsequently, the virtual audio device parses the received first audio data. For example, it can identify its original audio parameters by analyzing the data structure of the first audio data.

[0037] Specifically, the application layer may include third-party applications that can be used to transmit the first audio data to the system framework layer. For example, these third-party applications can be any application capable of playing audio, such as game applications, music playback applications, audio players, video applications, or social applications. For instance, the third-party application could be a music player.

[0038] S130: Based on the original audio parameters and the device parameters of the audio playback device, perform parameter adjustment processing on the first audio data to obtain second audio data that matches at least one audio playback device.

[0039] Specifically, the device parameters of an audio playback device are its own hardware parameters and playback capabilities, such as its supported sampling rate range, bit depth, channel configuration, and maximum output power. After obtaining the original audio parameters of the first audio data, the audio parameters of the first audio data can be adjusted according to the original audio parameters and the device parameters of the audio playback device so that the first audio data can be adapted to the playback capabilities of the audio playback device. For example, if the sampling rate of the original audio data is higher than the highest sampling rate supported by the playback device, downsampling processing is required to obtain the second audio data.

[0040] S140: Transmit the second audio data to the transport layer of the audio transmission system via a virtual audio device, and transmit the second audio data to the audio playback device via the transport layer.

[0041] Specifically, the transport layer is the module in the audio transmission system that sends processed audio data to audio playback devices. This transport layer handles data packet encapsulation, network protocols, and data distribution. After processing the second audio data, the transport layer can encapsulate the data and send it to one or more audio playback devices via a preset communication protocol. For example, data can be passed to a backend service via local inter-process communication (IPC), and then the backend service can send it to a remote audio playback device via a network protocol (such as Wi-Fi). This allows the audio application layer to be separated from the underlying multi-room transmission process. Music applications do not need to be aware of the complex processes of underlying device discovery, grouping, clock synchronization, audio buffering, network transmission, and playback control. The virtual audio device, as an intermediate adaptation layer, completes the interception, format adaptation, parameter adjustment, and transmission to the transport layer, enabling third-party audio applications to support multi-room playback seamlessly. For example, music applications do not need to modify code or integrate any multi-room SDK; they only need to output audio data according to the standard ALSA PCM interface, and then the virtual audio device intercepts this audio data at the underlying level.

[0042] For example, a user's electronic device (such as a smart speaker) is connected to two external audio playback devices: a wireless speaker and a smart TV located in different rooms. First, a virtual audio device is built into the system framework layer of the smart speaker's audio transmission system. This virtual audio device is configured as the system's default audio output interface. When the user plays music through a music application on the smart speaker, the music application will play music according to the standard ALSA... The PCM interface outputs raw PCM audio data to the system's default audio playback device. This virtual audio device intercepts and receives this first audio data, parses it, and identifies its original audio parameters as 44.1kHz sampling rate, 16-bit depth, and stereo. Simultaneously, the virtual audio device obtains the device parameters of the wireless speaker and smart TV. For example, the wireless speaker's device parameters are 48kHz sampling rate and 24-bit depth, while the smart TV's are 48kHz sampling rate and 16-bit depth. Based on these raw audio parameters and the device parameters of each audio playback device, the virtual audio device converts the 44.1kHz, 16-bit deep, stereo first audio data into 48kHz, 16-bit deep, stereo second audio data. Finally, the virtual audio device transmits the processed second audio data to the transmission layer of the smart speaker's internal audio transmission system. This transmission layer then transmits the second audio data to the wireless speaker and smart TV via a network (e.g., Wi-Fi).

[0043] In one possible implementation, based on the original audio parameters and the device parameters of the audio playback device, the first audio data is subjected to parameter adjustment processing to obtain second audio data that matches at least one audio playback device. This includes: obtaining the device parameters corresponding to each audio playback device in the audio playback device group through a virtual audio device; determining target transmission parameters based on the device parameters and the original audio parameters; wherein the target transmission parameters are audio parameters that match each audio playback device in the audio playback device group, and the value of the target transmission parameters is not higher than the lowest value of the corresponding parameters supported by each audio playback device in the audio playback device group; and performing format conversion processing on the first audio data according to the target transmission parameters to obtain the second audio data.

[0044] Specifically, the audio playback device group is a preset device group, which may include several audio playback devices. The virtual audio device can obtain the device parameters corresponding to each audio playback device, and determine an audio parameter set that is compatible with all audio playback devices based on the device parameters corresponding to different audio devices. Each original audio parameter is then adjusted to the audio parameters in this audio playback parameter set. For example, if one audio playback device supports a 48kHz sampling rate and another audio playback device supports a 44.1kHz sampling rate, then the 44.1kHz sampling rate is used as the determined audio sampling rate, and the sampling rate in the original audio parameters is adjusted to this sampling rate to unify the format of the first audio data, ensuring consistency and stability when multiple audio playback devices play synchronously.

[0045] For example, an electronic device connects to three audio playback devices: a high-fidelity speaker supporting 24-bit / 96kHz stereo, a Bluetooth headset supporting 16-bit / 48kHz stereo, and a smart speaker supporting only 16-bit / 44.1kHz mono. When the audio application layer transmits first audio data with original parameters of 24-bit / 96kHz stereo, the virtual audio device obtains the device parameters of these three audio playback devices and learns that the high-fidelity speaker supports the original parameters, the Bluetooth headset supports up to 16-bit / 48kHz stereo, and the smart speaker supports up to 16-bit / 44.1kHz mono. At this point, the virtual audio device can determine that the target transmission parameter is 16-bit / 44.1kHz mono. Subsequently, the virtual audio device performs format conversion processing on the first audio data to convert the 24-bit bit depth to 16-bit, the 96kHz sampling rate to 44.1kHz, and mix the stereo into mono, thereby obtaining the second audio data, making the format-converted second audio data compatible with all audio playback devices.

[0046] In one possible implementation, both the original audio parameters and the target transmission parameters include at least a sampling rate, bit depth, and number of channels. The first audio data is format-converted according to the target transmission parameters to obtain the second audio data. This includes: if the original audio parameters and the target transmission parameters are inconsistent, the first audio data is resampled and / or bit depth converted and / or number of channels is converted using a virtual audio device to obtain the second audio data; wherein the audio parameters of the second audio data are consistent with the target transmission parameters.

[0047] Specifically, sampling rate refers to the number of times the analog audio signal is sampled per second, bit depth refers to the number of binary bits used for each sample point in the audio data, and channel number refers to the number of channels in the audio signal, such as mono, stereo, or multi-channel surround sound. When the sampling rate, bit depth, or channel number of the first audio data is inconsistent with the target transmission parameters, the first audio data needs to be format converted and isolated. For example, the sampling rate of the first audio data can be changed through resampling processing, such as converting 48kHz audio to 44.1kHz; the bit depth of the first audio data can be changed through bit depth conversion processing; and the number of channels of the audio data can be changed through channel number conversion processing, such as converting 24-bit audio to 16-bit. The key audio parameters of the second audio data, including sampling rate, bit depth, and channel number, are completely consistent with the target transmission parameters, so that the processed audio data can match the playback parameters of each audio playback device.

[0048] In one possible implementation, before transmitting the second audio data to the transport layer of the audio transmission system via a virtual audio device, and then transmitting the second audio data to the audio playback device via the transport layer, the method further includes: constructing playback text corresponding to the second audio data via a virtual audio device based on the audio parameters of the second audio data; wherein the playback text includes at least one of a session identifier, a playback group identifier, an audio format, and an audio payload; and dividing the second audio data into multiple audio data blocks with the same playback duration based on the playback text; wherein the block number of the audio data block is associated with the division order of the division process.

[0049] Specifically, the playback text is a data structure used to represent audio data transmission and playback-related parameters. For example, the playback text can be a JSON-formatted description file, defining various playback attributes in key-value pairs. The session identifier is a string used to uniquely identify a specific audio playback session. In multi-user scenarios, different audio streams need to be controlled separately, and the session identifier can distinguish the playback process of these audio streams. The playback group identifier is used to associate multiple audio playback devices and can be a playback group ID set by the user during configuration. The audio format indicates the basic parameters of the audio data, such as the encoding method, sampling rate, bit depth, and number of channels. For example, the encoding format can be PCM. The audio payload is the data content carried in the audio data, which may include the frame length, frame duration, buffer target depth, playback start time, etc.

[0050] Specifically, the second audio data can be divided based on the playback text. For example, the second audio data can be divided into multiple audio data blocks according to a fixed time length (e.g., 20 milliseconds). Each audio data block has a corresponding block number, which is an index used to identify the order of the audio data block in the entire audio stream. This ensures that the audio playback device can reassemble the audio data blocks in the correct order, thereby achieving continuous and seamless playback. For example, the block number can be an integer sequence that increments from 0, allowing the transport layer to transmit audio data blocks in units, improving the flexibility and efficiency of transmission. Furthermore, the audio data blocks can be better recognized and sorted by the audio playback device, improving playback consistency.

[0051] For example, playback text is constructed based on the audio parameters of the second audio data. For instance, the session identifier in the playback text is a UUID generated by the system; the playback group identifier can be "living room speaker group," indicating that the audio stream will be sent to multiple speaker devices in the living room area for synchronized playback; the audio format field reflects that the current audio data is PCM encoded, with a 48kHz sampling rate, 24-bit bit depth, and stereo; the audio payload field represents the actual data stream in the second audio data. Then, the virtual audio device will divide the second audio data stream according to a preset playback duration (e.g., 50 milliseconds per block). If the total duration of the second audio data is 10 seconds, it can be divided into 200 audio data blocks. Each audio data block is assigned a corresponding block number, starting from 0 and incrementing. For example, the first block is numbered 0, the second block is numbered 1, and so on, until the last block is numbered 199. These audio data blocks are then transmitted to the transport layer to be transmitted to the audio playback device.

[0052] In one possible implementation, after dividing the second audio data into multiple audio blocks with the same playback duration based on the playback text, the method further includes: generating a target playback timestamp for each audio data block through a virtual audio device based on the playback duration and the original timestamp corresponding to the audio data block; and transmitting the audio data block to the audio playback device through a transport layer based on the target playback timestamp, so as to control the audio playback device to synchronously output the audio data block.

[0053] Specifically, the original timestamp is the time position mark of the audio data in the original audio stream before it is segmented. It is used to indicate the relative order and start time of the data block in the entire audio stream. Based on the original timestamp, a target playback timestamp can be generated for each audio data block. The target playback timestamp is the reference time for the data block to start playing on all audio playback devices. When sending audio data blocks through the transport layer, the transport layer can adjust the sending time of the audio data blocks according to the target playback timestamp. For example, audio data blocks that are about to expire can be sent first to ensure that they arrive at the audio playback devices before the target playback time. This allows all audio playback devices to output the same audio content accurately at the same time, avoiding problems such as delays and misalignments that may occur when multiple devices play the audio, and achieving audio playback synchronization.

[0054] For example, the second audio data can be divided into audio data blocks with a playback duration of 100 milliseconds each. If the target playback timestamp of the first audio data block is T0, then the target playback timestamp of the second audio data block is T0+100ms, and the third audio data block is T0+200ms.

[0055] Furthermore, the target playback timestamp of the audio data block can be calculated using the following formula:

[0056]

[0057]

[0058] in, For the first The target playback timestamp for each audio data block; The base time for the start of the playback session is the original timestamp corresponding to the audio data block; This is the block number of the audio data block; The playback duration for each audio data block, which can be measured in milliseconds (ms). The safety margin for the entire group transmission can be obtained directly. This safety margin can be related to the number of audio playback devices and their parameters. For the first The updated exponentially weighted moving average estimate of latency; This is the time delay smoothing coefficient, 0 < <1; For the first The network one-way transmission latency of a certain audio playback device detected this time can be measured in milliseconds (ms). This is the maximum time delay amplification factor. ≥ 1, usually The value can be in the range of 1.2-1.5; The maximum estimated latency value for all audio playback devices in the current audio playback device group can be calculated by examining the latency of each audio playback device within the group. And obtain the maximum value from it; The standard deviation of the latency estimates of each audio playback device in the current audio playback device group can be used to reflect the latency dispersion within the device group; This is the jitter compensation coefficient. ≥ 0, The value range is usually 1.0-2.0.

[0059] For example, the audio playback device group includes three audio playback devices: audio playback device A (living room speaker) with block number 0, audio playback device B (bedroom speaker) with block number 1, and audio playback device C (kitchen speaker) with block number 2, and is pre-configured. =0, =50 ms, =0.2, =1.3, =1.5, if after After round-robin detection, the estimated current latency values ​​for the three audio playback devices can be obtained respectively: =8ms; =20ms; =35ms, therefore the maximum estimated latency within the audio playback device group is: =35ms, and the standard deviation of the delay was calculated. =11.1ms, overall transmission safety margin =62.2ms, therefore the time to audio playback device A is calculated separately. =62.2ms, audio playback device B =112.2ms, audio playback setting C =162.2ms, if the first The measured latency of the audio playback device C was detected by the wheel. = 40 ms, then =36ms, and recalculate the relevant parameters. This method enables dynamic calculation of latency based on actual measured latency of various audio playback devices. This ensures that even the latest audio playback device receives data before the target time, thereby controlling multi-device synchronization errors to within milliseconds, eliminating perceptible audio misalignment between rooms, and making... It can adjust the margin in real time based on the measured latency. For example, it automatically reduces the margin to decrease system playback latency when network conditions are good, and automatically increases the margin to ensure synchronization stability when network jitter increases. When a new audio playback device is added, its latency estimate accumulates from the initial value. It can also be updated so that the target timestamp takes effect in the next playback batch, the playback progress of the original audio playback device is not affected, and when the original audio playback device leaves the audio playback device group, the maximum latency of the audio playback device group is reduced accordingly. This means that when an audio playback device joins or leaves the audio playback device group, there is no need to reset the playback session, which improves the flexibility of audio playback devices for synchronous playback. In one possible implementation, the method further includes: constructing a multi-level cache structure in the virtual audio device, the multi-level cache structure including at least an application receive cache, a format processing cache, a transmission send cache, and an exception recovery cache; receiving first audio data through the application receive cache, and performing parameter adjustment processing on the first audio data in the format processing cache to obtain second audio data; transmitting the second audio data to the transmission send cache, and dynamically adjusting the cache load of the multi-level cache structure based on the network transmission rate of the transport layer and the transmission speed of the audio application layer; transmitting the second audio data to the transport layer through the transmission send cache; when the transport layer detects a network connection failure in at least one audio playback device, retaining the audio data corresponding to the current playback progress through the exception recovery cache, and after the audio playback device restores its network connection, sending pre-buffered data to the audio playback device based on the playback progress data retained in the exception recovery cache, so that the audio playback device rejoins synchronous playback; and maintaining the continuous writing of audio data from the audio application layer to the virtual audio device throughout the above process.

[0060] Specifically, multiple memory regions can be configured within the virtual audio device, such as a circular buffer or a double buffer. The application receive buffer is used to receive the first audio data transmitted from the audio application layer; the format processing buffer is an intermediate buffer in a multi-level buffer structure, used to adjust the parameters of the first audio data in the format processing buffer to obtain the second audio data that matches the audio playback device; the transmission send buffer is used to transmit the parameter-adjusted second audio data to the transport layer. The transmission send buffer can smooth the data transmission rate, adapt to changes in the network transmission rate of the transport layer, and ensure data stability; the exception recovery buffer is used to retain the audio data corresponding to the current playback progress when the audio playback device experiences a network connection failure. This buffer can be a circular buffer used to store audio data that has been processed and sent in the most recent period of time. Buffer load refers to the ability to adjust the amount of data in each buffer area of ​​a multi-level buffer structure based on the difference between the network transmission rate of the transport layer and the transmission speed of the audio application layer. For example, when the amount of data in a certain buffer area is detected to be close to the upper limit, a backpressure signal is sent to the audio application layer to slow down the data transmission speed; when the amount of data is detected to be too low, the audio application layer can be controlled to increase the data transmission speed. When the transport layer detects a network connection failure of the audio playback device, it can retain the audio data corresponding to the current playback progress through the abnormal recovery buffer. When the audio playback device restores its network connection, the virtual audio device uses the audio data retained in the abnormal recovery buffer to send pre-buffered data to the audio playback device, enabling it to quickly rejoin synchronous playback, thereby ensuring the continuity and stability of the audio data stream.

[0061] In one possible implementation, the method further includes: when the virtual audio device receives an audio control command sent by the audio application layer, identifying the playback state corresponding to the audio control command, wherein the playback state includes at least one of pause, resume, stop, volume adjustment, and mute; generating a playback control command corresponding to the playback state through the virtual audio device, and sending the playback control command to the transport layer; and sending the playback control command to at least one audio playback device through the transport layer to control at least one audio playback device to switch the playback state according to the playback control command.

[0062] Specifically, when the audio application layer needs to control audio playback—for example, when a user clicks the pause button or adjusts the volume—it sends corresponding audio control commands to the virtual audio device in the system framework layer. When the virtual audio device recognizes the specific playback state corresponding to the command, it generates a standardized playback control command based on that state and sends it to the transmission layer of the audio transmission system. In this way, the audio playback device can receive and execute these commands, thereby achieving real-time switching of its playback state. This ensures that the audio application layer can uniformly control remote or local audio playback devices through the virtual audio device, enhancing the user's control over the audio playback device.

[0063] For example, when a user clicks the "Pause" button in a music playback application, the application generates an audio control command containing a "Pause" identifier and sends it to the virtual audio device in the system framework layer via IPC. Upon receiving this command, the virtual audio device parses it and identifies its corresponding playback state as "Pause." Subsequently, the virtual audio device generates a playback control command based on this playback state. This command might be a JSON message containing `"command": "pause"`, and sends this JSON message to the transport layer. The transport layer can then encapsulate this JSON message into a TCP packet and send it via Wi-Fi to the audio playback device connected to the electronic device. Upon receiving this TCP packet, the smart speaker parses the playback control command and switches its playback state from "Playing" to "Pause" accordingly, thus stopping audio output. When the user clicks the "Play" button again, a similar process triggers a "Resume" command, causing the smart speaker to resume playback from the paused point.

[0064] In one possible implementation, the method further includes: identifying the current playback mode of the audio transmission system through a virtual audio device; the current playback mode is a local playback mode or a multi-target playback mode; if the current playback mode is a local playback mode, transmitting the second audio data to the current playback sound card of the electronic device through the transport layer, so as to output the first audio data through the current playback sound card; if the current playback mode is a multi-target playback mode, distributing the second audio data to multiple audio playback devices in the audio playback device group through the transport layer.

[0065] Specifically, local playback mode refers to audio data being played through the electronic device's own audio output hardware (such as built-in speakers or headphone jacks). Multi-target playback mode refers to audio data being transmitted to one or more external audio playback devices, such as Bluetooth speakers, Wi-Fi speakers, or wired headphones. If the current playback mode is local playback mode, a second audio data can be transmitted to the electronic device's internal audio output hardware, ensuring that the audio can be played through the local sound card. This allows for flexible adaptation to different playback scenarios and optimizes the audio data transmission path and playback process.

[0066] In one possible implementation, the audio application layer includes multiple audio applications, and the method further includes: when the system framework layer receives third audio data transmitted by multiple audio applications respectively, constructing a playback session corresponding to each third audio data through a virtual audio device; determining the playback priority of each playback session based on the audio source type corresponding to the playback session; and performing mixing processing or volume reduction processing on the third audio data based on the playback priority.

[0067] Specifically, the audio source type corresponding to the playback session can be the category or purpose of the audio stream, such as media playback, navigation voice, call voice, system notification, etc. The playback priority of each playback session is determined according to the priority of the audio source type, so as to perform mixing or volume reduction processing on multiple third audio data. For example, the priority of call voice can be preset to be higher than that of navigation voice, and the priority of navigation voice to be higher than that of background music. Audio streams with higher priority can be played first; while audio streams with lower priority will be mixed or have their volume reduced by the virtual audio device to avoid audio conflicts and information loss.

[0068] For example, when a user is listening to music, a navigation app starts broadcasting turn information and initiates a voice call. After receiving these three audio streams at the system framework layer, the virtual audio device constructs independent playback sessions for each and identifies their audio source types as "media playback," "navigation voice," and "call voice," respectively. Based on preset priority rules, the virtual audio device determines that "call voice" has the highest priority, followed by "navigation voice," and then "media playback" has the lowest priority. When the call voice transmission begins, the virtual audio device reduces the volume of the music and navigation voice. Then, it mixes the processed music and navigation voice with the call voice to generate a mixed audio stream, allowing the user to clearly hear the call while keeping the background music and navigation voice low enough not to interfere with the call. When the call ends, the virtual audio device restores the volume of the music and navigation voice.

[0069] In one possible implementation, the method further includes: performing channel mapping processing on the second audio data based on the channel roles of each audio playback device in the audio playback device group; wherein the channel mapping processing includes at least one of the following: assigning the left and right channels of the second audio data to the corresponding audio playback devices respectively, mixing the multi-channel signal of the second audio data into a mono signal and assigning it to a mono audio playback device, and performing low-pass filtering processing on the second audio data to generate a subwoofer channel signal and assigning it to the subwoofer device.

[0070] Specifically, a channel role refers to a specific channel possessed by an audio playback device in the entire audio playback system, such as the left channel, right channel, center channel, surround channel, subwoofer channel, etc. Based on the channel of each audio playback device, the different channel information in the second audio playback device is allocated to the corresponding audio playback device. For example, the left channel part of the stereo signal is sent to the left channel speaker, and the right channel part is sent to the right channel speaker, thereby optimizing the overall audio playback effect.

[0071] This application discloses an audio transmission method applied to an electronic device. The electronic device includes an audio transmission system and is connected to at least one audio playback device. The method includes: constructing a virtual audio device in the system framework layer of the audio transmission system; wherein the interface type of the audio transmission interface of the virtual audio device is consistent with the interface type of the audio output interface of the audio application layer of the audio transmission system, so that the audio application layer can seamlessly access the virtual audio device through a standard audio output interface; when the system framework layer receives first audio data transmitted by the audio application layer, receiving the first audio data through the virtual audio device and parsing the first audio data to obtain the original audio parameters corresponding to the first audio data; based on the original audio parameters and the device parameters of the audio playback device, performing parameter adjustment processing on the first audio data to obtain second audio data matching at least one audio playback device; transmitting the second audio data to the transmission layer of the audio transmission system through the virtual audio device, and transmitting the second audio data to the audio playback device through the transmission layer. In this method, by constructing a virtual audio device in the system framework layer to realize audio transmission processing, the application layer and the transmission layer are decoupled, avoiding direct interconnection between the application layer and the audio playback device, improving the stability and quality of audio transmission, and improving the consistency and stability of simultaneous playback by multiple playback devices.

[0072] Corresponding to the aforementioned application function implementation method embodiments, this application also provides an audio transmission device, an electronic device, and corresponding embodiments. The entity executing the audio transmission method provided in this application embodiment can be an audio transmission device. Exemplarily, the audio transmission device can be an electronic device, or a functional component or entity within that electronic device.

[0073] Figure 2 This is a schematic diagram of the structure of an audio transmission device shown in an embodiment of this application.

[0074] See Figure 2 An audio transmission device 200, the device comprising: Module 210 is used to build a virtual audio device in the system framework layer of the audio transmission system; wherein, the interface type of the audio transmission interface of the virtual audio device is consistent with the interface type of the audio output interface of the audio application layer of the audio transmission system, so that the audio application layer can access the virtual audio device without being aware of it through the standard audio output interface. The parsing module 220 is used to receive the first audio data transmitted by the audio application layer through a virtual audio device when the system framework layer receives the first audio data, and to parse the first audio data to obtain the original audio parameters corresponding to the first audio data. Processing module 230 is used to perform parameter adjustment processing on the first audio data based on the original audio parameters and the device parameters of the audio playback device to obtain second audio data that matches at least one audio playback device; The transmission module 240 is used to transmit the second audio data to the transmission layer of the audio transmission system through a virtual audio device, and transmit the second audio data to the audio playback device through the transmission layer.

[0075] In one possible implementation, the parsing module 220 is further configured to obtain the device parameters corresponding to each audio playback device in the audio playback device group; determine the target transmission parameters based on the device parameters and the original audio parameters; wherein the target transmission parameters are audio parameters that are matched by each audio playback device in the audio playback device group, and the value of the target transmission parameters is not higher than the lowest value of the corresponding parameters supported by each audio playback device in the audio playback device group; and perform format conversion processing on the first audio data according to the target transmission parameters to obtain the second audio data.

[0076] In one possible implementation, both the original audio parameters and the target transmission parameters include at least a sampling rate, bit depth, and number of channels. The processing module 230 is further configured to perform resampling processing and / or bit depth conversion processing and / or channel number conversion processing on the first audio data through a virtual audio device when the original audio parameters and the target transmission parameters are inconsistent, to obtain second audio data; wherein the audio parameters of the second audio data are consistent with the target transmission parameters.

[0077] In one possible implementation, the processing module 230 is further configured to construct playback text corresponding to the second audio data through a virtual audio device based on the audio parameters of the second audio data; wherein the playback text includes at least one of a session identifier, a playback group identifier, an audio format, and an audio payload; and based on the playback text, divide the second audio data into multiple audio data blocks with the same playback duration; wherein the block number of the audio data block is associated with the division order of the division process.

[0078] In one possible implementation, the processing module 230 is further configured to generate a target playback timestamp for each audio data block through a virtual audio device based on the playback duration and the original timestamp corresponding to the audio data block; and based on the target playback timestamp, transmit the audio data block to the audio playback device through the transport layer to control the audio playback device to synchronously output the audio data block.

[0079] In one possible implementation, the processing module 230 is further configured to construct a multi-level cache structure in the virtual audio device, the multi-level cache structure including at least an application receive cache, a format processing cache, a transmission send cache, and an exception recovery cache; receive first audio data through the application receive cache, and perform parameter adjustment processing on the first audio data in the format processing cache to obtain second audio data; transmit the second audio data to the transmission send cache, and dynamically adjust the cache load of the multi-level cache structure based on the network transmission rate of the transport layer and the transmission speed of the audio application layer; transmit the second audio data to the transport layer through the transmission send cache; when the transport layer detects that at least one audio playback device has a network connection failure, retain the audio data corresponding to the current playback progress through the exception recovery cache, and after the audio playback device restores its network connection, send pre-buffered data to the audio playback device based on the playback progress data retained in the exception recovery cache, so that the audio playback device rejoins synchronous playback; and keep the audio data writing state from the audio application layer to the virtual audio device continuous throughout the above process.

[0080] In one possible implementation, the processing module 230 is further configured to, when the virtual audio device receives an audio control command sent by the audio application layer, identify the playback state corresponding to the audio control command, wherein the playback state includes at least one of pause, resume, stop, volume adjustment, and mute; generate a playback control command corresponding to the playback state through the virtual audio device, and send the playback control command to the transport layer; and send the playback control command to at least one audio playback device through the transport layer to control at least one audio playback device to switch the playback state according to the playback control command.

[0081] In one possible implementation, the transmission module 240 is further configured to identify the current playback mode of the audio transmission system via a virtual audio device; the current playback mode is either a local playback mode or a multi-target playback mode; if the current playback mode is a local playback mode, the second audio data is transmitted to the current playback sound card of the electronic device via the transmission layer, so as to output the first audio data via the current playback sound card; if the current playback mode is a multi-target playback mode, the second audio data is distributed to multiple audio playback devices in the audio playback device group via the transmission layer.

[0082] In one possible implementation, the processing module 230 is further configured to, when receiving third audio data transmitted by multiple audio applications at the system framework layer, construct a playback session corresponding to each third audio data through a virtual audio device; determine the playback priority of each playback session based on the audio source type corresponding to the playback session; and perform mixing processing or volume reduction processing on the third audio data based on the playback priority.

[0083] In one possible implementation, the construction module 210 is further configured to perform channel mapping processing on the second audio data based on the channel roles of each audio playback device in the audio playback device group; wherein the channel mapping processing includes at least one of the following: assigning the left and right channels of the second audio data to the corresponding audio playback devices respectively, mixing the multi-channel signal of the second audio data into a mono signal and assigning it to a mono audio playback device, and performing low-pass filtering processing on the second audio data to generate a subwoofer channel signal and assigning it to the subwoofer device.

[0084] This application discloses an audio transmission device, comprising: a construction module for constructing a virtual audio device in the system framework layer of an audio transmission system; wherein the interface type of the audio transmission interface of the virtual audio device is consistent with the interface type of the audio output interface of the audio application layer of the audio transmission system, so that the audio application layer can seamlessly access the virtual audio device through a standard audio output interface; a parsing module for receiving first audio data transmitted by the audio application layer through the virtual audio device and parsing the first audio data to obtain the original audio parameters corresponding to the first audio data when the system framework layer receives the first audio data transmitted by the audio application layer; a processing module for adjusting the parameters of the first audio data based on the original audio parameters and the device parameters of the audio playback device to obtain second audio data matching at least one audio playback device; and a transmission module for transmitting the second audio data to the transmission layer of the audio transmission system through the virtual audio device, and transmitting the second audio data to the audio playback device through the transmission layer. In this approach, by constructing a virtual audio device in the system framework layer to achieve audio transmission processing, the application layer and the transmission layer are decoupled, avoiding direct interconnection between the application layer and the audio playback device, improving the stability and quality of audio transmission, and enhancing the consistency and stability of simultaneous playback from multiple playback devices.

[0085] This application also provides an audio transmission system 300, which is applied to an electronic device connected to at least one audio playback device. The audio transmission system 300 includes: an audio application layer 310, a system framework layer 320, and a transmission layer 330. A virtual audio device 3201 is constructed in the system framework layer 320. The system framework layer 320 is used to receive the first audio data transmitted by the audio application layer 310 through the virtual audio device 3201, and parse the first audio data to obtain the original audio parameters corresponding to the first audio data. The virtual audio device 3201 is used to perform parameter adjustment processing on the first audio data based on the original audio parameters and the device parameters of the audio playback device to obtain second audio data matching at least one audio playback device, and transmit the second audio data to the transmission layer. The transmission layer 330 is used to receive the second audio data and transmit the second audio data to the audio playback device.

[0086] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated further here.

[0087] This application also provides an electronic device. Figure 4This is a schematic diagram of the hardware structure of an embodiment of the electronic device of this application. The electronic device includes a memory 420 and at least one processor 410. The memory 420 is electrically connected to the at least one processor 410. The memory 420 stores instructions. The at least one processor 410 calls the instructions in the memory 420 to cause the electronic device to execute the audio transmission method according to any of the foregoing embodiments of this application.

[0088] Specifically, the processor 410 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.

[0089] Memory 420 may include a large-capacity memory 420 for data or instructions. For example, and not limitingly, memory 420 may include a hard disk drive (HDD), a floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 420 may include removable or non-removable (or fixed) media. Where appropriate, memory 420 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 420 is non-volatile solid-state memory. In a particular embodiment, memory 420 includes read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or flash memory, or a combination of two or more of these.

[0090] In one example, the control device may also include a communication interface 430 and a bus 440. The processor 410, memory 420, and communication interface 430 are connected via the bus 440 and communicate with each other.

[0091] The communication interface 430 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.

[0092] Bus 440 includes hardware, software, or both, that couples components of an online data traffic metering device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 440 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, any suitable bus or interconnect is contemplated herein.

[0093] Furthermore, in conjunction with the audio transmission methods in the above embodiments, this application embodiment can provide a computer-readable storage medium for implementation. This computer-readable storage medium stores executable code, which, when executed by a processor, implements any of the audio transmission methods in the above embodiments.

[0094] This application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.

[0095] The functional blocks shown in the above block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.

[0096] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0097] The above are merely specific embodiments of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.

Claims

1. An audio transmission method, characterized in that, Applied to an electronic device, the electronic device including an audio transmission system, the electronic device being connected to at least one audio playback device, the method includes: A virtual audio device is constructed in the system framework layer of the audio transmission system; wherein the interface type of the audio transmission interface of the virtual audio device is consistent with the interface type of the audio output interface of the audio application layer of the audio transmission system, so that the audio application layer can seamlessly access the virtual audio device through the standard audio output interface. When the system framework layer receives the first audio data transmitted by the audio application layer, the first audio data is received through the virtual audio device, and the first audio data is parsed to obtain the original audio parameters corresponding to the first audio data. Based on the original audio parameters and the device parameters of the audio playback device, the first audio data is subjected to parameter adjustment processing to obtain second audio data that matches the at least one audio playback device; The second audio data is transmitted to the transmission layer of the audio transmission system via the virtual audio device, and then transmitted to the audio playback device via the transmission layer.

2. The method according to claim 1, characterized in that, The step of adjusting the parameters of the first audio data based on the original audio parameters and the device parameters of the audio playback device to obtain second audio data matching the at least one audio playback device includes: The device parameters corresponding to each audio playback device in the audio playback device group are obtained through the virtual audio device; Based on the device parameters and the original audio parameters, a target transmission parameter is determined; wherein, the target transmission parameter is an audio parameter that is matched by each audio playback device in the audio playback device group, and the value of the target transmission parameter is not higher than the lowest value of the corresponding parameter supported by each audio playback device in the audio playback device group; The first audio data is converted according to the target transmission parameters to obtain the second audio data.

3. The method according to claim 2, characterized in that, Both the original audio parameters and the target transmission parameters include at least sampling rate, bit depth, and number of channels. The step of performing format conversion processing on the first audio data according to the target transmission parameters to obtain the second audio data includes: If the original audio parameters are inconsistent with the target transmission parameters, the first audio data is resampled and / or bit-depth converted and / or channel number converted by the virtual audio device to obtain the second audio data; wherein the audio parameters of the second audio data are consistent with the target transmission parameters.

4. The method according to claim 1, characterized in that, Before transmitting the second audio data to the transport layer of the audio transmission system via the virtual audio device, and transmitting the second audio data to the audio playback device via the transport layer, the method further includes: Based on the audio parameters of the second audio data, a playback text corresponding to the second audio data is constructed through the virtual audio device; wherein, the playback text includes at least one of session identifier, playback group identifier, audio format, and audio payload; Based on the playback text, the second audio data is divided into multiple audio data blocks with the same playback duration; wherein, the block number of the audio data block is associated with the division order of the division process.

5. The method according to claim 4, characterized in that, After dividing the second audio data into multiple audio blocks with the same playback duration based on the playback text, the method further includes: Based on the playback duration and the original timestamps corresponding to the audio data blocks, the target playback timestamps corresponding to each audio data block are generated through the virtual audio device; Based on the target playback timestamp, the audio data block is transmitted to the audio playback device through the transport layer to control the audio playback device to synchronously output the audio data block.

6. The method according to claim 1, characterized in that, The method further includes: A multi-level cache structure is constructed in the virtual audio device, and the multi-level cache structure includes at least an application receive cache, a format processing cache, a transmission send cache, and an exception recovery cache. The application receives the first audio data through the receiving cache, and performs the parameter adjustment processing on the first audio data in the format processing cache to obtain the second audio data. The second audio data is transmitted to the transmission buffer, and the buffer load of the multi-level buffer structure is dynamically adjusted based on the network transmission rate of the transmission layer and the transmission speed of the audio application layer. The second audio data is transmitted to the transport layer via the transmission buffer. When the transport layer detects a network connection failure in at least one of the audio playback devices, it retains the audio data corresponding to the current playback progress through the failure recovery cache. After the audio playback device restores its network connection, it sends pre-buffered data to the audio playback device based on the playback progress data retained in the failure recovery cache, so that the audio playback device can rejoin synchronous playback. During the above process, the audio application layer keeps writing audio data to the virtual audio device continuously.

7. The method according to claim 1, characterized in that, The method further includes: When the virtual audio device receives an audio control command sent by the audio application layer, it identifies the playback state corresponding to the audio control command. The playback state includes at least one of pause, resume, stop, volume adjustment, and mute. The virtual audio device generates playback control commands corresponding to the playback state, and sends the playback control commands to the transport layer. The playback control command is sent to the at least one audio playback device through the transport layer to control the at least one audio playback device to switch playback states according to the playback control command.

8. The method according to claim 1, characterized in that, The method further includes: The virtual audio device identifies the current playback mode of the audio transmission system; the current playback mode is either a local playback mode or a multi-target playback mode. If the current playback mode is local playback mode, the second audio data is transmitted to the current playback sound card of the electronic device through the transmission layer, so as to output the first audio data through the current playback sound card; If the current playback mode is a multi-target playback mode, the second audio data is distributed to multiple audio playback devices in the audio playback device group through the transport layer.

9. The method according to claim 1, characterized in that, The audio application layer includes multiple audio applications, and the method further includes: When the system framework layer receives third audio data transmitted by multiple audio applications respectively, a playback session corresponding to each third audio data is constructed through the virtual audio device; Based on the audio source type corresponding to the playback session, determine the playback priority of each playback session; Based on the playback priority, the third audio data is mixed or its volume is reduced.

10. The method according to claim 2, characterized in that, The method further includes: Based on the channel roles of each audio playback device in the audio playback device group, the second audio data is subjected to channel mapping processing; wherein, the channel mapping processing includes at least one of the following: assigning the left and right channels of the second audio data to the corresponding audio playback devices respectively, mixing the multi-channel signal of the second audio data into a mono signal and assigning it to a mono audio playback device, and performing low-pass filtering processing on the second audio data to generate a subwoofer channel signal and assigning it to the subwoofer device.

11. An audio transmission device, characterized in that, The device includes: Modules for building virtual audio devices within the system framework layer of an audio transmission system; The parsing module is used to receive the first audio data transmitted by the audio application layer through the virtual audio device when the system framework layer receives the first audio data, and to parse the first audio data to obtain the original audio parameters corresponding to the first audio data; wherein, the interface type of the audio transmission interface of the virtual audio device is consistent with the interface type of the audio output interface of the audio application layer of the audio transmission system, so that the audio application layer can access the virtual audio device without being aware of it through the standard audio output interface. The processing module is used to perform parameter adjustment processing on the first audio data based on the original audio parameters and the device parameters of the audio playback device to obtain second audio data that matches at least one audio playback device. The transmission module is used to transmit the second audio data to the transmission layer of the audio transmission system through the virtual audio device, and transmit the second audio data to the audio playback device through the transmission layer.

12. An audio transmission system, characterized in that, An audio transmission system is applied to an electronic device connected to at least one audio playback device, and the audio transmission system includes: an audio application layer, a system framework layer, and a transmission layer; a virtual audio device is constructed in the system framework layer. The system framework layer is used to receive the first audio data transmitted by the audio application layer through the virtual audio device, and parse the first audio data to obtain the original audio parameters corresponding to the first audio data; wherein, the interface type of the audio transmission interface of the virtual audio device is consistent with the interface type of the audio output interface of the audio application layer of the audio transmission system, so that the audio application layer can access the virtual audio device without being aware of it through the standard audio output interface. The virtual audio device is used to perform parameter adjustment processing on the first audio data based on the original audio parameters and the device parameters of the audio playback device to obtain second audio data that matches the at least one audio playback device, and transmit the second audio data to the transport layer. The transport layer is used to receive the second audio data and transmit the second audio data to the audio playback device.

13. An electronic device, characterized in that, include: processor; as well as A memory having executable code stored thereon, which, when executed by the processor, causes the processor to perform the method as described in any one of claims 1-10.

14. A computer-readable storage medium, characterized in that, It stores executable code that, when executed by a processor of an electronic device, causes the processor to perform the method as described in any one of claims 1-10.

Citation Information

Patent Citations

  • Audio stream data processing method and device, cloud server and readable storage medium

    CN115801740A

  • Audio data processing method and device, equipment and storage medium

    CN120704634A