Audio data transmission method and system, chip, networking, storage medium and program product
By caching audio data in the network and dynamically adjusting the data frame interval, the problem of audio data transmission and playback when the node clock frequency is different from the audio source clock frequency is solved, achieving better audio data transmission and playback performance and synchronization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-10
- Publication Date
- 2026-03-13
AI Technical Summary
In a network, there is room for improvement in the audio data transmission and playback performance of nodes, especially when the node clock frequency is different from the audio source clock frequency, which can easily lead to audio data buffer overflow or underflow, resulting in audio data loss or playback stuttering.
By caching audio data in the first buffer unit and dynamically adjusting the data frame interval, the clock frequency of the audio source is adaptively matched according to the amount of cached audio data and the preset threshold information, so as to ensure that the transmission frequency of audio data and the clock frequency of the audio source reach a dynamic balance and avoid buffer overflow or underflow.
It improves the performance of audio data transmission and playback in the network, prevents audio data loss or playback stuttering, and enhances the synchronization and stability of audio synchronous playback.
Smart Images

Figure CN121664784A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of electronic communication technology, and in particular to an audio data transmission method, system, chip, network, storage medium, and program product. Background Technology
[0002] Currently, networking, including but not limited to ring networks and daisy-chain networks, requires audio data transmission and playback during use. For example, in vehicle networking, it is often necessary to transmit audio data from a sound source to multiple nodes in the network for synchronized playback to improve the in-vehicle audio experience. However, in current audio data transmission solutions, there is still room for improvement in the performance of audio data transmission and playback among nodes in the network. Summary of the Invention
[0003] In view of the above, embodiments of this disclosure provide an audio data transmission method, system, chip, network, storage medium, and program product to at least partially solve the above problems.
[0004] According to a first aspect of the present disclosure, an audio data transmission method is provided for a first node in a network of multiple nodes, the multiple nodes further including other nodes besides the first node, wherein the clock frequency of the first node is different from the clock frequency of the audio source, the method comprising: Audio data from the sound source is cached through the first cache unit; After reading the audio data of the i-th data frame from the first cache unit, the data frame interval between the i-th data frame and the (i+1)-th data frame is determined based on the first data volume of the audio data cached in the first cache unit, where i is a positive integer; After the data frame interval ends, the audio data of the (i+1)th data frame is read from the first buffer unit and transmitted to the corresponding node among the other nodes through the network.
[0005] According to a second aspect of the present disclosure, an audio data transmission method is provided for an audio playback node in a network of multiple nodes, the method comprising: Audio data from the sound source is cached through the second buffer unit; Based on the second data volume of the audio data cached in the second cache unit, the data reading frequency of the audio data is determined; Based on the data reading frequency, audio data is read from the second buffer unit to synchronously play audio data with at least one other node among the plurality of nodes at a predetermined time.
[0006] According to a third aspect of the present disclosure, an audio data transmission system is provided, comprising: a plurality of nodes connected in a network, wherein the plurality of nodes includes a first node and a second node, the second node being any one of the other nodes besides the first node, and the clock frequency of the first node being the same as the clock frequency of the audio source. The first node is configured to: send target data to the second node through the network, so that the second node synchronizes its clock frequency to the same frequency as the sound source based on the target data; receive audio data from the sound source, and transmit the audio data to the second node through the network; The second node is configured to: perform delayed playback on the received audio data based on a preset delay time corresponding to the second node, so as to achieve audio synchronization with at least one of the other nodes.
[0007] According to a fourth aspect of the present disclosure, a chip is provided, comprising: a processor and a memory, wherein the processor and the memory communicate with each other; the memory is used to store at least one executable instruction, which causes the processor to perform an operation corresponding to the method as described in any one of the first and second aspects.
[0008] According to a fifth aspect of the present disclosure, a network is provided, comprising: a plurality of connected nodes, at least one of the plurality of nodes including a chip as described in the fourth aspect.
[0009] According to a sixth aspect of the present disclosure, a computer storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method as described in any one of the first and second aspects.
[0010] According to a seventh aspect of the present disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method as described in any one of the first and second aspects.
[0011] The audio data transmission method in this embodiment can be used in a network of multiple nodes, specifically a first node whose clock frequency differs from that of the audio source. It can cache audio data from the audio source using a first buffer unit. After reading the audio data of the i-th data frame from the first buffer unit, it determines the data frame interval between the i-th and (i+1)-th data frames based on the first data volume of the audio data cached in the first buffer unit, where i is a positive integer. Then, after the data frame interval ends, it can read the audio data of the (i+1)-th data frame from the first buffer unit and transmit the audio data of the (i+1)-th data frame to the corresponding node in the other nodes through the network. Therefore, through the technical solution of this embodiment, the clock frequency of the first node and the audio source can be compared. In a network with different source clock frequencies, the data frame interval between two data frames can be dynamically adjusted by the first data amount of audio data cached in the first buffer unit. This allows audio data to be transmitted to the corresponding node through the network at an adaptive pace. In this optional method, the first node can adaptively match the clock frequency of the audio source, enabling the transmission frequency of the first node's audio data to achieve a better dynamic balance with the clock frequency of the audio source. This prevents audio data buffer overflow or underflow problems caused by the difference between the clock frequency of the first node and the clock frequency of the audio source, thereby avoiding audio data loss or audio playback stuttering at the corresponding node. This improves the audio data transmission and playback performance in various audio playback scenarios in the network. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings.
[0013] Figure 1A A schematic diagram of a network structure according to an embodiment of this disclosure is shown.
[0014] Figure 1B A schematic diagram of a network structure, representing another example of an embodiment of this disclosure, is shown.
[0015] Figure 1C A schematic diagram of a network structure, representing another example of an embodiment of this disclosure, is shown.
[0016] Figure 1D A schematic diagram of the network structure of yet another example of an embodiment of this disclosure is shown.
[0017] Figure 2 A schematic flowchart illustrating some examples of audio data transmission methods in embodiments of this disclosure is shown.
[0018] Figure 3 The format of some example data frames in embodiments of this disclosure is shown.
[0019] Figure 4A A schematic diagram showing the adjustment of the data frame interval between two data frames is shown.
[0020] Figure 4B Another schematic diagram showing the adjustment of the data frame interval between two data frames is shown.
[0021] Figure 5 This diagram illustrates the reading of audio data by the first node after phase adjustment of the data frame and audio frame.
[0022] Figure 6 Schematic flowcharts of audio data transmission methods for some other examples of embodiments of this disclosure are shown.
[0023] Figure 7A A schematic diagram showing the adjustment of the data reading frequency of audio data is shown.
[0024] Figure 7B Another schematic diagram showing the adjustment of the data reading frequency of audio data is shown.
[0025] Figure 8 A schematic diagram is shown showing the third and fourth threshold values corresponding to different audio playback nodes in a network.
[0026] Figure 9A A schematic diagram of an example where the sound source and the first node are connected to the same clock source is shown.
[0027] Figure 9B A schematic diagram of another example where the sound source and the first node are connected to the same clock source is shown.
[0028] Figure 9C A schematic diagram is shown in which the sound source and the first node are connected to the same clock source in another example.
[0029] Figure 10 A schematic diagram of a chip, representing some examples of embodiments of this disclosure, is shown. Detailed Implementation
[0030] To enable those skilled in the art to better understand the technical solutions in the embodiments of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and thoroughly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art should fall within the protection scope of this disclosure.
[0031] This disclosure provides an audio data processing scheme that can improve the performance of audio data transmission and playback of nodes in a network, which will be described in detail below.
[0032] First, some exemplary networking methods applicable to the technical solutions in the embodiments of this disclosure will be described. For example, refer to... Figure 1A , Figure 1B , Figure 1C , Figure 1D Some network topology diagrams from embodiments of this disclosure are shown. Figure 1A and Figure 1D The network shown is a ring network; while Figure 1B and Figure 1C The network shown is a daisy-chain network, in which, Figure 1B The example is a single chrysanthemum chain network. Figure 1C An example is a double daisy chain network (including a daisy branch formed by nodes 0, 1, and 2, and a daisy branch formed by nodes 0, 3, and 4).
[0033] Among them, Figure 1A , Figure 1B , Figure 1C , Figure 1D In this network, node 0 can be connected to a processing unit (such as, but not limited to, MCU (Microcontroller Unit), CPU (Central Processing Unit) and other processors). Node 0 forms a "one master and many slaves" network with other nodes (i.e., node 1, node 2, node 3, and node 4). Node 0 can be the master node, and nodes 1 to 4 can all be slave nodes.
[0034] For example, in Figure 1A In the example, the audio source is directly connected to node 0 (the main node), so node 0 can be the audio source node. Nodes 1 through 4 are all connected to the speakers, so nodes 1 through 4 can all be audio playback nodes. Figure 1A Similarly, in Figure 1B In the example, node 0 is the audio source node, and nodes 1 through 4 can all be audio playback nodes. Figure 1A Similarly, in Figure 1C In the example, node 0 is the sound source node, and nodes 1 through 4 can all be audio playback nodes. Figure 1DIn the example, since the audio source is directly connected to node 1 (child node) and node 0 (master node) is also connected to the speaker, node 1 can be the audio source node, and nodes 0 and 2-4 can all be audio playback nodes. Optionally, during audio data transmission and playback, the audio source can transmit the audio data to the audio source node, and then the audio source node transmits the audio data along the network to one or more corresponding audio playback nodes. The audio playback nodes can play the audio data through the speaker, thereby facilitating the implementation of various network audio functions, including but not limited to synchronous audio data playback and on-demand playback.
[0035] In the embodiments of this disclosure, a node can be a physical conceptual node. For example, in some embodiments, a node can be any module, device, or chip capable of implementing the solutions of the embodiments of this disclosure. For example, in some embodiments, each node includes at least one chip for data transmission in the network. In some embodiments, a node can also be a chip, that is, a chip can be directly used as a node. Optionally, all nodes can be physical layer chips. Optionally, nodes can directly implement data transmission based on the physical layer, thereby helping to reduce data transmission latency in the network.
[0036] It should be noted that the above Figures 1A-1D The network topology and number of nodes shown are for illustrative purposes only. The actual number of nodes can be more or less depending on the specific network requirements. Further details can be adapted accordingly in the following text. Figures 1A-1D The network structure is used to describe the audio data transmission method of the embodiments of this disclosure.
[0037] According to a first aspect of the present disclosure, an audio data transmission method is provided, which can be used as a first node in a network of multiple nodes. The multiple nodes also include other nodes besides the first node, and the clock frequency of the first node is different from the clock frequency of the audio source. The present disclosure does not specifically limit the clock frequency of the first node and the clock frequency of the audio source; hereinafter, an example can be given where the clock frequency of the audio source is 24.576MHz and the clock frequency of the first node is 25MHz. Figure 2 The flowchart shown illustrates that the audio data transmission method may include the following steps S102, S104, and S106, specifically: S102: Cache audio data from the sound source through the first buffer unit.
[0038] Optionally, the network can adopt any type and any network structure. For example, in some optional embodiments, the network in this disclosure can be a ring network, a daisy-chain network, or other types of networks. When it is a ring network, it can be combined with... Figure 1A and Figure 1D The above is for illustrative purposes. When using a daisy-chain network, it can be combined with... Figure 1B and Figure 1C The following diagram illustrates this concept. Optionally, the first node can be the master node in the network (e.g., ...). Figures 1A-1D In each example, node 0), while other nodes besides the first node can be child nodes (e.g., Figures 1A-1D (Nodes 1-4 in each example).
[0039] It should be understood that when the network topology used in the technical solution of this disclosure embodiment is a ring network or a daisy chain network, in a ring network or daisy chain network where the clock frequency of the first node and the clock frequency of the audio source are different, the first node can dynamically adjust the data frame interval between two data frames by using the first data amount of audio data cached in the first buffer unit. This allows the audio data to be transmitted to the corresponding node through the ring network or daisy chain network at an adaptive rhythm. Thus, through this optional method, the first node can adaptively match the clock frequency of the audio source, enabling the transmission frequency of the first node's audio data to achieve a better dynamic balance with the clock frequency of the audio source. This prevents audio data buffer overflow or underflow problems caused by the difference between the clock frequency of the first node and the clock frequency of the audio source, thereby avoiding audio data loss or audio playback stuttering at the corresponding node. This improves the audio data transmission and playback performance in various audio playback scenarios, including but not limited to synchronous audio playback of multiple nodes, in a ring network or daisy chain network.
[0040] The technical solutions in this disclosure can be used in any application scenario. For example, in some optional embodiments, the network can be an in-vehicle network, which can be used for in-vehicle data transmission and / or processing functions. Correspondingly, the audio data transmission scheme can be used for in-vehicle audio data transmission, thereby improving the audio data transmission and playback performance in in-vehicle audio playback scenarios.
[0041] In other examples, networking can also be used for data transmission and / or processing in scenarios including but not limited to home or security. Correspondingly, audio data transmission schemes can be used for audio data transmission in scenarios including but not limited to home or security. There is no unique limitation in the embodiments of this disclosure.
[0042] In this embodiment of the disclosure, the first cache unit can be a cache corresponding to the first node, and can be used to cache audio data. For example... Figure 1A , Figure 1B , Figure 1C As shown, audio data from a sound source can include data transmitted from a sound source directly connected to the first node to the first node and cached in the first cache unit. Alternatively, as... Figure 1D As shown, audio data from the sound source can also include audio data transmitted from the sound source directly connected to other nodes to the first node and cached in the first cache unit.
[0043] S104: After reading the audio data of the i-th data frame from the first buffer unit, the data frame interval between the i-th data frame and the (i+1)-th data frame is determined based on the first data amount of the audio data cached in the first buffer unit, where i is a positive integer.
[0044] In this embodiment of the disclosure, the first node can send data to other nodes via data frames. For example, the first node can read audio data that needs to be sent to a certain node from the first buffer unit, and then transmit the audio data of the data frame to the corresponding node through networking.
[0045] In the embodiments of this disclosure, the data frame can be implemented in any format. For example, it can be a data frame used only for transmitting audio data; or it can be a data frame that can transmit audio data and other types of data (such as MAC (Media Access Control) data, SPI (Serial Peripheral Interface) protocol data, CAN (Controller Area Network) protocol data, LIN (Local Interconnect Network) protocol data).
[0046] For example, Figure 3 The formats of some example data frames are shown. In some alternative embodiments, refer to Figure 3 As shown, a data frame may include at least a common header and multiple data blocks, each of which can be used to store data. For example, one or more data blocks in the data frame may be used to store audio data. This data frame format, as described in this embodiment, facilitates the network transmission of audio data and other types of data, and it also allows for direct transmission and processing through the physical layer. Optionally, the common header records some common information about the network and the data frame, such as a preamble, for identifying the protocol of the data frame. Protocol identification can be achieved through the preamble in the common header. Optionally, an ACK (Acknowledge Character) data field may also be included at the end of the data frame to ensure reliable data transmission.
[0047] Optionally, the i-th data frame can be any data frame sent after the first node powers on, and the (i+1)-th data frame can be the next data frame after the i-th data frame. In the following example, the i=3rd data frame can be used; correspondingly, the (i+1)-th data frame can be the 4th data frame.
[0048] In this embodiment of the disclosure, the first node can dynamically adjust the data frame interval between two data frames by using the first data amount of audio data cached in the first cache unit, thereby facilitating the subsequent transmission of audio data to the corresponding node through the network at an adaptive pace.
[0049] The specific implementation of step S104 is not limited in the embodiments disclosed herein. In some optional embodiments, step S104 may include: determining the data frame interval between the i-th data frame and the (i+1)-th data frame based on the first data volume of audio data cached in the first buffer unit and the first preset threshold information.
[0050] Based on this, in this embodiment of the present disclosure, by determining the first data volume of the audio data cached in the first cache unit, and based on the first data volume and the first preset threshold information, the data frame interval between the i-th data frame and the (i+1)-th data frame can be effectively and dynamically determined. This allows the audio data to be transmitted to the corresponding node through the network at an adaptive pace. As a result, the first node can adaptively match the clock frequency of the sound source, enabling the transmission frequency of the audio data of the first node to achieve a better dynamic balance with the clock frequency of the sound source. This prevents audio data buffer overflow or underflow problems caused by the difference between the clock frequency of the first node and the clock frequency of the sound source, thereby avoiding audio data loss or audio playback stuttering at the corresponding node. This improves the audio data transmission and playback performance in various audio playback scenarios in the network, including but not limited to synchronous audio playback of multiple nodes.
[0051] Optionally, the first preset threshold information may include one or more threshold values, which may be used to compare with the first data volume in order to determine the data frame interval between the i-th data frame and the (i+1)-th data frame.
[0052] In some optional embodiments, the data volume unit of any threshold value in the first preset threshold information is bits. Therefore, by comparing the threshold value in bits with the first data volume, the data frame interval between two data frames can be dynamically adjusted with higher precision. This facilitates the transmission of audio data to the corresponding nodes through the network with higher precision and an adaptive rhythm, thereby improving the audio data transmission and playback performance in various audio playback scenarios within the network, including but not limited to synchronized audio playback across multiple nodes.
[0053] In some optional embodiments, the first buffer unit buffers audio data in units of bits. This allows the determined first data volume of the audio data to be measured in bits. By using a first data volume in bits, the data frame interval between two data frames can be dynamically adjusted with higher precision. This facilitates the transmission of audio data to the corresponding nodes through the network with higher precision and an adaptive rhythm, thereby improving the audio data transmission and playback performance in various audio playback scenarios within the network, including but not limited to synchronized audio playback across multiple nodes.
[0054] This disclosure does not limit the specific method for determining the data frame interval. In some optional embodiments, the first preset threshold information may include a first threshold value and a second threshold value. The first threshold value is greater than the second threshold value; that is, the first threshold value can be a high threshold value, and the second threshold value can be a low threshold value.
[0055] In some optional embodiments, if the first data volume is greater than a first threshold value, a reduction adjustment is made based on the default frame interval to determine the data frame interval between the i-th data frame and the (i+1)-th data frame.
[0056] It should be understood that when the first data volume is greater than the first threshold (high threshold), it indicates that the amount of audio data cached in the first buffer unit is too large. Therefore, in this embodiment, the default frame interval is reduced to determine the data frame interval between the i-th data frame and the (i+1)-th data frame. This accelerates the audio data reading from the first buffer unit by the first node. Thus, in a network where the clock frequency of the first node and the clock frequency of the audio source are different, the data frame interval between two data frames can be effectively and dynamically adjusted. This allows for adaptive and faster reading of audio data, and transmission of the audio data to the corresponding node through the network. Through this optional method, the first node can adaptively match the clock frequency of the audio source, enabling a better dynamic balance between the audio data transmission frequency of the first node and the clock frequency of the audio source. This prevents audio data buffer overflow caused by the difference between the clock frequency of the first node and the clock frequency of the audio source, thereby avoiding audio data loss and improving the audio data transmission and playback performance in various audio playback scenarios within the network. For example, in scenarios where multiple nodes are playing audio synchronously, the synchronization of audio playback can be effectively improved.
[0057] The default frame interval can be set as needed. For example, optionally, the default frame interval may include one or more clock cycles of the first node. The reduction adjustment can be implemented in any way. In some optional embodiments, the default frame interval can be subtracted by a first preset number of clock cycles, thereby facilitating and accurately determining the data frame interval between the i-th data frame and the (i+1)-th data frame.
[0058] For example, Figure 4A A schematic diagram illustrating the adjustment of the data frame interval between two data frames is shown. Optionally, in this example, one clock cycle can be represented by one clock tick. Combined with... Figure 4A In the example shown, the default frame interval can be set to x beats, and the first preset number can be set to n. After reading the audio data of the i=2nd data frame from the first buffer unit, when the first data volume is greater than the first threshold value, the default frame interval (i.e., x beats) can be subtracted from the first preset number of clock cycles (i.e., n beats), and the data frame interval between the second and third data frames is determined to be "xn beats".
[0059] In other optional embodiments, when adjusting the reduction based on the default frame interval, the default frame interval can be multiplied by a preset value greater than 0 and less than 1 (which can be preset as needed) to determine the data frame interval between the i-th data frame and the (i+1)-th data frame; or, when adjusting the reduction based on the default frame interval, the default frame interval can be divided by a preset value greater than 1 (which can be preset as needed) to determine the data frame interval between the i-th data frame and the (i+1)-th data frame. Alternatively, other methods can be used; there is no single limitation in the embodiments disclosed herein.
[0060] In some optional embodiments, if the first data volume is less than the second threshold value, an increase adjustment is made based on the default frame interval to determine the data frame interval between the i-th data frame and the (i+1)-th data frame.
[0061] It should be understood that when the first data volume is less than the second threshold (lower threshold), it indicates that the amount of audio data cached in the first buffer unit is too small. In this embodiment, the default frame interval is increased to determine the data frame interval between the i-th data frame and the (i+1)-th data frame. This slows down the audio data reading from the first buffer unit by the first node. Therefore, in a network where the clock frequency of the first node and the clock frequency of the audio source are different, the data frame interval between the two data frames can be effectively and dynamically adjusted. This allows for adaptive, slower audio data reading and transmission to the corresponding node via the network. Through this optional method, the first node can adaptively match the clock frequency of the audio source, enabling a better dynamic balance between the audio data transmission frequency of the first node and the clock frequency of the audio source. This prevents audio data buffer underflow caused by the difference between the clock frequency of the first node and the clock frequency of the audio source, thus avoiding audio playback stuttering and improving audio data transmission and playback performance in various audio playback scenarios within the network. For example, in scenarios where audio is played synchronously across multiple nodes, it can effectively improve the synchronization of audio playback and avoid playback stuttering.
[0062] The adjustment can be implemented in any way. In some optional embodiments, a second preset number of clock cycles can be added to the default frame interval to facilitate convenient and accurate determination of the data frame interval between the i-th data frame and the (i+1)-th data frame. The second preset number can be set to be the same as or different from the first preset number.
[0063] For example, Figure 4B Another schematic diagram illustrating the adjustment of the data frame interval between two data frames is shown. Optionally, in this example, one clock cycle can be represented by one clock tick. Combined with... Figure 4B In the example shown, the default frame interval can be set to x beats, and the second preset number can be set to n (that is, the second preset number is set to the same as the first preset number). After reading the audio data of the i=2th data frame from the first buffer unit, when the first data volume is less than the second threshold value, the default frame interval (i.e., x beats) can be added to the second preset number of clock cycles (i.e., n beats), and the data frame interval between the second and third data frames is determined to be "x+n beats".
[0064] In other optional embodiments, when increasing the adjustment based on the default frame interval, the default frame interval can be multiplied by a preset value greater than 1 (which can be preset as needed) to determine the data frame interval between the i-th data frame and the (i+1)-th data frame; or, when increasing the adjustment based on the default frame interval, the default frame interval can be divided by a preset value greater than 0 and less than 1 (which can be preset as needed) to determine the data frame interval between the i-th data frame and the (i+1)-th data frame. Alternatively, other methods can be used; there is no unique limitation in the embodiments of this disclosure.
[0065] In some optional embodiments, if the first data volume is less than or equal to the first threshold value and greater than or equal to the second threshold value, the default frame interval is determined as the data frame interval between the i-th data frame and the (i+1)-th data frame.
[0066] It should be understood that when the first data volume is between the second threshold (low threshold) and the first threshold (high threshold), it indicates that the amount of audio data cached in the first buffer unit is normal. In this embodiment, the default frame interval can be directly determined as the data frame interval between the i-th data frame and the (i+1)-th data frame. This allows for effective dynamic adjustment of the data frame interval between two data frames in a network where the clock frequency of the first node and the clock frequency of the audio source are different. This enables the audio data to be read at an adaptive speed and transmitted to the corresponding node through the network. Through this optional method, the first node can adaptively match the clock frequency of the audio source, achieving a better dynamic balance between the audio data transmission frequency of the first node and the clock frequency of the audio source. This prevents audio data buffer overflow or underflow caused by the difference between the clock frequency of the first node and the clock frequency of the audio source, thereby avoiding audio data loss or audio playback stuttering at the corresponding node. This improves the audio data transmission and playback performance in various audio playback scenarios within the network.
[0067] Optionally, in Figure 4A , Figure 4B In the example, one clock cycle can be represented by one clock tick. Combined with... Figure 4A , Figure 4B In the example shown, the default frame interval can be set to x frames. When the first data volume is normal (i.e., between the second threshold and the first threshold), the default frame interval (i.e., x frames) can be used to determine the data frame interval between two adjacent data frames as "x frames".
[0068] It should be noted that, Figure 4A and Figure 4B ,as well as Figure 5For ease of understanding, the schematic diagram of the data frame only shows the data blocks corresponding to audio data and other data, and is not intended to limit the technical solutions of the embodiments of this disclosure.
[0069] S106: After the data frame interval ends, read the audio data of the (i+1)th data frame from the first buffer unit, and transmit the audio data of the (i+1)th data frame to the corresponding node among other nodes through the network.
[0070] After the data frame interval between the i-th data frame and the (i+1)-th data frame is determined, we can wait for the data frame interval to end, and then read the audio data of the (i+1)-th data frame from the first buffer unit, thereby transmitting the audio data of the (i+1)-th data frame to the node that needs to play the audio data for playback.
[0071] Based on this, through the audio data transmission scheme of steps S102 to S106 in the embodiments of this disclosure, in a network where the clock frequency of the first node and the clock frequency of the audio source are different, the data frame interval between two data frames can be dynamically adjusted by the first data amount of audio data cached in the first buffer unit. This allows the audio data to be transmitted to the corresponding node through the network at an adaptive rhythm. Thus, through this optional method, the first node can adaptively match the clock frequency of the audio source, enabling the transmission frequency of the audio data of the first node to achieve a better dynamic balance with the clock frequency of the audio source. This prevents audio data buffer overflow or underflow problems caused by the difference between the clock frequency of the first node and the clock frequency of the audio source, thereby avoiding audio data loss or audio playback stuttering at the corresponding node. This improves the audio data transmission and playback performance in various audio playback scenarios in the network.
[0072] In some optional embodiments, in step S106, the audio data of the (i+1)th data frame can be transmitted to the second node among the other nodes through networking, so that the second node can at least synchronize the playback of audio data with at least one node among the other nodes at a predetermined time.
[0073] Based on this, the technical solution in this embodiment can dynamically adjust the data frame interval between two data frames by using the first data amount of audio data cached in the first buffer unit, thereby enabling the audio data to be transmitted to the second node through the network at an adaptive rhythm. This can improve the audio data transmission and playback performance when the second node in the network and at least one other node synchronously execute audio data playback at a predetermined time, which is beneficial to improving the synchronization of multi-node audio synchronous playback.
[0074] The second node can be any of the other audio playback nodes, for example, using... Figure 1A , Figure 1B , Figure 1C In the example shown, any one of nodes 1 to 4 can be used, and the second node can be node 1. The audio data of the (i+1)th data frame can be transmitted to the second node (node 1) among the other nodes through networking, so that the second node can synchronously play the audio data with at least one other audio playback node (i.e., at least one of node 2, node 3, and node 4) at a predetermined time.
[0075] It should be understood that the relevant content regarding the synchronous execution of audio data playback by the audio playback node and other nodes at a predetermined time can be understood with reference to the embodiments in the second aspect below, and will not be elaborated here.
[0076] In some optional embodiments, i ≥ 2 in step S104 of the present disclosure. Optionally, the audio data transmission method in the present disclosure further includes: after the first node is powered on and before the audio data is read from the first buffer unit for the first time, performing phase adjustment with the audio frames including audio data of the sound source based on the first data frame to the j-th data frame, so that the data frames and audio frames have a fixed phase relationship, wherein 0 < j < i and j is an integer.
[0077] For example, if the initial time when the first node generates data frames and the initial time when the audio source generates audio frames are random after the first node is powered on, and the time interval between the two initial times is also random, this can easily complicate the audio noise reduction of the network. For example, when the network is applied to vehicle networking, the above-mentioned randomness can complicate the entire vehicle's ANC (Active Noise Control) or RNC (Road Noise Control).
[0078] Based on this, in the above optional embodiments, during the initialization stage from the first node power-on to the first reading of audio data from the first buffer unit, phase adjustment can be performed on the audio frames including audio data of the sound source based on the first data frame to the j-th data frame, so that the data frame and the audio frame have a fixed phase relationship, and the initial phase alignment can be achieved. In this way, a correct time starting point can be established for the stable transmission of subsequent audio data before the audio data transmission of the first node. It is understood that after the phase adjustment in this embodiment is completed, the audio data of the i-th data frame is read and transmitted. This can effectively reduce the complexity of audio noise reduction in the network and facilitate the dynamic adjustment of the data frame interval between two data frames based on the first data amount of audio data cached in the first buffer unit. This allows the audio data to be transmitted to the corresponding node through a ring network or daisy chain network at an adaptive rhythm. This enables the first node to more easily and accurately adaptively match the clock frequency of the sound source, and enables the transmission frequency of the audio data of the first node to achieve a better dynamic balance with the clock frequency of the sound source. This better prevents audio data buffer overflow or underflow caused by the difference between the clock frequency of the first node and the clock frequency of the sound source, thereby avoiding audio data loss or audio playback stuttering at the corresponding node. This is more conducive to improving the audio data transmission and playback performance in various audio playback scenarios, including but not limited to synchronous audio playback of multiple nodes, in a ring network or daisy chain network.
[0079] Optionally, j can be selected as needed. For example, j can be 1, 2, 3, etc. Here, we can take j=2 as an example, then i can be at least 3.
[0080] Optionally, from the first data frame to the j-th data frame, the first node does not read audio data, and the data segment storing audio data in the first data frame to the j-th data frame can be empty data.
[0081] The phase adjustment method described above can be selected according to actual needs. For example, in some optional embodiments, refer to... Figure 1A , Figure 1B , Figure 1C In the example of the network shown, the audio source can be directly connected to the first node (such as node 0). Phase adjustment can be performed as follows: determine one or more target data frames from the first data frame to the j-th data frame. For each target data frame, based on the received audio frame, adjust the data frame interval between the target data frame and the next data frame of the target data frame to achieve phase adjustment.
[0082] Based on this, using the aforementioned optional method, when the audio source is directly connected to the first node, by adjusting the target data frame and the data frame interval of its next frame, phase adjustment can be accurately and effectively achieved between the first and j-th data frames and the audio frame. This ensures a fixed phase relationship between the data frames and the audio frame, achieving initial phase alignment. Consequently, a correct time starting point can be established for the stable transmission of subsequent audio data before the audio data is transmitted at the first node. After the phase adjustment is completed in this embodiment, the audio data of the i-th data frame is read and transmitted, thereby effectively reducing the complexity of audio noise reduction in the network.
[0083] It should be noted that the target data frame can include one or more of the first to the j-th data frames. That is, all data frames from the first to the j-th data frames can be designated as target data frames to adjust the data frame interval between the target data frame and the next data frame, thereby achieving phase adjustment; alternatively, a portion of the data frames from the first to the j-th data frames can be designated as target data frames, while another portion can be left undesignated, to adjust the data frame interval between the target data frame and the next data frame, thereby achieving phase adjustment.
[0084] Optionally, when adjusting the data frame interval between the target data frame and the next data frame, the data frame interval can be increased or decreased as needed. For example, increasing the interval can be achieved by adding a preset number of clock cycles to a default data frame interval. Conversely, decreasing the interval can be achieved by subtracting a preset number of clock cycles from a default data frame interval.
[0085] For example, Figure 5 This diagram illustrates the reading of audio data from the first node after phase adjustment of the data and audio frames. (Refer to...) Figure 5 In the example shown, the initial time when the first node generates the first data frame is different from the initial time when the sound source generates the first audio frame including audio data (1), and the time interval between the two initial times is random. Figure 5In the example shown, using j=2, the target data frame can be selected as the j=2nd data frame. Before adjusting the data frame interval between two adjacent data frames, the default data frame interval can be x beats (i.e., x clock cycles). After adjusting the data frame interval between the target data frame (the j=2nd data frame) and its next data frame (i.e., the i=3rd data frame), the data frame interval between the target data frame and its next data frame is adjusted to "x+m beats" (that is, adding m clock cycles to the default data frame interval), thereby achieving phase adjustment and ensuring that the data frame and audio frame have a fixed phase relationship. Figure 5 As shown, after phase adjustment, the audio frame including audio data (1) can be read and stored in the i=3rd data frame. It should be understood that the above... Figure 5 The examples provided are for illustrative purposes only and are not intended to limit the scope of the embodiments disclosed herein.
[0086] In some alternative embodiments, refer to Figure 1D In the example of the network shown, the audio source can be directly connected to a target node (such as node 1) in other nodes besides the first node. Phase adjustment can be performed as follows: determine one or more target data frames from the first data frame to the j-th data frame, and for each target data frame, send the target data frame to the target node through the network so that the target node determines the phase offset based on the target data frame and the audio frame. Receive the phase offset transmitted by the target node and perform phase adjustment based on the phase offset.
[0087] Based on this, using the aforementioned optional method, even when the audio source is not directly connected to the first node but directly connected to the target node among other nodes, the phase offset can be determined at the target node based on the target data frame and the audio frame. This allows the first node to accurately and effectively adjust the phase with the audio frame based on the phase offset received from the target node, from the first data frame to the j-th data frame. This ensures a fixed phase relationship between the data frame and the audio frame, achieving initial phase alignment. Consequently, it can accurately and conveniently establish the correct time starting point for stable transmission of subsequent audio data before the first node transmits audio data. After the phase adjustment is completed in this embodiment, the audio data of the i-th data frame is read and transmitted, effectively reducing the complexity of audio noise reduction in the network.
[0088] It should be noted that the target data frame can include one or more of the first to the j-th data frames. That is, all data frames from the first to the j-th data frames can be identified as target data frames, and the target data frames can be transmitted to the target node to determine the phase offset, so as to achieve phase adjustment; or a portion of the data frames from the first to the j-th data frames can be identified as target data frames, while the other portion of data frames can be deferred as target data frames, and the target data frames can be transmitted to the target node to determine the phase offset, so as to achieve phase adjustment.
[0089] Optionally, the target node can receive the phase offset transmitted by the network and perform phase adjustment based on the phase offset. Alternatively, the target node can also transmit the phase offset in other ways, such as, but not limited to, wireless transmission outside the network.
[0090] In some optional embodiments, the phase offset transmitted by the target node can be received by receiving a data frame containing audio data returned by the target node through the network, and obtaining the phase offset from the common header of the data frame.
[0091] Alternatively, you can refer to the previous section about Figure 3 The common header of the data frame shown is used to understand its format. Optionally, after the first node transmits the target data frame to the target node via the network, the target node can store the audio data from the audio frame obtained from the audio source into the target data frame, store the determined phase offset in the common header of the target data frame, and then return the target data frame to the first node via the network. This allows the phase offset to be fed back through the common header of the data frame. The first node can extract the phase offset from the common header of the returned target data frame to perform phase adjustment based on the phase offset.
[0092] It should be understood that by storing the phase offset in the common header of the data frames (including audio data) returned by the target node through the network, the first node can more easily obtain the phase offset. This allows for accurate and efficient phase adjustment with the audio frame based on the first to j-th data frames, thus ensuring a fixed phase relationship between the data frames and the audio frame. Furthermore, since the target node directly connected to the audio source can return data frames to the first node to transmit audio data while simultaneously storing the phase offset in the common header of those data frames, this effectively improves data transmission efficiency and reduces data transmission overhead.
[0093] It is understood that the above content is only some optional implementations of the audio data transmission method in the first aspect, and is not intended to limit the embodiments of this disclosure.
[0094] According to a second aspect of the embodiments of this disclosure, an audio data transmission method is provided, which can be used as an audio playback node among multiple nodes in a network. For example... Figure 6 The flowchart shown illustrates that the audio data transmission method may include the following steps S202, S204, and S206, specifically: S202: Cache audio data from the sound source through the second buffer unit.
[0095] Optionally, the network can adopt any type and any network structure. For example, in some optional embodiments, the network in this disclosure can be a ring network, a daisy-chain network, or other types of networks. When it is a ring network, it can be combined with... Figure 1A and Figure 1D Please refer to the diagram for clarification. The audio playback node has already been explained previously and will not be repeated here.
[0096] It should be understood that when the network topology used in the technical solution of this embodiment is a ring network or a daisy-chain network, the audio playback node can dynamically adjust the data reading frequency of the audio data through the second data amount of the audio data cached in the second buffer unit. This allows the audio data to be read at an adaptive rhythm for playback. The audio playback node can adaptively match the clock frequency of the sound source, enabling the audio data reading frequency of the audio playback node to achieve a better dynamic balance with the clock frequency of the sound source. This prevents audio data buffer overflow or underflow problems caused by the difference between the clock frequency of the audio playback node and the clock frequency of the sound source, thereby avoiding audio data loss or audio playback stuttering of the corresponding node. This improves the audio data transmission and playback performance of multiple nodes in the ring network or daisy-chain network for synchronized audio playback.
[0097] The technical solutions in this disclosure can be used in any application scenario. For example, in some optional embodiments, the network can be an in-vehicle network, which can be used for in-vehicle data transmission and / or processing functions. Correspondingly, the audio data transmission scheme can be used for in-vehicle audio data transmission, thereby improving the audio data transmission and playback performance in in-vehicle audio playback scenarios.
[0098] In other examples, networking can also be used for data transmission and / or processing in scenarios including but not limited to home or security. Correspondingly, audio data transmission schemes can be used for audio data transmission in scenarios including but not limited to home or security. There is no unique limitation in the embodiments of this disclosure.
[0099] In this embodiment of the disclosure, the second buffer unit can be a buffer corresponding to an audio playback node, and can be used to cache audio data. For example... Figure 1A , Figure 1B , Figure 1C As shown, audio data from the audio source can be transmitted from the master node (node 0) directly connected to the audio playback node through a network to the audio playback nodes (i.e., each child node, nodes 1-4), and cached in the second cache unit. Alternatively, as... Figure 1D As shown, audio data from the audio source can also include audio data transmitted from the audio source directly connected to the audio playback node (such as node 1) to the audio playback node and cached in the second cache unit.
[0100] Optionally, the audio playback node can be a first node or a second node, wherein the clock frequency of the first node is different from the clock frequency of the audio source, and the second node can be any node among multiple nodes other than the first node.
[0101] For example Figure 1A , Figure 1B , Figure 1C In any of the network topologies shown, the audio playback nodes can be nodes 1 through 4 (each can be a second node), and the clock frequency of node 0 (which can be the first node) can be different from the clock frequency of the audio source. For example, Figure 1D In the network shown, the audio playback node can be node 0 (which can be the first node) or nodes 1 to 4 (which can all be the second node). The clock frequency of node 0 (which can be the first node) can be different from the clock frequency of the audio source.
[0102] It should be understood that, since the clock frequency of the first node is different from the clock frequency of the audio source, the second node can be any node other than the first node among multiple nodes, and the audio playback node can be either the first node or the second node. Therefore, when adopting the audio data transmission scheme of this embodiment, in a network where the clock frequency of the first node is different from the clock frequency of the audio source, any audio playback node among the first node and the second node can dynamically determine the data reading frequency of the audio data through the second data amount of the audio data cached in the second buffer unit. This allows the audio playback node to read the audio data at an adaptive rhythm for playback. Any audio playback node among the first node and the second node can adaptively match the clock frequency of the audio source, enabling the audio data reading frequency of the audio playback node among the first node and the second node to achieve a better dynamic balance with the clock frequency of the audio source. This prevents audio data buffer overflow or underflow problems caused by the difference between the clock frequency of the audio playback node and the clock frequency of the audio source, thereby avoiding audio data loss or audio playback stuttering of the corresponding node. This improves the audio data transmission and playback performance of audio synchronous playback of multiple nodes in the network.
[0103] Optionally, the node directly connected to the audio source can be a non-audio playback node (such as...). Figure 1A , Figure 1B, Figure 1C As shown in node 0). Alternatively, a node directly connected to the audio source can also be an audio playback node (such as...). Figure 1D (As shown in node 1). Settings can be configured as needed.
[0104] S204: Determine the data reading frequency of the audio data based on the second data volume of the audio data cached in the second buffer unit.
[0105] In this embodiment of the disclosure, the audio playback node can dynamically adjust the data reading frequency of the audio data by using the second data amount of the audio data cached in the second buffer unit, thereby facilitating the synchronous playback of audio data from multiple nodes in the future.
[0106] The specific implementation of step S204 is not limited in this embodiment. In some optional embodiments, step S204 may include: determining the data reading frequency of the audio data based on the second data volume of the audio data cached in the second buffer unit and the second preset threshold information.
[0107] Based on this, in this embodiment of the present disclosure, by determining the second data volume of the audio data cached in the second cache unit, and based on the second data volume and the second preset threshold information, the data reading frequency of the audio data can be effectively and dynamically determined. This allows the audio data to be read at an adaptive rhythm for playback. As a result, the audio playback node can adaptively match the clock frequency of the sound source, enabling the audio data reading frequency of the audio playback node to achieve a better dynamic balance with the clock frequency of the sound source. This prevents audio data cache overflow or underflow problems caused by the difference between the clock frequency of the audio playback node and the clock frequency of the sound source, thereby avoiding audio data loss or audio playback stuttering of the corresponding node. This improves the audio data transmission and playback performance of audio synchronous playback of multiple nodes in the network.
[0108] Optionally, the second preset threshold information may include one or more threshold values, which may be used to compare with the second data volume to facilitate the data reading frequency of the audio data.
[0109] In some optional embodiments, the data volume unit of any threshold value in the second preset threshold information is bits. This allows the determined second data volume of the audio data to be in bits. By using a second data volume in bits, the data reading frequency of the audio data can be dynamically adjusted with higher precision, facilitating more precise reading of the audio data at an adaptive rhythm for playback. This, in turn, can better improve the audio data transmission and playback performance for synchronized audio playback across multiple nodes in the network.
[0110] In some optional embodiments, the second buffer unit buffers audio data in bits. This allows the determined second data volume of the audio data to be measured in bits. By using a bit-based second data volume, the data reading frequency of the audio data can be dynamically adjusted with higher precision. This facilitates more precise reading of the audio data at an adaptive rhythm for playback, thereby improving the audio data transmission and playback performance for synchronized audio playback across multiple nodes in the network.
[0111] This disclosure does not limit the specific method for determining the data reading frequency of audio data. In some optional embodiments, the second preset threshold information may include a third threshold and a fourth threshold. The third threshold is greater than the fourth threshold; that is, the third threshold can be a high threshold, and the fourth threshold can be a low threshold.
[0112] In some optional embodiments, if the second data volume is greater than the third threshold, the data reading frequency of the audio data is determined by increasing the default reading frequency.
[0113] It should be understood that when the second data volume is greater than the third threshold (high threshold), it indicates that the amount of audio data cached in the second buffer unit is too large. In this embodiment, the default reading frequency is increased to determine the data reading frequency of the obtained audio data. This can accelerate the audio data reading from the second buffer unit by the audio playback node, thereby effectively and dynamically adjusting the audio data reading frequency. This allows for adaptive and faster audio data reading for better audio playback. Through this optional method, the audio playback node can adaptively match the clock frequency of the audio source, enabling a better dynamic balance between the audio data reading frequency of the audio playback node and the clock frequency of the audio source. This prevents audio data buffer overflow or underflow problems caused by differences between the clock frequency of the audio playback node and the clock frequency of the audio source, thus avoiding audio data loss or audio playback stuttering at the corresponding node. This improves the audio data transmission and playback performance for synchronized audio playback across multiple nodes in the network.
[0114] The default reading frequency can be set as needed and is not limited here. Increasing or adjusting it can be done in any way. In some optional embodiments, the default reading frequency can be adjusted to a first preset reading frequency greater than the default reading frequency, thereby facilitating and accurately determining the data reading frequency of the obtained audio data.
[0115] For example, Figure 7A A schematic diagram illustrating the adjustment of the audio data readout frequency is shown. Optionally, combined with... Figure 7AIn the example shown, the default read frequency can be set to p, and the first preset read frequency can be set to q1. When the second data volume corresponding to the second cache unit is greater than the third threshold value, the default read frequency (i.e., p) can be adjusted to the first preset read frequency (i.e., q1) which is greater than the default read frequency.
[0116] In other optional embodiments, when increasing the adjustment based on the default reading frequency, the default reading frequency can be multiplied by a preset value greater than 1 (which can be preset as needed) to determine the data reading frequency of the audio data; alternatively, when increasing the adjustment based on the default reading frequency, the default reading frequency can be divided by a preset value greater than 0 and less than 1 (which can be preset as needed) to determine the data reading frequency of the audio data; or, the default reading frequency can be added to a preset value (which can be preset as needed) to determine the data reading frequency of the audio data. Other methods can also be used; there is no single limitation in the embodiments disclosed herein.
[0117] In some optional embodiments, if the second data volume is less than the fourth threshold, a reduction adjustment is made based on the default reading frequency to determine the data reading frequency of the audio data.
[0118] It should be understood that when the second data volume is less than the fourth threshold (low threshold), it indicates that the amount of audio data cached in the second buffer unit is too small. In this embodiment, the default reading frequency is reduced to determine the data reading frequency of the obtained audio data. This slows down the audio data reading from the second buffer unit by the audio playback node, thereby effectively and dynamically adjusting the audio data reading frequency. This allows for adaptive, slower audio data reading for better audio playback. Through this optional method, the audio playback node can adaptively match the clock frequency of the audio source, achieving a better dynamic balance between the audio data reading frequency of the audio playback node and the clock frequency of the audio source. This prevents audio data buffer overflow or underflow caused by differences between the clock frequency of the audio playback node and the clock frequency of the audio source, thus avoiding audio data loss or audio playback stuttering at the corresponding node. This improves the audio data transmission and playback performance for synchronized audio playback across multiple nodes in the network.
[0119] The reduction adjustment can be implemented in any way. In some optional embodiments, the default reading frequency can be adjusted to a second preset reading frequency that is lower than the default reading frequency, thereby facilitating and accurately determining the data reading frequency of the obtained audio data. The second preset reading frequency can be set to be the same as or different from the first preset reading frequency.
[0120] For example, Figure 7BAnother schematic diagram illustrating the adjustment of the audio data readout frequency is shown. Optionally, combined with... Figure 7B In the example shown, the default reading frequency can be set to p, and the second preset reading frequency can be set to q2. When the second data volume is less than the fourth threshold, the default reading frequency (i.e., p) can be adjusted to the second preset reading frequency (i.e., q2), which is less than the default reading frequency.
[0121] In other optional embodiments, when adjusting the reduction based on the default reading frequency, the default reading frequency can be multiplied by a preset value greater than 0 and less than 1 (which can be preset as needed) to determine the audio data reading frequency; alternatively, the default reading frequency can be divided by a preset value greater than 1 (which can be preset as needed) to determine the audio data reading frequency; or, the default reading frequency can be subtracted from a preset value (which can be preset as needed) to determine the audio data reading frequency. Other methods may also be used; there is no single limitation in the embodiments disclosed herein.
[0122] In some optional embodiments, if the second data volume is less than or equal to the third threshold and greater than or equal to the fourth threshold, the default reading frequency will be determined as the upper or lower limit of the data reading frequency of the audio data.
[0123] It should be understood that when the second data volume is between the fourth threshold (low threshold) and the third threshold (high threshold), it indicates that the amount of audio data cached in the second buffer unit is normal. In this embodiment, the default reading frequency can be directly determined as the upper or lower limit of the audio data reading frequency, thereby effectively and dynamically adjusting the audio data reading frequency to read audio data at an adaptive speed for audio playback. Through this optional method, the audio playback node can adaptively match the clock frequency of the audio source, enabling the audio data reading frequency of the audio playback node to achieve a better dynamic balance with the clock frequency of the audio source. This prevents audio data buffer overflow or underflow problems caused by differences between the clock frequency of the audio playback node and the clock frequency of the audio source, thus avoiding audio data loss or audio playback stuttering at the corresponding node. This improves the audio data transmission and playback performance for synchronized audio playback across multiple nodes in the network.
[0124] Optionally, in Figure 7AIn the example, the default reading frequency can be set to p. Then, when the second data volume recovers from being greater than the third threshold to normal (i.e., between the fourth and third thresholds), the default reading frequency p can be determined as the lower limit of the audio data reading frequency, while the upper limit can be determined as the first data reading frequency q1. That is, as... Figure 7B As shown, if the data reading frequency is t, then p≤t<q1.
[0125] Optionally, in Figure 7B In the example, the default read frequency can be set to p. Then, when the second data volume recovers from below the fourth threshold to normal (i.e., between the fourth and third thresholds), the default read frequency p can be determined as the upper limit of the audio data read frequency, while the lower limit can be determined as the second data read frequency q2. That is, as... Figure 7B As shown, if the data reading frequency is t, then q2<t≤p.
[0126] S206: Based on the data reading frequency, read audio data from the second buffer unit to synchronously play audio data with at least one other node among multiple nodes at a predetermined time.
[0127] Once the audio data reading frequency is determined, audio data can be read from the second buffer unit based on that frequency, thus facilitating synchronized audio data playback with at least one other node among multiple nodes at a predetermined time, achieving synchronized audio playback across multiple nodes.
[0128] Based on this, the audio data transmission scheme of steps S202-S206 in the embodiments of this disclosure can be used for audio playback nodes among multiple nodes in a network. It can cache audio data from the sound source through a second buffer unit, and then determine the audio data reading frequency based on the second data volume of the audio data cached in the second buffer unit. Furthermore, based on the data reading frequency, audio data can be read from the second buffer unit to synchronously execute audio data playback with at least one other node among the multiple nodes (excluding the audio playback node) at a predetermined time. Therefore, through the technical solution of the embodiments of this disclosure, the audio data cached in the second buffer unit can be used to... The second data volume dynamically determines the audio data reading frequency, enabling the audio data to be read at an adaptive rhythm for playback. Through this optional method, the audio playback node can adaptively match the clock frequency of the audio source, achieving a better dynamic balance between the audio data reading frequency of the audio playback node and the clock frequency of the audio source. This prevents audio data buffer overflow or underflow problems caused by differences between the clock frequency of the audio playback node and the clock frequency of the audio source, thereby avoiding audio data loss or audio playback stuttering at corresponding nodes. This improves the audio data transmission and playback performance for synchronized audio playback across multiple nodes in the network.
[0129] In some optional embodiments, the third and fourth threshold values corresponding to the audio playback node are both determined based on the target number of nodes through which the audio data is transmitted between the node directly connected to the audio source and the audio playback node in the network; wherein, the smaller the target number, the larger the third threshold value and the larger the fourth threshold value; conversely, the larger the target number, the smaller the third threshold value and the smaller the fourth threshold value.
[0130] In this embodiment of the disclosure, different second preset threshold values can be used for different audio playback nodes. For example, the third threshold value corresponding to different audio playback nodes can be different, and the fourth threshold value corresponding to different audio playback nodes can also be different. In the above optional embodiments, the third and fourth threshold values corresponding to the audio playback node can both be determined based on the target number of nodes through which audio data is transmitted between the node directly connected to the audio source (i.e., the audio source node) and the audio playback node in the network. It can be understood that the smaller the target number, the closer the audio playback node is to the audio source node on the audio data transmission path, and the smaller the accumulated transmission delay of the audio data; conversely, the larger the target number, the farther the audio playback node is from the audio source node on the audio data transmission path, and the greater the accumulated transmission delay of the audio data.
[0131] In the optional embodiments of this disclosure, the smaller the target quantity, the larger the third threshold value and the larger the fourth threshold value are set; conversely, the larger the target quantity, the smaller the third threshold value and the smaller the fourth threshold value are set. Therefore, through this optional approach, larger third and fourth threshold values can be used for audio playback nodes closer to the audio source node, and smaller third and fourth threshold values can be used for audio playback nodes farther from the audio source node. This effectively improves the stability of multi-node audio synchronous playback and ensures that multiple audio playback nodes can keep up with the rhythm of audio synchronous playback, enabling audio playback nodes and at least one node to more accurately synchronize audio data playback at a predetermined time, thereby effectively reducing the playback latency of multi-node audio synchronous playback.
[0132] The third and fourth threshold values can be preset as needed, as long as the requirements are met. In some optional embodiments, the third threshold value A corresponding to the audio playback node satisfies the relationship: A = C + D + E; the fourth threshold value B corresponding to the audio playback node satisfies the relationship: B = C + D - E; where C is half of the total data capacity of the second buffer unit, D is a preset offset negatively correlated with the target quantity, and E is a preset threshold offset.
[0133] The total data capacity of the second cache unit depends on the actual selection and is not limited here. For example, if the total data capacity of the second cache unit is 512 bits, then half of the total data capacity C = 512 / 2 = 256 bits.
[0134] For the audio source node and each node, their positions in the network are fixed. Therefore, given a fixed audio data transmission path and direction, the latency of audio data transmission from the audio source to each node is also fixed. This latency is related to the target number of nodes through which the transmitted audio data passes between the directly connected audio source node and the audio playback node. A larger target number results in a larger latency, and a smaller target number results in a smaller latency. Therefore, a preset offset corresponding to this latency can be pre-set for the second buffer unit of each audio playback node. This preset offset is negatively correlated with the target number; that is, a smaller target number results in a larger preset offset, and a larger target number results in a smaller preset offset. Consequently, it can be effectively ensured that a smaller target number for the audio playback node results in a larger third threshold value and a larger fourth threshold value; conversely, a larger target number results in a smaller third threshold value and a smaller fourth threshold value. The preset offset D can be preset as needed.
[0135] The preset threshold offset E can be set as needed. By using C+D±E, the high threshold value (the third threshold value A) and the low threshold value (the fourth threshold value B) can be obtained respectively.
[0136] For example, Figure 8 A schematic diagram is shown showing the third and fourth threshold values corresponding to different audio playback nodes in a network. Figure 8 Can be combined to Figure 1A or Figure 1B To understand the network topology shown, for ease of explanation, we can use the following: Figure 1B This can be illustrated using a single chrysanthemum chain network. For example... Figure 1B As shown, node 0 is the audio source node directly connected to the audio source. Nodes 1 through 4 can all be audio playback nodes. The total data capacity of the second buffer unit of nodes 1 through 4 can be 512 bits (of course, it is not limited to this). Therefore, half of the total data capacity C of the second buffer unit of nodes 1 through 4 is 256 bits. The audio source can transmit audio data to node 0, and the audio data is written to the first buffer unit of node 0. If nodes 1 through 4 are made to synchronously perform audio playback, then node 0 can read audio data from the first buffer unit and adaptively send it to nodes 1, 2, 3, and 4 through network topology. And as... Figure 1BAs can be seen, in the audio data transmission path, for the four audio playback nodes, there are 0 nodes between node 1 and node 0 (target quantity is 0), 1 node between node 2 and node 0 (target quantity is 1), 2 nodes between node 3 and node 0 (target quantity is 2), and 3 nodes between node 4 and node 0 (target quantity is 3). Node 1 receives the audio data first, and node 4 receives the audio data last. Therefore, the preset offset D for node 1 can be set to 3×30 bits, the preset offset D for node 2 to 2×30 bits, the preset offset D for node 3 to 1×30 bits, and the preset offset D for node 4 to 0 bits. It is evident that the preset offset D for nodes 1 to 4 is negatively correlated with the target quantity; that is, the smaller the target quantity, the larger the preset offset, and vice versa. Assuming a preset threshold offset E = 50 bits, the following values can be obtained: Node 1: Third threshold A = (256 + 3 × 30 + 50) bits = 396 bits, Fourth threshold B = (256 + 3 × 30 - 50) bits = 296 bits; Node 2: Third threshold A = (256 + 2 × 30 + 50) bits = 366 bits, Fourth threshold B = (256 + 2 × 30 - 50) bits = 266 bits; Node 3: Third threshold A = (256 + 1 × 30 + 50) bits = 336 bits, Fourth threshold B = (256 + 1 × 30 - 50) bits = 236 bits; Node 4: Third threshold A = (256 + 0 + 50) bits = 306 bits, Fourth threshold B = (256 + 0 - 50) bits = 206 bits. It is understandable that the above combined... Figure 1B and Figure 8 The descriptions provided are merely illustrative and not intended to limit the scope of the embodiments disclosed herein. It should be understood that, for Figure 1A In a ring network, when the direction of audio data transmission (such as clockwise or counterclockwise) is fixed, for example, node 0 directly connected to the audio source... Figure 1A The clockwise audio data transmission shown above can be referred to in conjunction with the above. Figure 1B and Figure 8 To understand the explanation; if node 0 along Figure 1A The principle is similar for audio data transmission in the counter-clockwise direction shown in the diagram; simply understand it in reverse order.
[0137] Based on this, by adopting the optional embodiment of the relationship between the third threshold value A and the fourth threshold value B, the third threshold value and the fourth threshold value can be effectively obtained. This allows audio playback nodes closer to the sound source node to use larger third and fourth threshold values, while audio playback nodes farther from the sound source node can use smaller third and fourth threshold values. This effectively improves the stability of multi-node audio synchronous playback and ensures that multiple audio playback nodes can keep up with the rhythm of audio synchronous playback. This enables audio playback nodes and at least one node to more accurately synchronize audio data playback at a predetermined time, thereby effectively reducing the playback delay of multi-node audio synchronous playback.
[0138] This disclosure does not limit the specific implementation of step S206. In some optional embodiments, step S206 may include: reading audio data from a predetermined storage location in the second buffer unit based on the data reading frequency, so as to synchronously perform audio data playback with at least one other node among a plurality of nodes at a predetermined time, wherein the predetermined storage location is half of the total data capacity of the second buffer unit plus the storage location indicated by a preset offset, the preset offset being negatively correlated with a target number, and the target number being the number of nodes through which audio data is transmitted between the node directly connected to the audio source and the audio playback node in the network.
[0139] Typically, when an audio playback node reads audio data from the second buffer unit, it can read the audio data from the storage location indicated by half of the total data capacity of the second buffer unit, that is, from the middle position of the second buffer unit.
[0140] The optional method described above in this embodiment can read audio data from a predetermined storage location in the second cache unit based on a determined data reading frequency. The predetermined storage location is half of the total data capacity of the second cache unit plus the storage location indicated by a preset offset. The preset offset is negatively correlated with the target number, which is the number of nodes through which audio data is transmitted between the node directly connected to the audio source and the audio playback node in the network. This allows the audio playback node to achieve more precise synchronous playback of audio data with at least one other node (which can also be an audio playback node) at a predetermined time, thereby improving the audio data transmission and playback performance of audio synchronous playback of multiple nodes in the network.
[0141] It should be understood that the total data capacity of the second cache unit, the preset offset, and the target number have been explained in the previous text and can be understood in conjunction with the previous text, so they will not be repeated here.
[0142] In some alternative embodiments, step S206 may include: reading audio data from the second buffer unit based on the data reading frequency, and performing delayed playback on the read audio data based on the preset delay time corresponding to the audio playback node, so as to synchronously perform audio data playback with at least one other node among the multiple nodes at a predetermined time.
[0143] As explained earlier, the location of the audio source node and each node in the network is fixed. Therefore, given a fixed audio data transmission path and direction, the delay in audio data transmission from the source to each node is also deterministic. This delay is related to the target number of nodes through which the transmitted audio data passes between the node directly connected to the audio source and the audio playback node. A larger target number results in a larger delay, and a smaller target number results in a smaller delay. Therefore, in this embodiment, different preset delay times can be set for different audio playback nodes to facilitate synchronized audio playback across multiple nodes. For example, a smaller target number for an audio playback node allows for a longer preset delay time, while a larger target number allows for a shorter preset delay time. The specific settings can be configured as needed, and there is no single limitation.
[0144] Based on this, in this embodiment of the present disclosure, by setting a preset delay time for the audio playback node, and then reading audio data from the second buffer unit based on the data reading frequency, the audio data read is played with a delay based on the preset delay time corresponding to the audio playback node. This enables the audio playback node to achieve more precise synchronous playback of audio data at a predetermined time with at least one other node (which can also be an audio playback node) among multiple nodes, thereby improving the audio data transmission and playback performance of audio synchronization playback of multiple nodes in the network.
[0145] It is understood that the above content is only some optional implementations of the audio data transmission method in the second aspect, and is not intended to limit the embodiments of this disclosure.
[0146] According to a third aspect of the present disclosure, an audio data transmission system is provided. The audio data transmission system includes: multiple nodes connected in a network, including a first node and a second node, wherein the second node is any one of the other nodes besides the first node, and the clock frequency of the first node is the same as the clock frequency of the audio source. The first node is configured to: send target data to the second node through the network, so that the second node synchronizes its clock frequency to be the same as the clock frequency of the audio source based on the target data; receive audio data from the audio source, and transmit the audio data to the second node through the network. The second node is configured to: perform delayed playback on the received audio data based on a preset delay time corresponding to the second node, so as to achieve synchronized audio playback with at least one of the other nodes.
[0147] Based on this, the audio data transmission system described in this embodiment of the present disclosure, since the clock frequency of the first node in its network is the same as the clock frequency of the audio source, and the first node can send target data to the second node through the network, the second node can synchronize its clock frequency to be the same as the clock frequency of the audio source based on the target data. The first node can also receive audio data from the audio source and transmit the audio data to the second node through the network, thereby enabling the second node to perform delayed playback on the received audio data based on the preset delay time corresponding to the second node, so as to achieve highly synchronized and high-precision audio synchronous playback with at least one other node, thereby improving the audio data transmission and playback performance of audio synchronous playback of multiple nodes in the network.
[0148] Optionally, the first node can be the master node in the network, and the second node can be any node other than the first node. Therefore, each of the other nodes can be configured with a corresponding preset delay time, and the preset delay time is smaller the farther away from the sound source and longer the closer to the sound source.
[0149] Optionally, the preset delay time can be pre-set for the second node. For the audio source node and each node, their positions in the network are fixed. Therefore, given a fixed audio data transmission path and direction, the delay of audio data transmission from the audio source to each node is also fixed. This delay is related to the target number of nodes through which the transmitted audio data passes between the node directly connected to the audio source and the audio playback node in the network. A larger target number results in a larger delay, and a smaller target number results in a smaller delay. Therefore, in this embodiment, different preset delay times can be set for different audio playback nodes to facilitate synchronized audio playback across multiple nodes. For example, a smaller target number for an audio playback node allows for a longer preset delay time, while a larger target number allows for a shorter preset delay time. The specific settings can be configured as needed, and there is no single limitation.
[0150] Optionally, the network can adopt any type and any network structure. For example, in some optional embodiments, the network in this disclosure can be a ring network, a daisy-chain network, or other types of networks. When it is a ring network, it can be combined with... Figure 1A , Figure 1D , Figure 9A The above is for illustrative purposes. When using a daisy-chain network, it can be combined with... Figure 1B , Figure 1C , Figure 9B , Figure 9C The following diagram illustrates this concept. Optionally, the first node can be the master node in the network (e.g., ...). Figures 1A-1D , Figures 9A-9C In each example, node 0), while other nodes besides the first node can be child nodes (e.g., Figures 1A-1D , Figures 9A-9C (Nodes 1-4 in each example).
[0151] Optionally, in the audio data transmission system, the audio source can be directly connected to the first node, or the audio source can be directly connected to the second node, as long as the clock frequency of the first node is the same as the clock frequency of the audio source.
[0152] In some alternative embodiments, the second node is further configured to: based on the target data, use Ethernet CDR (Clock and Data Recovery) technology to make the clock frequency of the second node track the clock frequency of the first node, so that the clock frequency of the second node is synchronized to be the same as the clock frequency of the audio source.
[0153] It should be understood that by using Ethernet CDR technology to track the clock frequency of the second node to the first node, the clock frequency of the second node can be easily, quickly and effectively synchronized to the same frequency as the clock frequency of the audio source. This facilitates the clock synchronization of the network nodes of the entire audio data transmission system, thereby enabling highly synchronized and precise multi-node audio data synchronous playback.
[0154] Optionally, when tracking the clock frequency using Ethernet CDR technology based on target data, subsequent nodes can be made to track the frequency of the preceding node sequentially. This ensures that the clock frequency of each node in the other nodes (including the second node) is the same as the clock frequency of the first node, thereby synchronizing each node in the other nodes (including the second node) to the same clock frequency as the audio source. Furthermore, the process of other nodes tracking the clock frequency of the first node can begin and proceed from any direction. For example, as... Figure 9A The ring network shown can start in a clockwise direction, a counterclockwise direction, or both directions simultaneously; for example, Figure 9B The single daisy chain network shown can start from a single direction of the daisy chain; for example, such as... Figure 9B The double daisy-chain network shown can be started simultaneously from both directions of the two daisy branches, or they can be started at different times. It should be understood that the goal is to ensure that the clock frequency of each node in the other nodes is the same as the clock frequency of the first node and the clock frequency of the sound source.
[0155] Optionally, after the second node synchronizes to the same clock frequency as the first node and the audio source, the first node can use the audio data time-division multiplexing method to transmit data, thereby transmitting the audio data to the second node and at least one of the other nodes through networking; the second node receives the audio data and then performs delayed playback on the received audio data based on the preset delay time corresponding to the second node, so as to achieve highly synchronized and high-precision audio synchronous playback with at least one of the other nodes.
[0156] In some alternative embodiments, the first node and the sound source can be connected to the same clock source, so that the clock frequency of the first node can be stably the same as the clock frequency of the sound source.
[0157] You can select the connection method for the first node and the audio source and clock source as needed. For example, ... Figure 9A In the example shown, the clock source could be connected to both the first node (which could be node 0) and the sound source. For example, as shown below... Figure 9B For example, the clock source can also be directly connected to the sound source, and then connected to the first node (which can be node 0) through the sound source. Another example is... Figure 9CFor example, the clock source can also be directly connected to the first node (which can be node 0), and then connected to the sound source through the first node. This is understandable. Figure 9A , Figure 9B , Figure 9C These correspond to ring-shaped, single-chrysanthemum chain-shaped, and double-chrysanthemum chain-shaped network configurations, respectively. Figure 9A , Figure 9B , Figure 9C The example uses three different clock source connection methods, which can be used in various network structures and can be understood by analogy.
[0158] In some alternative embodiments, the first node and the sound source can be connected to different clock sources, and the clock sources have the same clock frequency. This can be chosen as needed, and no single embodiment is limited in this disclosure.
[0159] In some optional embodiments, the first node caches audio data from the sound source through a buffer unit, reads the audio data of the i-th data frame from the buffer unit and writes it into the i-th data frame, and transmits the i-th data frame through a network to transmit the audio data in the i-th data frame to the second node, where i ≥ 1 and i is an integer. Further optionally, i ≥ 2 and i is an integer.
[0160] Optionally, the buffer unit here can be the first buffer unit mentioned above, which can be understood by referring to the previous text. The relevant content regarding data frames has already been explained above, which can also be understood by referring to the previous text.
[0161] Optionally, the first node is further configured to: after the first node is powered on until the first audio data is read from the buffer unit for the first time, perform phase adjustment with the audio frames including audio data of the sound source based on the first data frame to the j-th data frame, so that the data frames and audio frames have a fixed phase relationship, where 0 < j < i and j is an integer.
[0162] For example, if the initial time when the first node generates data frames and the initial time when the audio source generates audio frames are random after the first node is powered on, and the time interval between the two initial times is also random, this can easily complicate the audio noise reduction of the network. For example, when the network is applied to vehicle networking, the above-mentioned randomness can complicate the entire vehicle's ANC (Active Noise Control) or RNC (Road Noise Control).
[0163] Based on this, in the above optional embodiments, during the initialization phase from the power-on of the first node to the first reading of audio data from the buffer unit, phase adjustment can be performed on the audio frames including audio data from the sound source based on the first to the j-th data frames. This ensures that the data frames and audio frames have a fixed phase relationship, achieving initial phase alignment. This establishes a correct time starting point for stable transmission of subsequent audio data before the first node transmits audio data. It is understood that after the phase adjustment in this embodiment is completed, reading and transmitting the audio data of the i-th data frame can effectively reduce the complexity of audio noise reduction in the network and facilitate the subsequent delayed playback of the received audio data by the second node based on the preset delay time corresponding to the second node. This enables highly synchronized and high-precision audio synchronous playback with at least one other node, thereby improving the audio data transmission and playback performance of multiple nodes in the network.
[0164] Optionally, j can be selected as needed. For example, j can be 1, 2, 3, etc. Here, we can take j=2 as an example, then i can be at least 3.
[0165] Optionally, from the first data frame to the j-th data frame, the first node does not read audio data, and the data segment storing audio data in the first data frame to the j-th data frame can be empty data.
[0166] The phase adjustment method described above can be selected according to actual needs. For example, in some optional embodiments, refer to... Figure 1A , Figure 1B , Figure 1C , Figure 9A , Figure 9B , Figure 9C In the example of the network shown, the audio source can be directly connected to the first node (such as node 0). Then, the first node is further configured to: determine one or more target data frames from the first data frame to the j-th data frame, and for each target data frame, adjust the data frame interval between the target data frame and the next data frame of the target data frame based on the received audio frame to achieve phase adjustment.
[0167] Based on this, using the aforementioned optional method, when the audio source is directly connected to the first node, by adjusting the target data frame and the data frame interval of its next frame, phase adjustment can be accurately and effectively achieved between the first and j-th data frames and the audio frame. This ensures a fixed phase relationship between the data frames and the audio frame, achieving initial phase alignment. Consequently, a correct time starting point can be established for the stable transmission of subsequent audio data before the audio data is transmitted at the first node. After the phase adjustment is completed in this embodiment, the audio data of the i-th data frame is read and transmitted, thereby effectively reducing the complexity of audio noise reduction in the network.
[0168] It should be noted that the target data frame can include one or more of the first to the j-th data frames. That is, all data frames from the first to the j-th data frames can be designated as target data frames to adjust the data frame interval between the target data frame and the next data frame, thereby achieving phase adjustment; alternatively, a portion of the data frames from the first to the j-th data frames can be designated as target data frames, while another portion can be left undesignated, to adjust the data frame interval between the target data frame and the next data frame, thereby achieving phase adjustment.
[0169] Optionally, when adjusting the data frame interval between the target data frame and the next data frame, the data frame interval can be increased or decreased as needed. For example, increasing the interval can be achieved by adding a preset number of clock cycles to a default data frame interval. Conversely, decreasing the interval can be achieved by subtracting a preset number of clock cycles from a default data frame interval.
[0170] Alternatively, it can be combined with the preceding text. Figure 5 The above optional embodiments are explained in the description, and will not be repeated here.
[0171] In some alternative embodiments, refer to Figure 1D In the example of the network shown, the audio source can be directly connected to a target node (such as node 1) among other nodes besides the first node. It should be understood that the target node here can be the second node or not. Then: the first node is further configured to: determine one or more target data frames from the first data frame to the j-th data frame; for each target data frame, send the target data frame to the target node through the network so that the target node can determine the phase offset based on the target data frame and the audio frame; receive the phase offset transmitted by the target node; and perform phase adjustment based on the phase offset.
[0172] Based on this, using the aforementioned optional method, even when the audio source is not directly connected to the first node but directly connected to the target node among other nodes, the phase offset can be determined at the target node based on the target data frame and the audio frame. This allows the first node to accurately and effectively adjust the phase with the audio frame based on the phase offset received from the target node, from the first data frame to the j-th data frame. This ensures a fixed phase relationship between the data frame and the audio frame, achieving initial phase alignment. Consequently, it can accurately and conveniently establish the correct time starting point for stable transmission of subsequent audio data before the first node transmits audio data. After the phase adjustment is completed in this embodiment, the audio data of the i-th data frame is read and transmitted, effectively reducing the complexity of audio noise reduction in the network.
[0173] It should be noted that the target data frame can include one or more of the first to the j-th data frames. That is, all data frames from the first to the j-th data frames can be identified as target data frames, and the target data frames can be transmitted to the target node to determine the phase offset, so as to achieve phase adjustment; or a portion of the data frames from the first to the j-th data frames can be identified as target data frames, while the other portion of data frames can be deferred as target data frames, and the target data frames can be transmitted to the target node to determine the phase offset, so as to achieve phase adjustment.
[0174] Optionally, the target node can receive the phase offset transmitted by the network and perform phase adjustment based on the phase offset. Alternatively, the target node can also transmit the phase offset in other ways, such as, but not limited to, wireless transmission outside the network.
[0175] In some optional embodiments, the first node is further configured to: receive a data frame containing audio data returned by the target node through the network, and obtain the phase offset from the common header of the data frame.
[0176] Alternatively, you can refer to the previous section about Figure 3 The common header of the data frame shown is used to understand its format. Optionally, after the first node transmits the target data frame to the target node via the network, the target node can store the audio data from the audio frame obtained from the audio source into the target data frame, store the determined phase offset in the common header of the target data frame, and then return the target data frame to the first node via the network. This allows the phase offset to be fed back through the common header of the data frame. The first node can extract the phase offset from the common header of the returned target data frame to perform phase adjustment based on the phase offset.
[0177] It should be understood that by storing the phase offset in the common header of the data frames (including audio data) returned by the target node through the network, the first node can more easily obtain the phase offset. This allows for accurate and efficient phase adjustment with the audio frame based on the first to j-th data frames, thus ensuring a fixed phase relationship between the data frames and the audio frame. Furthermore, since the target node directly connected to the audio source can return data frames to the first node to transmit audio data while simultaneously storing the phase offset in the common header of those data frames, this effectively improves data transmission efficiency and reduces data transmission overhead.
[0178] It is understood that the above description of the audio data transmission system in the embodiments of this disclosure is only some optional embodiments of this disclosure and is not a limitation on the embodiments of this disclosure.
[0179] According to a fourth aspect of the present disclosure, a chip is provided, comprising: a processor and a memory, wherein the processor and the memory communicate with each other; the memory is used to store at least one executable instruction, which causes the processor to perform an operation corresponding to the method as described in any one of the first aspects.
[0180] Figure 10 This is a schematic block diagram of a chip provided in an embodiment of this disclosure. Specific embodiments of this disclosure do not limit the specific implementation of the chip. Figure 10 As shown, the chip 1000 may include a processor 1002 and a memory 1006. Wherein: The processor 1002 and the memory 1006 communicate with each other.
[0181] The processor 1002 is used to execute program 1010, which can specifically execute the relevant steps in any of the aforementioned audio data transmission method embodiments.
[0182] Specifically, program 1010 may include program code that includes computer operation instructions.
[0183] The processor 1002 may be a CPU, an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present disclosure. The smart device may include one or more processors of the same type, such as one or more CPUs; or it may include processors of different types, such as one or more CPUs and one or more ASICs.
[0184] RISC-V is an open-source instruction set architecture based on the Reduced Instruction Set Computing (RISC) principle. It can be applied to various aspects of microcontrollers and FPGA chips, specifically in areas such as IoT security, industrial control, mobile phones, and personal computers. Because its design considers small size, speed, and low power consumption, it is particularly suitable for modern computing devices such as warehouse-scale cloud computers, high-end mobile phones, and tiny embedded systems. With the rise of AIoT (Artificial Intelligence of Things), the RISC-V instruction set architecture is receiving increasing attention and support and is expected to become the next generation of widely used CPU architecture.
[0185] The computer operation instructions in this embodiment can be computer operation instructions based on the RISC-V instruction set architecture. Correspondingly, the processor 1002 can be designed based on the RISC-V instruction set. Specifically, the chip provided in this embodiment can be a chip designed using the RISC-V instruction set. This chip can execute executable code based on the configured instructions, thereby implementing the audio data transmission method in the above embodiments.
[0186] Memory 1006 is used to store program 1010. Memory 1006 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0187] Specifically, program 1010 can be used to cause processor 1002 to execute the audio data transmission method in any of the foregoing embodiments.
[0188] The specific implementation of each step in program 1010 can be found in the corresponding steps and units described in any of the aforementioned audio data transmission method embodiments, and will not be repeated here. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the devices and modules described above can be referred to the corresponding process descriptions in the aforementioned method embodiments, and will not be repeated here.
[0189] According to a fifth aspect of the present disclosure, the present disclosure also provides a network comprising: a plurality of connected nodes, at least one of the plurality of nodes including the chip 1000 as described above.
[0190] This disclosure does not limit the network structure; any structure that meets the requirements is acceptable. For example, it can be... Figure 1A , Figure 1D The ring network in the diagram, or it can also be... Figure 1B or Figure 1C The chrysanthemum chain network in the middle ( Figure 1B It is a single chrysanthemum chain network. Figure 1C(This is a double chrysanthemum chain network; please refer to the previous text for details.)
[0191] According to a sixth aspect of the present disclosure, the present disclosure also provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the audio data transmission method as described in any of the foregoing embodiments.
[0192] For example, the computer storage media includes, but is not limited to: Compact Disc Read-Only Memory (CD-ROM), Random Access Memory (RAM), floppy disk, hard disk, or magneto-optical disk.
[0193] According to a seventh aspect of the present disclosure, the present disclosure also provides a computer program product including a computer program that, when executed by a processor, implements the audio data transmission method as described in any of the foregoing embodiments.
[0194] The chip 1000, network, computer storage medium, and computer program product embodiments in this disclosure have been described in detail in the foregoing audio data transmission method embodiments. Therefore, their related content and beneficial effects can be understood by referring to the above embodiments, and will not be repeated here.
[0195] Furthermore, it should be noted that the user-related information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to sample data used for training the model, data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this disclosure are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0196] It should be noted that, depending on the implementation needs, the various components / steps described in the embodiments of this disclosure can be broken down into more components / steps, or two or more components / steps or parts of the operations of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of this disclosure. It should be understood that the various technical features in the technical solutions of the embodiments of this disclosure can be combined and / or broken down in any suitable manner.
[0197] The methods described above according to embodiments of this disclosure can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code originally stored on a remote recording medium or a non-transitory machine-readable medium and to be stored on a local recording medium, downloaded over a network. Thus, the methods described herein can be stored on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an Application Specific Integrated Circuit (ASIC) or a Field Programmable Gate Array (FPGA)). It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components (e.g., Random Access Memory (RAM), Read-Only Memory (ROM), Flash Memory, etc.) capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods described herein. Furthermore, when a general-purpose computer accesses code used to implement the methods shown herein, the execution of the code transforms the general-purpose computer into a dedicated computer for performing the methods shown herein.
[0198] Those skilled in the art will recognize that the units and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for specific applications, but such implementations should not be considered beyond the scope of the embodiments disclosed herein.
[0199] The above embodiments are only used to illustrate the embodiments of this disclosure, and are not intended to limit the embodiments of this disclosure. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the embodiments of this disclosure. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of this disclosure, and the patent protection scope of the embodiments of this disclosure should be defined by the claims.
[0200] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". It should be noted that the concepts of "first", "second", etc., mentioned in the embodiments of this disclosure are only used to distinguish different devices, modules, or units, and are not used to limit the order of functions performed by these devices, modules, or units or their interdependencies. It should be noted that the modifications of "a" and "a plurality" mentioned in the embodiments of this disclosure are illustrative and not restrictive, and those skilled in the art should understand that unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0201] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of this disclosure, and are not intended to limit them. Although the embodiments of this disclosure have been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure.
Claims
1. An audio data transmission method for a first node in a network of multiple nodes, the multiple nodes further including other nodes besides the first node, wherein the clock frequency of the first node is different from the clock frequency of the audio source, the method comprising: Audio data from the sound source is cached through the first cache unit; After reading the audio data of the i-th data frame from the first cache unit, the data frame interval between the i-th data frame and the (i+1)-th data frame is determined based on the first data volume of the audio data cached in the first cache unit, where i is a positive integer; After the data frame interval ends, the audio data of the (i+1)th data frame is read from the first buffer unit and transmitted to the corresponding node among the other nodes through the network.
2. The method according to claim 1, wherein, Determining the data frame interval between the i-th data frame and the (i+1)-th data frame based on the first data volume of audio data cached in the first cache unit includes: Based on the first data volume of audio data cached in the first cache unit and the first preset threshold information, the data frame interval between the i-th data frame and the (i+1)-th data frame is determined.
3. The method according to claim 2, wherein, The first preset threshold information includes a first threshold value and a second threshold value, wherein the first threshold value is greater than the second threshold value; The step of determining the data frame interval between the i-th data frame and the (i+1)-th data frame based on the first data volume of audio data cached in the first cache unit and the first preset threshold information includes: If the first data volume is greater than the first threshold value, then a reduction adjustment is performed based on the default frame interval to determine the data frame interval between the i-th data frame and the (i+1)-th data frame; and / or, If the first data volume is less than the second threshold value, then an increase adjustment is made based on the default frame interval to determine the data frame interval between the i-th data frame and the (i+1)-th data frame; and / or, If the first data volume is less than or equal to the first threshold value and greater than or equal to the second threshold value, then the default frame interval is determined as the data frame interval between the i-th data frame and the (i+1)-th data frame.
4. The method according to claim 3, wherein, The default frame interval includes one or more clock cycles of the first node; wherein... The reduction adjustment based on the default frame interval includes: subtracting a first preset number of clock cycles from the default frame interval; or, The adjustment based on the default frame interval includes: adding a second preset number of clock cycles to the default frame interval.
5. The method according to claim 2, wherein, The data volume unit of any threshold value in the first preset threshold information is bits, and / or, the first buffer unit buffers audio data in units of bits.
6. The method according to claim 1, wherein, The step of transmitting the audio data of the (i+1)th data frame to the corresponding node among the other nodes through the network includes: The audio data of the (i+1)th data frame is transmitted to the second node among the other nodes through the network, so that the second node can synchronize with at least one of the other nodes to play the audio data at a predetermined time.
7. The method according to any one of claims 1-6, wherein, i≥2; the method further includes: After the first node is powered on until the first audio data is read from the first buffer unit for the first time, phase adjustment is performed on the audio frames including audio data of the sound source based on the first data frame to the j-th data frame, so that the data frames and audio frames have a fixed phase relationship, where 0 < j < i and j is an integer.
8. The method according to claim 7, wherein, If the sound source is directly connected to the first node, then: the step of performing phase adjustment with the audio frames including audio data of the sound source based on the first data frame to the j-th data frame includes: determining one or more target data frames from the first data frame to the j-th data frame, and for each target data frame, adjusting the data frame interval between the target data frame and the next data frame of the target data frame based on the received audio frame to achieve the phase adjustment; or, If the audio source is directly connected to a target node among the other nodes, then: the step of performing phase adjustment with the audio frames including audio data of the audio source based on the first data frame to the j-th data frame includes: determining one or more target data frames from the first data frame to the j-th data frame; for each target data frame, sending the target data frame to the target node through the network, so that the target node determines a phase offset based on the target data frame and the audio frame; receiving the phase offset transmitted by the target node; and performing the phase adjustment based on the phase offset.
9. The method according to claim 8, wherein, The receipt of the phase offset transmitted by the target node includes: Receive the data frame containing audio data returned by the target node through the network, and obtain the phase offset from the common header of the data frame.
10. The method according to any one of claims 1-6, wherein, The network can be a ring network or a daisy-chain network.
11. An audio data transmission method for an audio playback node in a network of multiple nodes, the method comprising: Audio data from the sound source is cached through the second buffer unit; Based on the second data volume of the audio data cached in the second cache unit, the data reading frequency of the audio data is determined; Based on the data reading frequency, audio data is read from the second buffer unit to synchronously play audio data with at least one other node among the plurality of nodes at a predetermined time.
12. The method according to claim 11, wherein, The audio playback node is either a first node or a second node, wherein the clock frequency of the first node is different from the clock frequency of the audio source, and the second node is any node among the plurality of nodes other than the first node.
13. The method according to claim 11, wherein, The determination of the audio data reading frequency based on the second data volume of the audio data cached in the second cache unit includes: Based on the second data volume of the audio data cached in the second cache unit and the second preset threshold information, the data reading frequency of the audio data is determined.
14. The method according to claim 13, wherein, The second preset threshold information includes a third threshold and a fourth threshold, wherein the third threshold is greater than the fourth threshold; The step of determining the audio data reading frequency based on the second data volume of the audio data cached in the second cache unit and the second preset threshold information includes: If the second data volume is greater than the third threshold, then an increase is made based on the default reading frequency to determine the audio data reading frequency; and / or, If the second data volume is less than the fourth threshold value, then the data reading frequency of the audio data is determined by reducing and adjusting based on the default reading frequency.
15. The method according to claim 14, wherein, The adjustment based on the default read frequency includes: adjusting the default read frequency to a first preset read frequency that is greater than the default read frequency; or... The reduction adjustment based on the default reading frequency includes: adjusting the default reading frequency to a second preset reading frequency that is lower than the default reading frequency.
16. The method of claim 14, wherein, The third and fourth threshold values corresponding to the audio playback node are both determined based on the target number of nodes through which audio data is transmitted between the node directly connected to the audio source and the audio playback node in the network. Wherein, the smaller the target quantity, the larger the third threshold value and the larger the fourth threshold value; conversely, the larger the target quantity, the smaller the third threshold value and the smaller the fourth threshold value.
17. The method according to claim 16, wherein, The third threshold value A corresponding to the audio playback node satisfies the following relationship: A = C + D + E; The fourth threshold value B corresponding to the audio playback node satisfies the following relationship: B = C + D - E; Where C is half of the total data capacity of the second cache unit, D is a preset offset that is negatively correlated with the target quantity, and E is a preset threshold offset.
18. The method according to claim 13, wherein, The data volume unit of any threshold value in the second preset threshold information is bits, and / or, the second buffer unit buffers audio data in units of bits.
19. The method according to any one of claims 11-18, wherein, The step of reading audio data from the second buffer unit based on the data reading frequency, so as to synchronously perform audio data playback with at least one other node among the plurality of nodes at a predetermined time, includes: Based on the data reading frequency, audio data is read from a predetermined storage location in the second cache unit to synchronously play audio data with at least one other node among the plurality of nodes at a predetermined time. The predetermined storage location is half of the total data capacity of the second cache unit plus the storage location indicated by a preset offset. The preset offset is negatively correlated with a target number, which is the number of nodes through which audio data is transmitted between the node directly connected to the audio source and the audio playback node in the network. or, Based on the data reading frequency, audio data is read from the second buffer unit, and based on the preset delay time corresponding to the audio playback node, the read audio data is played with a delay so that the audio data is played synchronously with at least one other node among the plurality of nodes at a predetermined time.
20. The method according to any one of claims 11-18, wherein, The network can be a ring network or a daisy-chain network.
21. An audio data transmission system, comprising: Multiple nodes connected to form a network, including a first node and a second node, wherein the second node is any node other than the first node, and the clock frequency of the first node is the same as the clock frequency of the sound source. The first node is configured to send target data to the second node through the network, so that the second node synchronizes its clock frequency to the same frequency as the sound source based on the target data; The system receives audio data from the sound source and transmits the audio data to the second node through the network. The second node is configured to: perform delayed playback on the received audio data based on a preset delay time corresponding to the second node, so as to achieve audio synchronization with at least one of the other nodes.
22. The system according to claim 21, wherein, The audio data transmission system satisfies at least one of the following conditions: The second node is further configured to: based on the target data, use Ethernet CDR technology to make the clock frequency of the second node track the clock frequency of the first node, so that the clock frequency of the second node is synchronized to be the same as the clock frequency of the sound source; The first node and the sound source are connected to the same clock source; The network can be a ring network or a daisy-chain network.
23. The system according to claim 21 or 22, wherein, The first node caches audio data from the sound source through a cache unit, reads the audio data of the i-th data frame from the cache unit and writes it into the i-th data frame, and transmits the i-th data frame through the network to transmit the audio data in the i-th data frame to the second node through the network, where i≥2 and i is an integer; The first node is further configured to: after the first node is powered on until the first audio data is read from the cache unit for the first time, perform phase adjustment with the audio frames including audio data of the sound source based on the first data frame to the j-th data frame, so that the data frames and audio frames have a fixed phase relationship, where 0 < j < i and j is an integer.
24. The system according to claim 23, wherein, If the audio source is directly connected to the first node, then the first node is further configured to: determine one or more target data frames from the first data frame to the j-th data frame, and for each target data frame, adjust the data frame interval between the target data frame and the next data frame of the target data frame based on the received audio frame, so as to achieve the phase adjustment; or, If the audio source is directly connected to the target node among the other nodes, then: the first node is further configured to: determine one or more target data frames from the first data frame to the j-th data frame; for each target data frame, send the target data frame to the target node through the network, so that the target node determines the phase offset based on the target data frame and the audio frame; receive the phase offset transmitted by the target node; and perform the phase adjustment based on the phase offset.
25. The system according to claim 24, wherein, The first node is further configured to: Receive the data frame containing audio data returned by the target node through the network, and obtain the phase offset from the common header of the data frame.
26. A chip, comprising: A processor and a memory, wherein the processor and the memory communicate with each other; The memory is used to store at least one executable instruction, which causes the processor to perform the operation corresponding to the method as described in any one of claims 1-20.
27. A network configuration, comprising: A plurality of connected nodes, at least one of which includes the chip as described in claim 26.
28. A computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as described in any one of claims 1-20.
29. A computer program product comprising a computer program that, when executed by a processor, implements the method as described in any one of claims 1-20.