Multi-room audio hierarchical encoding transmission method based on scene recognition and related device
By acquiring the decoding capabilities of audio devices and network quality parameters, the encoding format is dynamically determined and audio transmission packets are encapsulated. This solves the problem that fixed encoding formats in multi-room audio systems cannot simultaneously achieve low latency, stability, and sound quality, resulting in a better synchronization and sound quality experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LINKPLAY TECHNOLOGY INC NANJING
- Filing Date
- 2026-06-24
- Publication Date
- 2026-07-21
Smart Images

Figure CN122435935A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio transmission technology, and in particular to a multi-room audio hierarchical coding transmission method and related equipment based on scene recognition. Background Technology
[0002] With the rapid popularization of wireless smart audio devices and the continuous upgrading of home network environments, multi-room audio playback has become an important functional requirement for smart speakers, wireless amplifiers, streaming media players, soundbars, and home theater systems. Multi-room audio playback requires multiple audio devices within the same playback group to simultaneously receive the same audio stream in a local area network environment and complete decoding and output at similar times, thereby achieving synchronized listening throughout the house. To reduce system complexity, existing technologies typically use a unified, fixed encoding format for audio stream transmission across all devices within the playback group. Common solutions include using MP3, AAC, or uncompressed PCM formats, which have been widely adopted commercially under conditions of a small number of devices, a stable network environment, and a simple playback scenario.
[0003] However, in real-world multi-room audio systems, the performance requirements for different playback scenarios vary significantly: scenarios such as TV audio return, game audio, and real-time input are extremely sensitive to end-to-end latency; ordinary online music playback scenarios focus more on transmission stability and bandwidth utilization efficiency when multiple devices are connected concurrently; and local high-resolution audio source playback scenarios place stringent requirements on bitrate, dynamic range, and audio fidelity. Fixed encoding schemes have fundamental limitations under these diverse scenarios: if high bitrate or lossless encoding is uniformly used, under conditions of limited network bandwidth, simultaneous access by multiple devices, or signal attenuation across rooms, it can easily lead to insufficient buffering, playback stuttering, and synchronization misalignment between devices; if low bitrate lossy encoding is uniformly used, it cannot meet users' demands for high-quality audio reproduction. More importantly, when the network conditions, decoding capabilities, and buffer capacities of devices within the playback group differ, fixed encoding schemes cannot differentiate between devices, nor do they have the ability to dynamically adjust the encoding strategy according to changes in the playback scenario. Therefore, how to dynamically match the encoding method according to the playback scenario, device capabilities and network status in a multi-room audio system, and effectively balance low latency, transmission stability and sound quality fidelity, has become an urgent technical problem to be solved. Summary of the Invention
[0004] The main objective of this invention is to solve the problem that existing multi-room audio systems use a fixed encoding format for transmission, which cannot dynamically match the encoding method according to the playback scenario, resulting in a difficulty in simultaneously achieving low latency, transmission stability, and sound quality fidelity.
[0005] The first aspect of this invention provides a multi-room audio hierarchical coding and transmission method based on scene recognition, the multi-room audio hierarchical coding and transmission method based on scene recognition includes: Obtain the audio stream to be transmitted and the audio decoding capability parameters and network transmission quality parameters of each audio device in the multi-room playback group, and take the intersection of each audio decoding capability parameter to obtain a common decoding format set; Based on the playback task attribute parameters of the audio stream to be transmitted, the current playback task of the audio stream to be transmitted is classified into scenarios to obtain scenario type parameters; Based on the scenario type parameter, the common decoding format set, and the network transmission quality parameter, a target encoding format is determined from the preset encoding format level. When the network transmission quality parameter is lower than the transmission quality threshold corresponding to the target encoding format, the target encoding format is updated to an encoding format in the preset encoding format level whose transmission quality threshold is lower than the network transmission quality parameter. Based on the target encoding format, the audio stream to be transmitted is encoded and encapsulated into an audio transmission packet. The header of the audio transmission packet includes an encoding format identifier field and a playback timestamp field. The audio transmission packet is distributed to each audio device. Each audio device decodes the audio frame data in the audio transmission packet based on the encoding format identifier field to obtain PCM audio data. Based on the playback timestamp field, the local clock of each audio device is aligned and the PCM audio data is output synchronously.
[0006] A second aspect of the present invention provides a multi-room audio hierarchical coding and transmission apparatus based on scene recognition, the multi-room audio hierarchical coding and transmission apparatus based on scene recognition comprising: The parameter acquisition module is used to acquire the audio stream to be transmitted and the audio decoding capability parameters and network transmission quality parameters of each audio device in the multi-room playback group, and to take the intersection of each audio decoding capability parameter to obtain a common decoding format set; The scene classification module is used to classify the current playback task of the audio stream to be transmitted according to the playback task attribute parameters of the audio stream to be transmitted, and obtain the scene type parameter. The encoding decision module is used to determine a target encoding format from a preset encoding format hierarchy based on the scenario type parameter, the common decoding format set, and the network transmission quality parameter. When the network transmission quality parameter is lower than the transmission quality threshold corresponding to the target encoding format, the target encoding format is updated to an encoding format in the preset encoding format hierarchy whose transmission quality threshold is lower than the network transmission quality parameter. The encoding and encapsulation module is used to encode the audio stream to be transmitted and encapsulate it into an audio transmission packet based on the target encoding format. The header of the audio transmission packet includes an encoding format identifier field and a playback timestamp field. The distribution synchronization module is used to distribute the audio transmission packet to each audio device. Each audio device decodes the audio frame data in the audio transmission packet based on the encoding format identifier field to obtain PCM audio data, and aligns the local clock of each audio device with the playback timestamp field to synchronously output the PCM audio data.
[0007] A third aspect of the present invention provides a scene-recognition-based multi-room audio hierarchical coding and transmission device, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor invokes the instructions in the memory to cause the scene-recognition-based multi-room audio hierarchical coding and transmission device to perform the various steps of the scene-recognition-based multi-room audio hierarchical coding and transmission method described above.
[0008] A fourth aspect of the present invention provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the steps of the above-described multi-room audio hierarchical coding and transmission method based on scene recognition.
[0009] The above-described multi-room audio hierarchical encoding and transmission method and related equipment based on scene recognition are described above. In this embodiment of the invention, the audio stream to be transmitted and the audio decoding capability parameters and network transmission quality parameters of each audio device in the multi-room playback group are obtained. The intersection of each audio decoding capability parameter is taken to obtain a common decoding format set, and the current playback task is classified into scene types according to the playback task attribute parameters to obtain scene type parameters. Based on the scene type parameters, the common decoding format set, and the network transmission quality parameters, a target encoding format is determined from a preset encoding format hierarchy. When the network transmission quality parameter is lower than the transmission quality threshold corresponding to the target encoding format, the target encoding format is updated to an encoding format whose transmission quality threshold is lower than the current network transmission quality parameter. Based on the target encoding format, the audio stream to be transmitted is encoded and encapsulated into an audio transmission packet with a header containing an encoding format identifier field and a playback timestamp field. The audio transmission packet is distributed to each audio device. Each audio device decodes the audio frame data based on the encoding format identifier field to obtain PCM audio data, and synchronously outputs PCM audio data after aligning with the local clock based on the playback timestamp field. This application establishes a hierarchical encoding decision-making mechanism constrained by a common decoding format set for devices and driven by playback scenarios. It replaces the fixed encoding transmission architecture with scene recognition, format hierarchy mapping, network adaptive degradation, and a unified timestamp synchronization mechanism. This solves the technical problems of existing multi-room audio systems using fixed encoding formats, which result in a single encoding strategy, inability to adapt to the different needs of multiple scenarios, and playback stuttering and synchronization mismatch between devices caused by high bitrate encoding under weak network conditions. It achieves a dynamic balance between low latency, transmission stability, and sound quality fidelity within multi-room playback groups, effectively improving synchronization consistency and user listening experience in multi-device concurrent playback scenarios.
[0010] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.
[0011] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0012] Figure 1 This is a schematic diagram of the first embodiment of the multi-room audio hierarchical coding and transmission method based on scene recognition in this invention. Figure 2 This is a schematic diagram of an embodiment of the multi-room audio hierarchical coding and transmission device based on scene recognition in this invention. Figure 3This is a schematic diagram of an embodiment of a multi-room audio hierarchical coding and transmission device based on scene recognition in this invention. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0014] The terms "comprising" and "having," and any variations thereof, used in the embodiments of this invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the steps or units listed, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.
[0015] To facilitate understanding of this embodiment, the specific process of this embodiment is described below. Please refer to [link / reference]. Figure 1 The first embodiment of the multi-room audio hierarchical coding and transmission method based on scene recognition in this invention includes: 101. Obtain the audio stream to be transmitted and the audio decoding capability parameters and network transmission quality parameters of each audio device in the multi-room playback group, and take the intersection of each audio decoding capability parameter to obtain a common decoding format set; In this embodiment, the step of obtaining the audio stream to be transmitted and the audio decoding capability parameters and network transmission quality parameters of each audio device in the multi-room playback group, and taking the intersection of each audio decoding capability parameter to obtain a common decoding format set includes: obtaining the audio stream to be transmitted, sending a capability query request to each audio device in the multi-room playback group, and receiving the audio decoding capability parameters returned by each audio device; collecting various network status data of each audio device in the multi-room playback group to obtain the network transmission quality parameters of each audio device; and performing an intersection operation on the encoding format field in the audio decoding capability parameters returned by each audio device to obtain a common decoding format set. After distributing the audio transmission packet to each audio device, the method further includes: resending capability query requests to each audio device in the playback group according to a preset acquisition cycle; receiving the current audio decoding capability parameters returned by each audio device; storing the current audio decoding capability parameters as the audio decoding capability parameters for the current acquisition cycle; comparing the audio decoding capability parameters for the current acquisition cycle with the audio decoding capability parameters for the previous acquisition cycle; if the audio decoding capability parameters for the current acquisition cycle are inconsistent with the audio decoding capability parameters for the previous acquisition cycle, performing a new intersection operation on the audio decoding capability parameters of each audio device to obtain an updated common decoding format set; verifying whether the target encoding format is included in the updated common decoding format set; if the target encoding format is not included in the updated common decoding format set, re-determining the target encoding format from the updated common decoding format set.
[0016] In practical applications, before a multi-room playback group begins transmitting audio, the master device sends capability query request messages to each audio device in the playback group via unicast over the local area network. Upon receiving the request, each audio device returns the device capability information pre-stored in its firmware to the master device in the form of a response message. Audio decoding capability parameters refer to a collection of descriptions of the audio decoding processing capabilities possessed by the audio device at the hardware and firmware levels. Specifically, this includes the encoding format field supported by the device (i.e., the audio compression format types that the device can recognize and reproduce, such as LC3+, AAC, FLAC), the maximum decodeable sampling rate, the maximum decodeable bit depth, and the maximum decodeable bit rate. Taking a playback group consisting of a living room amplifier, bedroom speakers, and a study streaming media player as an example, the living room amplifier returns support for LC3+, AAC, and FLAC, with a maximum sampling rate of 192kHz / 24bit; the bedroom speakers return support for AAC and FLAC, with a maximum sampling rate of 96kHz / 24bit; and the study player returns support for LC3+, AAC, and FLAC, with a maximum sampling rate of 192kHz / 24bit. The master device writes the audio decoding capability parameters returned by each device into a local device capability table. Simultaneously, the master device continuously collects network status data from each audio device through the LAN management module. Specific collected metrics include packet loss rate, network round-trip latency, link jitter (i.e., the fluctuation range of network latency), link throughput, and Wi-Fi signal strength (RSSI). The master device weights and sums these metrics according to preset weights to obtain the current network transmission quality parameters for each audio device. Specifically, packet loss rate and network round-trip latency have the most direct impact on the stability of multi-room audio playback. Increased packet loss rate results in audio frame data not reaching the receiving device completely, leading to frequent compensation padding. Increased network round-trip latency reduces the clock alignment accuracy between the master device and each slave device; therefore, these two metrics are given higher weights. Link jitter, link throughput, Wi-Fi signal strength (RSSI), retransmission count, and local playback buffer change trends are used as auxiliary metrics and are given relatively lower weights. The weights of each metric are preset during the system firmware configuration phase, and the weighted sum yields a comprehensive network transmission quality parameter value reflecting the current link quality. For example, when the RSSI of the bedroom speaker drops from -55dBm to -75dBm and the packet loss rate increases significantly, the overall calculated result of its network transmission quality parameters will be significantly lower than that of the living room amplifier, thus triggering differentiated processing in the subsequent encoding format decision stage. The master device then extracts the encoding format field from the audio decoding capability parameters returned by each device, performs an intersection operation on the encoding format field sets of all devices, and retains only the encoding formats that appear in all device encoding format fields, obtaining a common decoding format set. Taking the three devices mentioned above as an example, the bedroom speaker does not support LC3+, therefore LC3+ is not included in the common decoding format set; all three devices support AAC and FLAC, so the common decoding format set is {AAC, FLAC}.The common decoding format set limits the candidate range for subsequent encoding format decisions to formats that all devices in the playback group have decoding capabilities, thus preventing the master device from choosing an encoding format that a slave device cannot decode, which would prevent that device from participating in playback.
[0017] During the continuous distribution of audio transmission packets to various audio devices, capability query requests are resent to each audio device in the playback group according to a preset acquisition cycle. The current audio decoding capability parameters returned by each device are received and written into the local device capability table, stored as the audio decoding capability parameters for this acquisition cycle. The preset acquisition cycle refers to the time interval for periodically polling and detecting device capabilities; this interval is preset during the configuration phase. Device firmware upgrades that add support for a certain encoding format, or temporary reductions in available decoding capabilities due to excessive DSP (Digital Signal Processor, a dedicated processing unit responsible for audio encoding and decoding operations), will cause changes in device capabilities during playback. By comparing the audio decoding capability parameters of the current acquisition cycle with the audio decoding capability parameters stored in the previous acquisition cycle field by field, if any device's encoding format field, maximum decodeable sampling rate, or maximum decodeable bit depth changes, the latest audio decoding capability parameters of all devices in the playback group are re-executed with an intersection operation of the encoding format fields to obtain an updated common decoding format set. The system then checks if the current target encoding format exists in the updated public decoding format set. If it does, it remains unchanged; otherwise, it selects the highest-level encoding format from the updated public decoding format set that matches the current scene type parameter as the new target encoding format. Taking the aforementioned playback group as an example, if the study room player firmware rollback causes FLAC to be removed from its encoding format field, the updated public decoding format set becomes {AAC}. The current target encoding format FLAC does not exist in this set, so the target encoding format is immediately updated to AAC, ensuring that all devices in the playback group can decode and play normally. This dynamic update mechanism enables the system to perceive and respond to changes in the capabilities of playback group members in real time, preventing some devices from continuously failing to decode and leaving the playback group due to changes in device capabilities, thus ensuring the continuity of multi-room playback.
[0018] 102. Based on the playback task attribute parameters of the audio stream to be transmitted, classify the current playback task of the audio stream to be transmitted into scenarios to obtain scenario type parameters; In this embodiment, the scene type parameters include low-latency scene type parameters, lossless playback scene type parameters, and normal multi-room playback scene type parameters. The step of classifying the current playback task of the audio stream to be transmitted into scene type parameters based on the playback task attribute parameters of the audio stream to be transmitted includes: determining the audio input source type of the audio stream to be transmitted; when the audio input source type belongs to a preset low-latency trigger condition, the scene type parameter of the current playback task is determined as a low-latency scene type parameter; determining the audio source quality parameter of the audio stream to be transmitted; when the audio source quality parameter belongs to a preset lossless trigger condition, the scene type parameter of the current playback task is determined as a lossless playback scene type parameter; when the audio input source type does not belong to a preset low-latency trigger condition and the audio source quality parameter does not belong to a preset lossless trigger condition, the scene type parameter of the current playback task is determined as a normal multi-room playback scene type parameter. After classifying the current playback task of the audio stream to be transmitted into scenarios and obtaining scenario type parameters, the method further includes: obtaining the user's playback mode setting parameters; when the playback mode setting parameters are low-latency mode, updating the scenario type parameters to low-latency scenario type parameters; when the playback mode setting parameters are lossless priority mode, updating the scenario type parameters to lossless playback scenario type parameters; when the user's playback mode setting parameters are neither low-latency mode nor lossless priority mode, determining the scenario type parameters of the current playback task based on the audio input source type and the audio source quality parameters.
[0019] In practical applications, after acquiring the audio stream to be transmitted, the audio input source type of the current audio stream is determined by reading the input interface identifier or protocol type identifier of the audio stream. The audio input source type refers to the physical interface or protocol channel type through which the current audio stream enters the main device, such as HDMI ARC / eARC interface, optical SPDIF interface, Bluetooth input, AirPlay protocol, DLNA protocol, or local file playback queue. Pre-configured low-latency trigger conditions are stored in a list of latency-sensitive input source types, specifically including HDMI ARC / eARC input, game audio input, and real-time voice interaction input. The currently read audio input source type is matched against each item in this list. If it matches any item in the list, a successful match is determined, and the scene type parameter of the current playback task is identified as the low-latency scene type parameter. For example, if a user connects a TV to the main audio device in the living room via an HDMI eARC interface and forms a multi-room playback group with living room and kitchen speakers, the audio input source type identifier is read as HDMI eARC, which matches the pre-configured low-latency trigger condition list, and the scene type parameter is determined as the low-latency scene type parameter. Audio input from interfaces such as HDMI eARC has a strict time correspondence with video. If the end-to-end audio latency is too high, users will intuitively perceive that the lip movements and sound are out of sync. By directly mapping such input sources to low-latency scene type parameters, we can ensure that subsequent encoding format decisions prioritize low-latency encoding formats and fundamentally eliminate the problem of audio-visual asynchrony.
[0020] When the audio input source type does not match any item in the preset low-latency trigger condition list, the audio source quality parameters of the audio stream to be transmitted are further extracted. Audio source quality parameters refer to a combined description of the original encoding format type, sampling rate, bit depth, and bitrate. For local file playback scenarios, these parameters are obtained by directly reading the audio file header. For example, the FLAC file header contains fields for sampling rate, bit depth, and encoding format type; parsing each field yields the complete audio source quality parameters. For network streaming media playback scenarios, the audio source quality parameters are obtained through the protocol description information transmitted when establishing a playback session using streaming media protocols (such as AirPlay and DLNA). This description information contains fields for audio encoding format type, sampling rate, and bit depth; parsing this description information yields the audio source quality parameters. Pre-configured lossless trigger conditions are configured at the firmware level. These conditions stipulate that the following requirements must be met simultaneously to be considered a lossless audio source: the value of the original encoding format type field belongs to the lossless compression format identifier set (containing identifier strings such as FLAC, WAV, and ALAC); the value of the sampling rate field is not less than 44100Hz; and the value of the bit depth field is not less than 16 bits. The extracted original encoding format type, sampling rate, and bit depth are sequentially compared with the preset lossless trigger conditions. If all three conditions are met, the scene type parameter of the current playback task is determined as the lossless playback scene type parameter. For example, if a user plays a FLAC music file with a sampling rate of 96kHz and a bit depth of 24bit from a local NAS, after parsing the file header, the original encoding format type is extracted as FLAC, the sampling rate is 96000Hz, and the bit depth is 24bit. All three conditions meet the preset lossless trigger conditions, and the scene type parameter is determined as the lossless playback scene type parameter. Conversely, when a 320kbps MP3 audio source is detected, the original encoding format type field does not belong to the lossless compression format identifier set, does not meet the preset lossless trigger conditions, and is not marked as a lossless playback scene, thus avoiding meaningless high-bandwidth transmission of low-quality audio sources. When the audio input source type does not meet the preset low-latency triggering condition and the audio source quality parameters do not meet the preset lossless triggering condition, the scene type parameter of the current playback task is determined as the ordinary multi-room playback scene type parameter, corresponding to daily playback scenarios such as online music streaming playback such as Spotify and background music for family gatherings.
[0021] After completing scene classification based on objective signals, the system reads the user's playback mode settings from the control application. These playback mode settings refer to the playback target mode actively selected by the user in the control application interface, specifically including low-latency mode, normal multi-room mode, high-quality mode, lossless priority mode, and automatic mode. When the playback mode setting parameter is set to low latency mode, the scene type parameter is forcibly updated to the low latency scene type parameter, and the status information of the currently used low latency encoding format is simultaneously displayed in the control application interface, such as "Currently using LC3+ low latency transmission", so that the user can perceive the current playback status in real time. When the playback mode setting parameter is set to lossless priority mode, the scene type parameter is forcibly updated to the lossless playback scene type parameter, and the status information of the currently used lossless encoding format, its sampling rate, and bit depth is displayed in the control application interface, such as "Lossless transmission: FLAC, 96kHz / 24bit". When the playback mode setting parameter is neither low latency mode nor lossless priority mode, the scene type parameter is determined based on the audio input source type and audio source quality parameters. If the actual encoding format used is inconsistent with the playback mode set by the user due to insufficient device capabilities or deteriorating network conditions, the control application outputs an encoding format switching prompt to the user, such as "Current network fluctuation, switched to AAC stable mode", to avoid the user mistakenly believing that the device is malfunctioning. Taking a user selecting lossless priority mode in the control application and playing FLAC music from their local NAS as an example, even if the current network conditions are average, the scene type parameter will be forcibly updated to the lossless playback scene type parameter, and subsequent attempts will prioritize FLAC transmission. Similarly, when a user selects low latency mode, even if the current network conditions are good and the audio source is a regular streaming media, the scene type parameter will be forcibly updated to the low latency scene type parameter, prioritizing low-latency encoding strategies over high-bitrate lossless encoding. This user-subjective selection mechanism for scene type parameters solves the problem that objective signals cannot fully reflect the user's actual playback intentions, ensuring that the system response matches the user's actual expectations.
[0022] 103. Based on the scenario type parameter, the common decoding format set, and the network transmission quality parameter, determine the target encoding format from the preset encoding format level. When the network transmission quality parameter is lower than the transmission quality threshold corresponding to the target encoding format, update the target encoding format to an encoding format in the preset encoding format level whose transmission quality threshold is lower than the network transmission quality parameter. In this embodiment, determining the target encoding format from the preset encoding format hierarchy based on the scene type parameter, the common decoding format set, and the network transmission quality parameter includes: mapping the scene type parameter to the encoding format with the corresponding priority in the preset encoding format hierarchy to obtain candidate encoding formats; verifying whether the candidate encoding format belongs to the common decoding format set; if the candidate encoding format belongs to the common decoding format set, then the candidate encoding format is determined as the target encoding format; if the candidate encoding format does not belong to the common decoding format set, then the highest-level encoding format with a lower level than the candidate encoding format in the preset encoding format hierarchy is selected from the common decoding format set and determined as the target encoding format. The step of mapping the scene type parameter to the corresponding priority encoding format in the preset encoding format hierarchy to obtain candidate encoding formats includes: determining latency requirement score and sound quality requirement score based on the scene type parameter, and determining network stability score based on the network transmission quality parameter; weighting and summing the latency requirement score, the network stability score, and the sound quality requirement score according to preset weights to obtain a comprehensive encoding format score; comparing the comprehensive encoding format score with the score intervals corresponding to each encoding format in the preset encoding format hierarchy, and selecting the encoding format corresponding to the score interval to which the comprehensive encoding format score belongs based on the comparison results to obtain candidate encoding formats.
[0023] In practical applications, three scores are determined based on the current scenario type parameters and network transmission quality parameters. The latency requirement score and audio quality requirement score are directly mapped from the scenario type parameters: when the scenario type parameter is a low-latency scenario, the latency requirement score is high and the audio quality requirement score is low, for example, a latency requirement score of 90 and an audio quality requirement score of 60; when the scenario type parameter is a lossless playback scenario, the latency requirement score is low and the audio quality requirement score is high, for example, a latency requirement score of 30 and an audio quality requirement score of 95; when the scenario type parameter is a normal multi-room playback scenario, both scores are taken as the median. The correspondence between the scenario type parameters and the score values is pre-configured in the system firmware as a score lookup table. The corresponding latency requirement score and audio quality requirement score are directly retrieved from the table based on the current scenario type parameter, without real-time calculation. The network stability score is derived from network transmission quality parameters through a linear mapping. For example, the specific conversion method could be to proportionally map the range of network transmission quality parameters to a scoring range of 0 to 100 points. A higher network transmission quality parameter value results in a higher network stability score; for instance, a high network transmission quality parameter corresponds to a network stability score of 90 points, while a decrease in network transmission quality due to a drop in Wi-Fi signal strength leads to a score of 70 points. All three scores are expressed on a percentage basis, ranging from 0 to 100 points. Specifically, the network transmission quality parameter Q is calculated using a non-linear weighted multiplication method, as shown in the following formula: ; The normalized degradation values of each indicator are obtained through the following mapping formula: ; The symbols in the above formulas have the following meanings: Q is the network transmission quality parameter, ranging from 0 to 100, with higher values indicating better link quality; i is the index number, ranging from 1 to 5, corresponding to packet loss rate, network round-trip delay, link jitter, RSSI degradation mapping, and retransmission count, respectively; x_i is the normalized degradation degree of the i-th index, ranging from 0 to 1; m_i is the currently collected raw value of the i-th index; M_i is the normalized upper limit benchmark value of the i-th index, pre-configured in the software; α_i is the penalty weight coefficient of the i-th index, ranging from 0 to 1, with higher values assigned to packet loss rate and network round-trip delay, and lower values assigned to retransmission count; β_i is the non-linear penalty exponent of the i-th index, β_i ≥ 1, when β_i > 1, the penalty is lighter in the early stage of degradation and increases sharply after the critical zone, when β_i = A value of 1 degrades to a linear penalty; ∏ is a multiplication operator, ensuring that the Q-value drops significantly when any metric deteriorates severely. Taking the kitchen equipment in the playback group as an example, the currently collected data shows a packet loss rate of 4%, network round-trip latency of 30ms, link jitter of 8ms, RSSI degradation mapping value of 20dB, and retransmissions of 3 times / second. The firmware pre-configured upper limit benchmark values for each metric are 10%, 100ms, 40ms, 50dB, and 20 times / second, respectively, corresponding to penalty weight coefficients. The values are 0.9, 0.8, 0.6, 0.5, and 0.4 respectively, representing non-linear penalty exponents. The body values are 2.0, 1.5, 1.2, 1.0, and 1.0. First, calculate each normalized value. , , , , Then, the penalty factor for each indicator was calculated. , , , , .final When this value falls below the transmission quality threshold corresponding to the FLAC encoding format (firmware default is 70), the system triggers a decision to downgrade from FLAC to AAC encoding format. Compared to the traditional linear weighted summation method, by adopting a multiplicative structure, a severe deterioration in any metric can independently lower the value. This value helps prevent a severely degraded indicator from being diluted by the mean of other good indicators, making coding degradation decisions more sensitive and accurate. Non-linear penalty index. This ensures the system doesn't trigger unnecessary degradation in low packet loss ranges, but the penalty increases sharply beyond the critical range, accurately reflecting the non-linear destructive effect of packet loss on audio frame integrity. Each indicator is tested through... Unified mapping to the 0 to 1 range, with an upper limit benchmark value. It can be configured independently for different network environments without modifying the overall system of weighting coefficients and nonlinear exponents, reducing the cost of firmware adaptation for multiple devices. Meanwhile, Value as network stability score By inputting the source of the data into the above coding format and comprehensive scoring formula, a more realistic score value can be obtained.
[0024] The comprehensive encoding format score is obtained by weighting and summing the latency requirement score, network stability score, and audio quality requirement score according to preset weights. The calculation formula is as follows: S = W1×S d + W2×S n + W3×S q ; Where S is the comprehensive score for the encoding format; S d Score the delayed demand; S n Score the network stability; q The system assigns a score to the audio quality requirement; W1 is the preset weight for the latency requirement score; W2 is the preset weight for the network stability score; and W3 is the preset weight for the audio quality requirement score, where W1 + W2 + W3 = 1. These preset weights are pre-configured in the system firmware. Taking a low-latency playback task as an example, S... d =90、S n =70、S q =60, the overall score of the encoding format after weighted summation of the three factors is relatively high, so LC3+ is selected as the candidate encoding format; taking a lossless playback scenario as an example, S d =30、S n =90、S q =95. The weighted sum of the three scores indicates a preference for high sound quality, leading to the selection of FLAC as the candidate encoding format. By integrating the three independent scores into a single comprehensive encoding format score through weighted summation, the system can quantitatively balance latency requirements, network stability, and sound quality requirements, avoiding the problem of a single dimension dominating the decision and causing severe performance degradation in other dimensions.
[0025] The overall score of the encoding format is then compared with the score ranges corresponding to each encoding format in the preset encoding format hierarchy to determine candidate encoding formats. The preset encoding format hierarchy refers to a system-defined priority ranking of encoding formats, divided into three levels: Level 1 is LC3+ encoding format, corresponding to low-latency scenarios; Level 2 is AAC encoding format, corresponding to normal multi-room playback scenarios; and Level 3 is FLAC encoding format, corresponding to lossless playback scenarios. The score ranges corresponding to each encoding format are pre-configured in the system firmware. For example, an encoding format with an overall score falling into a high score range corresponds to LC3+, falling into a middle score range corresponds to AAC, and falling into a low score range corresponds to FLAC. By comparing the current overall score of the encoding format with the score ranges corresponding to each encoding format one by one, the encoding format corresponding to the score range in which the overall score falls is selected as the candidate encoding format. The process then verifies whether the candidate encoding format belongs to the common decoding format set. If it does, the candidate encoding format is directly determined as the target encoding format. If it does not, the highest-level encoding format in the preset encoding format hierarchy below the candidate encoding format is selected from the common decoding format set as the target encoding format. Simultaneously, the application outputs a prompt to the user, indicating that there are devices in the current playback group that do not support the candidate encoding format. For example, it might prompt, "This device does not support FLAC lossless decoding and has switched to AAC multi-room mode." The user can choose to continue playback in the entire group or remove the unsupported device from the playback group. Taking the candidate encoding format as LC3+ but the common decoding format set as {AAC, FLAC} as an example, LC3+ does not belong to the common decoding format set. Therefore, the highest-level encoding format below LC3+ in the preset encoding format hierarchy, AAC, is selected from {AAC, FLAC}, and the target encoding format is determined as AAC. This ensures that all devices in the playback group can decode normally, and when device capabilities are limited, the highest-level available encoding format is selected to optimize playback quality within constraints.
[0026] 104. Based on the target encoding format, the audio stream to be transmitted is encoded and encapsulated into an audio transmission packet, wherein the header of the audio transmission packet includes an encoding format identifier field and a playback timestamp field; In this embodiment, encoding and encapsulating the audio stream to be transmitted into an audio transmission packet based on the target encoding format includes: decoding the audio stream to be transmitted into initial audio data, and performing sampling rate conversion and channel number adaptation on the initial audio data based on the maximum decodable sampling rate and the maximum decodable number of channels corresponding to the common decoding format set to obtain adapted audio data; performing frame sequence encoding on the adapted audio data based on the target encoding format to obtain an encoded audio frame sequence; encapsulating each encoded audio frame in the encoded audio frame sequence into an audio transmission packet, and filling the header of the audio transmission packet with an encoding format identifier field and a playback timestamp field. The step of encoding the adapted audio data into an encoded audio frame sequence based on the target encoding format includes: when the target encoding format is LC3+ encoding format, performing LC3+ frame sequence encoding on the adapted audio data using a preset short frame length parameter to obtain an encoded audio frame sequence; when the target encoding format is AAC encoding format, determining a target bitrate parameter based on the network transmission quality parameter, and performing AAC frame sequence encoding on the adapted audio data using the target bitrate parameter to obtain an encoded audio frame sequence; when the target encoding format is FLAC encoding format, performing lossless frame sequence encoding on the adapted audio data to obtain an encoded audio frame sequence.
[0027] In practical applications, the audio stream to be transmitted is first decoded into initial audio data. Initial audio data refers to uncompressed PCM format audio data obtained by restoring the original encoding format of the audio stream to be transmitted. PCM format is a common intermediate representation of audio data before encoding processing, and all encoders use PCM data as input. The specific decoding method is determined based on the original encoding format of the audio stream to be transmitted: if the original format is MP3, the MP3 decoder is called to restore it to PCM data; if the original format is AAC, the AAC decoder is called to restore it to PCM data; if the original format is FLAC, the FLAC decoder is called to restore it to PCM data. After decoding, the sampling rate and number of channels of the initial audio data are compared with the above upper limits by reading the maximum decodeable sampling rate and the maximum decodeable number of channels recorded in the common decoding format set. If the sampling rate of the initial audio data exceeds the maximum decodeable sampling rate, the resampling processor is called to perform downsampling processing on the PCM data. Specifically, this involves sampling points at equal intervals in the original PCM sampling sequence according to the target sampling rate, and using linear interpolation to complete the values between adjacent sampling points, so that the temporal resolution of the audio data matches the target sampling rate. If the number of channels in the initial audio data exceeds the maximum number of decodeable channels, channel downmixing is performed on the multi-channel PCM data. Specifically, the PCM sample values of each channel are weighted and superimposed according to fixed mixing coefficients specified in the ITU-R BS.775 standard to obtain PCM data that matches the target number of channels. For example, if the audio source is 192kHz / 24bit dual-channel FLAC, but a device in the playback group only supports 96kHz / 24bit decoding, the initial audio data is downsampled from 192kHz to 96kHz to obtain adapted audio data. Similarly, if the input is 5.1-channel audio and all playback group devices are dual-channel, the 5.1-channel PCM data is first downmixed before encoding and transmission. This ensures that the sampling rate and number of channels are uniformly adapted to the range that all devices in the playback group can decode, preventing some devices from failing to participate in playback due to parameters exceeding their own decoding limits.
[0028] Then, based on the target encoding format, frame sequence encoding is performed on the adapted audio data to obtain an encoded audio frame sequence. Frame sequence encoding refers to dividing continuous PCM audio data into several audio frames of fixed duration, compressing and encoding each frame independently, and finally outputting an encoded audio frame sequence composed of multiple encoded audio frames arranged sequentially. When the target encoding format is LC3+ encoding format, LC3+ frame sequence encoding is performed on the adapted audio data using a preset short frame length parameter. LC3+ is an encoding format designed specifically for low-latency real-time audio transmission. The short frame length parameter refers to the audio duration corresponding to each encoded audio frame. The shorter the frame length, the smaller the amount of encoded data per frame and the lower the encoder output latency. The short frame length parameter is pre-configured in the firmware, and shorter encoded frames are used when the local area network quality is good, so that the end-to-end encoding latency remains at a low level. Taking a user inputting TV audio via HDMI eARC, with the playback group including the main living room speakers and subwoofer, as an example, short-frame-long LC3+ encoding is used. The encoder completes encoding and outputs immediately upon receiving each frame of PCM data, effectively ensuring synchronized output of TV picture and audio. After LC3+ encoding, the corresponding audio transmission packet is placed in a high-priority transmission queue, sent to each playback device before ordinary music buffer data, ensuring that the TV sound will not experience significant delay due to ordinary music buffer packets competing for transmission resources. When the target encoding format is AAC, the target bitrate parameter is first determined based on the current network transmission quality parameters. The correspondence between the target bitrate parameter and the network transmission quality parameters is pre-configured as a bitrate lookup table in the system firmware. After reading the current network transmission quality parameters, the target bitrate parameter is directly looked up in the table. When the network transmission quality parameters are high, the target bitrate parameter takes a higher value to ensure sound quality; when the network transmission quality parameters are low, the target bitrate parameter takes a lower value to reduce transmission pressure. The AAC-LC encoder is then called with the target bitrate parameter to segment and encode the adapted audio data frame by frame, outputting an AAC encoded audio frame sequence. Taking a user connecting three devices in the living room, dining room, and bedroom to play online music, with a weaker network signal in the bedroom, the target bitrate parameter is determined by looking up the network transmission quality parameters of the bedroom device in a table and set to a lower value. AAC-LC frame sequence encoding is then performed at this bitrate to reduce transmission bandwidth requirements and ensure stable reception on the bedroom device. When the target encoding format is FLAC, if the original audio source is FLAC and its sampling rate and bit depth are within the range supported by the playback devices, the original FLAC frames are directly pass-through encapsulated without re-encoding, avoiding processing delays caused by repeated encoding. If the original audio source is PCM or WAV, the FLAC encoder is used to perform lossless frame sequence encoding on the adapted audio data before encapsulation and transmission. FLAC encoding uses a reversible compression algorithm; the encoded data can be completely restored to the original PCM data after decoding without losing any audio information, making it suitable for scenarios with stringent requirements for audio fidelity, such as high-resolution music playback on local NAS devices.
[0029] Finally, each encoded audio frame in the encoded audio frame sequence is sequentially encapsulated into an audio transmission packet. The header of each audio transmission packet is filled with an encoding format identifier field, a playback timestamp field, a frame sequence number field, and a playback scene identifier field. The frame sequence number field identifies the frame's position within the entire audio stream, allowing the receiving device to determine frame continuity and detect any dropped frames. The playback scene identifier field indicates the playback scene type of the current audio transmission packet (low-latency scene, normal multi-room scene, or lossless playback scene). Reading this field allows the receiving device to adjust its local buffering strategy accordingly; for example, a smaller buffer is used for priority playback in low-latency scenes, while a larger buffer is used in lossless scenes to improve stability. The encoding format identifier field in the audio transmission packet header indicates the encoding format type used by the current audio frame. Reading this field allows the receiving device to determine which decoder should be called to decode the audio frame data in the packet body. For example, a field value of LC3+ calls the LC3+ decoder, a field value of AAC calls the AAC decoder, and a field value of FLAC calls the FLAC decoder, achieving automatic decoder routing without requiring the upper-layer application to be aware of encoding format changes. The playback timestamp field in the audio transmission packet header indicates the absolute time value at which the audio frame should be output and played. It specifically includes two pieces of information: the frame sequence number and the playback timestamp. The frame sequence number identifies the frame's position within the entire audio stream, while the playback timestamp identifies the absolute playback time of that frame. Upon receiving the audio transmission packet, each audio device compares the playback timestamp field with its local clock and outputs the PCM data for that frame when the local clock reaches the time corresponding to the playback timestamp. For example, in a playback group consisting of a living room amplifier and bedroom speakers, the main device fills the 1000th frame of audio data with a frame sequence number of 1000 and a playback timestamp of 10.000 seconds. The living room amplifier and bedroom speakers each output the frame when their local clocks reach 10.000 seconds, achieving synchronized playback across multiple rooms.
[0030] 105. Distribute the audio transmission packet to each audio device. Each audio device decodes the audio frame data in the audio transmission packet based on the encoding format identifier field to obtain PCM audio data. Based on the playback timestamp field, it aligns the local clock of each audio device and synchronously outputs the PCM audio data.
[0031] In this embodiment, the audio devices decode the audio frame data in the audio transmission packet based on the encoding format identifier field to obtain PCM audio data. This includes: each audio device reading the encoding format identifier field in the header of the audio transmission packet, performing a consistency check between the header field and the audio frame data; when the check passes, decoding the audio transmission packet based on the encoding format identifier field to obtain PCM audio data; when the check fails, marking the current frame as a missing frame, and performing compensation padding on the missing frame based on the preceding PCM audio data in the local cache of each audio device, outputting the compensated PCM audio data; when the number of consecutive compensation padding exceeds a preset threshold, updating the target encoding format to an encoding format in a preset encoding format level that is lower than the current target encoding format level. The step of aligning the local clocks of each audio device with the playback timestamp field and synchronously outputting the PCM audio data includes: each audio device reading the playback timestamp field in the audio transmission packet header and subtracting the playback timestamp field from the current value of the local clock corresponding to each audio device to obtain a clock deviation value; when the clock deviation value is positive, reducing the local audio output sampling rate by a preset sampling rate step; when the clock deviation value is negative, increasing the local playback buffer depth by a preset buffer step; and when the clock deviation value is less than a preset synchronization deviation threshold, outputting the PCM audio data when the local clock reaches the time corresponding to the playback timestamp field. After outputting the PCM audio data when the local clock reaches the time corresponding to the playback timestamp field, the method further includes: subtracting the current value of the local clock corresponding to each audio device from the playback timestamp field to obtain a clock deviation value; when the clock deviation value is greater than a preset synchronization deviation threshold and is positive, reducing the local audio output sampling rate by a preset sampling rate step; when the clock deviation value is greater than the preset synchronization deviation threshold and is negative, increasing the local playback buffer depth by a preset buffer step; when the clock deviation value is greater than a preset maximum deviation threshold within a preset number of consecutive detection cycles, clearing the local playback buffer of the corresponding audio device, and rereading audio frame data from the audio transmission packet based on the playback timestamp field and writing it into the local playback buffer.After distributing the audio transmission packet to each audio device, the method further includes: collecting network transmission quality parameters of each audio device; when the network transmission quality parameters are lower than the transmission quality threshold corresponding to the target encoding format, updating the target encoding format to an encoding format in a preset encoding format layer whose transmission quality threshold is lower than the network transmission quality parameters, thus obtaining a switched target encoding format; generating a preset number of encoded audio frames based on the switched target encoding format and writing them into the buffer of each audio device; adding one frame duration to the playback timestamp field of the last frame of the target encoding format and assigning it to the playback timestamp field of the first frame of the switched target encoding format, thus obtaining an encoding format switching timestamp; when the local clock of each audio device reaches the encoding format switching timestamp, decoding the encoded audio frames in the buffer of each audio device based on the switched target encoding format, performing amplitude reduction processing on the PCM audio data before the encoding format switching timestamp, and performing amplitude increase processing on the PCM audio data after the encoding format switching timestamp before outputting it. After obtaining the encoding format switching timestamp, the process further includes: when the network transmission quality parameters of each audio device are higher than the transmission quality threshold corresponding to the target encoding format and the duration exceeds the preset network stability duration threshold, updating the target encoding format to the encoding format corresponding to the scene type parameter in the preset encoding format hierarchy to obtain the upgraded target encoding format; generating a preset number of encoded audio frames according to the upgraded target encoding format and writing them into the buffer of each audio device; adding one frame duration to the value of the playback timestamp field of the last frame of the target encoding format and assigning it to the playback timestamp field of the first frame of the upgraded target encoding format to obtain the encoding format upgrade timestamp; when the local clock of each audio device reaches the encoding format upgrade timestamp, decoding the encoded audio frames in the buffer of each audio device based on the upgraded target encoding format, performing amplitude reduction processing on the PCM audio data before the encoding format upgrade timestamp, and performing amplitude increase processing on the PCM audio data after the encoding format upgrade timestamp before outputting.After distributing the audio transmission packet to each audio device, the method further includes: collecting the network transmission quality parameters of each audio device; when the network transmission quality parameter of any audio device is lower than the transmission quality threshold corresponding to the lowest level encoding format in the preset encoding format hierarchy, invalidating the playback group identifier field of the audio device and stopping the distribution of audio transmission packets to the audio device; when the network transmission quality parameter of the audio device is higher than the transmission quality threshold corresponding to the lowest level encoding format in the preset encoding format hierarchy, resending a capability query request to the audio device, receiving the current audio decoding capability parameters returned by the audio device, re-performing the intersection operation on the audio decoding capability parameters of each audio device in the playback group to obtain an updated common decoding format set, and after aligning the local clock of the audio device based on the playback timestamp field, validating the playback group identifier field of the audio device, and resuming the distribution of audio transmission packets to the audio device; when the audio decoding capability parameters of an audio device in the playback group do not include the target encoding format, re-determining the target encoding format based on the updated common decoding format set and the scene type parameter, and re-encoding and encapsulating the audio stream to be transmitted according to the target encoding format before distributing it to each audio device.
[0032] In practical applications, after receiving an audio transmission packet, each audio device first reads the encoding format identifier field in the packet header to confirm the encoding format type used by the current audio frame. Then, it performs a consistency check between the packet header fields and the audio frame data. This consistency check verifies the correspondence between the frame sequence number, data length, and checksum recorded in the packet header and the actual audio frame data in the packet body. Specifically, the receiving device calculates the checksum value for the audio frame data in the packet body using the CRC (Cyclic Redundancy Check) algorithm and compares the result with the pre-stored checksum field in the packet header. If they match, the check passes; otherwise, it fails. When the check passes, the receiving device routes the audio frame data in the packet body to the corresponding decoder based on the value of the encoding format identifier field: LC3+ is used when the field value is LC3+, AAC when the field value is AAC, and FLAC when the field value is FLAC. After decoding, PCM audio data is output. When verification fails, the receiving device marks the current frame as a missing frame, retrieves the preceding PCM audio data from its local cache, and copies this data to fill the time position corresponding to the missing frame, ensuring continuous playback output and preventing pops or silences caused by missing single frames. When the number of consecutive compensation fills exceeds a preset threshold, it indicates that the current network condition can no longer support stable transmission of the current target encoding format. The receiving device sends an encoding degradation notification message to the master device via the local area network. Upon receiving the notification, the master device updates the target encoding format to a preset encoding format level lower than the current target encoding format level. The preset threshold is pre-configured in the system firmware. If the threshold is too low, even a brief network fluctuation will trigger encoding degradation, leading to frequent switching and impacting user experience. If the threshold is too high, the device's prolonged reliance on compensation fill output results in significant cumulative sound quality loss. Taking the example of a bedroom speaker experiencing multiple frame verification failures during FLAC lossless playback due to the Wi-Fi signal dropping from -55dBm to -75dBm, when the number of consecutive compensation padding attempts reaches a preset threshold, the bedroom speaker sends an encoding downgrade notification message to the main device. The main device then downgrades the target encoding format from FLAC to AAC, reducing the amount of data transmitted per unit time, thus enabling the bedroom speaker to restore stable reception and decoding under the current network conditions.
[0033] Each audio device reads the playback timestamp field from the audio transmission packet header and subtracts the value of the playback timestamp field from the current value of its local clock to obtain the clock deviation value. The clock deviation value is the difference between the expected playback time of the current audio frame and the current reading of the local clock. A positive value indicates that the current reading of the local clock has exceeded the expected playback time of the frame, meaning the local clock is too fast; a negative value indicates that the current reading of the local clock has not yet reached the expected playback time of the frame, meaning the local clock is too slow. When the clock deviation value is positive, the local clock is too fast. Directly reducing the buffer depth would cause audio frames already accumulated in the buffer to be skipped, resulting in missing playback content. Therefore, the local audio output sampling rate is reduced by a preset sampling rate step. The local audio output sampling rate refers to the number of sampling points output per second by the device's DAC (Digital-to-Analog Converter, responsible for converting PCM digital signals into analog audio signals). Reducing the output sampling rate slightly increases the actual duration of each sampling point, resulting in a slight slowdown in the audio output rhythm and a decrease in the effective advance speed of the local clock, gradually aligning the local clock with the time indicated by the playback timestamp field. The preset sampling rate step size is pre-configured in the system firmware to a small increment, ensuring that the human ear cannot perceive pitch changes during adjustment. When the clock deviation value is negative, the local clock is too slow. Directly increasing the sampling rate would accelerate the audio output, causing pitch distortion. Therefore, the system increases the local playback buffer depth by the preset buffer step size. The local playback buffer depth refers to the number of audio frames accumulated in the local cache before the device outputs PCM audio data. Increasing the buffer depth allows the device to wait for more frames to arrive before outputting, shifting the actual output time backward to align with the playback timestamp field, and also providing buffer space for the natural alignment of the subsequent sampling rate. When the absolute value of the clock deviation is less than the preset synchronization deviation threshold, the deviation between the local clock and the time indicated by the playback timestamp field is considered to be within an acceptable range. The device outputs PCM audio data when the local clock reaches the time corresponding to the playback timestamp field. The preset synchronization deviation threshold is pre-configured in the system firmware. This threshold is set based on the human ear's perception boundary of multi-room audio synchronization errors. When the error exceeds this boundary, the user will perceive echo effects from different rooms. Taking the example of a kitchen device whose playback clock gradually leads other devices by 10 milliseconds due to local crystal oscillator error, the clock deviation value is positive. The local audio output sampling rate of the kitchen device is slightly reduced, so that its playback rhythm is slightly slowed down and gradually aligned with the playback timestamp field.
[0034] After each audio device outputs PCM audio data, the playback synchronization status is continuously corrected by subtracting the current value of the local clock of each audio device from the playback timestamp field. When the clock deviation value is greater than the preset synchronization deviation threshold and is positive, the device reduces the local audio output sampling rate by a preset sampling rate step; when the clock deviation value is greater than the preset synchronization deviation threshold and is negative, the device increases the local playback buffer depth by a preset buffer step. When the clock deviation value is greater than the preset maximum deviation threshold within a preset number of consecutive detection cycles, it indicates that neither sampling rate fine-tuning nor buffer depth adjustment can eliminate the current clock deviation, and the device performs forced resynchronization. The specific implementation of forced resynchronization is as follows: the device clears all cached audio frame data in the local playback buffer and rereads the audio frame data corresponding to the current time position from the audio transmission packet based on the playback timestamp field and writes it into the local playback buffer. Specifically, the device sends a resynchronization request message to the master device, carrying the current local clock reading. Upon receiving the request, the master device searches for the frame number corresponding to that moment in the audio stream based on the local clock reading in the request message. Starting from that frame number, it resends audio transmission packets to the device. After receiving the resent audio transmission packets, the device writes the audio frame data into its local playback buffer and realigns its local clock according to the playback timestamp field before resuming output. The preset number of continuous detection cycles and the preset maximum deviation threshold are both pre-configured in the system firmware. The setting of the number of continuous detection cycles ensures that the system has given enough opportunities for fine-tuning before triggering forced resynchronization, avoiding brief playback interruptions caused by a single, occasional deviation triggering forced resynchronization. Taking the case where the clock deviation of the kitchen device continues to increase and cannot converge even after multiple sampling rate fine-tunings, when the deviation exceeds the preset maximum deviation threshold within the preset number of continuous detection cycles, the kitchen device clears its local playback buffer and sends a resynchronization request message to the master device. The master device resends audio transmission packets to the kitchen device starting from the corresponding frame number. The kitchen device rewrites the buffer, aligns its clock, and resumes synchronized output with other devices in the playback group.
[0035] During the continuous distribution of audio transmission packets to various audio devices, network transmission quality parameters of each audio device are periodically collected. Specific collected indicators include packet loss rate, network round-trip latency, link jitter, link throughput, Wi-Fi signal strength (RSSI), retransmission count, and local playback buffer change trends. When the network transmission quality parameter of an audio device falls below the transmission quality threshold corresponding to the current target encoding format, the target encoding format is updated to an encoding format in the preset encoding format layer whose transmission quality threshold is lower than the current network transmission quality parameter, thus obtaining the target encoding format for switching. The transmission quality threshold refers to the minimum value of network transmission quality parameters required to maintain stable transmission of the corresponding encoding format. The transmission quality threshold corresponding to the FLAC encoding format is the highest, followed by the AAC encoding format, and the transmission quality threshold corresponding to the LC3+ encoding format is the lowest. After determining the target encoding format for switching, the master device pre-generates a preset number of encoded audio frames according to the target encoding format and sends them to each audio device. Each device writes these frames into its local buffer, ensuring that there are sufficient audio frames in the target encoding format available for decoding and output before the switching time arrives. Taking the system's planned switch from FLAC to AAC at frame 20000 as an example, the master device begins generating AAC encoded frames at frame 19900 and sends them to the buffers of each device, ensuring that each device's buffer has sufficient AAC frames before the switch occurs. Then, the playback timestamp field of the last frame of the current target encoding format is incremented by one frame duration and assigned to the playback timestamp field of the first frame of the target encoding format to obtain the encoding format switch timestamp, ensuring that the two audio segments before and after the switch are strictly continuous and seamless in time. When the local clock of each audio device reaches the encoding format switch timestamp, each device stops decoding the old encoding format audio frames and instead decodes the encoded audio frames in the buffer based on the target encoding format. The PCM audio data before the encoding format switch timestamp undergoes a gradual decrease in amplitude processing, while the PCM audio data after the encoding format switch timestamp undergoes a gradual increase in amplitude processing. Amplitude reduction processing refers to gradually reducing the amplitude value of the PCM audio data from its original amplitude to zero linearly within approximately 20 milliseconds before the switching timestamp. Amplitude increase processing refers to gradually restoring the amplitude value of the PCM audio data from zero to its original amplitude linearly within approximately 20 milliseconds after the switching timestamp. These two processes together create a fade-in / fade-out effect, eliminating the abrupt change in sound caused by the difference in timbre characteristics between the two encoding formats at the encoding format switching point. Taking a balcony device where the network transmission quality parameters are lower than the corresponding FLAC transmission quality threshold due to weakened Wi-Fi signal as an example, the target encoding format is determined to be AAC. An AAC encoded audio frame is pre-generated and written to the balcony device's buffer. After performing fade-in / fade-out processing at the encoding format switching timestamp, the device switches to AAC decoding output, resulting in a smooth and seamless listening experience for the user.
[0036] When the network transmission quality parameters of each audio device are detected to be consistently higher than the transmission quality threshold corresponding to the target encoding format and the duration exceeds the preset network stability duration threshold, the target encoding format is upgraded to the encoding format corresponding to the current scene type parameter in the preset encoding format level. The preset network stability duration threshold is pre-configured in the system firmware. This threshold ensures that the network status remains stable for a sufficiently long time before the upgrade is performed, avoiding repeated switching caused by a brief network rebound triggering the upgrade and then dropping again. After determining the target encoding format, the master device pre-generates a preset number of encoded audio frames according to the target encoding format and sends them to the buffers of each audio device. The playback timestamp field value of the last frame of the target encoding format is added with one frame duration and assigned to the playback timestamp field of the first frame of the target encoding format, resulting in the encoding format upgrade timestamp. When the local clock of each audio device reaches the encoding format upgrade timestamp, the PCM audio data before the encoding format upgrade timestamp is processed by gradually decreasing amplitude, and the PCM audio data after the encoding format upgrade timestamp is processed by gradually increasing amplitude before output, thus achieving a smooth upgrade from a lower-level encoding format to a higher-level encoding format. Taking the example of a balcony device switching to AAC playback and the Wi-Fi signal continuously recovering and stabilizing for a duration exceeding a preset network stability threshold, the target encoding format for the upgrade is determined to be FLAC. After performing amplitude reduction and increase processing at the encoding format upgrade timestamp, the output switches back to FLAC decoding. Lossless audio quality is automatically restored when network conditions permit, without manual user intervention. In some embodiments, when there are significant differences in network quality among devices in the playback group, higher-level encoding formats (such as FLAC) are allowed to be transmitted to devices with acceptable network quality, while lower-level encoding formats (such as AAC) are transmitted to devices with unacceptable network quality. Both encoded streams share the same playback timestamp field, ensuring that devices with different encoding qualities can still synchronously output PCM audio data at the same time. Taking the living room amplifier and study player connected via wired network to receive FLAC lossless audio streams, and the balcony device connected via weaker Wi-Fi to receive AAC audio streams as an example, the main device generates both FLAC and AAC streams simultaneously and assigns the same playback timestamp field value to the corresponding frames of the two encoded streams. Each device outputs synchronously when its local clock reaches the timestamp, achieving synchronous playback in multiple rooms.
[0037] During playback, the aforementioned dynamic device capability update mechanism (polling device capabilities according to a preset collection cycle and re-performing the intersection operation) and the out-of-group / re-entry mechanism for network-abnormal devices together constitute a complete closed loop for dynamic management of playback groups. Finally, during audio packet distribution, network transmission quality parameters of each audio device are continuously collected. When the network transmission quality parameter of any audio device falls below the transmission quality threshold corresponding to the lowest level of the preset encoding format hierarchy, it indicates that the current network condition of that device can no longer support stable transmission of any encoding format. The playback group identifier field of that audio device is then invalidated, and audio packets are no longer distributed to that device. The playback group identifier field is the field in the audio packet header that identifies the playback group to which the current audio packet belongs. The receiving device only processes audio packets whose playback group identifier field matches the playback group to which it belongs. After invalidating this field, the audio packets sent by the master device no longer contain the playback group identifier of that device, and the device immediately stops receiving and processing them. Taking a disconnected balcony device as an example, invalidating the playback group identifier field of the balcony device allows the devices in the living room, bedroom, and kitchen to continue playing normally, unaffected by the balcony device's disconnection. When the network transmission quality parameters of the balcony device recover to a level higher than the transmission quality threshold corresponding to the lowest level encoding format, a capability query request is resent to the device. The device's returned current audio decoding capability parameters are received, and an intersection operation is re-performed on the audio decoding capability parameters of all devices in the playback group to obtain an updated common decoding format set. The device's local clock is aligned based on the playback timestamp field, and its playback group identifier field is re-enabled. Audio transmission packets are then distributed to the device again. If the current audio decoding capability parameters returned by a newly joined device do not include the current target encoding format, the target encoding format is re-determined based on the updated common decoding format set and the current scene type parameters. The audio stream to be transmitted is then re-encoded and encapsulated according to the new target encoding format before being distributed to all devices in the playback group, ensuring that the newly joined device can participate in decoding and playback normally.
[0038] In this embodiment of the invention, a hierarchical encoding decision mechanism is established, constrained by a common decoding format set for devices and driven by playback scenarios. This mechanism replaces the fixed encoding transmission architecture with scene recognition, format hierarchy mapping, network adaptive degradation, and a unified timestamp synchronization mechanism. This solves the technical problems of existing multi-room audio systems using fixed encoding formats, which result in a single encoding strategy, inability to adapt to the different needs of multiple scenarios, and playback stuttering and synchronization mismatch between devices caused by high bitrate encoding under weak network conditions. It achieves a dynamic balance between low latency, transmission stability, and sound quality fidelity within multi-room playback groups, effectively improving synchronization consistency and user listening experience in multi-device concurrent playback scenarios.
[0039] The above describes the multi-room audio hierarchical coding and transmission method based on scene recognition in the embodiments of the present invention. The following describes the multi-room audio hierarchical coding and transmission apparatus based on scene recognition in the embodiments of the present invention. Please refer to [link / reference]. Figure 2 One embodiment of the multi-room audio hierarchical coding and transmission device based on scene recognition in this invention includes: The parameter acquisition module 201 is used to acquire the audio stream to be transmitted and the audio decoding capability parameters and network transmission quality parameters of each audio device in the multi-room playback group, and to take the intersection of each audio decoding capability parameter to obtain a common decoding format set. The scene classification module 202 is used to classify the current playback task of the audio stream to be transmitted according to the playback task attribute parameters of the audio stream to be transmitted, and obtain the scene type parameter. The encoding decision module 203 is used to determine a target encoding format from a preset encoding format hierarchy based on the scene type parameter, the common decoding format set and the network transmission quality parameter. When the network transmission quality parameter is lower than the transmission quality threshold corresponding to the target encoding format, the target encoding format is updated to an encoding format in the preset encoding format hierarchy whose transmission quality threshold is lower than the network transmission quality parameter. The encoding and encapsulation module 204 is used to encode the audio stream to be transmitted and encapsulate it into an audio transmission packet based on the target encoding format. The header of the audio transmission packet includes an encoding format identifier field and a playback timestamp field. The distribution synchronization module 205 is used to distribute the audio transmission packet to each audio device. Each audio device decodes the audio frame data in the audio transmission packet based on the encoding format identifier field to obtain PCM audio data, and aligns the local clock of each audio device based on the playback timestamp field to synchronously output the PCM audio data.
[0040] In this embodiment of the invention, a hierarchical encoding decision mechanism is established, constrained by a common decoding format set for devices and driven by playback scenarios. This mechanism replaces the fixed encoding transmission architecture with scene recognition, format hierarchy mapping, network adaptive degradation, and a unified timestamp synchronization mechanism. This solves the technical problems of existing multi-room audio systems using fixed encoding formats, which result in a single encoding strategy, inability to adapt to the different needs of multiple scenarios, and playback stuttering and synchronization mismatch between devices caused by high bitrate encoding under weak network conditions. It achieves a dynamic balance between low latency, transmission stability, and sound quality fidelity within multi-room playback groups, effectively improving synchronization consistency and user listening experience in multi-device concurrent playback scenarios.
[0041] above Figure 2The multi-room audio hierarchical coding and transmission device based on scene recognition in this embodiment of the invention will be described in detail from the perspective of modular functional entities. The multi-room audio hierarchical coding and transmission device based on scene recognition in this embodiment of the invention will be described in detail from the perspective of hardware processing.
[0042] Figure 3 This is a schematic diagram of a scene-recognition-based multi-room audio hierarchical coding transmission device 300 provided in an embodiment of the present invention. The scene-recognition-based multi-room audio hierarchical coding transmission device 300 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 310 (e.g., one or more processors) and a memory 320, and one or more storage media 330 (e.g., one or more mass storage devices) for storing application programs 333 or data 332. The memory 320 and storage media 330 can be temporary or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the diagram), each module may include a series of instruction operations on the scene-recognition-based multi-room audio hierarchical coding transmission device 300. Furthermore, the processor 310 may be configured to communicate with the storage media 330 and execute the series of instruction operations in the storage media 330 on the scene-recognition-based multi-room audio hierarchical coding transmission device 300.
[0043] The scene-recognition-based multi-room audio hierarchical coding transmission device 300 may also include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input / output interfaces 360, and / or one or more operating systems 331, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will understand that... Figure 3 The illustrated multi-room audio hierarchical coding transmission device structure based on scene recognition does not constitute a limitation on the multi-room audio hierarchical coding transmission device based on scene recognition. It may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.
[0044] The present invention also provides a multi-room audio hierarchical coding and transmission device based on scene recognition. The computer device includes a memory and a processor. The memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor performs each step of the multi-room audio hierarchical coding and transmission method based on scene recognition in the above embodiments.
[0045] The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to perform various steps of the scene recognition-based multi-room audio hierarchical coding and transmission method.
[0046] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0047] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0048] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0049] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-room audio hierarchical coding and transmission method based on scene recognition, characterized in that, The scene-recognition-based multi-room audio hierarchical coding and transmission method includes: Obtain the audio stream to be transmitted and the audio decoding capability parameters and network transmission quality parameters of each audio device in the multi-room playback group, and take the intersection of each audio decoding capability parameter to obtain a common decoding format set; Based on the playback task attribute parameters of the audio stream to be transmitted, the current playback task of the audio stream to be transmitted is classified into scenarios to obtain scenario type parameters; Based on the scenario type parameter, the common decoding format set, and the network transmission quality parameter, a target encoding format is determined from the preset encoding format level. When the network transmission quality parameter is lower than the transmission quality threshold corresponding to the target encoding format, the target encoding format is updated to an encoding format in the preset encoding format level whose transmission quality threshold is lower than the network transmission quality parameter. Based on the target encoding format, the audio stream to be transmitted is encoded and encapsulated into an audio transmission packet. The header of the audio transmission packet includes an encoding format identifier field and a playback timestamp field. The audio transmission packet is distributed to each audio device. Each audio device decodes the audio frame data in the audio transmission packet based on the encoding format identifier field to obtain PCM audio data. Based on the playback timestamp field, the local clock of each audio device is aligned and the PCM audio data is output synchronously.
2. The multi-room audio hierarchical coding and transmission method based on scene recognition according to claim 1, characterized in that, The process of obtaining the audio decoding capability parameters and network transmission quality parameters of the audio stream to be transmitted and each audio device in the multi-room playback group, and taking the intersection of each of the audio decoding capability parameters to obtain a common decoding format set, includes: Acquire the audio stream to be transmitted, send capability query requests to each audio device in the multi-room playback group, and receive the audio decoding capability parameters returned by each audio device; Collect various network status data of each audio device in the multi-room playback group to obtain the network transmission quality parameters of each audio device; The common decoding format set is obtained by performing an intersection operation on the encoding format field in the audio decoding capability parameters returned by each audio device.
3. The multi-room audio hierarchical coding and transmission method based on scene recognition according to claim 1, characterized in that, The scene type parameters include low-latency scene type parameters, lossless playback scene type parameters, and normal multi-room playback scene type parameters. The scene type parameters are obtained by classifying the current playback task of the audio stream to be transmitted according to its playback task attribute parameters, resulting in the following: Determine the audio input source type of the audio stream to be transmitted. When the audio input source type belongs to a preset low-latency trigger condition, determine the scene type parameter of the current playback task as the low-latency scene type parameter. Determine the audio source quality parameters of the audio stream to be transmitted. When the audio source quality parameters belong to the preset lossless trigger conditions, determine the scene type parameters of the current playback task as lossless playback scene type parameters. When the audio input source type does not belong to the preset low latency triggering condition and the audio source quality parameter does not belong to the preset lossless triggering condition, the scene type parameter of the current playback task is determined to be the normal multi-room playback scene type parameter.
4. The multi-room audio hierarchical coding and transmission method based on scene recognition according to claim 3, characterized in that, After classifying the current playback task of the audio stream to be transmitted into a scene and obtaining the scene type parameter, the method further includes: Obtain the user's playback mode setting parameters. When the playback mode setting parameters are set to low latency mode, update the scene type parameter to the low latency scene type parameter. When the playback mode setting parameter is set to lossless priority mode, the scene type parameter is updated to the lossless playback scene type parameter; When the user playback mode setting parameter is not low latency mode and not lossless priority mode, the scene type parameter of the current playback task is determined based on the audio input source type and the audio source quality parameter.
5. The multi-room audio hierarchical coding and transmission method based on scene recognition according to claim 1, characterized in that, The step of determining the target encoding format from the preset encoding format hierarchy based on the scene type parameter, the common decoding format set, and the network transmission quality parameter includes: The scene type parameter is mapped to the corresponding priority encoding format in the preset encoding format hierarchy to obtain the candidate encoding format; Verify whether the candidate encoding format belongs to the common decoding format set. If the candidate encoding format belongs to the common decoding format set, then determine the candidate encoding format as the target encoding format. If the candidate encoding format does not belong to the public decoding format set, then the highest level encoding format with a lower level than the candidate encoding format is selected from the preset encoding format hierarchy in the public decoding format set and determined as the target encoding format.
6. The multi-room audio hierarchical coding and transmission method based on scene recognition according to claim 5, characterized in that, The step of mapping the scene type parameter to the corresponding priority encoding format in the preset encoding format hierarchy to obtain candidate encoding formats includes: The latency requirement score and audio quality requirement score are determined based on the scenario type parameters, and the network stability score is determined based on the network transmission quality parameters. The latency requirement score, the network stability score, and the audio quality requirement score are weighted and summed according to preset weights to obtain a comprehensive encoding format score. The comprehensive score of the encoding format is compared with the score intervals corresponding to each encoding format in the preset encoding format hierarchy. Based on the comparison results, the encoding format corresponding to the score interval to which the comprehensive score of the encoding format belongs is selected to obtain the candidate encoding format.
7. The multi-room audio hierarchical coding and transmission method based on scene recognition according to claim 1, characterized in that, The step of encoding and encapsulating the audio stream to be transmitted into an audio transmission packet based on the target encoding format includes: The audio stream to be transmitted is decoded into initial audio data, and based on the maximum decodable sampling rate and the maximum decodable number of channels corresponding to the common decoding format set, the initial audio data is sampled and adapted to the number of channels to obtain adapted audio data. Based on the target encoding format, the adapted audio data is frame sequence encoded to obtain an encoded audio frame sequence; Each encoded audio frame in the encoded audio frame sequence is encapsulated into an audio transmission packet, and the header of the audio transmission packet is filled with an encoding format identifier field and a playback timestamp field.
8. The multi-room audio hierarchical coding and transmission method based on scene recognition according to claim 7, characterized in that, The step of encoding the adapted audio data into a frame sequence based on the target encoding format to obtain an encoded audio frame sequence includes: When the target encoding format is LC3+ encoding format, the adapted audio data is encoded using LC3+ frame sequence using a preset short frame length parameter to obtain an encoded audio frame sequence. When the target encoding format is AAC encoding format, the target bit rate parameter is determined based on the network transmission quality parameter, and the target bit rate parameter is used to encode the adapted audio data into an AAC frame sequence to obtain an encoded audio frame sequence. When the target encoding format is FLAC encoding format, the adapted audio data is subjected to lossless frame sequence encoding to obtain an encoded audio frame sequence.
9. The multi-room audio hierarchical coding and transmission method based on scene recognition according to claim 1, characterized in that, Each audio device decodes the audio frame data in the audio transmission packet based on the encoding format identifier field to obtain PCM audio data, including: Each audio device reads the encoding format identifier field in the header of the audio transmission packet, performs a consistency check between the header field of the audio transmission packet and the audio frame data, and when the check passes, decodes the audio transmission packet based on the encoding format identifier field to obtain PCM audio data; When the verification fails, the current frame is marked as a missing frame, and the missing frame is compensated and filled based on the preceding PCM audio data in the local cache of each audio device, and the compensated PCM audio data is output. When the number of consecutive compensation padding exceeds a preset threshold, the target encoding format is updated to an encoding format in the preset encoding format hierarchy that is lower than the current target encoding format level.
10. The multi-room audio hierarchical coding and transmission method based on scene recognition according to claim 1, characterized in that, The step of aligning the local clocks of each audio device with the local clocks based on the playback timestamp field and synchronously outputting the PCM audio data includes: Each audio device reads the playback timestamp field in the audio transmission packet header and subtracts the playback timestamp field from the current value of the local clock corresponding to each audio device to obtain the clock deviation value; When the clock deviation value is positive, the local audio output sampling rate is reduced by a preset sampling rate step; when the clock deviation value is negative, the local playback buffer depth is increased by a preset buffer step. When the clock deviation value is less than the preset synchronization deviation threshold, the PCM audio data is output when the local clock reaches the time corresponding to the playback timestamp field.
11. The multi-room audio hierarchical coding and transmission method based on scene recognition according to claim 10, characterized in that, After outputting the PCM audio data when the local clock reaches the time corresponding to the playback timestamp field, the method further includes: The clock deviation value is obtained by subtracting the current value of the local clock corresponding to each audio device from the playback timestamp field. When the clock deviation value is greater than the preset synchronization deviation threshold and is positive, the local audio output sampling rate is reduced by a preset sampling rate step; when the clock deviation value is greater than the preset synchronization deviation threshold and is negative, the local playback buffer depth is increased by a preset buffer step. When the clock deviation value is greater than the preset maximum deviation threshold within a preset number of consecutive detection cycles, the local playback cache of the corresponding audio device is cleared, and the audio frame data is reread from the audio transmission packet based on the playback timestamp field and written to the local playback cache.
12. The multi-room audio hierarchical coding and transmission method based on scene recognition according to claim 1, characterized in that, After distributing the audio transmission packet to each audio device, the method further includes: According to the preset acquisition cycle, the system resends the capability query request to each audio device in the playback group, receives the current audio decoding capability parameters returned by each audio device, and stores the current audio decoding capability parameters as the audio decoding capability parameters for this acquisition cycle. The audio decoding capability parameters of the current acquisition cycle are compared with those of the previous acquisition cycle. If the audio decoding capability parameters of the current acquisition cycle are inconsistent with those of the previous acquisition cycle, the audio decoding capability parameters of each audio device are re-intersected to obtain an updated common decoding format set. Verify whether the target encoding format is included in the updated common decoding format set. If the target encoding format is not included in the updated common decoding format set, redetermine the target encoding format from the updated common decoding format set.
13. The multi-room audio hierarchical coding and transmission method based on scene recognition according to claim 1, characterized in that, After distributing the audio transmission packet to each audio device, the method further includes: Collect network transmission quality parameters of each audio device. When the network transmission quality parameters are lower than the transmission quality threshold corresponding to the target encoding format, update the target encoding format to an encoding format in the preset encoding format level whose transmission quality threshold is lower than the network transmission quality parameters, and obtain the target encoding format to be switched. Based on the target encoding format being switched, a preset number of encoded audio frames are generated and written to the buffers of each audio device. The playback timestamp field of the last frame of the target encoding format is added to the duration of one frame and then assigned to the playback timestamp field of the first frame of the target encoding format being switched, thus obtaining the encoding format switching timestamp. When the local clock of each audio device reaches the encoding format switching timestamp, the encoded audio frames in the buffer of each audio device are decoded based on the target encoding format to be switched, and the PCM audio data before the encoding format switching timestamp is subjected to amplitude reduction processing, and the PCM audio data after the encoding format switching timestamp is subjected to amplitude increase processing before being output.
14. A multi-room audio hierarchical coding and transmission device based on scene recognition, characterized in that, The scene recognition-based multi-room audio hierarchical coding and transmission device includes: The parameter acquisition module is used to acquire the audio stream to be transmitted and the audio decoding capability parameters and network transmission quality parameters of each audio device in the multi-room playback group, and to take the intersection of each audio decoding capability parameter to obtain a common decoding format set; The scene classification module is used to classify the current playback task of the audio stream to be transmitted according to the playback task attribute parameters of the audio stream to be transmitted, and obtain the scene type parameter. The encoding decision module is used to determine a target encoding format from a preset encoding format hierarchy based on the scenario type parameter, the common decoding format set, and the network transmission quality parameter. When the network transmission quality parameter is lower than the transmission quality threshold corresponding to the target encoding format, the target encoding format is updated to an encoding format in the preset encoding format hierarchy whose transmission quality threshold is lower than the network transmission quality parameter. The encoding and encapsulation module is used to encode the audio stream to be transmitted and encapsulate it into an audio transmission packet based on the target encoding format. The header of the audio transmission packet includes an encoding format identifier field and a playback timestamp field. The distribution synchronization module is used to distribute the audio transmission packet to each audio device. Each audio device decodes the audio frame data in the audio transmission packet based on the encoding format identifier field to obtain PCM audio data, and aligns the local clock of each audio device with the playback timestamp field to synchronously output the PCM audio data.
15. A multi-room audio hierarchical coding and transmission device based on scene recognition, characterized in that, The scene recognition-based multi-room audio hierarchical coding and transmission device includes: a memory and at least one processor, wherein the memory stores instructions; The at least one processor invokes the instructions in the memory to cause the scene-recognition-based multi-room audio hierarchical coding transmission device to perform the steps of the scene-recognition-based multi-room audio hierarchical coding transmission method as described in any one of claims 1-13.
16. A computer-readable storage medium storing instructions thereon, characterized in that, When the instructions are executed by the processor, they implement the various steps of the multi-room audio hierarchical coding and transmission method based on scene recognition as described in any one of claims 1-13.