Scene mode switching method and device of multichannel audio system, electronic equipment and system

By detecting the status and characteristic events of the audio input terminal, and combining the device capability map and rule engine, scene templates are automatically matched to achieve seamless switching of multi-channel audio systems. This solves the problem of low scene switching efficiency in existing technologies and improves the consistency of user experience and ease of operation.

CN121985262APending Publication Date: 2026-05-05LINKPLAY TECHNOLOGY INC NANJING
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LINKPLAY TECHNOLOGY INC NANJING
Filing Date
2025-12-16
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing multi-room/multi-channel audio systems suffer from low scene switching efficiency in multi-mode management and cross-device parameter coordination. Users need to manually group devices and adjust parameters, which makes operation cumbersome and prone to errors, making it difficult to guarantee immersive sound field and consistency of operation.

Method used

By detecting the signal status and characteristic events at the audio input end, combined with the device capability map, and using a rule engine to match scene templates, the system automatically determines the target group and audio parameters, enabling seamless switching of multi-channel audio systems, including virtual center channel processing and smooth parameter transition.

Benefits of technology

It improves the scene switching efficiency of multi-channel audio systems, lowers the user operation threshold, and realizes an intelligent and consistent listening experience of "turning on is like turning on a movie theater, and switching is like matching", avoiding the tedious work of manual mode management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121985262A_ABST
    Figure CN121985262A_ABST
Patent Text Reader

Abstract

The invention relates to a scene mode switching method and device of a multichannel audio system, electronic equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: detecting a signal state and a characteristic event of an audio input end; determining a device capability map of each audio playing device included in the multi-channel audio system; wherein the equipment capability map represents the audio playing capability of the audio playing equipment; based on the signal state and the feature event, a scene template matched with the playing scene to be switched is determined, and an equipment marshalling strategy and audio processing parameters are predefined in the scene template; and determining a target group and audio parameters of the audio playing device according to the scene template in combination with the device capability map, and switching the scene mode of the multichannel audio system based on the target group and audio parameters of the audio playing device. By adopting the method, the scene mode switching efficiency of the multi-channel audio system can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of media system switching technology, and in particular to a method, apparatus, electronic device and system for scene mode switching in a multi-channel audio system. Background Technology

[0002] In the current field of home audio technology, multi-room / multi-channel audio systems can support various listening modes such as watching movies, listening to music, and playing games. Through independent parameter configuration (such as channel grouping, EQ selection, delay adjustment, etc.), they provide basic sound presentation solutions for different scenarios, forming the existing multi-mode audio processing framework.

[0003] As users increasingly demand immersive audio experiences and ease of operation, especially in applications involving multi-device collaboration and rapid switching between multiple scenarios, related audio systems face new challenges in terms of user experience consistency, operational continuity, and system-level collaboration.

[0004] Currently, users face at least one problem: low efficiency in scene switching when managing multiple modes and coordinating parameters across devices in audio systems. Summary of the Invention

[0005] Therefore, it is necessary to provide a method, device, electronic device, and system for scene mode switching in a multi-channel audio system that can improve scene switching efficiency, in order to address the aforementioned technical problems.

[0006] Firstly, this application provides a method for scene mode switching in a multi-channel audio system, the method comprising: Detect the signal status and characteristic events at the audio input; Determine the device capability map of each audio playback device included in the multi-channel audio system; wherein, the device capability map represents the audio playback capability of the audio playback device; Based on signal status and characteristic events, a scene template matching the playback scene to be switched is determined. The scene template is predefined with device grouping strategy and audio processing parameters. Based on the scene template and the device capability map, the target group and audio parameters of the audio playback device are determined, and the scene mode of the multi-channel audio system is switched based on the target group and audio parameters of the audio playback device.

[0007] In one embodiment, detecting the signal state at the audio input includes: The physical link status of at least one of the following interfaces—the high-definition multimedia interface, the audio return channel, and the enhanced audio return channel—is detected as the signal status of the audio input.

[0008] In one embodiment, detecting a characteristic event includes: Detect at least one of the following characteristics of the audio stream encoding format, number of channels, sampling rate, and dynamic range at the audio input end; and / or, Device control commands received via consumer electronics control protocols are detected as characteristic events.

[0009] In one embodiment, the device control instructions include instructions to power on the television, enter a system audio mode, or identify the type of the source device; the method further includes: Upon detecting a device control command, the triggering conditions for scene switching are determined based on predefined entry and exit time thresholds.

[0010] In one embodiment, based on signal state and characteristic events, a scene template matching the playback scene to be switched is determined, including: The logical rules predefined by the rule engine will be applied to signal states and characteristic events; The logical rules map the combination of signal states and / or feature events to the corresponding scene templates.

[0011] In one embodiment, the logic rules are also applied to time information, combining and mapping the time information with signal states and / or characteristic events.

[0012] In one embodiment, the logical rules executed by the rule engine include: If the feature event contains a multi-channel encoding format, then the home theater scene template is triggered; If the feature event contains an identifier instruction from the preset game console, the game scene template is triggered.

[0013] In one embodiment, the scene template includes: At least one of the following: home theater template, music appreciation template, game mode template, nighttime movie viewing template, and news broadcast template.

[0014] In one embodiment, the home theater template's predefined device grouping strategy includes selecting a 5.1 channel group; The home theater template's predefined audio processing parameters include enabling the subwoofer, setting the crossover point to 80Hz, and boosting the gain of the 1kHz to 4kHz frequency band to enhance dialogue.

[0015] In one embodiment, the predefined audio processing parameters of the game mode template include: reducing audio processing buffer latency and disabling dynamic range compression.

[0016] In one embodiment, the predefined audio processing parameters for the nighttime viewing template include: enabling dynamic range compression, attenuating low frequencies, and limiting overall gain.

[0017] In one embodiment, the target grouping of audio playback devices is determined based on a scene template and a device capability map, including: Based on the device capability map, evaluate all channel group combinations supported by currently online audio playback devices; Select the target group from the supported grouping combinations based on the predefined grouping strategy in the scene template.

[0018] In one embodiment, the method further includes: If the key equipment in the target group is offline, the target group will be downgraded to the next level of grouping based on the equipment capability map and grouping strategy.

[0019] In one embodiment, downgrading the target group to the next level of grouping includes: When the physical center channel device is offline, the preset virtual center channel algorithm is activated to group the front left and right channel devices into a virtual multi-channel system.

[0020] In one embodiment, switching the scene mode of the multi-channel audio system includes: Using one audio device in a multi-channel audio system as the master clock source, the clock domains of all audio playback devices within the target group are synchronized. Based on the audio-visual latency characteristics of video display devices, global audio latency compensation is calculated and configured.

[0021] In one embodiment, the method further includes a mode fallback process: If the duration of the loss of a valid signal at the current audio input exceeds a first preset threshold, the multi-channel audio system will be switched to the preset default scene template or the last validly activated non-trigger scene template.

[0022] In one embodiment, switching the scene mode of the multi-channel audio system includes: The parameter switching is completed within a second preset time using crossfading to achieve a smooth parameter transition.

[0023] Secondly, this application also provides a scene mode switching device for a multi-channel audio system, the device comprising: The status and feature detection module is used to detect the signal status and feature events at the audio input terminal. The device capability determination module is used to determine the device capability map of each audio playback device included in the multi-channel audio system; wherein, the device capability map represents the audio playback capability of the audio playback device; The scene template determination module is used to determine the scene template that matches the playback scene to be switched based on the signal status and characteristic events. The scene template is predefined with device grouping strategy and audio processing parameters. The scene mode switching module is used to determine the target group and audio parameters of the audio playback device based on the scene template and the device capability map, and to switch the scene mode of the multi-channel audio system based on the target group and audio parameters of the audio playback device.

[0024] Thirdly, this application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method.

[0025] Fourthly, this application also provides a multi-channel audio system, comprising: Multiple audio playback devices; The electronic device mentioned above is used to control multiple audio playback devices and switch between scene modes of a multi-channel audio system.

[0026] Fifthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.

[0027] The scene mode switching method, device, electronic device, and system for multi-channel audio systems provided in this application can accurately capture the starting point of the user's playback intention by real-time detection of input signal status and characteristic events; by maintaining a dynamic device capability map, the system can clearly grasp the available audio resources; by matching input features with predefined scene templates according to rules, the system can achieve intelligent understanding of the user's complex scene needs; by combining template strategies with actual device conditions, the system can calculate and execute the optimal target grouping and audio parameters, enabling collaborative configuration and seamless switching across multiple audio devices. Compared with related technologies, this application can avoid the tedious, repetitive, and error-prone manual mode management work for users. Users no longer need to manually group devices, turn subwoofers on and off, select EQ presets, and adjust crossover points and dynamic ranges in different APP pages. The system can automatically match the appropriate sound field mode according to the playback content (such as movies and games), and automatically optimize the grouping according to the device's online status (such as enabling a virtual center channel if the center channel is offline), and can even automatically enable parameters that balance experience and user-friendliness at specific times such as nighttime. This improves the efficiency and smoothness of switching between multiple scenes, lowers the operational threshold for users, and achieves an intelligent and consistent auditory experience of "cinema upon startup and matching upon switching." Consequently, it can improve the efficiency of scene switching in terms of multi-mode management and cross-device parameter collaboration of the audio system. Attached Figure Description

[0028] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0029] Figure 1 A flowchart illustrating a scene mode switching method for a multi-channel audio system provided in this application embodiment; Figure 2 A flowchart illustrating the detection of signal state and characteristic events at an audio input terminal, provided as an embodiment of this application; Figure 3 A flowchart illustrating the execution logic rules of a rule engine, provided as an embodiment of this application; Figure 4 This application provides a schematic diagram of a process for determining a target group. Figure 5 A flowchart illustrating a scene mode switching method provided in an embodiment of this application; Figure 6 A schematic diagram of the structure of a scene mode switching device for a multi-channel audio system provided in this application embodiment; Figure 7 This is a schematic diagram of the internal structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0030] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0031] Existing multi-room / multi-channel audio systems often require users to manually perform the following steps in the app when switching between modes such as "movie watching," "music," "gaming," and "night mode": grouping / disbanding satellite or center channels, switching subwoofer output, selecting EQ and crossover, and adjusting gain / delay and dynamic range. Different brands have weak recognition of input sources (HDMI / ARC, optical, Bluetooth, WiFi) and scene linkage, lacking a unified rule engine. This results in slow switching, multiple steps, and a high risk of errors. Furthermore, inconsistent parameters across devices make it difficult to guarantee a consistent experience of immersive sound field and lip-sync.

[0032] In one exemplary embodiment, Figure 1 This is a flowchart illustrating a scene mode switching method for a multi-channel audio system provided in an embodiment of this application, as shown below. Figure 1As shown, a scene mode switching method for a multi-channel audio system is provided. This method is illustrated using a terminal as an example. It is understood that this method can also be applied to a server, or to a system including both a terminal and a server, and is implemented through interaction between the terminal and the server. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, and tablets. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. In this embodiment, the method includes the following steps S101 to S104: Wherein: S101, Detect the signal status and characteristic events of the audio input terminal.

[0033] Here, "audio input" refers to the physical or logical interface through which a channel audio system receives external audio signals, such as HDMI, optical fiber, Bluetooth, or network streaming interfaces. "Signal status" refers to the electrical connection and data link layer status of the interface, such as link lockout or carrier presence. "Characteristic events" refer to specific information or instructions that can be parsed during signal transmission, which may include the format characteristics of the audio stream itself (such as encoding and channel count) and device instructions transmitted through control protocols.

[0034] For example, the method can be executed by control logic deployed on a home theater host (such as WiiM Amp Pro) or a connected smart terminal app. It can continuously monitor the status of all enabled input ports. When the user turns on the TV and selects to output audio via HDMI ARC, the terminal can detect that the physical connection of the HDMI link has been established (signal status), further deduce that the currently transmitted audio stream is Dolby Digital 5.1 encoded, and receive the "System Audio Mode On" command (characteristic event) sent by the TV via the CEC protocol.

[0035] As an example, the TV is turned on and sends a CEC command, with the audio locked at 48kHz / 2.0 PCM. Based on the combination of these signal states and characteristic events, the system determines that it may be entering a cinema scene, thereby triggering subsequent decision-making processes.

[0036] By actively and continuously detecting multidimensional information at the input end, the system can accurately perceive user intent and content changes, providing reliable and timely triggering basis for subsequent automated scenario decisions, thereby improving the response speed or recognition comprehensiveness to changes in the input source.

[0037] S102. Determine the device capability map of each audio playback device included in the multi-channel audio system; wherein, the device capability map represents the audio playback capability of the audio playback device.

[0038] In this context, a multi-channel audio system refers to a system composed of multiple independent audio playback devices (speakers) designed to create a surround sound field. Audio playback devices can refer to specific physical units within the system, such as the left front speaker, center speaker, and subwoofer. A device capability map is a structured data model used to digitally describe the inherent attributes and adjustable parameters of each device, serving as the knowledge base for intelligent system grouping and parameter optimization.

[0039] For example, when a device first joins the network or starts up, the system can query or receive capability parameters reported by each speaker through inter-device communication protocols (such as Wi-Fi, Bluetooth, or proprietary wireless protocols). The terminal can aggregate this information to form a global device capability map, which records, for example: Device A can be used as a left / right front or surround channel, supporting crossover point adjustment range of 60Hz-120Hz; Device B is a dedicated center channel, without crossover function but supporting dialogue enhancement EQ; Device C is an active subwoofer, supporting phase 0 / 180 degree switching, etc.

[0040] As an example, the system can collect the available channel roles, crossover / phase adjustment range, delay, and EQ parameter capabilities of each speaker after the device is powered on or connected to the network, and generate a device capability map based on this to calculate feasible grouping schemes (such as 2.1, 3.1, 5.1).

[0041] By constructing a device capability map, the system can dynamically perceive the real capabilities and status of devices within the network, providing a foundation for subsequent optimal grouping and parameter adaptation based on actual resource conditions. This allows for flexible handling of audio grouping and playback with added or removed devices or mixed models.

[0042] S103. Based on signal status and characteristic events, determine a scene template that matches the playback scene to be switched. The scene template is predefined with device grouping strategy and audio processing parameters.

[0043] The playback scenario refers to the specific context in which a user consumes audio, such as watching a movie, listening to music, or playing a game. Different scenarios have different requirements for sound field effects and parameters. A scenario template can be a predefined set of complete configuration schemes for a specific playback scenario, serving as the fundamental strategy connecting user intent with system execution. Device grouping strategies define which devices should be activated in the aforementioned scenarios and how their channel roles should be assigned. Audio processing parameters can include specific sound effect processing settings such as EQ curves, dynamic range compression, crossover points, delay, and gain.

[0044] For example, when the input signal is detected to be Dolby Atmos encoded in step S101, the system can match this feature event with preset rules. The rule engine determines that the feature meets the triggering conditions of the "home theater" scene, and then calls the corresponding scene template. This template can predefine the priority to try to build a 5.1 or more channel system (grouping strategy), and preset specific audio processing parameters such as improving dialogue clarity (1-4kHz +3dB), enabling the subwoofer and setting the crossover point to 80Hz.

[0045] As an example, the logic for mapping signal status (HDMI) and characteristic events (specific codec formats) to a "home theater" scene template could include: IF input.type == HDMI AND input.codec in {PCM2.0, Dolby5.1} THENscene := THEATER.

[0046] Through predefined scene templates and rule matching mechanisms, the system can instantly transform complex audio signal features into clear scene intentions and specific configuration instructions, enabling the system execution process from "perception" to "understanding," saving users the tedious process of manually recalling and applying complex parameter sets in different scenes.

[0047] S104. Based on the scene template and the device capability map, determine the target group and audio parameters of the audio playback device, and switch the scene mode of the multi-channel audio system based on the target group and audio parameters of the audio playback device.

[0048] Target grouping refers to the optimal combination of devices and role allocation scheme calculated by the system based on the current scene requirements and actual equipment conditions. Audio parameters are the set of processing instructions finally issued to each device after being calculated by the system or selected from templates and fine-tuned for the specific capabilities of the current devices. Switching scene modes refers to the complete process by which the system controls relevant devices to establish communication links according to the target grouping and applies audio parameters to transition the entire system from one sound effect state to another.

[0049] For example, after determining that the system needs to switch to the "Home Theater" template, it queries the current device capability map. Assume the map shows that the left and right front speakers, center speaker, left surround speaker, right surround speaker, and subwoofer are currently online and available. Based on the template's "5.1 priority" strategy, the system assigns roles to these six devices and groups them into target groups. Simultaneously, it fine-tunes parameters in the template, such as dialogue enhancement EQ and subwoofer crossover points, taking into account the specific capabilities of each device (e.g., the subwoofer's actual supported crossover range), to form the final audio parameters. Subsequently, the system sends grouping instructions and parameters to these devices and synchronizes their clocks to complete the mode switch.

[0050] As an example, when the conditions are met, the system can perform "automatic grouping {L,R,C,LS,RS,Sub}, frequency division 80Hz, automatic phase, EQ dialogue +3dB".

[0051] In this way, by combining abstract scene templates with specific device capability maps, a feasible optimal configuration can be generated and all devices can be coordinated to execute synchronously, thus achieving the effect of "one-click triggering, with the system automatically completing all complex settings." This completely avoids the repetitive operations of manually grouping and adjusting parameters one by one, thereby improving the efficiency and accuracy of multi-channel audio scene switching.

[0052] In this embodiment, by real-time detection of input signal status and characteristic events (S101), the system can accurately capture the starting point of the user's playback intention; by maintaining a dynamic device capability map (S102), the system can clearly grasp the available audio resources; by matching input features with predefined scene templates according to rules (S103), the system can achieve intelligent understanding of the user's complex scene needs; by combining template strategies and device status, calculating and executing the optimal target grouping and audio parameters (S104), the system can complete collaborative configuration and seamless switching across multiple audio devices. Compared with related technologies, this application can avoid the tedious, repetitive, and error-prone manual mode management work for users. Users no longer need to manually group devices, turn subwoofers on and off, select EQ presets, and adjust crossover points and dynamic ranges in different APP pages. The system can automatically match the appropriate sound field mode according to the playback content (such as movies and games), and automatically optimize the grouping according to the device's online status (such as enabling a virtual center channel if the center channel is offline), and can even automatically enable parameters that take into account both experience and user-friendliness at specific times such as night. This improves the efficiency and smoothness of switching between multiple scenes, lowers the operational threshold for users, and achieves an intelligent and consistent auditory experience of "cinema upon startup and matching upon switching." Consequently, it can improve the efficiency of scene switching in terms of multi-mode management and cross-device parameter collaboration of the audio system.

[0053] In one exemplary embodiment, such as Figure 2 As shown, Figure 2 This application provides a flowchart illustrating the detection of signal state and characteristic events at an audio input terminal, which can be implemented in various ways. Figure 1 Based on this, the steps for scene mode switching in a multi-channel audio system are illustrated below. Specifically, step S101, detecting the signal state at the audio input terminal, may include: S201. Detect the physical link status of at least one of the following interfaces: the high-definition multimedia interface, the audio return channel, and the enhanced audio return channel, as the signal status of the audio input terminal.

[0054] Among them, the High-Definition Multimedia Interface (HDMI), Audio Return Channel (ADR), and Enhanced Audio Return Channel (ERC) are industry-standard interfaces used for transmitting high-definition audio and video signals between devices such as televisions, players, and amplifiers. Physical link status refers to the connectivity, synchronization, and locking status of these interfaces at the hardware connection level, and can be the fundamental basis for determining whether a signal is being transmitted effectively.

[0055] For example, the system's main controller or dedicated HDMI chip can continuously monitor the hot-plug detection pin level of the HDMI / ARC / eARC port, the clock lock status of the TMDS data channel, and the carrier detection of ARC. When audio is detected being sent back from the TV via the eARC interface, the system can confirm that the physical link of eARC has been successfully established and locked. The eARC interface establishment state is a stable signal state, indicating the possibility of high-quality audio signal input.

[0056] As an example, it can detect the status of HDMI / ARC / eARC signals and monitor carrier lock and EDID handshake on the HDMI port.

[0057] In this embodiment, by accurately detecting the physical link status of key audio and video interfaces, different situations such as "no signal input", "unstable signal" and "stable signal input" can be reliably distinguished, providing a basis for subsequent scene judgment based on signal characteristics, avoiding false triggering caused by link flickering or incompetence, thereby improving the accuracy of scene switching.

[0058] In one exemplary embodiment, such as Figure 2 As shown, step S101, detecting feature events, can specifically include: S202, Detect at least one of the following characteristics of the audio stream encoding format, number of channels, sampling rate, and dynamic range at the audio input end; and / or, S203. Detect device control commands received through the consumer electronics control protocol as characteristic events.

[0059] Among these, audio stream encoding format, number of channels, sampling rate, and dynamic range characteristics can refer to metadata describing the content attributes of the audio stream itself. Consumer electronics control protocol refers to a communication standard that allows control commands to be sent between connected devices via HDMI cable. Device control commands refer to standardized messages transmitted via the CEC protocol for controlling device behavior (such as power on / off, switching input sources).

[0060] For example, after confirming that the physical link is working, the system can further parse the transmitted data packets. The audio decoder or analysis module will extract metadata from the stream, such as determining that it is "Dolby Digital Plus 7.1" encoded and has a 48kHz sampling rate. Simultaneously, the CEC controller can parse commands transmitted over the cable, such as receiving a "..." message.<Image ViewOn> The command indicates that the television is turned on.

[0061] As an example, the detection of characteristic events includes: detecting "audio format and number of channels" and "CEC commands (such as TV_ON, System Audio Mode)".

[0062] In this embodiment, by deeply analyzing the audio stream content and the smart device's linkage commands, it can not only know that "there is a signal," but also "what kind of signal" and "what the signal source device is doing." This enriches the basis for scene determination, enabling the system to distinguish between different content such as movies, concerts, and games, and respond to user behaviors such as turning the TV on and off, thereby achieving more accurate and user-friendly automatic scene switching.

[0063] In one exemplary embodiment, the device control instructions include instructions to power on the television, enter a system audio mode, or identify the type of the source device; the method further includes: Upon detecting a device control command, the triggering conditions for scene switching are determined based on predefined entry and exit time thresholds.

[0064] The entry time threshold refers to the minimum time required for the system to maintain a stable state for a certain duration after detecting a trigger signal before executing a switching action; this is used for debouncing. The exit time threshold refers to the minimum time required for the system to wait for a certain period of time after detecting the disappearance of a trigger signal before executing an exit or rollback action; this is used to prevent frequent mode switching caused by brief interruptions.

[0065] For example, when the system detects a television sending a message such as "" via CEC<System Audio Mode On> "When a command is issued, switching to cinema mode does not have to be performed immediately. A timer (e.g., 800ms) can be started, and the trigger can be confirmed to be valid after the command state has been held for more than the entry time threshold. Similarly, when the TV is turned off and the HDMI signal is lost, the system can wait for an exit time threshold (e.g., 2 seconds) to confirm that the signal is not a brief interruption before performing operations such as reverting to music mode."

[0066] As an example, entry thresholds (lock ≥ 800 ms) and exit thresholds (unlock ≥ 2 s) can be used to prevent accidental touches during mode switching, thereby achieving anti-interference for mode switching.

[0067] In this embodiment, a confirmation method based on time thresholds is introduced, which can effectively filter out false detections and false triggers caused by device handshakes, signal transients, or accidental interference, thereby improving the stability and reliability of the system's automatic switching.

[0068] In an exemplary embodiment, based on signal state and characteristic events, a scene template matching the playback scene to be switched is determined, including: The logical rules predefined by the rule engine will be applied to signal states and characteristic events; The logical rules map the combination of signal states and / or feature events to the corresponding scene templates.

[0069] In this context, a rule engine can refer to a software component that is responsible for executing a set of "if-then" logical rules. These rules can associate input conditions (a specific combination of signal states and feature events) with output actions (selecting a scene template).

[0070] For example, the rule engine can store a rule like this: IF (input_interface == HDMI_ARC) AND (audio_format == Dolby_Atmos) THEN (load_template == Home_Theater). This means that when the data reported by the input sensing module meets the requirements of "HDMI ARC interface" and has "Dolby Atmos audio stream", the rule engine is triggered, and the output decision is to load the "Home Theater" template.

[0071] In this embodiment, by introducing a rule engine, the system can modularize and configure complex scene judgment logic. This makes the scene matching strategy clear, maintainable, and easily extensible. Scene switching logic can be optimized or enhanced by adding, deleting, or modifying rules without altering the core program code, thereby improving the system's flexibility and scalability.

[0072] In one exemplary embodiment, the logic rules can also be applied to time information, combining and mapping the time information with signal states and / or characteristic events.

[0073] The time information refers to the system's current clock information, which can include specific times of day. Including this as part of the rule conditions allows for automated behavior based on time periods.

[0074] For example, a rule can be defined in the rule engine: IF (current_time BETWEEN 22:00 AND 06:00) AND (active_template == Home_Theater) THEN (apply_subtemplate == Night_Mode). This means that even in home theater mode, if the system time is during nighttime, parameters that enable the "nighttime subtemplate" (such as compressing dynamic range and reducing low frequencies) can be automatically applied.

[0075] As an example, when night mode is enabled automatically, the preset condition can be 22:00-06:00, or the user can manually switch it, and the time information (22:00-06:00) can be used as one of the trigger conditions.

[0076] In this embodiment, time information can be incorporated into the decision-making dimension of the rule engine, enabling the system to make more intelligent adjustments based on the context. This allows it to automatically adapt to specific user needs at different times (such as not disturbing neighbors at night), achieving a progression from "content-based" automation to "context-based" intelligence. Furthermore, it can reduce the need for manual user intervention and improve scene switching efficiency.

[0077] In one exemplary embodiment, such as Figure 3 As shown, Figure 3 This application provides a flowchart illustrating the execution logic rules of a rule engine, wherein the logic rules executed by the rule engine may specifically include: S301. If the feature event contains a multi-channel encoding format, then the home theater scene template is triggered. S302. If the feature event contains an identification instruction from the preset game console, then the game scene template is triggered.

[0078] Multichannel encoding formats refer to audio encoding that includes information from more than two independent channels, such as Dolby Digital 5.1 and DTS-HD MA 7.1, and are often characteristics of movie content. Preset game console identification instructions refer to specific device manufacturer IDs or behavioral patterns that can be identified through protocols such as CEC.

[0079] For example, when the audio analysis module detects that the input stream is in "DTS:X" format (a multi-channel immersive audio format), the rule engine can match the "multi-channel encoding" condition and trigger the loading process of the home theater template. When the CEC module parses the source device's manufacturer ID as "a certain brand" and the behavior pattern matches the characteristics of a game console, the rule engine triggers the game scene template.

[0080] In this embodiment, the two rules mentioned above can be scene judgment rules in the rule engine, which can accurately indicate the most unique signal characteristics of the two core scenes of "watching movies" and "playing games," enabling fast and accurate scene recognition. This ensures that when a user plays a movie or starts a game console, the system can instantly enter the optimal sound effect state, improving the user's immersion and experience consistency, thereby effectively improving scene switching efficiency.

[0081] In one exemplary embodiment, the scene template includes: At least one of the following: home theater template, music appreciation template, game mode template, nighttime movie viewing template, and news broadcast template.

[0082] The Home Theater template is optimized for movie watching, emphasizing surround sound, clear dialogue, and powerful bass. The Music Appreciation template focuses on high-fidelity stereo reproduction, pursuing original timbre and a wide soundstage. The Gaming Mode template prioritizes ultra-low audio latency and accurate sound positioning. The Nighttime Movie Viewing template, a derivative of the Home Theater template, preserves key auditory qualities while compressing dynamics and limiting low frequencies to reduce the risk of disturbing neighbors. The News Broadcast template drastically enhances speech clarity and simplifies channels to highlight human voices.

[0083] For example, the system can pre-configure the five commonly used templates mentioned above. Each template can be a complete data structure, containing a full set of preset parameters, from device grouping preferences, EQ curves, frequency division points, dynamic range processing, delay strategies to CEC control strategies. When the rules engine decides to switch to a certain scenario, it calls the corresponding template data.

[0084] As an example, specific strategies and parameters can be defined for these five scene templates (Home Theater, Music, Game, Night, News), including: The system maintains multiple scenario templates, each containing automatic grouping strategies, EQ, dynamic range (DRC), delay parameters, etc.

[0085] (1) Home Theater Sense Template: Automatic grouping: Prioritize 5.1, then 3.1, and then 2.1; Subwoofer enabled, initial crossover 80Hz, automatic phase; EQ: Dialogue enhancement +2~+4dB (1–4kHz); Delay: Automatic calibration based on lip-sync.

[0086] (2) Music Templates: Primarily based on 2.0 or 2.1 grouping; EQ can use a balanced or user-defined curve; The CEC control is disabled, and the main features are high fidelity and low noise.

[0087] (3) Game template: Prioritize low latency and shorten the filter tail and buffer; Turn off DRC and long reverb; If CEC recognizes the PlayStation / Xbox vendor ID, it will be enabled automatically.

[0088] (4) Nighttime template: Enable DRC compression and 3dB low-frequency attenuation; Maintain dialogue enhancement; Automatically lowers the overall volume to prevent disturbing neighbors.

[0089] (5) News Template: Enhance the voice range of 1–4kHz, and limit low frequencies to <150Hz; The EQ uses a "clear voice" curve; Turn off the surround output and keep only the front channels on.

[0090] Rule engine example: IF input.type == HDMI AND input.codec in {PCM2.0, Dolby5.1} THEN scene := THEATER, enable_subwoofer := true,EQ := "MovieVoice+" The system can dynamically switch templates based on input signals and user preferences to achieve automation in multiple scenarios.

[0091] In this embodiment, by pre-setting these five categories of templates covering mainstream home audio consumption scenarios, the system can meet the daily needs of the vast majority of users. The template-based design allows each experience to be carefully tuned, enabling users to obtain near-professional-level listening optimization, avoiding the operational difficulties caused by numerous complex parameters, achieving a professional experience, and thus further improving the scene switching effect of the multi-channel audio system.

[0092] In one exemplary embodiment, the home theater template's predefined device grouping strategy includes selecting a 5.1 channel group; The home theater template's predefined audio processing parameters include enabling the subwoofer, setting the crossover point to 80Hz, and boosting the gain of the 1kHz to 4kHz frequency band to enhance dialogue.

[0093] In this context, a 5.1 channel configuration can refer to a standard home theater speaker layout that includes a left front speaker, a right front speaker, a center speaker, left surround speakers, right surround speakers, and a low-frequency effects channel. The crossover point refers to the critical frequency that determines which frequencies are played by the main speakers and which by the subwoofer. The 1kHz to 4kHz frequency band refers to the main area where the energy of human speech (dialogue) is concentrated.

[0094] For example, when the "Home Theater" template is loaded, the system can attempt to group available devices according to a 5.1 configuration. Simultaneously, the subwoofer can be instructed to turn on, and the crossover frequency between the main speakers and the subwoofer can be set to 80Hz. In terms of EQ processing, the system can apply a wideband boost in this frequency band, making vocals more prominent, clearer, and easier to hear movie dialogue.

[0095] As an example, for the "Home Theater Template", the automatic grouping can prioritize 5.1, then 3.1, and then 2.1; enable the subwoofer, initial crossover 80Hz, automatic phase; EQ: dialogue enhancement +2~+4dB (1–4 kHz).

[0096] In this embodiment, the template definition transforms the core elements of a home theater experience (surround sound field, solid low frequencies, and clear dialogue) into specific, executable system configurations. When a user starts a movie, the system can apply these optimized parameters, eliminating the need for the user to manually set the subwoofer, adjust the crossover, or search for dialogue enhancement options, thus achieving an immersive viewing experience in one step and improving efficiency from the start of watching to achieving a satisfactory listening experience.

[0097] In one exemplary embodiment, the predefined audio processing parameters of the game mode template include: reducing audio processing buffer latency and disabling dynamic range compression.

[0098] Audio processing buffer latency refers to the time lag introduced when audio signals are processed within the system (such as decoding, effects processing, and resampling). Dynamic range compression is an audio processing technique used to reduce the difference between the loudest and quietest parts of a sound. It is often used in movies to ensure clear dialogue while minimizing the loudness of explosions, but this can result in a loss of impact and realism.

[0099] For example, in game mode, the system can adjust the settings of the audio processing pipeline, such as reducing the length of the resampling filter and using smaller data buffer blocks, to minimize the total latency from the sound emitted from the game console to its playback from the speakers. Simultaneously, DRC processing can be disabled to ensure that gunshots, explosions, and ambient sounds in the game have their original dynamic impact, enhancing the sense of presence and the clarity of responsive cues.

[0100] As an example, for the "game template", low latency can be prioritized, the filter tail and buffer shortened, and DRC and long reverb turned off.

[0101] In this embodiment, by specifically optimizing latency and dynamic range, the game template can address two key metrics in game scenarios: audio-visual synchronization and the realistic impact of sound. Players no longer need to search for a "low latency mode" in complex audio settings menus or worry about compressed explosion sounds; the system can provide the most suitable audio configuration for the game, thereby further improving the scene switching efficiency of the multi-channel audio system.

[0102] In one exemplary embodiment, the predefined audio processing parameters for the nighttime movie-watching template include: enabling dynamic range compression, attenuating low frequencies, and limiting overall gain.

[0103] Enabling dynamic range compression can reduce the volume difference between sudden, high-dynamic-range sound effects (such as explosions or impacts) and quiet dialogue in movies, preventing disturbance to the user. Attenuating low frequencies can reduce the energy of ultra-low-frequency effects, as low-frequency sounds have strong penetrating power and easily propagate through building structures. Limiting overall gain can set a maximum volume limit to prevent excessive noise caused by users unintentionally increasing the volume.

[0104] For example, when night mode is activated, the system can enable the DRC processor to appropriately reduce large peak volumes. Simultaneously, in the EQ settings, roll-off attenuation can be applied to frequencies below, for example, 80Hz or 100Hz. Furthermore, the system can override the user's volume settings to ensure that the final output maximum sound pressure level does not exceed a preset safety value.

[0105] As an example, for the "Night Template", DRC compression and low-frequency attenuation of 3dB can be enabled; dialogue enhancement can be maintained; and the overall volume can be automatically reduced to prevent disturbing others.

[0106] In this embodiment, the nighttime movie-watching template intelligently implements a noise-free strategy while preserving the core auditory experience of the movie (such as clear dialogue). Users do not need to manually lower the volume or subwoofer while watching movies late at night, or constantly worry about sudden loud noises. The system can automatically maintain a balanced and reassuring auditory level, allowing users to relax and enjoy the content without being distracted by the potential impact of sound. This improves the reliability of scene mode switching in multi-channel audio systems.

[0107] In one exemplary embodiment, such as Figure 4 As shown, Figure 4 This application provides a flowchart illustrating the process of determining a target group, which can be implemented in accordance with embodiments of the present application. Figure 1 Based on this, the steps for scene mode switching in a multi-channel audio system are illustrated below. In step S104, the target group of the audio playback devices is determined according to the scene template and the device capability map. Specifically, this may include: S401. Based on the device capability map, evaluate all channel group combinations supported by currently online audio playback devices.

[0108] S402. Select the target group from the supported grouping combinations according to the predefined grouping strategy in the scene template.

[0109] Evaluating all channel grouping combinations refers to the system calculating all possible device-channel allocation schemes based on the role each device can play (e.g., a speaker can be used as a front or surround speaker) recorded in the device capability map. Predefined grouping strategies can be part of a scene template, specifying grouping preferences for that scene, such as "5.1 preferred, 3.1 second, 2.1 third." Selecting the target group is the process by which the system sorts and selects feasible schemes according to the strategy to arrive at the final execution plan.

[0110] For example, the current system has left and right front speakers (A, B), a center speaker (C), a subwoofer (Sub), and two surround speakers (D, E). The device capability map shows that A, B, C, D, E, and Sub are all online. The system calculates feasible combinations including: 5.1 (A, B, C, D, E, Sub), 3.1 (A, B, C, Sub), 2.1 (A, B, Sub), etc. If the currently active template is "Home Theater," and its policy is "5.1 preferred," then the system selects 5.1 as the target group.

[0111] As an example, the system can select the highest-scoring combination from the feasible set based on the current template and DCG.

[0112] In this embodiment, by dynamically calculating and selecting groups according to a strategy, the system can achieve optimal resource allocation. It can fully utilize all currently online devices to provide users with the best channel configuration under current conditions. This avoids the severe performance degradation that occurs in fixed-group systems when some devices are temporarily offline, or the slow response and complex operation of manual regrouping by users, ensuring the continuity and flexibility of mode switching. Therefore, it can improve the efficiency and reliability of scene mode switching in multi-channel audio systems.

[0113] In one exemplary embodiment, the system can quantitatively evaluate different grouping schemes using a device grouping optimization scoring algorithm. This algorithm comprehensively considers device capabilities, spatial location, and scene matching to calculate a comprehensive score and select the optimal grouping. For example, different grouping schemes can be quantitatively evaluated using the following formula: (1); Where S is the overall score (0-100); C is the equipment capability score; P is the spatial location adaptability; M is the scene matching degree; α, β, and γ are all weight coefficients (summing to 1); Equipment capability rating: (2); Where F is the frequency response score, D is the dynamic range score, and L is the delay performance score.

[0114] Spatial location adaptability: (3); Where θᵢ is the actual angle, θᵢ* is the ideal angle, and σ is the tolerance parameter (typically 15°).

[0115] Scene matching degree: (4); Where, N match To meet the required number of devices, N ideal This represents the ideal number of devices.

[0116] As an example, the system selects a 5.1 channel group for "home theater" with 6 devices currently online. Set α=0.4, β=0.3, γ=0.3. According to the above formulas (1)-(4), we can calculate: C≈2 points (overall device capability); P≈80.1 points (right surround deviation 10°); M=100 points (6 / 6 devices meet the requirements); S=0.4×82+0.3×80.1+0.3×100=86.83 points; In this embodiment, subjective judgments are transformed into objective values, enabling automated decision-making; multi-dimensional comprehensive evaluation ensures full optimization of the grouping scheme; it can dynamically adapt to equipment changes and achieve system self-healing; and it achieves millisecond-level scoring, significantly improving scene switching efficiency.

[0117] In one exemplary embodiment, the method further includes: If the key equipment in the target group is offline, the target group will be downgraded to the next level of grouping based on the equipment capability map and grouping strategy.

[0118] Among them, preset key equipment refers to equipment that is crucial to the experience in a specific scenario. For example, in a home theater scenario, the center speaker (responsible for dialogue) and the subwoofer (responsible for low-frequency effects) can be considered key equipment. Degradation refers to the system automatically switching to a suboptimal alternative grouping scheme that can still provide the core experience of the scenario when the optimal grouping scheme cannot be implemented due to equipment shortage.

[0119] For example, if the system is already running in 5.1 grouping, and the left surround speaker E goes offline due to a power outage, the system can detect this change and query the device capability map and the grouping strategy of the "Home Theater" template (5.1 preferred, followed by 3.1, etc.). Since 5.1 is no longer feasible, the system can automatically perform a downgrade, switching the target group to 3.1 (A, B, C, Sub) and reconfiguring the audio signal splitting and parameters.

[0120] As an example, when a critical device is offline, the system can automatically downgrade to a virtual hub or a 2.1 combination, or, if the Sub is offline, downgrade to 3.1.

[0121] In this embodiment, the introduction of an automatic degradation mechanism enhances system robustness and user experience continuity. Temporary device failures, network fluctuations, or power issues prevent system-wide shutdowns or the need for emergency user intervention. The system can smoothly and automatically adapt to changes in device status, providing the best possible auditory experience within available resources, achieving effective degradation and thus improving system reliability and user satisfaction.

[0122] In one exemplary embodiment, downgrading a target group to the next level of grouping includes: When the physical center channel device is offline, the preset virtual center channel algorithm is activated to group the front left and right channel devices into a virtual multi-channel system.

[0123] The virtual center channel algorithm refers to an algorithm that uses digital signal processing technology and the acoustic interference principle of the left and right front speakers to "virtually" create a stable center channel image at the listening position. A virtual multi-channel system refers to a system that uses such algorithms to simulate a multi-channel surround sound effect with fewer physical speakers than the standard number.

[0124] For example, when the physical center speaker is detected to be offline, and the current scenario requires a center channel (such as a home theater), the system does not simply revert to 2.1 stereo. Instead, the system can enable a virtual center channel algorithm. This algorithm performs special amplitude and phase processing on the signals sent to the left and right front speakers, making human dialogue sound as if it is still coming from the center of the screen, thus achieving a virtual 3.0 or 3.1 listening experience on the basis of a 2.0 or 2.1 physical system.

[0125] As an example, the system can automatically downgrade to a virtual center, enable the virtual center algorithm, and adaptively divide the frequency between 70 and 90 Hz.

[0126] In this embodiment, the virtualization algorithm is enabled when key equipment is missing, which avoids the problem of the entire surround sound collapsing due to the absence of a single speaker, and preserves the core auditory features of the scene (such as dialogue positioning) to the greatest extent. This allows users to still obtain an immersive experience far exceeding that of ordinary stereo when they do not have the conditions to configure a complete physical multi-channel system or encounter temporary equipment problems, thereby improving the practical value of the system and user satisfaction.

[0127] In one exemplary embodiment, Figure 5 This is a flowchart illustrating a scene mode switching method provided in an embodiment of this application, as shown below. Figure 5 As shown, it is possible to Figure 1 Based on this, the steps for switching scene modes in a multi-channel audio system are illustrated below. Specifically, step S104, switching the scene mode of the multi-channel audio system, may include: S501: Use one audio device in a multi-channel audio system as the master clock source to synchronize the clock domains of all audio playback devices in the target group. S502: Calculate and configure global audio delay compensation based on the audio-visual delay characteristics of the video display device.

[0128] In this context, the master clock source refers to the device selected as the system time base, and other devices need to maintain clock synchronization with it to prevent audio playback stuttering or noise. Clock domain synchronization refers to the process of ensuring that all audio devices in the network operate under a unified time base. Audio-visual latency characteristics refer to the inherent latency introduced by video display devices (such as televisions) in processing image signals (such as frame interpolation and motion compensation). Global audio latency compensation refers to artificially adding a corresponding delay to the entire audio path to match video latency and achieve lip-sync.

[0129] For example, after determining the target group, the system can designate the host or one of the speakers supporting master clock functionality as the master clock source and send synchronization signals to other devices in the group via a network protocol (such as PTP). Simultaneously, the system can determine the current audio-visual latency of the television (e.g., 40 milliseconds) from a pre-stored database or through user calibration. Subsequently, the system inserts a 40-millisecond delay into the audio processing link to ensure that sound and picture arrive at the user simultaneously.

[0130] As an example, the host acts as the clock master, and all speakers can be synchronized according to a unified clock domain. The system can calculate global latency compensation based on the television video delay.

[0131] In this embodiment, precise clock synchronization and global delay compensation are implemented to improve sound synchronization and lip-sync. This ensures that all sounds, whether surround sound or dialogue, are played in perfect synchronization and perfectly match the visuals. It eliminates the asynchrony caused by wireless transmission and television processing, providing a stable and immersive audiovisual experience comparable to or even better than wired systems.

[0132] In one exemplary embodiment, the system may employ an adaptive audio latency dynamic compensation algorithm to calculate compensation values ​​in real time based on the transmission path, device processing latency, and network jitter, ensuring accurate audio-visual synchronization.

[0133] For example, global delay compensation can be performed using the following formula (5): (5); Among them, T comp For global compensation (ms); T video For video delay; T net For network latency; T proc For audio processing delay; T buffer This is for buffering delay.

[0134] Network latency dynamics are estimated using formula (6): (6); Where RTT is the round-trip time, α emaThis is the filter coefficient (typical value 0.8).

[0135] Compensation for equipment differences is achieved through formula (7): (7); Among them, T trans(i) For transmission delay, T render(i) This is for rendering delay.

[0136] The jitter buffer is adjusted using formula (8): (8); Where, μ delay Let σ be the mean of the delay. delay k represents the standard deviation. σ This is the confidence level coefficient (typical value 2.5).

[0137] As an example, users can watch 4K movies via HDMI eARC, where T is known to be the key. video =45ms, T proc =8ms, T buffer =5ms, estimate T net ≈12.3ms.

[0138] Calculate global compensation: T comp =45+12.3-(8+5)=44.3ms; Differential compensation was calculated for each device: front speaker 28.3ms, center speaker 26.3ms, surround speakers 22.3ms, subwoofer 25.3ms.

[0139] Jitter buffer (μ=12.5ms, σ=3.2ms): B_jitter=12.5+2.5×3.2=20.5ms; In this embodiment, the "lip-syncing" problem can be accurately eliminated, improving the viewing experience; real-time adaptive network status eliminates the need for manual calibration; multi-device differentiated processing achieves high-precision synchronization; anti-network jitter improves system stability; and fully automated compensation reduces operating costs and improves scene switching efficiency.

[0140] In one exemplary embodiment, the method further includes a mode fallback process: If the duration of the loss of a valid signal at the current audio input exceeds a first preset threshold, the multi-channel audio system will be switched to the preset default scene template or the last validly activated non-trigger scene template.

[0141] The "mode fallback process" refers to the process by which the system automatically returns to a normal or user-preferred state after the signal that triggered the automatic scene disappears. "Loss of valid signal" refers to the audio input returning to a state without a stable signal lock. The "first preset threshold" refers to the aforementioned exit time threshold, used to confirm that the signal has truly disappeared rather than being a brief interruption. "Non-triggered scene templates" refer to scenes that are not triggered by a specific external signal (such as an HDMI movie signal), but can be manually selected by the user or loaded by the system by default, such as the "music" mode.

[0142] For example, after watching TV (HDMI ARC input), the user turns off the TV. The system can detect the loss of the HDMI signal and start a timer. After more than 2 seconds (a first preset threshold), it confirms that the TV is off. At this point, the system can perform a mode rollback process: automatically disbanding the surround sound group assembled in Cinema mode, turning off the subwoofer, and switching the system back to the user's most frequently used "Music Appreciation" template (the preset default scene template).

[0143] As an example, mode switching and rollback strategies could include: "Automatically rollback to the previous active template, such as music or news mode, if HDMI is unlocked for more than 2 seconds".

[0144] In this embodiment, the mode rollback process enhances the automation experience. It allows the system to intelligently switch modes not only when the user starts watching a movie, but also when the user finishes watching, automatically reverting to normal mode. Users no longer need to manually switch back to music mode or disable the speakers after watching a movie, achieving seamless automation throughout the entire process and further improving the scene mode switching efficiency of multi-channel audio systems.

[0145] In one exemplary embodiment, switching the scene mode of a multi-channel audio system includes: The parameter switching is completed within a second preset time using crossfading to achieve a smooth parameter transition.

[0146] Crossfading can refer to a technique used in audio processing for smooth transitions. During a switch, the effect of the old parameter gradually weakens while the effect of the new parameter gradually strengthens, with the two overlapping over a period of time to avoid abrupt changes. The second preset time refers to the duration of this gradual transition, usually very short (tens of milliseconds), so that the human ear is not easily aware of the switch process, only perceiving the change in the final result.

[0147] For example, when the system switches from "Music" mode to "Cinema" mode, the EQ curves, dynamic range processing, and other parameters of the two modes may differ significantly. A direct switch would cause a sudden and drastic change in timbre and loudness, potentially producing audible pops or sound image jumps. By using crossfading, the system can linearly decay the effect of the "Music" EQ to zero within 80 milliseconds, while simultaneously linearly boosting the effect of the "Cinema" EQ from zero to its full value, achieving a seamless and smooth auditory transition.

[0148] As an example, the system can perform a smooth transition of parameters during switching (crossfade ≤ 80 ms) to avoid pops or sound image drift.

[0149] In this embodiment, by employing cross-fading technology for smooth parameter transitions, the system can eliminate any auditory "click" sounds, sudden pitch changes, or sound field jumps that may occur during scene switching. This makes automatic mode switching extremely smooth and imperceptible, like a fade-in / fade-out effect, thereby improving the overall system completeness and, consequently, enhancing the reliability of scene mode switching in a multi-channel audio system.

[0150] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0151] In this application, through the aforementioned embodiments, a complete closed-loop process of "detection—decision-grouping—optimization—synchronization—rollback" is constructed through input signal detection, device capability mapping, and multi-scene template linkage. For HDMI home theater and multimedia playback scenarios, the system can automatically switch between modes such as "home theater, music, game, night, and news," achieving adaptive grouping, EQ, and latency optimization of satellite / center channels and subwoofers. This significantly improves the automation and consistency of home audio systems, truly realizing an intelligent experience of "theater on startup, matching upon switching," and possesses significant patent protection and product commercialization value.

[0152] The following describes the scene mode switching device for a multi-channel audio system provided in the embodiments of this application. The scene mode switching device for a multi-channel audio system has the same inventive concept as the scene mode switching method for a multi-channel audio system described above. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations of one or more scene mode switching device embodiments for a multi-channel audio system provided below can be referred to the limitations of the scene mode switching method for a multi-channel audio system described above. The scene mode switching device for a multi-channel audio system described below and the scene mode switching method for a multi-channel audio system described above can be referred to each other, and will not be repeated here.

[0153] In one exemplary embodiment, Figure 6 This application provides a schematic diagram of the structure of a scene mode switching device for a multi-channel audio system, as shown in the embodiments of this application. Figure 6 As shown, the scene mode switching device 60 of the multi-channel audio system includes: a status and feature detection module 610, a device capability determination module 620, a scene template determination module 630, and a scene mode switching module 640, wherein: The status and feature detection module 610 is used to detect the signal status and feature events of the audio input terminal.

[0154] The device capability determination module 620 is used to determine the device capability map of each audio playback device included in the multi-channel audio system; wherein, the device capability map represents the audio playback capability of the audio playback device.

[0155] The scene template determination module 630 is used to determine the scene template that matches the playback scene to be switched based on the signal state and characteristic events. The scene template is predefined with device grouping strategy and audio processing parameters.

[0156] The scene mode switching module 640 is used to determine the target group and audio parameters of the audio playback device based on the scene template and the device capability map, and to switch the scene mode of the multi-channel audio system based on the target group and audio parameters of the audio playback device.

[0157] In an exemplary embodiment, the status and feature detection module 610 is used to detect the status of at least one of the high-definition multimedia interface, the audio return channel, and the enhanced audio return channel, as the signal status of the audio input terminal.

[0158] In an exemplary embodiment, the status and feature detection module 610 is used to detect at least one of the following features of the audio stream encoding format, number of channels, sampling rate, and dynamic range of the audio input terminal; and / or, to detect device control commands received through a consumer electronics control protocol as feature events.

[0159] In an exemplary embodiment, the device control command includes instructions to turn on the TV, enter the system audio mode, or identify the type of the source device; the scene template determination module 630 is used to confirm the triggering conditions for scene switching based on predefined entry time thresholds and exit time thresholds when the device control command is detected.

[0160] In an exemplary embodiment, the scene template determination module 630 is used to apply predefined logical rules through a rule engine to signal states and feature events; wherein the logical rules map the feature combination of signal states and / or feature events to the corresponding scene template.

[0161] In an exemplary embodiment, the logical rules in the scene template determination module 630 are also applied to time information, combining and mapping the time information with signal states and / or feature events.

[0162] In an exemplary embodiment, the scene template determination module 630 is used to trigger a home theater scene template if the feature event contains a multi-channel encoding format, and to trigger a game scene template if the feature event contains an identification instruction from a preset game console.

[0163] In one exemplary embodiment, the scene template includes at least one of the following: home theater template, music appreciation template, game mode template, nighttime movie viewing template, and news broadcast template.

[0164] In one exemplary embodiment, the home theater template's predefined device grouping strategy includes selecting a 5.1 channel group; the home theater template's predefined audio processing parameters include enabling the subwoofer, setting the crossover point to 80Hz, and boosting the gain of the 1kHz to 4kHz frequency band to enhance dialogue.

[0165] In one exemplary embodiment, the predefined audio processing parameters of the game mode template include: reducing audio processing buffer latency and disabling dynamic range compression.

[0166] In one exemplary embodiment, the predefined audio processing parameters for the nighttime movie-watching template include: enabling dynamic range compression, attenuating low frequencies, and limiting overall gain.

[0167] In an exemplary embodiment, the scene mode switching module 640 is used to evaluate all channel grouping combinations supported by the currently online audio playback device based on the device capability map; and select a target group from the supported grouping combinations according to the predefined grouping strategy in the scene template.

[0168] In an exemplary embodiment, the scene mode switching module 640 is used to downgrade the target group to the next level of grouping combination when the preset key equipment in the target group is offline, based on the equipment capability map and grouping strategy.

[0169] In an exemplary embodiment, the scene mode switching module 640 is used to enable a preset virtual center channel algorithm when the physical center channel device is offline, and group the front left and right channel devices into a virtual multi-channel system.

[0170] In an exemplary embodiment, the scene mode switching module 640 is used to use an audio device in a multi-channel audio system as the master clock source to synchronize the clock domains of all audio playback devices in the target group; and to calculate and configure global audio delay compensation based on the audio-visual delay characteristics of the video display device.

[0171] In an exemplary embodiment, the scene mode switching module 640 is used to switch the multi-channel audio system to a preset default scene template or the last validly activated non-trigger scene template when the duration of the loss of a valid signal at the current audio input terminal exceeds a first preset threshold.

[0172] In an exemplary embodiment, the scene mode switching module 640 is used to complete the parameter switching within a second preset time using crossfading to achieve a smooth parameter transition.

[0173] Each module in the scene mode switching device of the aforementioned multi-channel audio system can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of the electronic device in hardware form or independent of it, or stored in the memory of the electronic device in software form, so that the processor can call and execute the corresponding operations of each module.

[0174] In one exemplary embodiment, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the scene mode switching method of any of the multi-channel audio systems described above.

[0175] In one exemplary embodiment, this application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the scene mode switching method of any of the multi-channel audio systems described above.

[0176] In one exemplary embodiment, a multi-channel audio system is provided, comprising: Multiple audio playback devices; The electronic device mentioned above is used to control multiple audio playback devices and switch between scene modes of a multi-channel audio system.

[0177] In one exemplary embodiment, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the scene mode switching method of any of the multi-channel audio systems described above.

[0178] Indicatively, such as Figure 7 As shown, Figure 7 This is a schematic diagram of the internal structure of an electronic device 700 provided in an embodiment of this application. The electronic device 700 can be provided as a server. (Refer to...) Figure 7 The electronic device 700 includes a processor 702, which further includes one or more processors, and memory resources represented by a memory 701 for storing instructions executable by the processor 702, such as a computer program. The computer program stored in the memory 701 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processor 702 is configured to execute instructions to perform a scene mode switching method for a multi-channel audio system according to any of the above embodiments. The electronic device 700 can operate on an operating system stored in the memory 701, such as Windows Server™, Mac OS X™, Unix™, Linux™, Free BSD™, or similar.

[0179] The electronic device 700 may further include a power supply component 703 configured to perform power management of the electronic device 700, a wired or wireless network interface 704 configured to connect the electronic device 700 to a network, and an input / output (I / O) interface 705. Wireless operation can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When executed by a processor, the computer program implements a scene mode switching method for a multi-channel audio system. The display unit 707 of the electronic device is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an e-ink display screen. The input device 706 of the electronic device can be a touch layer covering the display screen, or buttons, a trackball, or a touchpad mounted on the casing of the electronic device, or an external keyboard, touchpad, or mouse, etc.

[0180] Those skilled in the art will understand that Figure 7The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0181] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0182] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0183] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0184] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A scene mode switching method for a multi-channel audio system, characterized in that, The method includes: Detect the signal status and characteristic events at the audio input; Determine the device capability map of each audio playback device included in the multi-channel audio system; wherein, the device capability map characterizes the audio playback capability of the audio playback device; Based on the signal state and the feature event, a scene template matching the playback scene to be switched is determined, wherein the scene template is predefined with device grouping strategy and audio processing parameters; Based on the scene template and the device capability map, the target group and audio parameters of the audio playback device are determined, and the scene mode of the multi-channel audio system is switched based on the target group and audio parameters of the audio playback device.

2. The method according to claim 1, characterized in that, Detecting the signal status at the audio input, including: The physical link status of at least one of the following interfaces—the high-definition multimedia interface, the audio return channel, and the enhanced audio return channel—is detected as the signal status of the audio input terminal.

3. The method according to claim 1 or 2, characterized in that, Detecting the characteristic event includes: Detect at least one of the following characteristics of the audio stream encoding format, number of channels, sampling rate, and dynamic range at the audio input terminal; and / or, The device control command received via the consumer electronics control protocol is detected as the characteristic event.

4. The method according to claim 3, characterized in that, The device control commands include commands instructing the television to power on, enter system audio mode, or identify the type of source device; the method further includes: Upon detecting the device control command, the triggering conditions for scene switching are determined based on predefined entry and exit time thresholds.

5. The method according to claim 1, characterized in that, The step of determining a scene template matching the playback scene to be switched based on the signal state and the feature events includes: The logical rules predefined by the rule engine are applied to the signal state and the feature event; The logical rules map the signal state and / or the feature combination of the feature event to the corresponding scene template.

6. The method according to claim 5, characterized in that, The logical rules are also applied to time information, combining and mapping the time information with the signal state and / or the feature events.

7. The method according to claim 5 or 6, characterized in that, The logical rules executed by the rule engine include: If the feature event contains a multi-channel encoding format, then the home theater scene template is triggered; If the feature event contains an identification instruction from a preset game console, then the game scene template is triggered.

8. The method according to claim 1, characterized in that, The scene template includes: At least one of the following: home theater template, music appreciation template, game mode template, nighttime movie viewing template, and news broadcast template.

9. The method according to claim 8, characterized in that, The home theater template's predefined device grouping strategy includes selecting 5.1 channel grouping; The home theater template has predefined audio processing parameters including enabling the subwoofer, setting the crossover point to 80Hz, and boosting the gain of the 1kHz to 4kHz frequency band to enhance dialogue.

10. The method according to claim 8, characterized in that, The predefined audio processing parameters in the game mode template include: reducing audio processing buffer latency and disabling dynamic range compression.

11. The method according to claim 8, characterized in that, The predefined audio processing parameters for the nighttime movie viewing template include: enabling dynamic range compression, attenuating low-frequency bands, and limiting overall gain.

12. The method according to claim 1, characterized in that, Based on the scenario template and the device capability map, the target grouping of audio playback devices is determined, including: Based on the device capability map, evaluate all channel group combinations supported by currently online audio playback devices; Based on the predefined grouping strategy in the scenario template, a target group is selected from the supported grouping combinations.

13. The method according to claim 12, characterized in that, The method further includes: If the preset key equipment in the target group is offline, the target group will be downgraded to the next level of grouping based on the equipment capability map and the grouping strategy.

14. The method according to claim 13, characterized in that, The grouping that downgrades the target group to the next level includes: When the physical center channel device is offline, the preset virtual center channel algorithm is activated to group the front left and right channel devices into a virtual multi-channel system.

15. The method according to claim 1, characterized in that, The scene mode for switching multi-channel audio systems includes: Using one audio device in the multi-channel audio system as the master clock source, the clock domains of all audio playback devices in the target group are synchronized. Based on the audio-visual latency characteristics of video display devices, global audio latency compensation is calculated and configured.

16. The method according to claim 1, characterized in that, The method also includes a mode fallback process: If the duration of the loss of a valid signal at the current audio input terminal exceeds a first preset threshold, the multi-channel audio system is switched to a preset default scene template or the last validly activated non-trigger scene template.

17. The method according to claim 1, characterized in that, The scene mode for switching multi-channel audio systems includes: The parameter switching is completed within a second preset time using crossfading to achieve a smooth parameter transition.

18. A scene mode switching device for a multi-channel audio system, characterized in that, The device includes: The status and feature detection module is used to detect the signal status and feature events at the audio input terminal. The device capability determination module is used to determine the device capability map of each audio playback device included in the multi-channel audio system; wherein, the device capability map represents the audio playback capability of the audio playback device; The scene template determination module is used to determine a scene template that matches the playback scene to be switched based on the signal state and the feature event, wherein the scene template is predefined with device grouping strategy and audio processing parameters; The scene mode switching module is used to determine the target group and audio parameters of the audio playback device based on the scene template and the device capability map, and to switch the scene mode of the multi-channel audio system based on the target group and audio parameters of the audio playback device.

19. An electronic device comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 17.

20. A multi-channel audio system, characterized in that, include: Multiple audio playback devices; The electronic device as described in claim 19; the electronic device is used to control the plurality of audio playback devices to realize the switching of scene modes of a multi-channel audio system.

21. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 17.