Intelligent audio control system and method
Through the intelligent audio control system of the smart audio host, acquisition equipment, and speakers, the system dynamically controls the speakers to play target audio in different zones, solving the problem that public place audio systems cannot adjust according to changes in environment and people, and achieving efficient and energy-saving audio playback.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANSONG NANJING TECH LTD
- Filing Date
- 2026-01-23
- Publication Date
- 2026-04-28
AI Technical Summary
Existing audio systems in public places cannot dynamically adjust audio playback according to changes in the environment and people, resulting in unsatisfactory listening effects.
The intelligent audio control system, which uses a smart audio host, acquisition equipment, and speakers, dynamically controls the speakers to play target audio by collecting personnel information and environmental data, thereby achieving precise audio playback.
It improves the listening experience of audio playback, reduces energy waste, ensures audio clarity and comfort, and enhances the user experience.
Smart Images

Figure CN121934458A_ABST
Abstract
Description
Technical Field
[0001] This manual relates to the field of audio equipment technology, and in particular to an intelligent audio control system and method. Background Technology
[0002] For audio systems installed in public places (such as museums, exhibition halls, and airports), their main function is to manage the broadcasting of background music, announcements, or pre-recorded looped explanations through preset zones. The control logic of existing audio systems is typically based on a fixed schedule or manual triggering, uniformly playing audio to pre-defined physical areas. However, due to the varying environmental structures of different public places, coupled with changes in the location and number of people, audio playback based on preset zones, with its coarse control unit and lack of awareness of the environment and people, often fails to achieve the desired listening effect.
[0003] Therefore, it is desirable to provide a method for dynamic partitioning of distributed audio devices, which can dynamically partition devices with fixed installation locations to meet the needs of different scenarios and improve the listening experience. Summary of the Invention
[0004] This specification provides one or more embodiments of an intelligent audio control method, which is executed by a smart audio host connected to a data acquisition device and a switching device connected to multiple speakers. The method includes: determining a target area based on personnel information within a spatial area; wherein the personnel information is acquired by the data acquisition device; and controlling target speakers within the target area to play target audio based on target playback parameters.
[0005] This specification provides one or more embodiments of an intelligent audio control system, the system comprising: a data acquisition device configured to acquire personnel information within a spatial area; a speaker configured to play audio; a switching device configured to connect a smart audio host to the speaker for signal transmission; and a smart audio host connected to the data acquisition device and the switching device, the smart audio host being configured to: determine a target area based on the personnel information; and control a target speaker within the target area to play target audio based on target playback parameters.
[0006] This specification provides one or more embodiments of a computer-readable storage medium that stores computer instructions. When a computer reads the computer instructions from the storage medium, the computer executes the intelligent audio control method. Attached Figure Description
[0007] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein: Figure 1 This is a schematic diagram illustrating an application scenario of the intelligent audio control method according to some embodiments of this specification; Figure 2 This is a schematic diagram of an intelligent audio control system according to some embodiments of this specification; Figure 3 This is an exemplary flowchart of an intelligent audio control method according to some embodiments of this specification; Figure 4 This is a schematic diagram of an intelligent audio control method according to some embodiments of this specification; Figure 5 This is a schematic diagram illustrating auxiliary volume adjustment according to some embodiments of this specification. Detailed Implementation
[0008] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of this specification. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.
[0009] It should be understood that the terms “system,” “device,” “unit,” and / or “module” used herein are one way to distinguish different components, elements, parts, sections, or assemblies at different levels. However, if other terms can achieve the same purpose, they may be replaced by other expressions.
[0010] As indicated in this specification and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of expressly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.
[0011] Flowcharts are used in this specification to illustrate the operations performed by the system according to embodiments of this specification. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the steps can be processed in reverse order or simultaneously. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.
[0012] Figure 1 This is a schematic diagram illustrating an application scenario of the intelligent audio control method according to some embodiments of this specification.
[0013] In some embodiments, the application scenario 100 of the intelligent audio control system can be applied to scenarios such as museums, exhibition halls, and airports that require dynamic audio guidance.
[0014] In some embodiments, the application scenario 100 of the intelligent audio control system may include a data acquisition device 110, a processing device 120, an audio device 130, a network 140, and a storage device 150.
[0015] Data acquisition device 110 refers to a device that collects relevant data within a spatial area. For example... Figure 1 As shown, the acquisition device 110 may include a camera 110-1, a sensor 110-2, a sound pickup device 110-3, a radar 110-4, or any combination thereof.
[0016] Camera 110-1 is configured to acquire image data within a spatial region. For example, camera 110-1 may include a point cloud camera for acquiring three-dimensional spatial data within the spatial region.
[0017] Sensor 110-2 is configured to collect environmental data, personnel data, etc., within a spatial area. Sensor 110-2 may include infrared sensors, distance sensors, etc. For more related information, please refer to [link to relevant documentation]. Figure 2 Corresponding description.
[0018] The sound pickup device 110-3 is configured to collect sound information from a spatial area, such as ambient noise information and human voice noise information. The sound pickup device 110-3 may include an embedded microphone or a microphone array.
[0019] Radar 110-4 is configured to collect personnel data (such as pedestrian flow data) in a spatial area. Radar 110-4 can be a millimeter-wave radar, etc.
[0020] The processing device 120 can be used to process information and / or data related to the application scenario 100 of the intelligent audio control system. For example, pedestrian flow information, location information of the audio device 130, etc. In some embodiments, the processing device 120 can process data, information, and / or processing results obtained from other devices or system components, and execute program instructions based on this data, information, and / or processing results to perform one or more functions described in this specification. For example, the processing device 120 can acquire information and / or data collected by the acquisition device 110, and determine spatial distribution information and pedestrian flow information of a spatial area based on the information and / or data.
[0021] Audio device 130 refers to a hardware device used for playing audio. Audio device 130 may include multiple speakers. For example, audio device 130 may include, but is not limited to, AoIP (Audio over IP) speaker 130-1, loudspeaker 130-2, ... speaker cabinet 130-n, etc. In some embodiments, audio device 130 may be pre-installed at a fixed location within a spatial area to play audio related to the application scenario 100 of the intelligent audio control system, such as background music, service prompt audio, explanatory audio, emergency evacuation prompt audio, etc.
[0022] Network 140 may include any suitable network capable of facilitating the exchange of information and / or data. In some embodiments, one or more components of application scenario 100 of the intelligent audio control system (e.g., acquisition device 110, processing device 120, audio device 130, and storage device 150, etc.) may exchange information and / or data with one or more components of application scenario 100 of the intelligent audio control system via network 140.
[0023] Storage device 150 can store data, instructions, and / or any other information. Storage device 150 may include one or more storage components, each of which may be a separate device or part of another device. In some embodiments, storage device 150 may include random access memory (RAM), read-only memory (ROM), removable memory, and any combination thereof. In some embodiments, storage device 150 may be connected to network 140 to communicate with one or more components (acquisition device 110, processing device 120, audio device 130, etc.) in application scenario 100 of the intelligent audio control system.
[0024] It should be noted that the above description of the application scenario 100 of the intelligent audio control system is for ease of description only and should not be construed as limiting this specification to the scope of the illustrated embodiments. It is understood that those skilled in the art, after understanding the principles of the system, may arbitrarily combine the various component modules or construct subsystems connected to other modules without departing from these principles. In some embodiments, Figure 1 The acquisition device 110, processing device 120, audio device 130, and storage device 150 disclosed herein can be different modules within a single system, or a single module can implement the functions of two or more of the aforementioned modules. For example, the modules can share a single storage module, or each module can have its own separate storage module. Such variations are all within the scope of protection of this specification.
[0025] Figure 2 This is a schematic diagram of an intelligent audio control system according to some embodiments of this specification. For example... Figure 2As shown, the intelligent audio control system 200 may include a data acquisition device 110, an intelligent audio host 210, a network speaker 220, and a switching device 230.
[0026] Network speaker 220 may be one or more speakers included in audio device 130. For example... Figure 2 As shown, the first speaker 221 and the second speaker 222 included in the network speaker 220 can be composed of the AoIP speaker 130-1 and the loudspeaker 130-2 in the audio device 130, respectively.
[0027] In some embodiments, the data acquisition device 110 is configured to collect personnel information within a spatial area.
[0028] For more information about the data acquisition equipment, please see [link / details]. Figure 1 The corresponding description.
[0029] A spatial area refers to the area requiring intelligent audio control. A spatial area can be pre-divided into multiple sub-areas. For example, in a museum audio guidance scenario, the spatial area could be multiple areas within the exhibition hall, including the exhibit area and the guided tour area.
[0030] Personnel information refers to information such as the location and density of people within a spatial area. For example, personnel information includes the location, density characteristics, flow direction, and proximity of people to the exhibition booth within the spatial area.
[0031] Density characteristics refer to the number of people and population density in each sub-region of a spatial area.
[0032] The direction of movement refers to the direction of people's movement.
[0033] Proximity to the booth can be represented by the distance between the crowd gathering center and the booth. The crowd gathering center can be represented by the center of the sub-area with the largest number of people.
[0034] The smart audio host 210 acquires personnel information through the acquisition device 110. For example, the acquisition device 110 is configured as a millimeter-wave radar, and the smart audio host 210 acquires real-time density characteristics, movement, and location information of personnel through the millimeter-wave radar. For more information about the acquisition device 110, please refer to [link to relevant documentation]. Figure 1 Corresponding description.
[0035] In some embodiments, the smart audio host 210 can also collect ambient noise information via the acquisition device 110. For example, ambient noise information can be collected via the microphone 110-3. For more information on ambient noise information, please refer to [link to documentation]. Figure 3 The corresponding description.
[0036] In some embodiments, the network speaker 220 is configured to play audio.
[0037] In some embodiments, different application scenarios require different audio files. For example, in a museum audio-guided tour scenario, the audio may include background music during the tour, audio explanations of exhibits, and background information about the tour areas. More information about audio can be found in [link to relevant documentation]. Figure 1 Corresponding description.
[0038] The intelligent audio host 210 is the control center of the intelligent audio control system 200. The intelligent audio host 210 can integrate a processing device 120 to process information and / or data related to the intelligent audio control system 200.
[0039] like Figure 2 As shown, the smart audio host 210 is connected to the acquisition device 110 and the switching device 230. The smart audio host 210 can receive data from the acquisition device 110, perform data analysis through the built-in algorithm, and send control commands to the network speaker 220 through the switching device 230 based on the analysis results.
[0040] In some embodiments, the smart audio host 210 is configured to determine a target area based on personnel information; control the target speakers within the target area to play target audio based on target playback parameters.
[0041] For more information on how the Smart Audio Host 210 determines the target area, controls the playback of the target audio, and adjusts the target playback parameters, please refer to the corresponding instructions below.
[0042] Switching device 230 refers to a network device used for forwarding and managing Ethernet frames, and is a forwarding and management node for audio networks.
[0043] In some embodiments, the switching device 230 is configured to connect the smart audio host 210 and the network speaker 220 for signal transmission.
[0044] For example, the switching device 230 can be configured to perform functions such as audio / control data forwarding, multicast distribution, PoE power management, QoS scheduling and time synchronization pass-through between the smart audio host 210 and the network speaker 220.
[0045] In some embodiments, the network speaker 220 is connected to the smart audio host 210 via a star-chain or daisy-chain connection method through the switching device 230.
[0046] In some embodiments, the speakers in the intelligent audio control system 200 support the AOIP (Audio over IP) audio protocol and can receive and play digital audio streams sent by the intelligent audio host through the switching device 230.
[0047] In some embodiments, the intelligent audio system 200 encapsulates audio and control information into network packets using the AOIP audio protocol. These network packets include a destination address, enabling the intelligent audio host 210 to precisely send audio and control packets to the target speaker or speaker group via the switching device 230 on demand, similar to data transmission over an IP network, thereby achieving fine-grained audio scheduling and control.
[0048] In some embodiments, speakers supporting the AOIP audio protocol encapsulate the transmitted audio and control information into network data packets and transmit them via switching device 230. The smart audio host 210 assigns a unique network address (e.g., an IP address) to each speaker. When the smart audio host 210 sends a data packet labeled with a target address and playback instructions (e.g., "Target IP address P1, please play 'Explanation A.mp3' at 70% volume"), the switching device 230 or a link forwarding mechanism routes the data packet to the target speaker (e.g., the target speaker with target IP address P1).
[0049] In some embodiments, the smart audio system 200 uses the AOIP audio protocol to transmit audio and control commands, enabling the smart audio host 210 to accurately send audio streams and parameter commands to the target speaker based on more accurate audio / positioning information.
[0050] In some embodiments, the network speaker 220 can be connected in series via a single network cable and then connected to the smart audio host 210 via a switching device 230.
[0051] Single-wire cascading refers to the connection between the smart audio host and the switching device based on a single network cable. After the switching device is connected to the smart audio host, it is then connected to the next speaker based on another network cable. The two speakers are also connected in series via a single network cable, thereby realizing the cascading transmission of data and control.
[0052] For example, the smart audio host 210 and the switching device 230 are connected in series via a single network cable; the switching device 230 and the first speaker 221 are connected in series via a single network cable; and the first speaker 221 and the second speaker 222 are connected in series via a single network cable; the smart audio host 210 sends the audio data stream to the switching device 230 in the order of the network cable connection, and then the switching device 230 transmits the audio data to the first speaker 221 and the second speaker 222 in sequence.
[0053] In some embodiments of this specification, speakers are connected in series to a switching device via a single network cable, allowing for the cascading of a larger number of speakers. This solution not only enables intelligent control of speaker playback and volume but also reduces power consumption and the average power consumption of each speaker. Simultaneously, it significantly reduces the complexity of wiring and construction, shortens the construction period, and reduces material and labor costs.
[0054] In some embodiments, the network speaker 220 can also be connected to the smart audio host 210 via the switching device 230.
[0055] The smart audio host 210 and each network speaker 220 are connected to different ports of the switching device 230 via their respective independent network cables; audio data is sent from the smart audio host 210 to the switching device 230, and then the switching device 230 distributes the audio data to one or more designated speakers.
[0056] For example, the switching device 230 is configured with multiple ports (e.g., a first port, a second port, etc.); the first speaker 221 is connected to the first port of the switching device 230 via a separate network cable, and the second speaker 222 is connected to the second port of the switching device 230 via a separate network cable. The smart audio host 210 sends audio data streams to the switching device 230, which then transmits them to the first speaker 221 and / or the second speaker 222 according to pre-configured software.
[0057] In some embodiments, to adapt to scenarios of varying scale and complexity, the intelligent audio control system 200 supports multiple component arrangement and expansion strategies. These include, for example, a distributed multi-host architecture (deploying multiple intelligent audio hosts by area within the exhibition hall, which then collaborate through an upper-level scheduler), a hybrid sensor topology (combining millimeter-wave radar and cameras to enhance robustness), and a cloud-based analytics and operations platform (for historical data analysis, model training, and remote configuration). The intelligent audio control system 200 can also reserve standardized interfaces for integration with third-party exhibition hall management systems, navigation apps, or security platforms.
[0058] In some embodiments, the layout and connection method of components such as the acquisition device 110, the smart audio host 210, and the network speaker 220 can be adjusted and expanded according to the needs of the actual application scenario to achieve better coverage and control effects. For example, multiple acquisition devices 110 can be set up in a spatial area and connected to the smart audio host 210 through the network 140.
[0059] In one or more embodiments of this specification, by collecting multi-dimensional information (such as personnel information and environmental information), playback parameters are intelligently adjusted to achieve real-time and efficient intelligent audio playback control, reducing energy waste caused by unnecessary audio playback in uninhabited areas, and ensuring the clarity and comfort of audio in different environments with different numbers of people and noise levels.
[0060] Figure 3 This is an exemplary flowchart illustrating an intelligent audio control method according to some embodiments of this specification. Figure 3 As shown, process 300 includes the following steps. In some embodiments, process 300 may be executed by the smart audio host 210.
[0061] Step 310: Determine the target area based on the personnel information within the spatial area.
[0062] For details regarding spatial areas and personnel information, please refer to [link / reference]. Figure 2 Corresponding content.
[0063] The target area refers to the spatial region where audio needs to be played or audio playback parameters need to be adjusted. For example, the target area could be a region where the foot traffic exceeds a preset threshold.
[0064] In some embodiments, the smart audio host 210 can determine the target area based on personnel information.
[0065] In some embodiments, when the number of people in a certain spatial area exceeds a preset threshold, that spatial area can be identified as a target area.
[0066] In some embodiments, the target area may be determined based on other feasible methods, such as the exhibition hall’s preset guided tour route or preset key exhibit areas.
[0067] Step 320: Control the target speaker within the target area to play the target audio based on the target playback parameters.
[0068] Target loudspeakers refer to some or all of the loudspeakers within the target area. For example, loudspeakers within the target area that are in normal working order.
[0069] In some embodiments, the smart audio host can determine the target speaker in a variety of ways. For example, the smart audio host can use all speakers in the target area or any preset number of speakers in the target area as the target speaker.
[0070] In some embodiments, the target loudspeaker may further consist of a main loudspeaker and auxiliary loudspeakers. For more information about target loudspeakers, see [link to relevant documentation]. Figure 4 The corresponding content.
[0071] Target playback parameters refer to the playback parameters corresponding to the target speaker.
[0072] The target playback parameters may include at least one of the target volume, target sound pressure level, and target frequency.
[0073] Target volume and target sound pressure level refer to the volume and sound pressure level of the audio played by the target speaker, respectively.
[0074] The target frequency refers to the frequency range in the audio that needs to be adjusted when the target speaker plays the audio, such as the mid-to-high frequency range that needs to be enhanced.
[0075] In some embodiments, the smart audio host uses preset playback parameters as target playback parameters.
[0076] The smart audio host can also generate or adjust the target playback parameters based on other feasible methods.
[0077] In some embodiments, the smart audio host 210 can adjust the target playback parameters of the target speakers in the target area in real time based on the collected personnel information.
[0078] Taking adjusting target playback parameters based on personnel density in personnel information as an example, when the personnel density in the target area suddenly increases during playback, such as exceeding a preset threshold, the smart audio host 210 will immediately send a command to control the target speakers in the target area to increase the target volume or target sound pressure level in the target playback parameters by a preset amount. Conversely, when personnel begin to disperse, such as when the personnel density decreases below the preset threshold, the target volume or target sound pressure level in the target playback parameters will be reduced by a preset amount.
[0079] In some embodiments, the smart audio host 210 adjusts the target playback parameters based on ambient noise information.
[0080] Environmental noise information refers to noise-related information for a spatial area. Environmental noise information includes the volume, spectral distribution, and energy distribution of environmental noise.
[0081] In some embodiments, the smart audio host 210 collects environmental noise information in real time through the acquisition device 110.
[0082] The volume of ambient noise is usually referred to as the real-time sound pressure level (in decibels (dB)). In some embodiments, the smart audio host 210 collects the volume of ambient noise through the pickup device 110-3, converts the collected analog sound signal into a digital electrical signal, and finally quantifies it into an average sound pressure level value per unit time to determine the volume of ambient noise.
[0083] Spectral distribution refers to the distribution of noise intensity across different frequency ranges in collected ambient audio data. For example, spectral distribution characterizes how the intensity of a noise signal varies with frequency across the entire audible frequency range (e.g., 20 Hz to 20 kHz). In some embodiments, the smart audio host 210 collects noise signals through the pickup device 110-3 and performs spectral analysis using a built-in algorithm to determine the spectral distribution. For example, the smart audio host 210 employs signal processing techniques such as Fourier transform (FFT) to analyze the collected noise signals, converting the time-domain signal into a frequency-domain signal to obtain the spectral distribution.
[0084] Energy distribution refers to the proportion of total acoustic energy (power) of noise in different frequency ranges within the collected environmental audio data.
[0085] In some embodiments, the smart audio host 210 collects noise signals through the pickup device 110-3 and determines the energy distribution based on the spectral distribution obtained after spectral analysis. For example, based on the spectral distribution obtained after spectral analysis, the smart audio host 210 calculates the noise energy and total noise energy in each frequency range by integrating or summing the signal strength in each frequency range; and divides the noise energy in each frequency range by the total noise energy to determine the energy distribution corresponding to different frequency ranges.
[0086] In some embodiments, the smart audio host 210 determines the target playback parameters based on ambient noise information and a first preset table.
[0087] The first preset table is constructed based on the correspondence between the volume of ambient noise, the spectral distribution of ambient noise, the energy distribution of ambient noise, and the target playback parameters. The first preset table can be set based on experience.
[0088] For example, when the volume of ambient noise is higher, the frequency that needs to be enhanced in the target playback parameters (i.e., the target frequency) is lower, and when the energy distribution of ambient noise corresponds to the energy concentrated in the low-frequency part (e.g., the frequency range is between 20Hz and 200Hz), the target frequency in the target playback parameters can be in the mid-to-high frequency part (e.g., the frequency range is between 2kHz and 6kHz) to enhance speech clarity.
[0089] In some embodiments, the smart audio host 210 adjusts playback parameters based on ambient noise information and adaptively increases mid-to-high frequencies or gain according to real-time noise, ensuring the clarity of the target audio playback in noisy environments.
[0090] In some embodiments, the smart audio host 210 may also determine the target playback parameters through other feasible methods. For example, the smart audio host 210 determines the adjusted target playback parameters based on changes in ambient noise using a built-in algorithm.
[0091] Target audio refers to the audio that needs to be played. Target audio is associated with a target area. For example, if the target area is an exhibition area, the target audio is the explanatory audio corresponding to the exhibits in that area. Or, if the target area is a guided tour area, the target audio is the background information audio corresponding to that area.
[0092] In some embodiments, the smart audio host 210 can control target speakers within a target area to play target audio based on target playback parameters. For example, after determining the target area, the smart audio host 210 identifies the target speakers located within that area and sends control commands to them. These commands include the target speaker's identification information (such as a target IP address), target playback parameters (such as target volume, target sound pressure level, and target frequency), and target audio associated with the area (such as exhibit explanations or background information). Upon receiving the control commands, the target speakers play the corresponding target audio according to the target playback parameters.
[0093] In some implementations, the playback of the target audio in step 320 can also be achieved through steps 321 and 322.
[0094] Step 321: Based on the personnel information, determine whether the personnel have entered the viewing area.
[0095] The viewing area refers to the space surrounding the exhibits for people to view them. The size of the viewing area is determined based on experience.
[0096] In some embodiments, the smart audio host 210 determines whether a person has entered the viewing area based on their location information. For example, if the location information of a person collected by the acquisition device is within a preset viewing area, it is determined that the person has entered the viewing area.
[0097] In some embodiments, the smart audio host 210 determines whether a person has entered the viewing area based on the proximity of the person to the booth. For example, if the calculated distance between the center of the crowd and the booth is less than a preset distance to the boundary of the viewing area, it can be determined that the person (or crowd) has entered the viewing area.
[0098] In some embodiments, the smart audio host 210 can determine whether a person has entered the viewing area based on any other feasible method.
[0099] In some embodiments, the smart audio host 210 can optimize the determination of whether a person has entered the viewing area by combining the dwell time of the person. For example, only when a person is in the viewing area and their dwell time exceeds a preset dwell time threshold (e.g., 3 seconds) is it confirmed that they have entered the viewing area, thereby avoiding unnecessary audio playback triggered by people passing by quickly.
[0100] In some embodiments, the smart audio host 210 can optimize the determination of whether people have entered the viewing area by taking into account the density of people. For example, when the density of people in the viewing area reaches a preset minimum density requirement, it is determined to be a valid entry, and the guided tour service for the people is initiated.
[0101] Step 322: In response to personnel entering the viewing area, control the target speaker corresponding to the viewing area to play the target audio.
[0102] In some embodiments, the smart audio host 210 determines that a person has entered the viewing area and triggers a playback action; it sends an instruction to the target speaker corresponding to the space area where the viewing area is located, and starts playing the target audio (such as the exhibit explanation) associated with the exhibit with preset or dynamically determined target playback parameters (such as target volume and target sound pressure).
[0103] In some embodiments, the smart audio host 210 can control the target speaker corresponding to the viewing area to play target audio through any other feasible means.
[0104] For example, the smart audio host 210 can control the target speakers corresponding to the viewing area to play target audio. The playback of the target audio can be started in a fade-in manner, that is, after entering the viewing area, the volume gradually increases from zero to the target volume, so as to avoid abrupt auditory stimulation and improve the user experience.
[0105] In some embodiments of this specification, the entry of personnel is used as a precise playback trigger condition, which enables accurate positioning and timely response of audio services, improves the efficiency of information transmission, and further enhances the energy-saving effect by strictly limiting the playback range, avoiding unnecessary audio coverage outside the viewing area.
[0106] In some embodiments, the smart audio host 210 can control the target speaker to play target audio in any feasible manner.
[0107] For example, when the data acquisition of the acquisition device 110 is interrupted or cannot provide accurate personnel information, the smart audio host 210 can control the target speaker to play using preset playback parameters (such as low-volume background music) to ensure that the system can provide basic audio services under any circumstances.
[0108] According to one or more embodiments in this specification, the intelligent audio control system plays audio only in areas where people are present and intelligently adjusts the volume, thus avoiding unnecessary energy consumption and conforming to the concept of energy conservation and emission reduction. The intelligent adjustment avoids audio interference and poor sound quality caused by multiple speakers playing at the same time. Combined with the environmental noise adaptive adjustment function, it makes the spatial audio clearer and of higher quality, improving audio quality and user experience. It realizes full automation from playback area, content, volume to sound quality.
[0109] It should be noted that the above description of process 300 is for illustrative purposes only and does not limit the scope of this specification. Those skilled in the art can make various modifications and changes to process 300 under the guidance of this specification. However, these modifications and changes remain within the scope of this specification.
[0110] In some embodiments, the smart audio host 210 determines the acoustic center of the target area based on personnel information and exhibit information; determines the main speaker and auxiliary speaker in the target area based on the acoustic center; controls the main speaker to play at the main volume; determines the auxiliary volume based on the distance between the auxiliary speaker and the acoustic center and the main volume, and controls the auxiliary speaker to play at the auxiliary volume.
[0111] In some embodiments, the smart audio host 210 determines the acoustic center of the target area based on personnel information and exhibit information.
[0112] Exhibit information refers to attribute information related to the exhibits. For example, exhibit information includes the exhibit's location, its identification ID, the path to the corresponding audio guide file, and its category and brief description.
[0113] In some embodiments, the smart audio host 210 can obtain exhibit information by calling data pre-stored in the storage device 150. For example, it can obtain the location information of the exhibits by pre-measuring fixed coordinate data (e.g., a coordinate system based on the exhibition hall floor plan) and inputting it into the storage device 150.
[0114] In some embodiments, the smart audio host 210 can obtain exhibit information by receiving data collected by the acquisition device 110. For example, the smart audio host 210 can obtain exhibit information in real time by receiving exhibit images uploaded by the camera 110-1.
[0115] The acoustic center refers to the central location within the target area where clear audio playback is most needed.
[0116] For example, the acoustics center established the origin point that theoretically satisfies the best listening effect in order to create the best listening experience.
[0117] The determination of the acoustic center can be a dynamic process, updated in real time as people move and their gathering status changes.
[0118] In some embodiments, the smart audio host 210 can determine the crowd gathering center based on the density characteristics of the people; and use the midpoint between the location of the exhibit and the crowd gathering center as the acoustic center.
[0119] For example, the intelligent audio host 210, according to preset rules, converts the number of people in each sub-region collected by the acquisition device 110 into the weight corresponding to each sub-region, and calculates the weighted average of the center coordinates of each sub-region and the corresponding weight to determine the real-time centroid position, which is then used as the acoustic center. The preset rules can be preset based on experience; the more people in a sub-region, the greater its weight, indicating that the sub-region is more decisive for the acoustic center.
[0120] In some embodiments, the smart audio host 210 can determine the acoustic center based on other feasible methods. For example, if personnel information collection fails or the crowd density is extremely low, the smart audio host 210 can use the preset center point of the target area as the preset acoustic center.
[0121] In some embodiments, the smart audio host 210 can determine the main speaker and auxiliary speaker in a target area based on the acoustic center.
[0122] The main loudspeaker refers to the target loudspeaker that is closest to the acoustic center.
[0123] Auxiliary loudspeakers refer to target loudspeakers other than the main loudspeakers.
[0124] In some embodiments, the smart audio host 210 determines the main and auxiliary speakers within a target area based on the acoustic center, thereby enabling sound focusing towards the acoustic center.
[0125] The intelligent audio host 210 sets the speaker (usually one or a few) that is closest to the acoustic center or that is less than a preset distance threshold from the acoustic center as the main speaker to provide primary and clearest audio coverage.
[0126] The Smart Audio Host 210 sets up other speakers in the target area, excluding the main speaker, as auxiliary speakers to provide additional sound energy coverage and environmental sound field supplementation.
[0127] In some embodiments, the smart audio host 210 may also determine the main speaker and auxiliary speaker in the target area through other feasible means.
[0128] For example, a smart audio host can determine the main speaker and auxiliary speakers in a target area by acquiring manual input or accessing historical data stored in a storage device, and then using the speaker with the highest historical usage rate in the target area as the main speaker and the other speakers in the target area as auxiliary speakers.
[0129] In some embodiments, the smart audio host 210 controls the main speaker to play at the master volume.
[0130] Master volume refers to the base playback volume of the main speaker.
[0131] In some embodiments, the smart audio host 210 determines the target volume in the target playback parameters as the main volume.
[0132] In some embodiments, the smart audio host 210 determines the main volume based on the characteristics of the exhibition hall, the content characteristics of the target audio, and the characteristics of human behavior.
[0133] The characteristics of an exhibition hall include its size, number of exhibits, and type.
[0134] Exhibition hall type refers to the functional attributes of an exhibition hall. For example, a history exhibition hall or a children's science exhibition hall.
[0135] In some embodiments, the smart audio host 210 obtains exhibition hall features by calling data such as exhibition hall floor plans and asset lists pre-stored in the storage device 150. The smart audio host 210 may also obtain exhibition hall features through any other feasible means, such as by obtaining relevant data of exhibition hall features input by the user.
[0136] Content characteristics refer to the type of audio content. For example, content characteristics include explanations of exhibits, summaries of the exhibition hall background, etc.
[0137] In some embodiments, the smart audio host 210 obtains content features by calling audio data pre-stored in the storage device 150. The smart audio host 210 may also obtain content features in any other feasible way, such as by obtaining relevant data of content features input by the user.
[0138] Human behavior characteristics refer to the behavioral features of people in a spatial area. For example, human behavior characteristics include static listening and dynamic passing by.
[0139] Static listening refers to the movement speed of people in the area being slow (e.g., less than a preset speed threshold) and their orientation being basically the same (e.g., facing the exhibits, and the difference in orientation angles being less than a preset angle threshold).
[0140] Dynamic passing refers to people quickly (e.g., moving at a speed greater than a preset speed threshold) passing through a target area.
[0141] In some embodiments, the smart audio host 210 obtains human point cloud data from multiple historical time periods based on personnel information within historical time periods through cluster analysis and Kalman filtering algorithms.
[0142] Human point cloud data refers to a collection of point clouds representing the same person.
[0143] In some embodiments, the smart audio host 210 performs clustering analysis on single-frame point cloud data to identify point cloud sets of the same person. For example, the smart audio host 210 may use a density-based clustering algorithm (such as the DBSCAN clustering algorithm) to group the point cloud, remove outlier noise points, and group spatially closely distributed point sets into clusters. For each cluster, the smart audio host 210 extracts the cluster center (centroid), number of points, and boundary information as an estimate of the person's position and size at that moment, and filters out non-human targets (e.g., cylinders) based on the size and shape of the cluster.
[0144] The parameters of the clustering process (e.g., neighborhood scale and minimum number of points) can be adjusted according to the resolution and installation height of the acquisition device 110 or obtained by on-site calibration; the clustering results can be jointly judged with scene geometric information to improve the reliability of single-frame recognition.
[0145] In some embodiments, the smart audio host 210 uses a Kalman filter algorithm to acquire point cloud data that locates the same target individual (like the same person) in the point cloud time series data. For example, the smart audio host 210 performs time series observations of the point cloud and applies tracking algorithms such as Kalman filtering to associate detections at different times with the trajectory of the same target individual.
[0146] For example, at each moment, the intelligent audio host 210 performs cluster analysis on the current point cloud data to identify independent target individuals (e.g., one cluster corresponds to one target individual) and determines their center position, velocity, and other motion states. Based on a Kalman filter, using the motion state (e.g., position and velocity) of the target individual at time t-1, it calculates the predicted position of the target individual at the current time t. The predicted position is compared with the positions of all new clustered targets actually acquired by the acquisition device 110 at the current time t. If the position of a new clustered target individual matches the predicted position of that target individual... If the locations are within a preset distance threshold (i.e., spatially high proximity), then the point cloud detections of these two clusters at different times are determined to belong to the point cloud data of the same target individual. After the association is determined, the Kalman filter uses the new observation value (the actual position at the current time) to correct the prediction state, obtain a more accurate filtering state, and update the motion trajectory of the target individual for prediction in the next time step. Through this association mechanism based on temporal prediction and spatial proximity, the smart audio host 210 "connects" the point cloud data collected at different time points, thereby locating and locking continuous data belonging to the same target individual.
[0147] In some embodiments, the smart audio host 210 determines the behavioral characteristics of individuals based on the trajectory of each person's point cloud data within a historical time period.
[0148] Historical time periods refer to a period of time before the current moment. The duration is set based on experience, such as 3 minutes, 5 minutes, etc.
[0149] The trajectory includes the movement speed and position of each person's point cloud data.
[0150] In some embodiments, the smart audio host 210 analyzes the average velocity of each human point cloud data based on the individual human point cloud data. For example, if the average velocity of a human point cloud data is close to 0 within a preset static time (e.g., 5 minutes, 10 minutes, etc.), it is determined as "static listening" or "standing still"; if the average velocity is greater than a preset velocity threshold and the trajectory is linear, it is determined as "dynamic passing by". The preset velocity threshold is set based on experience.
[0151] In some embodiments, the smart audio host 210 determines the master volume based on a first vector to be matched and through a first vector database.
[0152] The first matching vector is constructed based on the exhibition hall features, the content features of the target audio, and the behavioral features of the people.
[0153] The first vector database includes multiple first reference vectors. Each first reference vector has a corresponding first label.
[0154] The first reference vector is constructed based on the characteristics of the historical exhibition hall, the characteristics of historical content, and the characteristics of historical behavior.
[0155] The first label includes the preferred master volume corresponding to the first reference vector.
[0156] The first tag is obtained based on historical data, practical experience, etc.
[0157] In some embodiments, the smart audio host 210 selects the historical main volume with the longest average dwell time of the crowd when playing audio from among multiple historical main volumes corresponding to the first reference vector, as the first label corresponding to the first feature vector.
[0158] The average dwell time of a crowd refers to the average dwell time of multiple people in an area while audio is playing. The Smart Audio Host 210 can obtain the average dwell time of a crowd by accessing data stored in the storage device or by obtaining user input.
[0159] In some embodiments, the smart audio host 210 can, based on the similarity between the first vector to be matched and multiple first reference vectors in the first vector database, select the first reference vector whose similarity meets a first preset condition as the first target vector, determine the first label corresponding to the first target vector, and use the preferred main volume corresponding to the first label as the main volume corresponding to the first vector to be matched. The first preset condition can be set according to the situation. For example, maximum similarity or similarity greater than a threshold, etc.
[0160] In one or more embodiments of this specification, by achieving precise adaptation of the master volume to multiple dimensions of context such as physical environment, playback content and human behavior, more accurate sound field control and more effective energy consumption control can be further achieved.
[0161] In some embodiments, the smart audio host 210 may also determine the master volume based on other feasible methods.
[0162] For example, the intelligent audio host 210 can use a preset master volume as a base, determine one gain based on the total number of people, and determine another gain based on the ambient noise level. These are then superimposed or weighted together to calculate a real-time optimized master volume for playback. This ensures that, in different environments and with varying crowd sizes, the main speaker's audio playback clearly conveys information while avoiding energy waste or noise interference caused by over-playing.
[0163] In some embodiments, the smart audio host 210 determines the auxiliary volume based on the distance between the auxiliary speaker and the acoustic center and the main volume, and controls the auxiliary speaker to play at the auxiliary volume.
[0164] Auxiliary volume refers to the playback volume of the auxiliary speaker.
[0165] In some embodiments, the smart audio host 210 determines the auxiliary volume of each auxiliary speaker based on the distance between the auxiliary speaker and the acoustic center and the main volume using a second preset table.
[0166] In some embodiments, the second preset table includes the relationship between the distance between the auxiliary speaker and the acoustic center, the master volume, and the auxiliary volume. The higher the master volume and the greater the distance between the auxiliary speaker and the acoustic center, the lower the auxiliary volume. The second preset table can be set based on experience.
[0167] In some embodiments, the smart audio host 210 may also determine the auxiliary volume based on other feasible methods.
[0168] For example, the intelligent audio host 210 can determine the auxiliary volume based on the dispersion of people near the auxiliary speakers. For instance, if there are a few people in the space near the auxiliary speakers, the volume attenuation can be reduced appropriately. If the distance calculation fails, all auxiliary speakers can uniformly play at a volume lower than a preset percentage of the main volume (e.g., 80% of the main volume) to maintain the consistency of the sound field.
[0169] For more information on determining the auxiliary volume, please refer to [link / reference]. Figure 5 Corresponding description.
[0170] The intelligent audio host 210 sends commands to control the auxiliary speakers, which then play the target audio at their respective determined auxiliary volumes.
[0171] Figure 4 This is a schematic diagram of an intelligent audio control method according to some embodiments of this specification.
[0172] like Figure 4 As shown, the intelligent audio control system 200 is deployed in a space such as an exhibition hall, which is divided into multiple sub-areas. Three example areas are shown in the figure: area A, area B and area C.
[0173] Area A includes exhibits A-3, sensors A-11 and A-12, and speakers A-21 and A-22. The number of personnel in Area A is 2.
[0174] Area B includes exhibit B-3, sensor B-11, and sensor B-12. The loudspeakers in Area B are divided into loudspeaker B-21 and loudspeaker B-22. The number of personnel in Area B is 5.
[0175] Area C includes sensor C-11 and speaker C-21. The number of personnel in Area C is 1.
[0176] Compared to areas A and B, area B has a larger number of people and a more concentrated population; therefore, area B will be selected as the target area. Based on the personnel and exhibit information of area B, the acoustic center O of the target area will be determined. B According to the Acoustics Center O B Determine the distance from the acoustic center O within the target area. B The nearest loudspeaker, B-22, is the main loudspeaker, and B-21 is the auxiliary loudspeaker. Loudspeaker B-22 is controlled to play at the main volume. Based on the distance of loudspeaker B-21 from the acoustic center and the main volume, the auxiliary volume is determined, and loudspeaker B-21 is controlled to play at the auxiliary volume to direct sound towards the acoustic center O. B The focus.
[0177] One or more embodiments of this specification, by determining the acoustic center and distinguishing between the main speaker and the auxiliary speaker, can effectively concentrate the sound energy of the audio playback at the acoustic center, thereby improving the clarity of the target audio, reducing the interference of audio on the periphery and adjacent areas of the target area, further improving the quality of the auditory environment, and optimizing the user experience.
[0178] Figure 5 This is a schematic diagram illustrating auxiliary volume adjustment according to some embodiments of this specification.
[0179] like Figure 5 As shown, based on the crowd density 511 in the corresponding area of the auxiliary speaker, the human voice noise 512 in the corresponding area of the auxiliary speaker, and the distance between the auxiliary speaker and the exhibit 513, the volume adjustment factor 520 is determined; based on the volume adjustment factor 520, the auxiliary volume 530 is adjusted.
[0180] The corresponding area for the auxiliary speaker refers to the area within a preset range around the auxiliary speaker. For example, the area within a preset radius (e.g., 1m) centered on the auxiliary speaker.
[0181] In some embodiments, the smart audio host 210 can determine the crowd density based on the personnel information collected by the acquisition device 110.
[0182] Human noise refers to environmental noise generated by the sounds of people talking, shouting, or engaging in activities within a spatial area.
[0183] In some embodiments, the smart audio host 210 can acquire human voice noise by receiving sound data collected by the acquisition device 110 (e.g., the pickup device 110-3). If the intensity of human voice noise around the auxiliary speaker is high, it may indicate that the population density in the area is relatively high, suggesting that there may be many listeners around the auxiliary speaker, and the volume needs to be increased appropriately.
[0184] Since the energy of human voice noise is mainly concentrated in a specific frequency range (e.g., between 300 Hz and 3400 Hz, i.e., the frequency of non-mechanical, non-ventilation noise), in some embodiments, the smart audio host 210 performs spectrum analysis on the collected sound data, and through bandpass filtering and energy calculation, quantifies only the sound energy within the frequency range corresponding to the human voice noise, thereby obtaining the intensity of the human voice noise.
[0185] The distance between the auxiliary speaker and the exhibit refers to the straight-line spatial distance between the auxiliary speaker and the exhibit. The closer the auxiliary speaker is to the exhibit, the more important the audio content it plays, and the higher the volume adjustment factor may be required to ensure that the sound can effectively transition to the central area covered by the main speaker.
[0186] The volume adjustment factor refers to the amount of auxiliary volume adjustment. For example, if there is a lot of noise in the human voice, the volume adjustment factor needs to be increased or set to a positive value to ensure audio clarity. Conversely, if the environment becomes quieter, the volume adjustment factor needs to be decreased or set to a negative value to intelligently achieve energy conservation and environmental protection.
[0187] In some embodiments, the volume adjustment factor can be a percentage in the range [-1, 1].
[0188] In some embodiments, the smart audio host 210 determines the volume adjustment factor based on the second vector to be matched through a second vector database.
[0189] The second matching vector is constructed based on the current crowd density, current human voice noise in the corresponding area of the auxiliary speaker, and the distance between the auxiliary speaker and the exhibit.
[0190] The second vector database includes multiple second reference vectors. Each second reference vector has a corresponding second label.
[0191] The second reference vector is constructed based on the historical crowd density, historical human voice noise, and historical distance between the auxiliary loudspeaker and the exhibit in the corresponding area of the historical auxiliary loudspeaker.
[0192] The second label includes the preferred volume adjustment factor corresponding to the second reference vector.
[0193] The second tag is obtained based on historical data, practical experience, etc.
[0194] In some embodiments, the smart audio host 210 uses the historical volume adjustment factor with the best subsequent feedback effect among multiple historical volume adjustment factors corresponding to the second reference vector as the second label corresponding to the second reference vector. The best feedback effect means the fewest negative feedback responses from the user. The smart audio host 210 can obtain the feedback effect by accessing data stored in a storage device or by obtaining user input.
[0195] In some embodiments, the smart audio host 210 can, based on the similarity between the second target vector and multiple second reference vectors in the second vector database, select the second reference vector whose similarity meets a second preset condition as the second target vector, determine the second label corresponding to the second target vector, and use the preferred volume adjustment factor corresponding to the second label as the volume adjustment factor corresponding to the second target vector (i.e., current crowd density, current distance, and current human voice noise). The second preset condition can be set according to the situation. For example, maximum similarity or similarity greater than a threshold, etc.
[0196] In some embodiments, the smart audio host 210 may also determine the volume adjustment factor in any other feasible manner. For example, based on a built-in multivariable function model, the required volume adjustment factor can be calculated according to three input parameters (crowd density, human voice noise, and distance) obtained in real time.
[0197] For more information on auxiliary volume, please refer to [link / reference]. Figure 4 The corresponding content.
[0198] In some embodiments, the smart audio host 210 adjusts the current auxiliary volume based on a volume adjustment factor to determine the adjusted auxiliary volume. The smart audio host 210 then sends a command to the corresponding auxiliary speaker to control it to play at the adjusted auxiliary volume.
[0199] In some embodiments, the smart audio host 210 can adjust the auxiliary volume based on the positive correlation between the volume adjustment factor and the auxiliary volume. For example, the auxiliary volume can be determined based on the following formula.
[0200]
[0201] In the formula, This refers to the adjusted auxiliary volume. This refers to the volume adjustment factor ( ∈[-1, 1]). This refers to the auxiliary volume before adjustment.
[0202] In some embodiments, the smart audio host 210 can adjust the auxiliary volume in any feasible manner. Further details can be found in [link to relevant documentation]. Figure 3 and Figure 4 Corresponding content.
[0203] In some embodiments of this specification, a volume adjustment factor is introduced and adjusted in real time to achieve intelligent and adaptive audio playback, thereby ensuring a natural and smooth transition of the overall sound field within the target area, and improving the audio coverage quality and user satisfaction throughout the target area.
[0204] The basic concepts have been described above. Obviously, for those skilled in the art, the detailed disclosure above is merely illustrative and does not constitute a limitation of this specification. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and corrections to this specification. Such modifications, improvements, and corrections are suggested in this specification and therefore remain within the spirit and scope of the exemplary embodiments described herein.
[0205] Furthermore, this specification uses specific terms to describe embodiments thereof. For example, "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Moreover, certain features, structures, or characteristics in one or more embodiments of this specification can be appropriately combined.
[0206] Furthermore, unless expressly stated in the claims, the order of processing elements and sequences, the use of numbers and letters, or other names described in this specification are not intended to limit the order of the processes and methods described herein. Although various examples have been discussed in the foregoing disclosure of some embodiments of the invention that are currently considered useful, it should be understood that such details are for illustrative purposes only, and the appended claims are not limited to the disclosed embodiments; rather, the claims are intended to cover all modifications and equivalent combinations that conform to the spirit and scope of the embodiments described herein. For example, while the system components described above can be implemented using hardware devices, they can also be implemented solely using software solutions, such as installing the described system on existing servers or mobile devices.
[0207] Similarly, it should be noted that, in order to simplify the description disclosed herein and thus aid in the understanding of one or more embodiments of the invention, the foregoing description of embodiments in this specification may sometimes combine multiple features into a single embodiment, drawing, or description thereof. However, this method of disclosure does not imply that the subject matter of this specification requires more features than those mentioned in the claims. In fact, the embodiments contain fewer features than all the features of a single embodiment disclosed above.
[0208] In some embodiments, numbers describing the quantity of components and attributes are used. It should be understood that such numbers used in the description of embodiments are modified in some examples with the terms "approximately," "approximately," or "generally." Unless otherwise stated, "approximately," "approximately," or "generally" indicates that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, which may be changed depending on the characteristics required by individual embodiments. In some embodiments, numerical parameters should take into account specified significant digits and employ a general method of digit reservation. Although the numerical ranges and parameters used to confirm their breadth of range in some embodiments of this specification are approximate values, in specific embodiments, such values are set as precisely as feasible.
[0209] For each patent, patent application, patent application publication, and other material such as articles, books, specifications, publications, and documents referenced in this specification, the entire contents of which are incorporated herein by reference. This excludes historical application documents that are inconsistent with or conflict with the content of this specification, as well as documents that limit the broadest scope of the claims in this specification (currently or subsequently appended to this specification). It should be noted that in the event of any inconsistency or conflict between the descriptions, definitions, and / or terminology used in the supplementary materials to this specification and the content of this specification, the descriptions, definitions, and / or terminology used in this specification shall prevail.
[0210] Finally, it should be understood that the embodiments described in this specification are merely illustrative of the principles of the embodiments described herein. Other variations may also fall within the scope of this specification. Therefore, alternative configurations of the embodiments described herein are intended to be illustrative rather than limiting, and should be considered consistent with the teachings of this specification. Accordingly, the embodiments described herein are not limited to those explicitly introduced and described herein.
Claims
1. An intelligent audio control system, comprising: The data collection device is configured to collect personnel information within a spatial area; The speaker is configured to play audio. A switching device is configured to connect the smart audio host and the speaker for signal transmission; The smart audio host is connected to the acquisition device and the switching device, and the smart audio host is configured as follows: Based on the personnel information, the target area is determined; as well as Control the target speaker within the target area to play the target audio based on the target playback parameters.
2. The system according to claim 1, wherein the acquisition device is further configured to acquire ambient noise information; and the smart audio host is further configured to: Adjust the target playback parameters based on the ambient noise information.
3. The system according to claim 1, wherein the smart audio host is further configured as follows: Based on the personnel information, determine whether the personnel have entered the viewing area; and In response to the person entering the viewing area, the target speaker corresponding to the viewing area is controlled to play the target audio.
4. The system according to claim 1, wherein the smart audio host is further configured as follows: Based on the personnel and exhibit information, determine the acoustic center of the target area; Based on the acoustic center, determine the main loudspeaker and auxiliary loudspeaker within the target area; Control the main speaker to play at the master volume; as well as The auxiliary volume is determined based on the distance between the auxiliary speaker and the acoustic center and the main volume, and the auxiliary speaker is controlled to play at the auxiliary volume.
5. In the system according to claim 1, the speaker is connected in series via a single network cable and connected to the smart audio host via the switching device.
6. A smart audio control method, the method being executed by a smart audio host, the smart audio host being connected to a data acquisition device and a switching device connected to multiple speakers, the method comprising: The target area is determined based on personnel information within the spatial area; wherein, the personnel information is collected by the acquisition device; and Control the target speaker within the target area to play the target audio based on the target playback parameters.
7. The method according to claim 6, further comprising: Adjust the target playback parameters based on the ambient noise information.
8. The method according to claim 6, wherein controlling the target loudspeaker within the target area to play target audio based on target playback parameters comprises: Based on the personnel information, determine whether the personnel have entered the viewing area; as well as In response to the person entering the viewing area, the target speaker corresponding to the viewing area is controlled to play the target audio.
9. The method according to claim 6, further comprising: Based on the personnel and exhibit information, determine the acoustic center of the target area; Based on the acoustic center, determine the main loudspeaker and auxiliary loudspeaker within the target area; Control the main speaker to play at the master volume; as well as The auxiliary volume is determined based on the distance between the auxiliary speaker and the acoustic center and the main volume, and the auxiliary speaker is controlled to play at the auxiliary volume.
10. A computer-readable storage medium storing computer instructions, wherein when a computer reads the computer instructions in the storage medium, the computer executes the intelligent audio control method as described in any one of claims 6 to 9.