A method and system for intelligent deployment of sound

By acquiring pedestrian flow and audio data, the system intelligently adjusts the configuration parameters of the speakers, solving the problems of wasted audio equipment resources and uneven sound effects, and achieving a better audio experience and resource utilization.

CN115967876BActive Publication Date: 2026-06-16HANSONG NANJING TECH LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310041601.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-12
Publication Date
2026-06-16
Estimated Expiration
2043-01-12

Smart Images

  • Figure CN115967876B_ABST
    Figure CN115967876B_ABST
Patent Text Reader

Abstract

The embodiment of the specification provides a sound intelligent deployment method and system, the method comprises: acquiring people flow distribution information and audio data of a target area; determining configuration parameters of at least one target sound among a plurality of sounds located in the target area based on the people flow distribution information and the audio data, the configuration parameters at least comprising a target position of the at least one target sound on a track of the target area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of intelligent speaker technology, and in particular to a method and system for intelligent speaker deployment. Background Technology

[0002] When playing music in commercial venues, there are too many fixed-location sound systems. This wastes sound resources, and the sound quality produced by these fixed-location systems is only good in certain areas of the venue. The sound quality deteriorates significantly outside these locations, resulting in a poor listening experience for users in the venue.

[0003] Therefore, we hope to propose a method and system for intelligent deployment of audio equipment, which can reduce the waste of audio equipment resources and intelligently improve the listening experience for users in the venue. Summary of the Invention

[0004] This specification provides one or more embodiments of an intelligent speaker deployment method, the method comprising: acquiring pedestrian flow distribution information and audio data in a target area; and determining, based on the pedestrian flow distribution information and audio data, configuration parameters of at least one target speaker among a plurality of speakers located in the target area, wherein the configuration parameters include at least the target position of at least one target speaker on a track in the target area.

[0005] This specification provides one or more embodiments of an intelligent audio deployment system, the system comprising: an acquisition module for acquiring pedestrian flow distribution information and audio data in a target area; and a determination module for determining, based on the pedestrian flow distribution information and audio data, configuration parameters of at least one target speaker among a plurality of speakers located in the target area, the configuration parameters including at least one target speaker's target position on a track in the target area.

[0006] One or more embodiments of the specification provide an intelligent sound system deployment device. The device includes a track set in a target area, multiple speakers located on the track, an information acquisition device, and a controller. The multiple speakers slide relative to the track. The controller is used to: determine the configuration parameters of at least one target speaker among the multiple speakers located in the target area based on pedestrian distribution information and audio data obtained by the information acquisition device. The configuration parameters include at least the target position of at least one target speaker on the track in the target area. The controller also controls at least one target speaker based on the configuration parameters.

[0007] This specification provides one or more embodiments of a computer-readable storage medium that stores computer instructions. When a computer reads the computer instructions in the storage medium, the computer executes the intelligent audio deployment method as described in any of the above embodiments. Attached Figure Description

[0008] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein:

[0009] Figure 1 These are schematic diagrams illustrating application scenarios of an intelligent audio deployment system based on some embodiments of this specification;

[0010] Figure 2 These are exemplary schematic diagrams of an intelligent audio deployment device according to some embodiments of this specification;

[0011] Figure 3 This is an exemplary flowchart of a method for intelligently deploying audio equipment according to some embodiments of this specification;

[0012] Figure 4 This is an exemplary schematic diagram illustrating the determination of configuration parameters based on a configuration model according to some embodiments of this specification;

[0013] Figure 5 This is an exemplary schematic diagram illustrating the determination of parameter configuration according to some embodiments of this specification;

[0014] Figure 6A This is an exemplary schematic diagram of the 2.0 mode according to some embodiments of this specification;

[0015] Figure 6B This is an exemplary schematic diagram of the 2.1 pattern shown according to some embodiments of this specification;

[0016] Figure 6C This is an exemplary schematic diagram of the 3.1 pattern shown according to some embodiments of this specification;

[0017] Figure 7A , Figure 7B This is an exemplary schematic diagram of the 5.1 pattern shown according to some embodiments of this specification;

[0018] Figure 7C This is an exemplary schematic diagram of the 7.1 pattern shown according to some embodiments of this specification;

[0019] Figure 7D This is an exemplary schematic diagram of the 9.1 pattern shown according to some embodiments of this specification;

[0020] Figure 8A , Figure 8B , Figure 8C These are exemplary schematic diagrams illustrating surround modes according to some embodiments of this specification. Detailed Implementation

[0021] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of this specification. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.

[0022] It should be understood that the terms “system,” “device,” “unit,” and / or “module” used herein are one way to distinguish different components, elements, parts, sections, or assemblies at different levels. However, if other terms can achieve the same purpose, they may be replaced by other expressions.

[0023] As indicated in this specification and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of expressly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.

[0024] Flowcharts are used in this specification to illustrate the operations performed by the system according to embodiments of this specification. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the steps can be processed in reverse order or simultaneously. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.

[0025] Figure 1 These are schematic diagrams illustrating application scenarios of an intelligent audio deployment system based on some embodiments of this specification. For example... Figure 1 As shown, the intelligent audio deployment system 100 may include a controller 110, a storage device 120, speakers 130, tracks 140, and an information acquisition device 150. In some embodiments, the components in the intelligent audio deployment system 100 may be connected and / or communicate with each other via a network (e.g., wireless connection, wired connection, or a combination thereof). In some embodiments, some components in the intelligent audio deployment system 100 may be directly connected.

[0026] The controller 110 can process information and / or data related to the intelligent sound system deployment system 100 to perform one or more functions described in this specification. For example, the controller 110 can acquire pedestrian flow distribution information and audio data in a target area; based on the pedestrian flow distribution information and audio data, it can determine the configuration parameters of at least one target sound among a plurality of sound systems located in the target area. In some embodiments, the controller 110 can process data and / or information acquired from the storage device 120 and the information acquisition device 150. For example, the controller 110 can process the pedestrian flow distribution information and audio data acquired by the information acquisition device 150.

[0027] In some embodiments, controller 110 may be a single server or a group of servers. In some embodiments, controller 110 may be local or remote. For example, controller 110 may be deployed near the speaker or at a distance from the speaker.

[0028] Storage device 120 can be used to store data, instructions, and / or any other information. In some embodiments, storage device 120 can store data and / or information obtained from controller 110, speaker 130, and / or information acquisition device 150, etc. For example, storage device 120 can store pedestrian distribution information and audio data of a target area obtained from information acquisition device 150. In some embodiments, storage device 120 can store data and / or instructions used by controller 110 to perform or use in order to complete the exemplary methods described herein. In some embodiments, storage device 120 can communicate with at least one other component (e.g., controller 110, speaker 130, information acquisition device 150) in the intelligent speaker deployment system 100. In some embodiments, storage device 120 can be part of controller 110.

[0029] Storage device 120 may include one or more storage components, each of which may be a separate device or part of another device. For example, storage device 120 may be located inside controller 110 or speaker 130.

[0030] The speaker 130 can be a hardware device for playing audio. In some embodiments, the speaker 130 can be mounted on a track 140 arranged in a target area and slide relative to the track. In some embodiments, the speaker 130 can include multiple speakers. The multiple speakers may be located in the same or different positions on the track 140. In some embodiments, the speaker 130 can be connected to a network to communicate with at least one other component (e.g., controller 110, storage device 120, information acquisition device 150) in the intelligent speaker deployment system 100. For more information about the speaker 130, see [link to relevant documentation]. Figure 2 And related explanations.

[0031] Track 140 refers to the cable channel through which the speaker 130 can move. In some embodiments, track 140 may include a traction mechanism and a drive sliding mechanism. The traction mechanism can be used to control the movement of the speaker 130 on track 140. The drive sliding mechanism can be disposed at the contact position between the speaker 130 and track 140, and is used to drive the speaker 130 to move on track 140. For more details about track 140, please refer to [link to relevant documentation]. Figure 2 And related explanations.

[0032] Information acquisition device 150 refers to a device or equipment for acquiring relevant information about a target area. Information acquisition device 150 may include a microphone, an infrared thermal imager, and a camera, etc. In some embodiments, information acquisition device 150 can communicate with at least one other component (e.g., controller 110, storage device 120, speaker 130) in the intelligent audio deployment system 100. For example, information acquisition device 150 can send the acquired information and / or data to controller 110 for processing, or it can send the acquired information and / or data to storage device 120 for storage. For more details about information acquisition device 150, please refer to [link to relevant documentation]. Figure 2 And related explanations.

[0033] It should be noted that the intelligent audio deployment system 100 is provided for illustrative purposes only and is not intended to limit the scope of this specification. Various modifications or variations can be made by those skilled in the art based on the description in this specification. For example, the intelligent audio deployment system 100 can be implemented on other devices to achieve similar or different functions. However, such variations and modifications will not depart from the scope of this specification.

[0034] Figure 2 This is an exemplary schematic diagram of an intelligent audio deployment device according to some embodiments of this specification. In some embodiments, the intelligent audio deployment device 200 may include a track disposed in a target area, a plurality of speakers located on the track, an information acquisition device, and a controller.

[0035] Tracks can be installed at multiple locations within the target area and are interconnected, allowing speakers to move along the tracks. Each target area can include multiple tracks, such as... Figure 2 As shown, the target area may include orbits a, b, c, d, e, f, and g.

[0036] In some embodiments, multiple tracks in each target area can be pre-installed in the target area according to a certain layout to form a track layout. The track layout can be adapted to the shape, size, and other characteristics of the target area. More information about track layouts can be found here. Figure 5 And its related descriptions.

[0037] In some embodiments, each track may be pre-set with a traction mechanism, a drive sliding mechanism, and multiple speakers. The traction mechanism can drive the drive sliding mechanism to move, thereby changing the position of the speakers in the track.

[0038] In some embodiments, one or more speakers may be installed on at least one track in the track layout. For example... Figure 2 As shown, a speaker 130-1 can be installed on track a, speakers 130-2 and 130-3 can be installed on track c, and speaker 130-4 can be installed on track e.

[0039] Multiple speakers on the track can move relative to the track. In some embodiments, the speakers can move and / or rotate relative to the track.

[0040] Multiple speakers on the track can adaptively determine their configuration parameters based on changes in the target area, resulting in better audio playback. In some embodiments, the configuration parameters of at least one target speaker can be adjusted based on pedestrian flow distribution information and audio information in the target area. Further details on adjusting configuration parameters can be found in the related description below.

[0041] An information acquisition device refers to a device or equipment used to collect relevant information about a target area. In some embodiments, the information acquisition device may include a microphone, an infrared thermal imager, and a camera. The microphone can be used to collect sound data of the target area; the infrared thermal imager can be used to collect a thermal imaging distribution map of the target area; and the camera can be used to collect regional images of the target area. In some embodiments, multiple information acquisition devices can be installed in a track, for example, such as... Figure 2 As shown, information acquisition devices 150-1, 150-2, 150-3, and 150-4 can be installed at the four corners of the track layout, and information acquisition device 150-5 can be installed in the center of the track layout. More information regarding the sound pickup equipment, thermal imaging distribution map, and area image can be found in step 310 and its related description.

[0042] The controller can be used to determine the configuration parameters of at least one target speaker among multiple speakers located within a target area, based on pedestrian distribution information and audio data acquired through an information acquisition device, and to control at least one target speaker based on the configuration parameters. For example, the controller can use a traction device to control the positional movement of at least one target speaker among multiple speakers on a track. As another example, the controller can control the number of target speakers currently playing, their rotation angle, acoustic parameters (e.g., volume, sound effects), etc. More information on determining configuration parameters can be found in [link to relevant documentation]. Figure 3 And its related descriptions.

[0043] Figure 3 This is an exemplary flowchart of a method for intelligently deploying audio systems according to some embodiments of this specification. In some embodiments, process 300 may be executed by a controller. Figure 3 As shown, process 300 includes the following steps:

[0044] Step 310: Obtain pedestrian flow distribution information and audio data for the target area. In some embodiments, step 310 may be performed by the acquisition module.

[0045] The target area refers to the region where audio needs to be played. For example, the target area could be a hotel lobby, a shopping mall, or similar areas.

[0046] Pedestrian distribution information refers to information that reflects the distribution of people in a target area. For example, pedestrian distribution information includes the location and number of each person in the target area.

[0047] Crowd distribution information can be characterized in various ways. In some embodiments, crowd distribution information may include thermal imaging distribution maps acquired by a thermal imager and / or area images captured by a camera. Crowd distribution information can also be represented by specific numerical values, such as the number of people and the location coordinates of each person.

[0048] A thermal imaging distribution map is an image used to record the heat or temperature radiated by people at different locations within a target area. Locations with a high concentration of people within the target area radiate higher heat or temperature, while sparsely populated areas radiate lower heat or temperature. Thermal imaging distribution maps use different colors to indicate the heat or temperature radiated by people at different locations within the target area. By mapping colors to temperatures, the distribution of people at different locations within the thermal imaging distribution map can be represented.

[0049] A region image is an image captured by a camera that contains a target area. For example, a region image could be an image captured by a camera that contains a lobby scene.

[0050] Audio data refers to sound-related data within a target area. In some embodiments, audio data includes ambient sounds within the target area. If audio is already playing from a speaker, the audio data also includes the audio played by the speaker. Ambient sounds may include other sounds within the target area besides the audio played by the speaker. For example, ambient sounds may include, but are not limited to, human voices, vehicle horns, and equipment operating sounds.

[0051] In some embodiments, audio data may include the sound type, volume, frequency, etc. of the sound in the target area.

[0052] In some embodiments, the audio data may include noise distribution information.

[0053] Noise distribution information refers to the distribution of sound within a target area. These sounds are considered noise relative to the audio played by a speaker. For example, noise distribution information can refer to the distribution of ambient sound within the target area.

[0054] In some embodiments, noise distribution information may include at least one of noise type and volume at multiple locations in the target area.

[0055] Noise type refers to the category of noise. In some embodiments, noise can be classified in various ways. For example, it can be classified according to noise frequency, and noise types can include low-frequency noise (dominant frequency below 300Hz), mid-frequency noise (dominant frequency between 300 and 800Hz), and high-frequency noise (dominant frequency above 800Hz), etc. Another example is that it can be classified according to noise source, and noise types can include: human voices, vehicle horns, equipment operating sounds, etc.

[0056] Volume refers to the loudness of noise. For example, volume can be 60dB.

[0057] In some embodiments, noise distribution information can be represented in the form of a noise distribution map. The noise distribution map may include at least one of the noise type and volume level at multiple locations within the target area. More information about noise distribution maps can be found in [link to documentation]. Figure 5 And its related descriptions.

[0058] In some embodiments, the controller can acquire audio data of the target area in various ways. For example, the controller can acquire audio data of the target area using multiple microphones deployed at different locations within the target area. A microphone is a device that receives and amplifies sound vibrations to obtain the ambient sound. For example, a microphone may include, but is not limited to, a microphone. More information about microphones can be found at [link to relevant documentation]. Figure 1 , Figure 2 And its related descriptions.

[0059] The controller can analyze the sound data collected by the microphone to determine the audio data or noise distribution information of the target area. In some embodiments, the controller can process the sound data collected by microphones at multiple locations using a sound recognition model to determine the noise type and volume. For example, the controller can process the sound data collected by the microphone at location A to determine the noise type and volume at location A. In some embodiments, a noise distribution map can be generated based on the noise type and volume at each location.

[0060] Voice recognition models can include one or any combination of neural network (NN) models, convolutional neural network (CNN) models, etc.

[0061] In some embodiments, the input to a sound recognition model may include sound data at a specific location. In some embodiments, the output of a sound recognition model may include the type and volume of noise at that location.

[0062] In some embodiments, the voice recognition model can be trained using multiple first training samples with first labels. For example, multiple first training data with first labels can be input into the initial voice recognition model. A loss function is constructed using the first labels and the output of the initial voice recognition model. The parameters of the initial voice recognition model are then iteratively updated based on the loss function. When the loss function of the initial voice recognition model satisfies a preset condition, the model training is complete, and a trained voice recognition model is obtained. The preset condition may be that the loss function converges, the number of iterations reaches a threshold, etc.

[0063] In some embodiments, the first training sample may include historical sound data from multiple locations collected by a sound pickup device. The first label may include the historical noise type and historical volume level corresponding to the historical sound data. The first label may be obtained through manual annotation.

[0064] Step 320: Based on pedestrian flow distribution information and audio data, determine the configuration parameters of at least one target speaker among multiple speakers located within the target area. In some embodiments, step 320 may be performed by a determining module.

[0065] A target speaker refers to a speaker that plays audio in a target area. In some embodiments, multiple speakers may be deployed on a track in the target area, and the target speaker may be at least one of the multiple speakers.

[0066] Configuration parameters refer to parameters related to audio playback by a target speaker in the target area. For example, configuration parameters could be parameters related to the location where the target speaker plays audio. Alternatively, configuration parameters could be parameters related to the sound effects and volume of the audio played by the target speaker.

[0067] In some embodiments, the configuration parameters include at least one target location of a target sound on a track in the target area.

[0068] The target position refers to the location on the track where the target speaker will play audio. Each target speaker can have a corresponding target position, and the positions of different target speakers can be the same or different.

[0069] The target location can be represented in several ways. For example, it can be represented by the track number of the target speaker and its distance from the leftmost or rightmost end of that track. As an example, the target location of speaker A could be a position on track a 2 meters away from the leftmost end of track a. Alternatively, the target location can be represented by coordinates, constructing a two-dimensional coordinate system to represent the target location.

[0070] In some embodiments, the configuration parameters may also include at least one of the following: the number of target speakers at the target location, the rotation angle, the volume, and the sound effect.

[0071] In some embodiments, the number of target speakers at the target location may include one or more. That is, there may be one or more target speakers at the same target location.

[0072] The rotation angle refers to the angle formed when a target speaker (e.g., the speaker's horn) is rotated relative to a specific direction. In some embodiments, the rotation angle may be the angle between the target speaker and the track.

[0073] Sound effects refer to the effects created by the sound played by the target speaker. For example, sound effects can include, but are not limited to, soothing, lyrical, and exciting sounds.

[0074] In some embodiments, configuration parameters can be represented in vector form. For example, configuration parameters can be represented as {([a, 1], 1, 90, 60, x), ([a, 10], 2, 75, 70, x)}. Here, [a, 1] represents the target position of target speaker 1 on track a in the target area, 1m from the leftmost end; 1 represents the number of target speakers 1 deployed at this target position is 1; 90 represents the rotation angle of target speaker 1 is 90°; 60 represents the volume of target speaker 1 is 60dB; and x represents the sound effect of target speaker 1. [a, 10] represents the target position of target speaker 2 on track a in the target area, 10m from the leftmost end; 2 represents the number of target speakers 2 deployed at this target position is 2; 75 represents the rotation angle of target speaker 2 is 75°; 70 represents the volume of target speaker 2 is 70dB; and x represents the sound effect of target speaker 2.

[0075] In some embodiments, the controller may determine the configuration parameters of the target speaker in a variety of ways.

[0076] In some embodiments, the number of target speakers at the target location can be set in various ways. For example, it can be set based on the volume of the noise. When the volume of the noise at the target location falls within a certain volume range, the number of target speakers corresponding to that volume range can be determined as the number of target speakers at the target location. The volume range and the number of target speakers corresponding to that volume range can be preset manually based on prior knowledge or historical data.

[0077] In some embodiments, the rotation angle of the target speaker at the target location can be set in various ways. For example, the rotation angle of the target speaker at the target location can be adjusted according to the distribution of pedestrian traffic, such as rotating the target speaker towards the direction with a higher pedestrian traffic.

[0078] In some embodiments, the volume of the target speaker at the target location can be adjusted according to the noise level. For example, when the noise level at the target location is high, the volume of the target speaker at that location can be increased.

[0079] In some embodiments, the controller can determine the configuration parameters of the target speaker based on pedestrian flow distribution information and audio data using a preset lookup table. In some embodiments, the preset lookup table includes multiple different reference pedestrian flow distribution information, multiple different reference audio data, and the correspondence between reference configuration parameters. In some embodiments, the preset lookup table can be obtained by constructing multiple different reference pedestrian flow distribution information, multiple different reference audio data, and the correspondence between reference configuration parameters based on prior knowledge or historical data (e.g., historical configuration data of the target area under different historical pedestrian flow distributions and different historical audio data).

[0080] In some embodiments, the controller can process pedestrian distribution information and audio data using a configuration model to determine configuration parameters. For more information on determining configuration parameters based on a configuration model, please refer to [link to relevant documentation]. Figure 4 And its related descriptions.

[0081] In some embodiments, the controller may acquire additional features and determine configuration parameters based on pedestrian distribution information, audio data, and additional features. More information on determining configuration parameters based on pedestrian distribution information, audio data, and additional features can be found at [link to relevant documentation]. Figure 5 And its related descriptions.

[0082] In some embodiments, after determining the configuration parameters of at least one target speaker, the controller can control the at least one target speaker to move to the corresponding target position via a traction mechanism. In some embodiments, when controlling the movement of the at least one target speaker, the controller can move the at least one target speaker according to a preset movement rule. An exemplary preset movement rule could be to minimize the total movement distance of the at least one target speaker. This movement method can effectively improve the deployment efficiency of the target speakers and save time and other costs.

[0083] In some embodiments described herein, the configuration parameters of the target speakers in the target area are determined by the pedestrian distribution information and audio data of the target area. The deployment position of the speakers can be automatically adjusted in a timely manner according to the distribution location and number of people, which can not only meet the needs of different scenarios, but also enable users in the target area to have a better experience.

[0084] Figure 4 This is an exemplary schematic diagram illustrating the determination of configuration parameters based on a configuration model according to some embodiments of this specification.

[0085] In some embodiments, the controller can process pedestrian flow distribution information and audio data through a configuration model to determine configuration parameters.

[0086] The configuration model can be a machine learning model. For example, the configuration model can include one or any combination of neural network (NN) models, convolutional neural network (CNN) models, etc.

[0087] In some embodiments, the input to the configuration model may include pedestrian flow distribution information and audio data. More information regarding pedestrian flow distribution information and audio data can be found in step 310 and its related description. In some embodiments, the output of the configuration model may include configuration parameters for the target speaker. More information regarding the configuration parameters can be found in step 320 and its related description.

[0088] In some embodiments, the input to the configuration model may also include music genre. Music genre refers to the style of music. For example, music genre may include, but is not limited to, classical music, pop music, jazz, rock music, etc.

[0089] In some embodiments, the configuration model can be obtained by training multiple second training samples with second labels. For example, multiple second training data with second labels can be input into the initial configuration model, and a loss function can be constructed using the second labels and the output of the initial configuration model. The parameters of the initial configuration model are then iteratively updated based on the loss function. When the loss function of the initial configuration model satisfies a preset condition, the model training is complete, and a trained configuration model is obtained. The preset condition may be that the loss function converges, the number of iterations reaches a threshold, etc.

[0090] In some embodiments, the second training samples may include sample pedestrian distribution information and sample audio data. The second label may include the actual configuration parameters corresponding to the sample pedestrian distribution information and sample audio data. In some embodiments, the second training samples may also include sample music genres.

[0091] In some embodiments, the second training sample can be determined based on historical pedestrian flow distribution information and historical audio data in the historical target area.

[0092] In some embodiments, the second label can be determined based on the boundary points of the historical thermal imaging distribution map corresponding to the historical target area. Boundary points refer to the locations of the boundary points of the pedestrian flow distribution in the thermal imaging distribution map. In some embodiments, the controller can process the boundary points of the historical thermal imaging distribution map using preset determination rules to determine the second label. In some embodiments, the preset determination rules may be: based on the point layout corresponding to the playback mode set by the user, determining the intersection or proximity of the boundary points of the thermal imaging distribution map with multiple tracks in the target area as the target location. More information about the user-set playback modes can be found in [link to relevant documentation]. Figure 5 And its related descriptions. For example, such as Figure 7A As shown, the user-set playback mode is 5.1 mode. 710 is the thermal imaging distribution map of the target area X. The thermal imaging distribution map 710 spans tracks b, c, and d. Therefore, three positions (i.e., positions 721, 722, and 723) where multiple boundary points in the thermal imaging distribution map 710 intersect or are close to track b in the target area X can be determined as the target position. Similarly, two positions (i.e., positions 725 and 726) where multiple boundary points in the thermal imaging distribution map 710 intersect or are close to track d in the target area X can be determined as the target position. In some embodiments, the second tag can be determined based on the center point of the historical thermal imaging distribution map corresponding to the historical target area. The center point refers to the isocenter point of the thermal imaging distribution map or a position close to the isocenter point of the thermal imaging distribution map. In some embodiments, the controller can determine the target position as the position where the center point of the thermal imaging distribution map intersects or is close to a track in the target area. For example, as... Figure 7AAs shown, the target location is determined by the position of the center point 711 in the thermal imaging distribution map 710 that is close to the position of track c in the target area X (i.e., the position corresponding to 724). In some embodiments, the second label can be determined by manual annotation.

[0093] In some embodiments of this specification, the configuration model is used to process the pedestrian distribution information and audio data to determine the configuration parameters, which not only improves the efficiency of the determination process, but also makes the final determined configuration parameters more accurate.

[0094] In some embodiments, the controller may perform at least one round of iterative updates on the output of the configuration model to determine the final configuration parameters. This at least one round of iterative updates may include: determining candidate configuration parameters; processing the candidate configuration parameters based on the effect model to determine effect parameters; adjusting the candidate configuration parameters based on the effect parameters; inputting the adjusted candidate configuration parameters into the effect model for processing to determine the updated effect parameters; and ending the iteration when the updated effect parameters meet preset conditions.

[0095] In the first iteration, the controller determines the initial candidate configuration parameters based on the configuration model output. It then processes these initial candidate configuration parameters using the effect model to determine the effect parameters. Finally, it adjusts the initial candidate configuration parameters based on the effect parameters to obtain the adjusted candidate configuration parameters. In the nth iteration (n ≥ 1), the controller processes the adjusted candidate configuration parameters from the previous iteration using the effect model to determine the updated effect parameters. It then adjusts the adjusted candidate configuration parameters based on the updated effect parameters to obtain the adjusted candidate configuration parameters. This process is repeated until the updated effect parameters meet preset conditions, at which point the iteration ends.

[0096] Effect parameters can be used to reflect the playback effect of audio playback according to candidate configuration parameters. For example, effect parameters can be used to reflect the volume and clarity effects of audio playback according to candidate configuration parameters. In some embodiments, effect parameters can be represented by real numbers between 0 and 1. The larger the value, the better the volume and clarity effects of audio playback according to candidate configuration parameters.

[0097] In some embodiments, the effect model can be a machine learning model. For example, the effect model can include one or any combination of neural network (NN) models, convolutional neural network (CNN) models, etc.

[0098] In some embodiments, the input to the effect model may include candidate configuration parameters and scene parameters, and the output may include effect parameters corresponding to the candidate configuration parameters. Scene parameters refer to scene-related parameters of the target region. For more information on scene parameters, please refer to [link to relevant documentation]. Figure 5 And its related descriptions.

[0099] In some embodiments, the performance model can be obtained by training multiple third training samples with third labels. For example, multiple third training data with third labels can be input into the initial performance model, and a loss function can be constructed using the third labels and the output of the initial performance model. The parameters of the initial performance model are then iteratively updated based on the loss function. When the loss function of the initial performance model satisfies a preset condition, the model training is complete, and the trained performance model is obtained. The preset condition may be that the loss function converges, the number of iterations reaches a threshold, etc.

[0100] In some embodiments, the third training sample may include candidate configuration parameters for the target region and scene parameters for the target region. In some embodiments, the third training sample may be determined based on historical configuration data of the target region. The third label may include sample effect parameters corresponding to the candidate configuration parameters.

[0101] In some embodiments, the controller can direct a sound-collecting robot to move to the target area of ​​the sample to collect the playback audio corresponding to the candidate configuration parameters of the sample, and analyze the playback audio to generate sample effect parameters as a third label. For example, the sample effect parameters can be determined by analyzing the playback audio using an algorithm or machine learning model. Alternatively, the sample effect parameters can be determined manually based on the playback audio. The sound-collecting robot can be an intelligent robot equipped with a sound-collecting device.

[0102] By combining intelligent robots, a large amount of training data can be obtained, ensuring that the training effect and model accuracy are high enough, while reducing the human cost caused by the collection of third training samples.

[0103] In some embodiments, the controller can adjust the candidate configuration parameters based on the difference between the effect parameters and the standard effect parameters to determine the adjusted candidate configuration parameters. The standard effect parameters can be used to determine whether the candidate configuration parameters need adjustment. The standard effect parameters can be system default values, empirical values, manually preset values, or any combination thereof, and can be set according to actual needs; this specification does not impose any restrictions on this.

[0104] In some embodiments, the controller can determine the adjustment method by comparing the effect difference between the effect parameter and the standard effect parameter with a preset table. The preset table contains various reference effect differences and their corresponding reference adjustment methods. During the comparison, the actual effect difference is matched with the reference effect difference, and the reference effect difference that meets the preset conditions (e.g., the same or closest) and its corresponding reference adjustment method are taken as the final adjustment method.

[0105] In some embodiments, the controller can process the effect difference based on a preset algorithm to determine the corresponding adjustment method. The preset algorithm includes an algorithm that converts the effect difference into adjustment values ​​for number, volume, target position, and rotation angle.

[0106] In some embodiments of this specification, the final configuration parameters are determined by iteratively updating the configuration parameters output by the configuration model, so that the final configuration parameters are more suitable for the scene of the target area, allowing users to have a better sound effect experience.

[0107] Figure 5 This is an exemplary flowchart illustrating the determination of configuration parameters according to some embodiments of this specification. In some embodiments, process 500 may be executed by a controller. Figure 5 As shown, process 500 includes the following steps:

[0108] Step 510: Obtain additional features.

[0109] Additional features may refer to feature data used to determine configuration parameters. In some embodiments, additional features may include at least one of scene parameters and user-set playback modes.

[0110] Scene parameters refer to parameters related to the scene of the target area. In some embodiments, scene parameters may include scene size (e.g., scene size could be 10m*10m, etc.), scene type (e.g., scene type could be KTV, bar, etc.), scene shape (e.g., scene shape could be square, circular, irregular, etc.), track layout, etc. The track layout refers to the arrangement of tracks within the target area. The track layout may include track length, number of tracks, spacing between tracks, etc. In some embodiments, different target areas may have different track layouts.

[0111] In some embodiments, the track layout can be set according to the scene size, scene type, and scene shape of the target area. For example, assuming that the scene size of the target area A is 10m*10m, the scene shape is square, and the scene type is KTV, the track layout of the target area A can be set to (10, 5, 2), which means that the length of the track in the track layout is 10m, the number of tracks is 5, and the tracks are equally spaced with a spacing of 2m.

[0112] The user-defined playback mode refers to the layout pattern of multiple target speakers pre-arranged by the user. In some embodiments, the user-defined playback mode may include 2.0 mode, 2.1 mode, 3.1 mode, 5.1 mode, 7.1 mode, 9.1 mode, surround mode, etc.

[0113] In each playback mode, a certain number of points can be selected and arranged according to the point layout corresponding to each playback mode. Among them, the points can be used as the target locations of the target speakers.

[0114] In some embodiments, the number of points selected in mode 2.0 is two. The point layout in mode 2.0 (also called a two-point layout) involves placing one point at each of the intersections of any two vertices (or near two vertices) on the thermal imaging distribution map with the track, for a total of two points. As an example, such as... Figure 6A As shown, the two-point layout involves placing one point at each of the two vertices (or near the vertices) of the thermal imaging distribution map 610 where it intersects with track b. These two points are designated as point 621 and point 622. Any two vertices can be selected based on the thermal imaging distribution map, specifically in areas with high pedestrian density. A vertex can refer to the vertex of the circumscribed polygon of the thermal imaging distribution map (e.g., a circumscribed rectangle, circumscribed hexagon, etc.), and the circumscribed polygon can be determined based on the shape of the thermal imaging distribution map.

[0115] In some embodiments, the number of points selected in mode 2.1 is three. The point layout in mode 2.1 (also called a three-point layout) involves setting one point at each of the intersections of any two vertices (or near the vertices) on the thermal imaging distribution map with the track, and setting one point between these two points according to a preset rule, for a total of three points. For example, the preset rule could be: when two points are on the same track, a point can be set at any location on that track (e.g., a location with high pedestrian traffic or the midpoint). Another example could be: when two points are on different tracks and the tracks are not adjacent, a point can be set at any location on the track in the middle position. As an example, such as... Figure 6B As shown, the three-point layout is as follows: one point is set at each of the two vertices (or near the vertices) of the thermal imaging distribution map 630 where it intersects with the track b. The two points are point 641 and point 643, and a point 642 is set at the midpoint between point 641 and point 643.

[0116] In some embodiments, four points are selected in mode 3.1. The point layout in mode 3.1 (also known as a four-point layout) involves setting one point at each of the intersections of any two vertices (or near the vertices) on the thermal imaging distribution map with the track, setting one point between these two points according to a preset rule, and setting one point based on the center point of the thermal imaging distribution map, for a total of four points. More information on setting the target position based on the center point can be found in [link to relevant documentation]. Figure 4 And its related description. For example, such as... Figure 6C As shown, the four-point layout can be configured by setting one point at each of the two vertices (or near the vertices) of the thermal imaging distribution map 650 where it intersects with the track b, with the two points being point 661 and point 663 respectively; setting a point 662 at the midpoint between point 661 and point 663; and setting a point 664 according to the center point of the thermal imaging distribution map 650.

[0117] In some embodiments, the number of points selected in mode 5.1 is six. The point layout in mode 5.1 (also called a six-point layout) involves setting one point at each of the four vertices (or near the vertices) of the thermal imaging distribution map where they intersect with the track, setting one point between any two points according to preset rules (e.g., between two points with high pedestrian traffic), and setting one point based on the center point of the thermal imaging distribution map, for a total of six points. As an example, such as... Figure 7A As shown, the six-point layout can be configured by setting one point at each of the four vertices (or near the vertices) of the thermal imaging distribution map 710 where they intersect with tracks b and d. The four points are point 721, point 723, point 725, and point 726. A point 722 is set at the midpoint between point 721 and point 723. A point 724 is set according to the center point 711 of the thermal imaging distribution map 710.

[0118] In some embodiments, the number of points selected in mode 7.1 is eight. The point layout in mode 7.1 (also known as an eight-point layout) involves setting one point at each of the four vertices (or near the vertices) of the thermal imaging distribution map where they intersect with the track; setting one point at the center point of the thermal imaging distribution map; setting one point at each of the two locations where the track of that center point intersects with or is close to the thermal imaging distribution map; and setting one point between two points on the same track according to a preset rule. As an example, such as... Figure 7CAs shown, the eight-point layout can be configured by setting one point at each of the four vertices (or near the vertices) of the thermal imaging distribution map 750 where they intersect with tracks b and d. The four points are point 761, point 763, point 767, and point 768. A point is set as point 765 based on the center point of the thermal imaging distribution map 750. Two points are set at the two locations where track c, where point 765 is located, intersects with or is close to the thermal imaging distribution map 750, namely point 764 and point 766. A point is set as point 762 at the midpoint between point 761 and point 763.

[0119] In some embodiments, the number of points selected in the 9.1 mode is ten. The points in the 9.1 mode can be distributed across four or more tracks. The point layout in the 9.1 mode (also known as a ten-point layout) involves setting one point at each of the four vertices (or near the vertices) of the thermal imaging distribution map where they intersect with the tracks; setting one point at the center point of the thermal imaging distribution map; setting one point at each of the two points where the track containing that point intersects with the thermal imaging distribution map; setting one point at each of the two points where the middle track (different from the track near the center point) intersects with the thermal imaging distribution map among the multiple tracks spanned by the four vertices; and setting one point between two points on the same track according to a preset rule. As an example, such as... Figure 7D As shown, the ten-point layout can be configured as follows: one point is set at each of the four vertices (or near the vertices) of the thermal imaging distribution map 770 where they intersect with tracks a and d, and the four points are point 781, point 783, point 789, and point 790. A point is set as point 787 based on the center point of the thermal imaging distribution map 770. Two points are set at the two locations where track c, where point 787 is located, intersects with or is close to the thermal imaging distribution map 770, and points are set as points 786 and 788, respectively. Two points are set at the two locations where track b (which is different from track c, which is close to the center point) intersects with the thermal imaging distribution map 770, and points are set as points 784 and 785, respectively. A point is set as point 782 at the midpoint between point 781 and point 783.

[0120] It should be noted that the order in which the above-mentioned points are set is not limited in some embodiments of this specification.

[0121] In some embodiments, the number of points and the layout of points (i.e., the two-point, three-point, four-point, five-point, eight-point, and ten-point layouts mentioned above) in modes 2.0, 2.1, 3.1, 5.1, 7.1, and / or 9.1 can remain fixed, while the size of the point layout can be adjusted according to actual conditions. For example, when the thermal imaging distribution map occupies a large proportion of the target area (i.e., when there is a large flow of people), the size of the point layout can be larger; when the thermal imaging distribution map occupies a small proportion of the target area (i.e., when there is a small flow of people), the size of the point layout can be smaller. As an example, such as Figure 7A As shown, when the thermal imaging distribution map 710 occupies a small proportion of the target area X, the size of its point layout can be small: points 721, 722, and 723 can be set on track b, and points 725 and 726 can be set on track d. As an example, such as... Figure 7B As shown, when the thermal imaging distribution map 730 occupies a large proportion of the target area X, the size of its point layout can be large: points 741, 742 and 743 can be set on track a, point 744 can be set on track c, and points 745 and 746 can be set on track e.

[0122] In some embodiments, the points selected in the surround mode can surround the thermal imaging distribution map, forming a surround layout. This specification does not limit the number of points selected in the surround mode. For example, one point can be set at each of the four vertices in the thermal imaging distribution map, and one or more points can be set between every two vertices according to preset rules. As an example, such as... Figure 8A As shown, when the user sets the playback mode to surround mode, a point can be set at each of the four vertices of the thermal imaging distribution map 810. The four points are point 821, point 823, point 826 and point 828, and a point 822 can be selected between point 821 and point 823, a point 824 can be selected between point 821 and point 826, a point 825 can be selected between point 823 and point 828, and a point 827 can be selected between point 826 and point 828.

[0123] In some embodiments, the layout of the points in the surround mode (i.e., the aforementioned four-sided surround layout) can remain fixed, while the number of points and the size of the point layout can be adjusted according to actual conditions. For example, when the thermal imaging distribution map occupies a large proportion of the target area (i.e., when there is a lot of foot traffic), the size of the point layout can be larger; when the thermal imaging distribution map occupies a small proportion of the target area (i.e., when there is a lot of foot traffic), the size of the point layout can be smaller. As an example, such as Figure 8AAs shown, when the thermal imaging distribution map 810 occupies a small proportion of the target area Y, the size of its point layout can be small: points 821, 822, and 823 can be set on track b, points 824 and 825 can be set on track c, and points 826, 827, and 828 can be set on track d. Figure 8B As shown, when the thermal imaging distribution map 830 occupies a large proportion of the target area Y, the size of its point layout can be larger: points 841, 842, and 843 can be set on track a, points 844 and 845 on track c, and points 846, 847, and 848 on track e. For example, when the thermal imaging distribution map occupies a large proportion of the target area (i.e., when there is a lot of foot traffic), the number of points can be larger; when the thermal imaging distribution map occupies a small proportion of the target area (i.e., when there is a lot of foot traffic), the number of points can be smaller. As an example, such as... Figure 8A As shown, when the thermal imaging distribution map 810 occupies a small proportion of the target area Y, the number of its points can be small, typically eight. For example... Figure 8C As shown, when the thermal imaging distribution map 830 occupies a large proportion of the target area Y, the number of its points can be relatively large, which can be achieved by... Figure 8B Based on the existing eight points, eight new points will be added: point 851 to point 858.

[0124] In some embodiments, the number of speakers at each location in the various playback modes described above can be adjusted according to actual conditions. The number of speakers at each location can be changed by the same multiple. For example, if there was originally one speaker at each location, it can be changed to two speakers at each location.

[0125] In some embodiments, the user-set playback mode can also be a static mode. In static mode, the target speaker locations can be selected based on the track layout. For example, target speaker locations can be selected at each vertex of the track layout (hereinafter referred to as vertex layout). Assuming the track layout is square, four locations can be selected at the four vertices. The number of locations, the location layout, and the number of speakers at each location remain unchanged in static mode, while the rotation angle, volume, and sound effects of the target speakers can be adjusted according to actual conditions. When the user-set playback mode is static mode, the controller can determine the target location based on the vertex layout and control at least one target speaker from multiple speakers to move to the target location.

[0126] By setting a static mode, when there is too little time to arrange the sound system in a target area due to rapid changes in pedestrian traffic or scattered pedestrian flow, the sound system can still provide a good sound experience for most locations in the target area, thus improving the user experience.

[0127] In some embodiments, the controller can determine additional features in various ways. For example, the controller can process information collected by the information acquisition device to determine scene parameters of the target area. The controller can perform image recognition on area images captured by the camera to determine the scene size, scene type, scene shape, track distribution map, etc. of the target area. As another example, the playback mode set by the user can be stored in a storage device (e.g., storage device 120), and the controller can obtain the playback mode set by the user by accessing the storage device.

[0128] Step 520: Determine configuration parameters based on pedestrian flow distribution information, audio data, and additional features.

[0129] In some embodiments, the controller can process pedestrian distribution information, audio data, and additional features through a configuration model to determine configuration parameters. Accordingly, the input to the configuration model may include pedestrian distribution information, audio data, and additional features, and the output may be configuration parameters.

[0130] In some embodiments, the configuration model can determine the approximate location distribution of the target speakers and the initial number of target speakers at each location based on the playback mode set by the user in the additional features. For example, when the user sets the playback mode to 5.1 mode, the approximate location distribution of the target speakers can be determined to be a six-point layout, with one target speaker at each location initially. When the user sets the playback mode to surround mode, the approximate location distribution of the target speakers can be determined to be a four-sided surround layout, with one target speaker at each location initially. When the user sets the playback mode to static mode, the approximate location distribution of the target speakers can be determined to be a vertex layout, with one target speaker at each location initially.

[0131] In some embodiments, the configuration model can further determine the target location, rotation angle, number of target speakers at each target location, volume, sound effects, etc., based on pedestrian distribution information, noise distribution information in audio data, and scene parameters in additional features.

[0132] In some embodiments, the second training samples used to train the configuration model may further include sample-added features of the target region. More information on training configuration models can be found at [link to relevant documentation]. Figure 4 The details and related descriptions will not be repeated here.

[0133] In some embodiments of this specification, by acquiring additional features and then using pedestrian flow information, audio data, and additional features to determine configuration parameters, the position of the speakers can be dynamically adjusted according to the actual situation on site, so that users can have the best sound experience in any location on site.

[0134] In some embodiments, the controller may divide the target area into multiple target sub-regions based on at least one of a noise distribution map and a thermal imaging distribution map; determine sub-configuration parameters corresponding to the multiple target sub-regions based on the sub-noise distribution map, sub-thermal imaging distribution map, sub-scene parameters, and user-set playback mode corresponding to the target sub-regions; and determine the configuration parameters of the target area based on the sub-configuration parameters of the multiple target sub-regions.

[0135] A noise distribution map is a graph that reflects the distribution of noise at various locations within a target area. Noise refers to any sound other than that emitted by an audio system. For example, a noise distribution map can reflect at least one of the noise type and volume level at various locations within the target area. In some embodiments, different noise types and volume levels can be labeled on the noise distribution map. For example, L, M, and H can be used as labels to represent low-frequency noise, mid-frequency noise, and high-frequency noise, respectively. Alternatively, specific decibel values ​​can be used as labels to represent volume levels.

[0136] In some embodiments, the controller can acquire the noise distribution map of the target area in various ways. For example, the controller can process the noise distribution information using acoustic algorithms or acoustic simulation software to acquire the noise distribution map of the target area.

[0137] A target sub-region refers to a portion of the target region.

[0138] The target sub-region can be divided in various ways. In some embodiments, the controller can divide the target region into multiple target sub-regions based on at least one of a noise distribution map and a thermal imaging distribution map, using preset division conditions. Exemplary preset division conditions may include dividing the target region according to a pre-set volume range in the noise distribution map and a color in the thermal imaging distribution map. For example, the portion of the noise distribution map corresponding to target region A in the volume range of 0-20dB can be divided into target sub-region A1, the portion corresponding to target region A in the volume range of 20-40dB can be divided into target sub-region A2, and the portion corresponding to target region A in the volume range of 40-60dB can be divided into target sub-region A3, etc. As another example, the portion of the thermal imaging distribution map corresponding to target region B that is displayed in blue can be divided into target sub-region B1, the portion corresponding to target region B that is displayed in yellow can be divided into target sub-region B2, and the portion corresponding to target region B that is displayed in red can be divided into target sub-region B3, etc. Exemplary preset division conditions may also include dividing by a specific area average. For example, a 10m*10m target area can be divided into 25 target sub-areas with a specific area of ​​2m*2m.

[0139] In some embodiments, the controller can also process the noise distribution map and thermal imaging distribution map through a segmentation model to determine multiple target sub-regions.

[0140] In some embodiments, the segmentation model can be a machine learning model. For example, the segmentation model can be an object detection model (You Only Look Once, YOLO).

[0141] In some embodiments, the input to the segmentation model may include a noise distribution map and a thermal imaging distribution map, and the output may include multiple target sub-regions.

[0142] In some embodiments, the input to the segmentation model may include a noise distribution map, a thermal imaging distribution map, and a track distribution map, and the output may include multiple target sub-regions. The track distribution map may be an image representing the track layout. In some embodiments, the track distribution map may be acquired by capturing images of the track layout using a camera.

[0143] In some embodiments, the segmentation model can be trained using multiple third training samples with third labels. For example, multiple third training samples with third labels can be input into the initial segmentation model, and a loss function can be constructed using the third labels and the results of the initial segmentation model. The parameters of the segmentation model are then iteratively updated based on the loss function. When the loss function of the initial segmentation model satisfies a preset condition, the model training is complete, and a trained segmentation model is obtained. The preset condition could be loss function convergence, the number of iterations reaching a threshold, etc.

[0144] In some embodiments, the third training sample may include a sample noise distribution map and a sample thermal imaging distribution map of the sample target region. In some embodiments, the third training sample may include a sample noise distribution map, a sample thermal imaging distribution map, and a sample orbit distribution map of the sample target region. In some embodiments, historical noise distribution maps and historical thermal imaging distribution maps of historical target regions may be used as the third training sample. In some embodiments, historical noise distribution maps, historical thermal imaging distribution maps, and historical orbit distribution maps of historical target regions may be used as the third training sample.

[0145] In some embodiments, the third label can be multiple target sub-regions of the sample target region. The third label can be obtained through manual annotation.

[0146] In some embodiments of this specification, multiple target sub-regions are determined by segmentation models. Areas with large differences in thermal imaging distribution or noise distribution in the target region can be divided into different target sub-regions, thereby intelligently determining the target sub-regions and making the division of target sub-regions more accurate.

[0147] After dividing the target area into multiple target sub-regions, the noise distribution map of the target area can be divided into multiple sub-noise distribution maps using the same method, and the thermal imaging distribution map of the target area can be divided into multiple sub-thermal imaging maps using the same method. That is, the sub-noise distribution map and the sub-thermal imaging map corresponding to a certain target sub-region in the target area are corresponding.

[0148] A sub-noise distribution map is a map that reflects the noise distribution within a target sub-region. Similar to a standard noise distribution map, more information about sub-noise distribution maps can be found in the preceding description of noise distribution maps.

[0149] A sub-thermal imaging distribution map is a map that reflects the location and number of people in a target sub-region. Sub-thermal imaging distribution maps are similar to thermal imaging distribution maps; for more information on sub-thermal imaging distribution maps, please refer to the previous description of thermal imaging distribution maps.

[0150] Sub-scene parameters refer to scene-related parameters within the target sub-region. Similar to scene parameters, more information about sub-scene parameters can be found in the preceding description of scene parameters.

[0151] Sub-configuration parameters refer to parameters related to the playback of target audio in a target sub-region. In some embodiments, sub-configuration parameters include at least one target position of a target audio on the track of the target sub-region. Sub-configuration parameters are similar to configuration parameters; for more information about sub-configuration parameters, please refer to the relevant description of configuration parameters above.

[0152] Sub-configuration parameters can be determined using methods specific to configuration parameters. For example, the controller can process the sub-noise distribution map, sub-thermal imaging distribution map, sub-scene parameters, and user-set playback mode of each target sub-region through a configuration model to determine the sub-configuration parameters for each target sub-region. For details on determining sub-configuration parameters, please refer to the previous description of determining configuration parameters.

[0153] In some embodiments, the controller can determine the configuration parameters of a target region through sub-configuration parameters of multiple target sub-regions. For example, the controller can merge the sub-configuration parameters of different target sub-regions to determine the configuration parameters of the target region. As an example, the sub-configuration parameters of target sub-region A1 are {([a, 1], 1, 90, 60, y), ([b, 1], 1, 75, 60, y)}, and the sub-configuration parameters of target sub-region A2 are {([a, 10], 1, 90, 60, y), ([b, 10], 1, 75, 60, y)}. Target region A is composed of target sub-region A1 and target sub-region A2, so the configuration parameters of target region A can be determined as {([a, 1], 1, 90, 60, y), ([b, 1], 1, 75, 60, y), ([a, 10], 1, 90, 60, y), ([b, 10], 1, 75, 60, y)}. For more explanation of the meaning of vectors, please refer to [link / reference]. Figure 3 And its related descriptions.

[0154] In some embodiments, the controller can perform fusion processing on the sub-configuration parameters of multiple target sub-regions using preset fusion rules to determine the configuration parameters of the target region.

[0155] Preset fusion rules refer to the conditions that must be met when fusion processing is performed on the sub-configuration parameters of the target sub-region.

[0156] In some embodiments, when the sub-configuration parameters of target speakers on a shared track differ in adjacent target sub-regions, the controller can determine the configuration parameters of the target region based on a preset fusion rule. Here, a shared track can refer to a track that is simultaneously divided into two target sub-regions.

[0157] In some embodiments, the preset fusion rule may be: to fuse the sub-configuration parameters of target speakers on a shared track in adjacent target sub-regions to determine the configuration parameters of the shared track. Exemplary fusion processing methods may include averaging, etc. For example, if track a is simultaneously divided into adjacent target sub-regions A1 and A2 within target region A, and the sub-configuration parameters of track a in target sub-region A1 are ([a, 6], 1, 90, 60, y), while the sub-configuration parameters of track a in target sub-region A2 are ([a, 8], 1, 100, 80, y), then the configuration parameters of track a in target region A can be determined as ([a, 7], 1, 95, 70, y). Exemplary fusion processing methods may also include weighted fusion, etc. The weights can be determined based on the pedestrian flow distribution information and noise distribution information of the target sub-regions. For example, the target sub-regions with concentrated pedestrian flow and higher noise levels have larger weights.

[0158] In some embodiments, the preset fusion rule may be: arbitrarily selecting one of the sub-configuration parameters of the target audio on the shared track in adjacent target sub-regions as the configuration parameter of the shared track. For example, in the example above, the sub-configuration parameter of track a in target sub-region A2 can be arbitrarily selected as ([a, 8], 1, 100, 80, y) as the final configuration parameter of track a.

[0159] In some embodiments of this specification, by dividing multiple target sub-regions to determine their respective sub-configuration parameters, and by processing the sub-configuration parameters of multiple target sub-regions according to preset merging rules to determine the configuration parameters of the target region, the most suitable sub-configuration parameters for each target sub-region can be determined through targeted processing of multiple target sub-regions. Then, by processing the sub-configuration parameters of each target sub-region through certain preset merging rules, the configuration parameters of the target region can achieve better regional sound effects while saving configuration costs.

[0160] It should be noted that the above descriptions of processes 300 and 500 are for illustrative purposes only and do not limit the scope of this specification. Those skilled in the art can make various modifications and changes to processes 300 and 500 under the guidance of this specification. However, these modifications and changes remain within the scope of this specification.

[0161] Some embodiments of this specification also provide an intelligent sound system deployment system, which includes an acquisition module for acquiring pedestrian distribution information and audio data in a target area; and a determination module for determining, based on the pedestrian distribution information and audio data, the configuration parameters of at least one target sound among a plurality of sound systems located in the target area, wherein the configuration parameters include at least the target position of at least one target sound on the track in the target area.

[0162] Some embodiments of this specification also provide a computer-readable storage medium for storing computer instructions, which, when executed by a processor, implement the intelligent audio deployment method as described in this specification.

[0163] The basic concepts have been described above. Obviously, for those skilled in the art, the detailed disclosure above is merely illustrative and does not constitute a limitation of this specification. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and corrections to this specification. Such modifications, improvements, and corrections are suggested in this specification and therefore remain within the spirit and scope of the exemplary embodiments described herein.

[0164] Furthermore, this specification uses specific terms to describe embodiments thereof. For example, "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Moreover, certain features, structures, or characteristics in one or more embodiments of this specification can be appropriately combined.

[0165] Furthermore, unless expressly stated in the claims, the order of processing elements and sequences, the use of numbers and letters, or other names described in this specification are not intended to limit the order of the processes and methods described herein. Although various examples have been discussed in the foregoing disclosure of some embodiments of the invention that are currently considered useful, it should be understood that such details are for illustrative purposes only, and the appended claims are not limited to the disclosed embodiments; rather, the claims are intended to cover all modifications and equivalent combinations that conform to the spirit and scope of the embodiments described herein. For example, while the system components described above can be implemented using hardware devices, they can also be implemented solely using software solutions, such as installing the described system on existing servers or mobile devices.

[0166] Similarly, it should be noted that, in order to simplify the description disclosed herein and thus aid in the understanding of one or more embodiments of the invention, the foregoing description of embodiments in this specification may sometimes combine multiple features into a single embodiment, drawing, or description thereof. However, this method of disclosure does not imply that the subject matter of this specification requires more features than those mentioned in the claims. In fact, the embodiments contain fewer features than all the features of a single embodiment disclosed above.

[0167] In some embodiments, numbers describing the quantity of components and attributes are used. It should be understood that such numbers used in the description of embodiments are modified in some examples with the terms "approximately," "approximately," or "generally." Unless otherwise stated, "approximately," "approximately," or "generally" indicates that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, which may be changed depending on the characteristics required by individual embodiments. In some embodiments, numerical parameters should take into account specified significant digits and employ a general method of digit reservation. Although the numerical ranges and parameters used to confirm their breadth of range in some embodiments of this specification are approximate values, in specific embodiments, such values ​​are set as precisely as feasible.

[0168] For each patent, patent application, patent application publication, and other material, such as articles, books, specifications, publications, and documents, referenced in this specification, the entire contents of which are incorporated herein by reference. This excludes historical application documents that are inconsistent with or conflict with the content of this specification, as well as documents that limit the broadest scope of the claims in this specification (currently or subsequently appended to this specification). It should be noted that in the event of any inconsistency or conflict between the descriptions, definitions, and / or terminology used in the supplementary materials to this specification and the content of this specification, the descriptions, definitions, and / or terminology used in this specification shall prevail.

[0169] Finally, it should be understood that the embodiments described in this specification are merely illustrative of the principles of the embodiments described herein. Other variations may also fall within the scope of this specification. Therefore, alternative configurations of the embodiments described herein are intended to be illustrative rather than limiting, and should be considered consistent with the teachings of this specification. Accordingly, the embodiments described herein are not limited to those explicitly introduced and described herein.

Claims

1. A method for intelligent deployment of audio equipment, characterized in that, The method includes: Acquire additional features, pedestrian distribution information and audio data of the target area. The additional features include at least one of the scene parameters of the target area and the playback mode set by the user. The scene parameters include at least one of the scene size, scene type, scene shape and track layout. The playback mode set by the user refers to the layout mode of multiple target speakers located in the target area that the user has pre-arranged. Based on the additional features, the pedestrian distribution information, and the audio data, the configuration parameters of at least one target speaker among the plurality of speakers are determined by a configuration model. The audio data includes noise distribution information, which includes at least one of the noise type and volume level at multiple locations in the target area. The configuration parameters include the target position of the at least one target speaker on the track in the target area, the number of target speakers at each target position, the rotation angle, the volume, and the sound effect. The configuration model is a machine learning model.

2. The method according to claim 1, characterized in that, The pedestrian distribution information includes thermal imaging distribution maps acquired by thermal imagers and / or regional images captured by cameras.

3. The method according to claim 1, characterized in that, The noise distribution information is determined in the following way: The sound data collected by the sound pickup devices at the multiple locations is processed by a sound recognition model to determine the type and volume of the noise. The sound recognition model is a machine learning model.

4. An intelligent audio deployment system, characterized in that, The system includes: The acquisition module is used to acquire additional features, pedestrian distribution information and audio data of the target area. The additional features include at least one of the scene parameters of the target area and the playback mode set by the user. The scene parameters include at least one of the scene size, scene type, scene shape and track layout. The playback mode set by the user refers to the layout mode of multiple target speakers located in the target area that the user has pre-arranged. The determination module is used to determine the configuration parameters of at least one target speaker among the plurality of speakers based on the additional features, the pedestrian distribution information, and the audio data, through a configuration model. The configuration parameters include the target position of the at least one target speaker on the track of the target area, the number of target speakers at each target position, the rotation angle, the volume, and the sound effect. The configuration model is a machine learning model. The audio data includes noise distribution information, which includes at least one of the noise type and volume level at multiple locations in the target area.

5. An intelligent audio deployment device, characterized in that, include: A track is set in the target area, multiple speakers are located on the track, an information acquisition device and a controller are provided, wherein the multiple speakers are slidably arranged relative to the track; The controller is used for: Based on the additional features, pedestrian distribution information, and audio data of the target area acquired by the information acquisition device, a configuration model is used to determine the configuration parameters of at least one target speaker among the multiple speakers located within the target area. The additional features include at least one of the scene parameters of the target area and the playback mode set by the user. The scene parameters include at least one of the scene size, scene type, scene shape, and track layout. The playback mode set by the user refers to the layout mode of the multiple target speakers pre-arranged by the user. The audio data includes noise distribution information, which includes at least one of the noise type and volume level at multiple locations in the target area. The configuration parameters include the target position of the at least one target speaker on the track in the target area, the number of target speakers at each target position, the rotation angle, the volume, and the sound effect. The configuration model is a machine learning model. The system also controls the sliding of the at least one target speaker on the track based on the configuration parameters.

6. A computer-readable storage medium, characterized in that, The storage medium stores computer instructions, which, when executed by a processor, implement the intelligent audio deployment method as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Optimal configuration method for sliding sound box

    CN113473354A

  • Broadcast system for automatically controlling volume

    KR1020150034534A