Loudspeaker arrangement identification based on directivity index
By determining speaker placement based on microphone-recorded vocal data and machine learning models, the problem of poor speaker placement in entertainment systems has been solved, improving sound quality and listening experience.
Patent Information
- Application Number
- CN202480050309.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-07-24
- Filing Date
- 2024-07-31
- Publication Date
- 2026-03-03
AI Technical Summary
The speaker placement in existing entertainment systems often deviates from recommended values, resulting in a poor listener experience and making it difficult to achieve optimized sound quality.
By determining the Directional Index (DI) pattern data of the person based on the vocal data recorded by the microphone, the speaker placement relative to the listener is determined using a trained machine learning model (such as a neural network), and audio playback settings are adjusted to optimize the sound field.
It enables dynamic adjustment of speaker placement based on the listener's location, improving sound quality and listening experience, and bringing it closer to recommended speaker placement guidelines or the listening experience desired by creators.
Smart Images

Figure CN121605656A_ABST
Abstract
Description
Technical Field
[0001] This application generally relates to speaker placement recognition based on human directional index. Background Technology
[0002] A loudspeaker converts electrical audio signals into corresponding sound. Loudspeakers are used to play music, listen to audio content corresponding to video content (e.g., audio from television programs or movies), and so on. Entertainment systems typically involve multiple loudspeakers that play audio. For example, an entertainment system may include a pair of left and right stereo speakers, a subwoofer, a center speaker, a pair of left and right surround speakers, and / or a pair of left and right rear surround speakers. The number of loudspeakers in a system is usually represented by the xy notation, where x is the number of loudspeakers used in the system and y refers to the number of subwoofers used in the system.
[0003] To optimize sound quality, speakers in entertainment systems are designed with a specific arrangement relative to the listener. For example, the ideal angle and distance from each speaker to the listener can be specified, for instance, by the recommendations described in the ITU-R BS.2159-4 standard. Summary of the Invention
[0004] Technical solution According to one aspect of this disclosure, a method includes: determining person directionality index (DI) pattern data corresponding to the position of a listener based on sound recordings from microphones at each of a plurality of loudspeakers; extracting a set of DI features from the DI pattern data; providing the set of DI features to a trained machine learning model; and determining the arrangement of each of the plurality of loudspeakers relative to the listener using the trained machine learning model and based on the set of DI features.
[0005] The method may further include: adjusting the audio playback settings of one or more of the plurality of speakers based on the determined arrangement of each of the plurality of speakers relative to the listener.
[0006] The method may further include: wherein the audio playback settings include spatial awareness correction.
[0007] The method may further include: wherein the arrangement includes the distance of each of the plurality of speakers relative to the listener.
[0008] The method may further include: wherein the arrangement includes the angle of each of the plurality of speakers relative to the listener.
[0009] The method may further include: wherein the trained machine learning model includes a trained neural network.
[0010] The method may further include: wherein the trained neural network is specific to a number of the plurality of loudspeakers.
[0011] The method may further include: wherein the DI pattern data includes DI values in each of a plurality of frequency bands for each of the plurality of loudspeakers.
[0012] The method may further include: detecting a predetermined command uttered by the listener; and in response to detecting the predetermined command utterance, determining the character DI pattern data.
[0013] The method may further include: wherein the trained machine learning model includes a first neural network, wherein the first neural network is trained to determine the relative distance between each of the plurality of speakers and the listener; and providing the set of DI features to a trained second neural network, wherein the trained second neural network is trained to determine the relative angle between each of the plurality of speakers and the listener; and determining the angle of each of the plurality of speakers relative to the listener using the trained second neural network and based on the set of DI features.
[0014] The method may further include: wherein the trained machine learning model is trained on training data, wherein the training data includes a set of training DI pattern data and a corresponding set of speaker distances; and the method further includes: determining a set of speaker distances between the plurality of speakers; extracting a set of speaker distance features from the set of speaker distances between the plurality of speakers; and determining the arrangement of each of the plurality of speakers relative to the listener based on the set of DI features and the set of speaker distance features using the trained machine learning model. Attached Figure Description
[0015] Figure 1 An example method for determining the arrangement of speakers relative to the listener's position in an entertainment system is shown.
[0016] Figure 2 Show Figure 1 Examples of certain steps in the method.
[0017] Figure 3 An example training method involving simulated data is shown.
[0018] Figure 4 An example randomized arrangement of a 4-speaker system relative to the listener is shown.
[0019] Figure 5 An example of analog DI data from a simulated 4-speaker system is shown.
[0020] Figure 6 An example computing system is shown. Detailed Implementation
[0021] To optimize sound quality, speakers in entertainment systems are designed with a specific arrangement relative to the listener. For example, ideal angles and distances from each speaker to the listener can be specified using recommendations outlined in ITU-R BS.2159-4. However, in practice, speaker placement and orientation often differ from these recommendations. For instance, room design and dimensions may limit speaker placement to specific locations different from those suggested by the recommendations. Furthermore, speaker setups are often imperfect, for example, because users may not precisely orient the speakers as specified in the recommendations. Moreover, recommended speaker placement and orientation are relative to a specific listener's location; therefore, listeners positioned relative to the entertainment system will experience a different relative speaker placement and orientation than the recommendations suggest.
[0022] Figure 1 An example method for determining the arrangement of speakers in an entertainment system relative to a listener's position is illustrated. As explained below, the speaker arrangement may include the speaker's position relative to the listener, the speaker's orientation relative to the listener, or both the speaker's position and orientation relative to the listener. As described below, once the arrangement of each speaker is determined, specific embodiments may use the aggregate arrangement information to adjust the system's sound field (e.g., the volume levels of one or more speakers) and / or apply spatial corrections, such as making the listening experience closer to the listening experience provided by recommended speaker arrangement guidelines or closer to the listening experience expected by the audio creator. For example, if the speaker system is not properly arranged relative to the listener, a delay corresponding to the propagation distance from the nearest speaker to the most distant speaker may be introduced, such that the sound from each speaker reaches the listener as if the listener were in a recommended position relative to the aggregate speaker system. In specific embodiments, based on the determined speaker arrangement relative to the listener, the listener may be notified to adjust their listening position or to adjust the arrangement of one or more speakers to improve their listening experience.
[0023] Figure 1Step 110 of the example method includes determining person directional index pattern data corresponding to the listener's location based on sound recordings made by microphones at each of the multiple loudspeakers (each microphone may be part of, mounted on, or otherwise located in the same position as one of the multiple loudspeakers). Human voices do not radiate sound uniformly around the head; instead, due to the position of the mouth within the head, more energy is radiated forward than backward, thus giving human voices a specific directional pattern that depends on the frequency of the sound, the speaker's angle / direction, and the distance from the speaker. The Directional Index (DI) is a measure of this directional pattern.
[0024] The vocalization can be a specific, predetermined word or phrase. In a particular embodiment, Figure 1 The method can begin with a vocal command that activates the microphone at each speaker when detected by the microphone, or otherwise initiate the process. Figure 1 The command can be the sound on which the DI measurement referenced in step 110 is based, or the command can be given before the sound. The user can initiate the command based on a command (whether verbal or non-verbal) (e.g., via a wirelessly connected device or a UI on a wired device such as a TV, smartphone, etc.). Figure 1 The speaker arrangement method. In a particular embodiment, when setting up an entertainment system, it can be used during the initialization phase. Figure 1 The method. In a particular embodiment, it can be repeated periodically. Figure 1 The method determines whether a listener frequently stays outside the recommended listening position relative to the speaker. In a particular embodiment, it is possible to store the information implemented... Figure 1 The method of determining previously used or frequently used listener locations allows users or the system to subsequently select specific previously determined listener locations without repeating the process. Figure 1 The method.
[0025] The sound can be recorded by a microphone located at each speaker. Each speaker is located in the same position as at least one microphone. The microphone can be a near-field microphone, and the microphone can be located near the main speaker driver.
[0026] The DI value represents the ratio of the acoustic energy measured in a specific direction to the acoustic energy output by the source in all directions. For example, DI can be calculated as: [Mathematical Expression 1]
[0027] Where H is the sound pressure (in this example, H_0 refers to the sound pressure at 0 degrees); w is the angular frequency w=2πf, where f refers to the discrete frequency band; and N is the total number of directions being measured.
[0028] In certain embodiments, signal processing may be performed on the recorded vocal data before determining the DI data from the vocal output. For example, the recorded audio may be passed through a high-pass filter (e.g., a 100 Hz high-pass filter) to remove unwanted low-frequency noise. As another example, the vocal data captured by each microphone may be segmented into multiple frequency range bands spanning typical frequencies used by human speech. For example, the vocal data may be segmented into bands centered at 250 Hz, 500 Hz, 1 kHz, 2 kHz, 4 kHz, and 8 kHz, but this disclosure contemplates the use of more or fewer bands, and the bands may be centered at frequencies different from those described in the foregoing examples.
[0029] In a particular embodiment, a filter may be applied to the recorded data on each microphone / speaker. For example, a 1 / 3 octave-band filter with a center frequency may be applied to the recorded audio, and this disclosure contemplates the use of other filters. In a particular embodiment, the average energy in each filtered band may then be determined for each recording. In a particular embodiment, if a speaker has more than one microphone, recordings occurring simultaneously at the microphones of that speaker may be combined, for example, the average energy in a particular filtered band may be averaged among the microphones in the speaker.
[0030] In a particular embodiment, DI data can be obtained for each frequency band and for each microphone at each speaker. For example, if the system includes 4 speakers and 6 frequency bands, 24 DI values are obtained, where one DI value is obtained for each channel for each speaker. In a particular embodiment, DI data can be normalized by frequency band, for example, by normalizing the highest DI value for a particular frequency band to 0 dB and then adjusting each other DI value in that particular frequency band accordingly relative to the 0 dB frequency band. An example table of DI values (in dB) for a system including 4 speakers and 6 example frequency bands is shown below: [Table 1]
[0031] The top row identifies each frequency band in Hz, and the leftmost column identifies each speaker. In this example, L is the left speaker, R is the right speaker, RB is the right rear speaker, and LB is the left rear speaker. As mentioned above, each column (each frequency band) contains a normalized 0 dB DI value, and the remaining values for that frequency band are normalized accordingly. DI data can be represented graphically, for example, as shown below. Figure 5 As shown, it illustrates an example of analog DI data collected from a simulated 4-speaker system. Figure 5 The example DI data shown corresponds to the DI data presented in the table above.
[0032] Figure 1 Step 120 of the example method includes extracting a set of DI features from the DI pattern data, and Figure 1 Step 130 of the example method includes providing the set of DI features to the trained machine learning model. Figure 1 Step 140 of the example method includes determining the arrangement of each of the plurality of speakers relative to the listener using a trained machine learning model and based on the set of DI features. Steps 120 to 140 and step 110, as contemplated in this disclosure, can be performed by any suitable computing device or combination of computing devices. For example, steps 110 to 140 can be performed by one of the speakers in an entertainment system, by a connected local device or server device, or by a combination of devices (e.g., one device may perform one or more steps while another device performs other steps, etc.).
[0033] Figure 2 Show Figure 1 Examples of steps 120 to 140 of the method. Figure 2 In this process, the DI value 202 undergoes principal component analysis 204, which essentially removes redundant information to identify a subset of the most meaningful speaker arrangement features. For example, as explained more fully below, a 40-dimensional input vector can be reduced to 36 features. In steps 120 through 140, at least DI features are used; in certain embodiments, additional data and corresponding features may also be used as input to a machine learning model. For example, and as described below, in certain embodiments, a machine learning model can be trained based on both DI values and speaker distances, and for such a model, step 130 includes extracting features from the DI pattern data and speaker distances, and step 140 includes determining the arrangement of each of the plurality of speakers based on the extracted DI features / speaker distance features. Although Figure 2 The example uses principal component analysis to extract features from DI values, but other feature extraction techniques can be used. In a particular embodiment, the DI values themselves (e.g., channelized DI values) can be features extracted in step 120; for example, the DI values can be directly placed into an M-dimensional vector, where M is the number of features subsequently input into the machine learning model. Figure 2 In this process, principal component analysis 204 is applied to the DI value 202 to generate a set of extracted input features 206.
[0034] Feature 206 is input into machine learning model 210, and machine learning model 210 in Figure 2The example is a neural network. The neural network includes an M-dimensional input layer (where M is typically the number of features extracted from the DI pattern data or from the DI pattern data and additional data such as speaker distances), two hidden layers 214 and 216, and an output layer 218. The output layer 218 of the machine learning model 210 results in a set of distance values 220, where each distance value specifies an estimated distance between the listener and the corresponding speaker.
[0035] Figure 2 The examples include two different machine learning models: Model 210, which outputs the estimated distance between the listener and each speaker; and Model 230, which outputs the estimated angle of each speaker relative to the listener. Therefore, Figure 2 An embodiment is shown in which two separate machine learning models are trained and subsequently used to identify specific arrangement attributes of each speaker. However, a particular embodiment may train and use a single machine learning model to determine both speaker distance and speaker angle, although this approach may have reduced accuracy compared to using a separate model.
[0036] In certain embodiments, individual models can be trained and subsequently used in entertainment systems with different numbers of speakers. For example, a 4-speaker system may have a pair of dedicated models to predict the arrangement of the four speakers, while a 5-speaker system may have a separate pair of dedicated models, and so on. Figure 2 An example of this embodiment is shown, wherein models 210 and 230 have four neurons in the output layer, thus specifically corresponding to a 4-speaker system. In a particular embodiment, more generally, one or more models can be trained to predict the arrangement of a series of speakers (e.g., 4 to 7 speakers) or any number of speakers, but this approach may have reduced performance compared to embodiments that use dedicated models for each specific number of speakers.
[0037] As mentioned above, Figure 2 An example of step 140 is shown, wherein the arrangement data includes a distance estimate 220 from model 210 and an orientation (angle) estimate from model 230. To obtain the angle estimate, M features 206 are input to an input layer 232 (typically M-dimensional). Model 230 includes three hidden layers 234, 236, and 238 and a four-dimensional output layer 240, wherein the four-dimensional output layer 240 outputs the angle 242: for the angle as... Figure 2 The example focuses on each speaker in a 4-speaker system, outputting an angle relative to the listener.
[0038] Figure 2The example shows a machine learning model 210 for estimating distance, comprising two hidden layers: a first hidden layer 214 containing 13 neurons and a second hidden layer 216 containing 95 neurons. This particular architecture of two hidden layers (each with a specified number of neurons) is particularly well-suited for estimating distance values for a 4-speaker system when using a neural network as a machine learning model. Similarly, Figure 2 An example of a model 230 for estimating angles is shown, comprising three hidden layers: a first hidden layer 234 comprising 5 neurons, a second hidden layer 236 comprising 42 neurons, and a third hidden layer 238 comprising 88 neurons. This particular architecture of three hidden layers (each with a specified number of neurons) is particularly well-suited for estimating the angles of a 4-speaker system when using a neural network as a machine learning model. However, this disclosure contemplates the use of other numbers of hidden layers and neurons, both through a 4-speaker neural network model and neural network models for other numbers of speakers. For example, an n-speaker system may have dedicated distance neural network models and angle neural network models, and these models may have multiple hidden layers and / or multiple constituent neurons specific to the n-speaker system. For example, the number of layers and neurons for a neural network model for a specific n-speaker arrangement can be determined using Bayesian optimization on a model-by-model basis. Finally, although in Figure 2 The example uses a neural network, but this disclosure is intended to train other machine learning models and use them to determine placement values based on DI data determined by the listener's vocalizations.
[0039] As mentioned above Figure 1 Example methods and Figure 2 The example architectures discussed herein involve one or more machine learning models used to determine placement values being trained before being deployed for runtime inference. Training can be based on simulated training data, real training data, or a combination thereof. For example, the following discussion provides a detailed example of training using simulated data.
[0040] Figure 3An example training method involving simulated data is illustrated. Step 302 includes defining room characteristics and sound characteristics for sound emanating from a simulated room. Step 302 includes simulating a receiver (microphone) 304, wherein simulating the receiver (microphone) 304 includes specifying a directional room impulse response (RIR) and location for each microphone. The simulated one or more microphones may match the characteristics of real microphones used in a particular speaker or entertainment system, making the simulated data specific to that system or speaker. Thus, in a particular embodiment, a model can be trained on data specific to one or more microphones in a particular real system, and the model can be deployed for that particular system or for similar systems containing similar microphones. In other embodiments, the model can be trained using data from a range of microphone types, and more generally, the model can be deployed for a system; in a particular embodiment, the speaker or microphone brand / model can then be specified by the user before inference.
[0041] exist Figure 3 In the example of acoustic room simulation, step 302 includes specifying a source 306 for simulation, wherein specifying a source 306 for simulation includes specifying a female or male directional impulse response. Step 302 also includes a room generator 308 for specifying relevant room parameters for a set of simulated rooms. Rooms can be of various sizes, dimensions, and shapes to simulate a wide range of room dynamics. For example, room sizes can be randomly selected from the following parameters: (1) height = 2.7 to 3.0 meters; (2) width = 4.0 to 8.0 meters; and (3) length = 6.0 to 12.0 meters. In a particular embodiment, room dimensions can be simulated such that there are approximately equal numbers of simulated rooms in a particular room size category (e.g., small (80 to 150 cubic meters), medium (150 to 220 cubic meters), and large (220 to 300 cubic meters)). Room parameters may include specifying the building materials used for the room, such as: drywall (with a 100mm cavity) mounted on a frame; mineral wool filled in the cavity and painted on the surface; double-glazed windows with 2 to 3mm glass and a 10mm air gap; plywood or hardwood boards with a 25mm air gap underneath; joist-style wood flooring; rubber floor tiles; thin carpet laid on top of a thin felt layer on wood flooring; drywall (with a 100mm cavity) mounted on a frame. The foregoing details regarding specific room dimensions, size categories, and characteristics provide examples of parameters that can be simulated, and other parameters or different values of these parameters can be used for room simulation. Additionally, simulated (or real) data may be focused on specific sizes or parameters to train speaker arrangement models specific to those sizes and / or parameters. For example, real or simulated data may be focused on a small room with carpet to train and deploy models specific to those characteristics.
[0042] Step 302 includes specifying the arrangement of each speaker and source listener in each simulated room configuration. Figure 4 An example randomized arrangement of four speakers (left speaker L, right speaker R, left rear speaker LB, and right rear speaker RB) relative to the listener / source 410 is shown. In a particular embodiment, the arrangement of each speaker may be restricted to a specific range; for example, each speaker may be positioned at its recommended location relative to the user and then adjusted by a specific random offset between a maximum and minimum offset. As another example, each speaker may be rotated from its recommended location by a randomized offset within a predetermined clockwise / counterclockwise angular range.
[0043] Figure 3 Step 310 of the example includes convolving the RIR of each microphone with the audio of female and male audio recordings 312 of previously recorded speech commands (words, phrases, etc.) under echo-free conditions. In the example of simulated training data, the result of convolution step 310 is simulated recorded audio at each microphone. Step 314 then includes extracting DI values from this simulated recorded audio, as described above. Although Figure 3 Examples involve simulating room conditions, speaker arrangements, and microphone properties to simulate audio received at each simulated microphone. However, training data can also include real-world training data obtained by setting up real rooms with real speakers and real microphones and recording the vocalization of real male and female voices in those rooms. However, generating real-world data can be more labor-intensive, especially when generating a large number of training samples under a wide range of room conditions. For example, training data could include 1000 or more training samples (e.g., 500 simulated rooms, each with male and female vocalizations), and generating such a large number of real-world training samples can be more labor-intensive than generating simulated training samples.
[0044] Whether using real or simulated training data, DI values are extracted from recorded audio (real or simulated) for a specific room / speaker setup. DI values can be processed, for example, as described above with respect to Table 1. The resulting DI values can be represented as an S×C matrix, where S is the number of speakers and C is the number of frequency bands, for example as... Figure 3 Step 316 is shown. Additionally, certain embodiments may use the inter-speaker distance during both the training and subsequent inference phases. For example, the inter-speaker distance may be represented as an S×S matrix, where each element represents the distance between the i-th speaker and the j-th speaker, where i and j both traverse all speaker indices. The distance value may refer to the distance between the driver coordinates of the i-th speaker and the microphone coordinates of the j-th speaker, for example, to account for the offset between the microphone position (which records training audio) and the speaker driver position.
[0045] Training data can be combined into a single N-dimensional vector, where, for example, in an embodiment training the model on both DI data and speaker distances, N equals S×C plus S×S. For example, in a 4-speaker setup using 6 frequency bands, N equals 40. As described above, feature selection can be performed on this vector to extract features, and the resulting features are fed into a machine learning model, such as... Figure 3 The neural network model 318. Real or synthetic training data is divided into training, testing, and validation parts, and the machine learning model is trained accordingly.
[0046] exist Figure 2 In the example, neural network model 210 can be trained using the following parameters: maximum epochs: 20,000; maximum failures: 5,000; optimizer: scaled conjugate gradient (SCG); objective: 0; minimum gradient: 10⁻⁶; μ: 0.005; σ: 5 × 10⁻⁵; λ: 5 × 10⁻⁷; and performance function: sum of squared errors. Neural network model 230 can be trained using the following parameters: maximum epochs: 20,000; maximum failures: 5,000; optimizer: scaled conjugate gradient (SCG); objective: 0; minimum gradient: 10⁻⁶; μ: 0.005; σ: 5 × 10⁻⁵; λ: 5 × 10⁻⁷; and performance function: sum of squared errors. However, these specific training parameters are merely one example of how to train these machine learning models to infer speaker placement, and any suitable training parameters can be used depending on the desired resolution, computational resources, and the properties of one or more models being trained.
[0047] In a particular embodiment, the relative propagation delay between speakers can be estimated by recording the sound emitted from one or more microphones to each speaker from the user, and then a cross-correlation algorithm is used to obtain the delay difference between the speakers. A geometric model and least squares method can be used to obtain the absolute distance between the user and the speakers, and then the angle of incidence from the speakers to the user can be analytically obtained. These distance / angle determinations can be performed in addition to or as an alternative to the ML-based methods described above (e.g., as a check on the ML-based methods described above). Furthermore, in a particular embodiment, the distance from each speaker to the listener can be obtained by other methods, such as direct measurement by the user, or by using a mobile phone containing a gyroscope, or by measuring the impulse response using an external microphone (e.g., a microphone included in a mobile phone). In this case, these measurements can be used in place of, or in addition to, the recordings made by the microphones within the speakers.
[0048] Figure 6 An example computer system 600 is illustrated. In a particular embodiment, one or more computer systems 600 perform one or more steps of one or more methods described or illustrated herein. In a particular embodiment, one or more computer systems 600 provide the functionality described or illustrated herein. In a particular embodiment, software running on one or more computer systems 600 performs one or more steps of one or more methods described or illustrated herein, or provides the functionality described or illustrated herein. Specific embodiments include one or more portions of one or more computer systems 600. Throughout this document, references to computer systems may include computing devices and vice versa, where appropriate. Furthermore, references to computer systems may include one or more computer systems, where appropriate.
[0049] This disclosure contemplates any suitable number of computer systems 600. This disclosure contemplates computer systems 600 in any suitable physical form. By way of example and not by limitation, computer system 600 may be an embedded computer system, a system-on-a-chip (SOC), a single-board computer system (SBC) (such as, for example, a computer-on-module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive self-service machine, a mainframe, a grid of computer systems, a mobile phone, a personal digital assistant (PDA), a server, a tablet computer system, or a combination of two or more of these. Where appropriate, computer system 600 may include one or more computer systems 600; may be single or distributed; span multiple locations; span multiple machines; span multiple data centers; or reside in a cloud, wherein the cloud may include one or more cloud components of one or more networks. Where appropriate, one or more computer systems 600 may perform one or more steps of the methods described or illustrated herein without substantial spatial or temporal limitations. By way of example and not limitation, one or more computer systems 600 may execute one or more steps of the methods described or illustrated herein in real time or in batch mode. Where appropriate, one or more computer systems 600 may execute one or more steps of the methods described or illustrated herein at different times or in different locations.
[0050] In a particular embodiment, computer system 600 includes a processor 602, memory 604, storage device 606, input / output (I / O) interface 608, communication interface 610, and bus 612. Although this disclosure describes and illustrates a particular computer system having a particular number of particular components in a particular arrangement, this disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable arrangement.
[0051] In a particular embodiment, processor 602 includes hardware for executing instructions, such as those constituting a computer program. By way of example, and not limitation, to execute instructions, processor 602 may retrieve (or fetch) instructions from internal registers, internal caches, memory 604, or storage device 606; decode and execute them; and then write one or more results to internal registers, internal caches, memory 604, or storage device 606. In a particular embodiment, processor 602 may include one or more internal caches for data, instructions, or addresses. Where appropriate, this disclosure contemplates that processor 602 may include any suitable number of suitable internal caches. By way of example, and not limitation, processor 602 may include one or more instruction caches, one or more data caches, and one or more translation back lookup buffers (TLBs). Instructions in the instruction cache may be copies of instructions in memory 604 or storage device 606, and the instruction cache may accelerate the retrieval of those instructions by processor 602. The data in the data cache may be a copy of the data in memory 604 or storage device 606 for operation of instructions executed at processor 602; the result of a previous instruction executed at processor 602 for access by a subsequent instruction executed at processor 602 or for writing to memory 604 or storage device 606; or other suitable data. The data cache can accelerate read or write operations of processor 602. The TLB can accelerate virtual address translation for processor 602. In a particular embodiment, processor 602 may include one or more internal registers for data, instructions, or addresses. Where appropriate, this disclosure contemplates processor 602 including any suitable number of suitable internal registers. Where appropriate, processor 602 may include one or more arithmetic logic units (ALUs); is a multi-core processor; or includes one or more processors 602. Although this disclosure describes and illustrates specific processors, this disclosure contemplates any suitable processor.
[0052] In a particular embodiment, memory 604 includes main memory for storing instructions for execution by processor 602 or data for operation of processor 602. By way of example and not limitation, computer system 600 may load instructions from storage device 606 or another source (such as, for example, another computer system 600) into memory 604. Processor 602 may then load the instructions from memory 604 into internal registers or internal caches. To execute the instructions, processor 602 may retrieve and decode the instructions from internal registers or internal caches. During or after instruction execution, processor 602 may write one or more results (which may be intermediate or final results) into internal registers or internal caches. Processor 602 may then write one or more of these results into memory 604. In a particular embodiment, processor 602 executes only instructions in one or more internal registers or internal caches or memory 604 (memory 604 or elsewhere, as opposed to storage device 606), and operates only on data in one or more internal registers or internal caches or memory 604 (memory 604 or elsewhere, as opposed to storage device 606). One or more memory buses (each of which may include an address bus and a data bus) couple processor 602 to memory 604. Bus 612 may include one or more memory buses, as described below. In a particular embodiment, one or more memory management units (MMUs) reside between processor 602 and memory 604 and facilitate access to memory 604 requested by processor 602. In a particular embodiment, memory 604 includes random access memory (RAM). Where appropriate, the RAM may be volatile memory. Where appropriate, the RAM may be dynamic RAM (DRAM) or static RAM (SRAM). Furthermore, where appropriate, the RAM may be single-port or multi-port RAM. This disclosure contemplates any suitable RAM. Where appropriate, memory 604 may include one or more memories 604. Although this disclosure describes and illustrates specific memories, this disclosure contemplates any suitable memory.
[0053] In a particular embodiment, storage device 606 includes mass storage for data or instructions. By way of example and not limitation, storage device 606 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk drive, magneto-optical disk drive, magnetic tape drive, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, storage device 606 may include removable or non-removable (or fixed) media. Where appropriate, storage device 606 may be internal or external to computer system 600. In a particular embodiment, storage device 606 is a non-volatile solid-state memory. In a particular embodiment, storage device 606 includes read-only memory (ROM). Where appropriate, the ROM may be a mask-programmable ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), electrically rewritable ROM (EAROM), or flash memory, or a combination of two or more of these. This disclosure contemplates mass storage device 606 in any suitable physical form. Where appropriate, storage device 606 may include one or more storage control units facilitating communication between processor 602 and storage device 606. Where appropriate, storage device 606 may include one or more storage devices 606. Although this disclosure describes and illustrates specific storage devices, this disclosure contemplates any suitable storage device.
[0054] In a particular embodiment, I / O interface 608 includes hardware, software, or both, providing one or more interfaces for communication between computer system 600 and one or more I / O devices. Where appropriate, computer system 600 may include one or more of these I / O devices. One or more of these I / O devices enable communication between a person and computer system 600. By way of example and not limitation, I / O devices may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, tablet computer, touchscreen, trackball, camera, another suitable I / O device, or a combination of two or more of these. I / O devices may include one or more sensors. This disclosure contemplates any suitable I / O device and any suitable I / O interface 608 for them. Where appropriate, I / O interface 608 may include one or more device or software drivers that enable processor 602 to drive one or more of these I / O devices. Where appropriate, I / O interface 608 may include one or more I / O interfaces 608. Although this disclosure describes and illustrates specific I / O interfaces, this disclosure considers any suitable I / O interface.
[0055] In a particular embodiment, communication interface 610 includes hardware, software, or both, to provide one or more interfaces for communication (such as packet-based communication) between computer system 600 and one or more other computer systems 600 or one or more networks. By way of example, and not limitation, communication interface 610 may include a network interface controller (NIC) or network adapter for communicating with Ethernet or other wired networks, or a wireless NIC (WNIC) or wireless adapter for communicating with wireless networks (such as Wi-Fi networks). This disclosure contemplates any suitable network and any suitable communication interface 610 used therefor. By way of example, and not limitation, computer system 600 may communicate with one or more portions of an ad hoc network, a personal area network (PAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), or the Internet, or a combination of two or more of these. One or more of these networks may be wired or wireless. As an example, computer system 600 may communicate with wireless PAN (WPAN) (such as, for example, Bluetooth WPAN), Wi-Fi network, Wi-MAX network, cellular telephone network (such as, for example, GSM network), or other suitable wireless network, or a combination of two or more of these. Where appropriate, computer system 600 may include any suitable communication interface 610 for any of these networks. Where appropriate, communication interface 610 may include one or more communication interfaces 610. Although specific communication interfaces are described and shown in this disclosure, any suitable communication interface is contemplated in this disclosure.
[0056] In a particular embodiment, bus 612 includes hardware, software, or both, that couple components of computer system 600 to each other. By way of example and not limitation, bus 612 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infiniband interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 612 may include one or more buses 612. Although this disclosure describes and illustrates specific buses, this disclosure contemplates any suitable bus or interconnect.
[0057] In this document, where appropriate, one or more computer-readable non-transitory storage media may include one or more semiconductor-based or other integrated circuits (ICs) (such as, for example, field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs)), hard disk drives (HDDs), hybrid hard disk drives (HHDs), optical disks, optical disk drives (ODDs), magneto-optical disks, magneto-optical drives, floppy disks, floppy disk drives (FDDs), magnetic tape, solid-state drives (SSDs), RAM drives, secure digital cards or drives, any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these. Where appropriate, computer-readable non-transitory storage media may be volatile, non-volatile, or a combination of volatile and non-volatile.
[0058] In this document, unless otherwise expressly stated or the context otherwise requires, “or” is inclusive rather than exclusive. Therefore, in this document, unless otherwise expressly stated or the context otherwise requires, “A or B” means “A, B, or both”. Furthermore, unless otherwise expressly stated or the context otherwise requires, “and” is joint and separate. Therefore, in this document, unless otherwise expressly stated or the context otherwise requires, “A and B” means “A and B together or separately”.
[0059] The scope of this disclosure covers all changes, substitutions, variations, alterations, and modifications to the exemplary embodiments described or illustrated herein that will be understood by those skilled in the art. The scope of this disclosure is not limited to the exemplary embodiments described or illustrated herein. Furthermore, although this disclosure describes and illustrates corresponding embodiments herein as including specific components, elements, features, functions, operations, or steps, any of these embodiments may include any combination or arrangement of any component, element, feature, function, operation, or step described or illustrated anywhere herein that will be understood by those skilled in the art.
Claims
1. A method comprising: Based on the sound recorded by the microphone at each of the multiple loudspeakers, determine the person directionality index (DI) pattern data corresponding to the listener's position; Extract a set of DI features from the DI pattern data; The set of DI features is provided to the trained machine learning model; and The arrangement of each of the plurality of loudspeakers relative to the listener is determined using the trained machine learning model and based on the set of DI features.
2. The method according to claim 1, further comprising: Based on the determined arrangement of each of the plurality of speakers relative to the listener, the audio playback settings of one or more of the plurality of speakers are adjusted.
3. The method according to claim 2, wherein, The audio playback settings include spatial awareness correction.
4. The method according to claim 1, wherein, The arrangement includes the distance of each of the plurality of speakers relative to the listener.
5. The method according to claim 1, wherein, The arrangement includes the angle of each of the plurality of speakers relative to the listener.
6. The method according to claim 1, wherein, The trained machine learning model includes a trained neural network.
7. The method according to claim 6, wherein, The trained neural network is specific to a number of the plurality of loudspeakers.
8. The method according to claim 1, wherein, The DI pattern data includes DI values for each of the multiple frequency bands for each of the plurality of loudspeakers.
9. The method according to claim 1, further comprising: Detect the listener's predetermined command to be spoken; and In response to the detection of the predetermined command sound, the character DI pattern data is determined.
10. The method according to claim 1, wherein, The trained machine learning model includes a first neural network, wherein the first neural network is trained to determine the relative distance between each of the plurality of speakers and the listener; and The method further includes: The set of DI features is provided to a trained second neural network, wherein the second neural network is trained to determine the relative angle between each of the plurality of speakers and the listener; The angle of each of the plurality of loudspeakers relative to the listener is determined using the trained second neural network and based on the set of DI features.
11. The method according to claim 1, wherein, The trained machine learning model is trained on training data, wherein the training data includes a set of training DI pattern data and a corresponding set of speaker distances; and The method further includes: Determine a set of speaker distances among the plurality of speakers; Extract a set of speaker distance features from the set of speaker distances among the plurality of speakers; and The trained machine learning model determines the arrangement of each of the plurality of speakers relative to the listener based on the set of DI features and the set of speaker distance features.
12. A non-transitory computer-readable storage medium containing one or more storage instructions, wherein, When executed, the instruction is capable of performing the following steps: Based on the sound recorded by the microphone at each of the multiple loudspeakers, determine the person directionality index (DI) pattern data corresponding to the listener's position; Extract a set of DI features from the DI pattern data; The set of DI features is provided to the trained machine learning model; as well as The arrangement of each of the plurality of loudspeakers relative to the listener is determined using the trained machine learning model and based on the set of DI features.
13. A system comprising: Multiple speakers; Multiple microphones, each microphone being located in the same position as at least one of the multiple speakers such that each of the multiple speakers is located in the same position as at least one of the multiple microphones; as well as One or more non-transitory computer-readable storage media storing instructions; as well as One or more processors, coupled to the one or more non-transitory computer-readable storage media and operable to execute the instructions to perform the following steps: Based on the sounds recorded by the multiple microphones, determine the person directionality index (DI) pattern data corresponding to the listener's position; Extract a set of DI features from the DI pattern data; The set of DI features is provided to the trained machine learning model; as well as The arrangement of each of the plurality of loudspeakers relative to the listener is determined using the trained machine learning model and based on the set of DI features.
14. The system of claim 13, further comprising one or more processors, wherein, The one or more processors are configured to: The instructions are executed to adjust the audio playback settings of one or more of the plurality of speakers based on the determined arrangement of each speaker relative to the listener.
15. The system according to claim 14, wherein, The audio playback settings include spatial awareness correction.