Acoustic speaker system

The acoustic speaker system uses portable devices with position sensors and a host device to calculate correction gain coefficients, ensuring balanced sound field generation and accurate sound localization despite changes in speaker count and placement.

JP2025115927APending Publication Date: 2025-08-07CEAR INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024098800
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-26
Filing Date
2024-06-19
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Existing acoustic speaker systems struggle to accurately simulate the original sound field when the number of speakers is increased or decreased and their positions are freely set, leading to imbalances in volume between channels and incorrect sound localization.

Method used

An acoustic speaker system comprising a plurality of portable speaker devices and a host device, where each speaker device includes a position sensor, distance estimation unit, and area estimation unit to determine its position and area, and the host device calculates correction gain coefficients to equalize output levels and adjust delay times, ensuring balanced sound field generation regardless of speaker arrangement.

Benefits of technology

The system generates a balanced sound field without losing volume balance between channels, even with varying numbers of speakers and positions, by equalizing outputs and synchronizing sound wave arrival times, enhancing sound image clarity and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025115927000001_ABST
    Figure 2025115927000001_ABST
Patent Text Reader

Abstract

To provide an acoustic speaker system for generating an excellent sound field with respect to the increase / decrease of the number of speakers and free setting of positions.SOLUTION: An acoustic speaker system includes a plurality of portable speaker devices 2 and a host device 3. The portable speaker device 2 includes: a distance estimation section for obtaining position information and distance of the portable speaker device 2; and an area estimation section for determining a belonging area from pre-held area information and the position information. The host device 3 includes a correction gain calculation section for calculating a gain coefficient which equalizes the total sum of outputs to be issued from the portable speaker device 2 in the respective belonging areas. The portable speaker device 2 includes: a level adjustment section for multiplying the correction gain coefficient by an acoustic signal; and a sound issuing section for converting the acoustic signal into a sonic wave so as to output the signal.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an acoustic speaker system having a portable speaker device that can be easily carried around. [Background technology]

[0002] A listener perceives the direction of a sound image by detecting the time difference, sound pressure difference, and reverberation of sound waves arriving at the left and right ears. If the head-related transfer function (HRTF) from the sound source to both ears matches well with the original sound source in the reproduced sound field, the listener can perceive a sound image in the reproduced sound field that simulates the original sound field.

[0003] Furthermore, sound waves undergo a specific change in sound pressure level for each frequency as they travel through space, the head, and the ears to the eardrum. This specific change in sound pressure level for each frequency is called the transfer characteristic. If the head-related transfer functions of the original sound field and the listening sound field match well, the listener will be able to hear the same timbre as the original sound due to similar transfer characteristics.

[0004] However, in most cases, the head-related transfer functions (HRTFs) of the original sound field and the listening sound field are different. For example, it is difficult to reproduce the sound field space of a real or virtual concert hall in a living room. Therefore, the positional relationship between the speaker and the listening point in the listening space differs in distance and angle from the positional relationship between the sound source and the listening point in the original sound field space. As a result, the HRTFs do not match, and the listener perceives a sound image position and timbre that are different from the sound source position and timbre of the original sound.

[0005] Therefore, in recording studios or mixing studios, acoustic processing is generally applied to recorded or artificially created audio signals to simulate the acoustic effects of the original sound in a predetermined listening environment. For example, in a studio, a mixer assumes a certain speaker arrangement and sound receiving point, and intentionally corrects the time difference and sound pressure difference for the multi-channel audio signals output from each speaker so that a sound image that mimics the sound source position of the original sound is perceived, and also changes the sound pressure level for each frequency to match the timbre of the original sound.

[0006] The International Telecommunication Union - Radio sector (ITU-R) provides specific recommendations for speaker placement such as 5.1ch, and THX, for example, has established standards for speaker placement in movie theaters, sound volume, theater size, etc. If mixers and listeners follow these recommendations and standards, the source position and timbre of the original sound will be accurately simulated when the acoustic signal reaches the listener's eardrum in the listening environment, even if the listening environment differs from the original sound field.

[0007] However, even though there is no longer a need to match the original sound field with the listening environment, it is still difficult to conveniently adapt a listening room to the above recommendations and standards. Also, if you want to make a 2-channel audio speaker system installed in a listening room compatible with a 5.1-channel source, for example, you will have to buy a new audio speaker system. In other words, listeners cannot freely increase or decrease the number of speakers.

[0008] Therefore, an acoustic system has been proposed that automatically detects the relative positions of multiple speaker devices placed at arbitrary locations and generates audio signals to obtain appropriate sound image localization (see, for example, Patent Document 1). This system does not place a burden on the listener and is suitable not only for building a new acoustic system, but also for adding speaker devices and changing their placement. [Prior art documents] [Patent documents]

[0009] [Patent Document 1] Japanese Patent Application Laid-Open No. 2005-198249 Summary of the Invention [Problem to be solved by the invention]

[0010] In Patent Document 1, on the premise that each speaker device is arranged as evenly as possible in each direction around the listener, the channel synthesis ratio of the audio signals is adjusted based on the number of speaker devices and the arrangement relationship after the position change, thereby generating a sound field that provides the desired sound image localization.

[0011] However, the speakers are not necessarily arranged evenly around the listener. Depending on the speaker arrangement, two or more speakers may be used for a single channel signal, which can result in an imbalance in the volume between channels. When the volume between channels is unbalanced, for example, the listener may perceive a particular sound source as being closer than expected, or other sound sources as being farther away or blurred than expected, making it difficult to achieve an acoustic speaker system that accurately simulates the position of the original sound source.

[0012] The present invention has been made to solve the above-mentioned problems, and its purpose is to provide an acoustic speaker system that can generate a good sound field even if the number of speakers is increased or decreased and the positions of the speakers can be freely set. [Means for solving the problem]

[0013] In order to achieve the above object, the acoustic speaker system of this embodiment is an acoustic speaker system comprising a plurality of portable speaker devices and a host device, wherein the portable speaker devices each comprise a position sensor that detects a position change of a housing, a distance estimation unit that determines position information and a distance in a three-dimensional space in accordance with the position change detected by the position sensor, a region estimation unit that determines a home region by comparing pre-stored region information obtained by dividing the three-dimensional space with the position information estimated by the distance estimation unit, and a speaker-side transmission unit that transmits the distance determined by the distance estimation unit and the home region determined by the region estimation unit to the host device, and the host device comprises a receiving unit that receives the distances and the home regions of all the portable speaker devices, and a correction gain coefficient that makes the sum of the outputs of the portable speaker devices emitting sounds in each of the home regions equal. and a transmitting unit that transmits the correction gain coefficient to the portable speaker device, wherein the correction gain calculating unit searches for the farthest distance among all of the portable speaker devices, counts the number M of the portable speaker devices that belong to the same belonging area, calculates a ratio L2 / L1 that is the ratio of distance L2 that is the distance of the portable speaker device to distance L1 that is the farthest distance, calculates 1 / M that is the reciprocal of the number M of the portable speaker devices that belong to the same belonging area, and calculates L2 / L1*1 / M, which is the gain coefficient obtained by multiplying the ratio L2 / L1 by the reciprocal 1 / M, for each of the portable speaker devices, and the portable speaker device further comprises a level adjusting unit that multiplies the correction gain coefficient by an acoustic signal, and a sound emitting unit that converts the acoustic signal that has passed through the level adjusting unit into a sound wave and outputs it.

[0014] The distance estimation unit of the portable speaker device may further determine the separation distance from the origin to the position of the portable speaker device itself based on the position change detected by the position sensor, and the speaker-side transmission unit may further transmit the separation distance to the host device, and the host device may further include a maximum required transmission time calculation unit that calculates the maximum required transmission time for a sound wave to reach the origin from the farthest portable speaker device within the same belonging area based on the separation distance transmitted from each of the portable speaker devices, and the transmission unit of the host device may further transmit the maximum required transmission time to the portable speaker device, and the portable speaker device may further include a delay correction calculation unit that delays sound emission based on the maximum required transmission time and the separation distance so that the sound wave from the farthest portable speaker device arrives at the origin at the same time.

[0015] The portable speaker device may include an acoustic signal receiving unit that receives the acoustic signal, and the level adjusting unit may multiply the acoustic signal received by the acoustic signal receiving unit by the correction gain coefficient.

[0016] The distance estimation unit may store distance upper limit information indicating a predetermined distance in advance, and compare the distance calculated in response to the position change with the predetermined distance indicated by the distance upper limit information, and the speaker-side transmission unit may stop transmitting the distance and the area to which the speaker belongs when the distance exceeds the predetermined distance indicated by the distance upper limit information. [Effects of the Invention]

[0017] According to the present invention, a sound field can be generated without losing the balance of volume between channels regardless of the number of portable speaker devices. [Brief explanation of the drawings]

[0018] [Figure 1] FIG. 1 is a block diagram showing a configuration of an acoustic speaker system. [Figure 2] FIG. 1 is a schematic diagram showing the appearance of a portable speaker device. [Figure 3]FIG. 2 is a block diagram showing the internal configuration of the portable speaker device. [Figure 4] FIG. 2 is a block diagram showing the internal configuration of a host device. [Figure 5] FIG. 2 is a block diagram showing the functions of the portable speaker device. [Figure 6] FIG. 2 is a schematic diagram showing the relationship between the portable speaker device and the origin. [Figure 7] 10A and 10B are schematic diagrams showing an example of a method for turning on the power of a portable speaker device. [Figure 8] FIG. 10 is a schematic diagram showing an example of dividing an area. [Figure 9] FIG. 10 is a schematic diagram showing an area map superimposed on a real space in which a portable speaker device is placed. [Figure 10] FIG. 10 is a schematic diagram showing an example of region setting information when a monaural sound field is constructed by remixing. [Figure 11] FIG. 10 is a schematic diagram showing an example of area setting information when a stereo sound field is constructed by remixing. [Figure 12] FIG. 10 is a schematic diagram showing an example of area setting information when a surround sound field is constructed by remixing. [Figure 13] 10 is a schematic diagram showing an example of area setting information when a sound field is constructed by extracting and arranging specific acoustic signals by remixing. FIG. [Figure 14] FIG. 10 is a schematic diagram showing an example of a method for obtaining a correction filter. [Figure 15] FIG. 2 is a block diagram showing the functions of the host device. [Figure 16] FIG. 10 is a schematic diagram showing an example of a method for calculating correction gain coefficients for portable speaker devices arranged in two areas. [Figure 17] FIG. 10 is a schematic diagram showing an example of setting parameters of an area map through GUI operations. [Figure 18] 10 is a flowchart showing an example of operations from when a program is started to when a portable speaker device is installed. [Figure 19] 10 is a flowchart illustrating an example of an initial setting operation related to the acoustic setting information of the host device. [Figure 20] 10 is a flowchart illustrating an example of an initial setting operation related to acoustic setting information of the portable speaker device. [Figure 21] 10 is a flowchart showing an example of an operation performed by the portable speaker device until the portable speaker device reproduces an acoustic signal. [Figure 22] FIG. 10 is a schematic diagram showing an example of redividing an area map in accordance with the structure of a room. DETAILED DESCRIPTION OF THE INVENTION

[0019] (First embodiment) An acoustic speaker system 1 according to a first embodiment will be described in detail with reference to the drawings.

[0020] Fig. 1 is a block diagram showing the configuration of an acoustic speaker system 1. As shown in Fig. 1, the acoustic speaker system 1 includes three or more portable speaker devices 2 and a host device 3. The portable speaker device 2 includes a sound emitting unit 224, and is a device that processes an acoustic signal 304 received from the host device 3, converts the signal into sound waves, and emits the sound. The host device 3 is a device that controls the portable speaker device 2 and transmits the acoustic signal 304 to the portable speaker device 2. The host device 3 is a computer that is portable by a listener and has a communication function, such as a smartphone, tablet terminal, or laptop PC.

[0021] FIG. 2 is a schematic diagram showing the appearance of a portable speaker device 2. As shown in FIG. 2, the portable speaker device 2 has a portable size that allows a listener to freely change the installation position and installation orientation. The portable speaker device 2 has, for example, a cubic shape. The portable speaker device 2 has sound emitting units 224 distributed on two opposing sides of its six faces. The sound emitting units 224 are, for example, dynamic, cone, or dome type, and have a diaphragm that converts an acoustic signal 304 input as an electrical signal into sound waves through physical vibrations. The sound emitting units 224 face in opposite directions and emit sound in 180-degree opposite directions.

[0022] Fig. 3 is a block diagram showing the internal configuration of the portable speaker device 2. As shown in Fig. 3, the portable speaker device 2 includes a memory 204, a CPU 205, an acceleration sensor 201, a gyro sensor 202, a data communication unit 214, a D / A converter 222, an amplifier 223, and a sound emitting unit 224, some or all of which may be integrated on a single SoC (System on a chip).

[0023] The memory 204 stores programs and data. The CPU 205 operates in accordance with the programs stored in the memory 204. The CPU 205 operates in accordance with the programs to perform acoustic processing on the acoustic signal 304 received from the host device 3.

[0024] The acceleration sensor 201 is a position displacement sensor for measuring the acceleration of an object, and simultaneously measures the acceleration in three axes along the X, Y, and Z axes. The gyro sensor 202 is a position displacement sensor for measuring the angular velocity when an object rotates, and simultaneously measures the angular velocity in three axes along the X, Y, and Z axes.

[0025] The data communication unit 214 is a modem that transmits and receives data in accordance with a wireless communication protocol defined by IEEE802.11, such as WiFI, LAN, or Bluetooth. The data communication unit 214 transmits data output from the CPU 205 to the host device 3. The data communication unit 214 also receives data sent from the host device 3.

[0026] The D / A converter 222 is a digital-to-analog converter, and converts the digital acoustic signal 304 to be sent to the sound emitting unit 224 into analog. The amplifier 223 amplifies the analog acoustic signal 304 and outputs it to the sound emitting unit 224.

[0027] 4 is a block diagram showing the internal configuration of the host device 3. As shown in FIG.

[0028] The memory 301 stores programs and data. The CPU 302 operates according to the programs stored in the memory 301. The CPU 302 processes operations of the listener, acoustic signals 304, and information received from the data communication unit 306.

[0029] The GUI 303 includes a monitor such as an LCD or organic EL monitor that visually displays the screen, and a touch panel that accepts operations by the listener using a pressure-sensitive or capacitance method. The GUI 303 displays various information to the listener, guides the listener, and also accepts operations by the listener and outputs operation information to the CPU 302.

[0030] Data communication unit 306 is a modem that transmits and receives data in accordance with a wireless communication protocol defined by IEEE802.11, such as WiFI, LAN, or Bluetooth. Data communication unit 306 receives data sent from portable speaker device 2, or transmits data from host device 3 to portable speaker device 2. Data communication unit 306 also transmits acoustic signal 304 to portable speaker device 2. Acoustic signal 304 may be stored in a memory in advance, or may be saved in storage on the Internet or LAN.

[0031] In such an acoustic speaker system 1, the portable speaker device 2 will be described in further detail with reference to Fig. 5. By the CPU 205 executing the program in the memory 204, the portable speaker device 2 is provided with an acoustic signal receiving unit 216, a distance estimation unit 206, an area estimation unit 207, a delay correction filter calculation unit 211, an attitude estimation unit 208, a correction filter memory unit 209, a correction filter determination unit 210, a channel selection unit 215, a sound source separation processing unit 217, an acoustic signal selection unit 218, a delay correction calculation unit 219, a correction filter convolution calculation unit 220, a level adjustment unit 221, a data transmission unit 212, and a data reception unit 213.

[0032] The distance estimation unit 206 includes a CPU 205, and uses the acceleration (x, y, z) output by the acceleration sensor 201 and the rotation angles (roll angle φ, pitch angle θ, yaw angle ψ) output by the gyro sensor 202 to determine the position information and separation distance of the portable speaker device 2.

[0033] Fig. 6 shows the relationship between the portable speaker device 2 and the origin 4. Fig. 6 shows a state in which the portable speaker device 2 is placed at an arbitrary position in a three-dimensional space of the x, y and z axes with the origin 4 at the center. The origin 4 is a reference point for measuring the positional relationship between a plurality of portable speaker devices 2, and by initializing the position displacement sensors of the plurality of portable speaker devices 2 at the same position, position changes are handled in a common coordinate space.

[0034] FIG. 7 is a schematic diagram showing an example of turning on the power of the portable speaker device 2. As shown in FIG. 7, the origin 4 is aligned with the position of the listener. That is, the portable speaker device 2 is turned on at the same position as the listener. Specifically, the portable speaker devices 2 are arranged horizontally and vertically with the logo mark 8 facing the listener, so that multiple portable speaker devices 2 are turned on at the same position, in the same direction, and with the same inclination, and their initial positions are aligned. The same range refers to the range within reach of the listener.

[0035] Distance estimation unit 206 determines position information (x, y, z) of the portable speaker device 2 itself from the acceleration of acceleration sensor 201. Distance estimation unit 206 periodically samples the acceleration and accumulates it in the position information to determine the latest position information. Distance estimation unit 206 also calculates the distance from origin 4 to the coordinates indicated by the position information as the separation distance. In this way, the position information is position coordinates in a three-dimensional space. The separation distance is the distance between origin 4 and portable speaker device 2 in the three-dimensional space.

[0036] Area estimation unit 207 includes CPU 205 and memory 204, and determines the area to which portable speaker device 2 belongs based on the position information obtained by distance estimation unit 206. The area to which portable speaker device 2 belongs is the area in which portable speaker device 2 exists when the space in which portable speaker device 2 exists is divided into a plurality of areas. The area to which portable speaker device 2 belongs is determined by comparing the area map stored in advance in memory 204 with the position information. The area map is made up of thresholds indicating the boundaries of each area. Area estimation unit 207 determines whether each coordinate component of the position information is larger or smaller than each coordinate component of each threshold, and identifies the area that includes the coordinates indicated by the position information, and sets the identified area as the area to which portable speaker device 2 belongs.

[0037] FIG. 8 shows an example of an area map in which the space in which the portable speaker device 2 actually plays back is divided into 27 sections. FIG. 8 is a schematic diagram showing a total of 27 divided areas, with the space centered on the origin 4 and divided into three sections in the vertical direction, three sections in the front-rear direction, and three sections in the left-right direction. In FIG. 8, the upper section, divided in the horizontal direction, is represented as U11 to U19, the middle section as M01 to M09, and the lower section as D21 to D29. The belonging area is information indicating in which of these 27 sections the portable speaker device 2 is located. FIG. 9 shows an example of an area divided into eight sections. As shown in FIG. 9, the space centered on the host device 3 can be divided into eight sections, with the two sections being 45° apart and equally spaced around the circumference. The belonging area is information indicating in which of these eight sections the portable speaker device 2 is located. The area map may be divided such that only certain areas are larger than others, and the sizes of the areas do not have to be uniform. The shape of the division of the region map may be not only horizontal and vertical, but also diagonal, such as at 45° in the elevation angle direction, or may be circular or elliptical.

[0038] The data transmission unit 212 is configured to include a data communication unit 214 and a CPU 205, and transmits to the host device 3 the distance of the portable speaker device 2 generated by the distance estimation unit 206 and the area to which the portable speaker device 2 belongs generated by the area estimation unit 207.

[0039] The data receiving unit 213 includes a data communication unit 214 and a CPU 205, and receives the maximum required propagation time, the area setting information, and the correction gain coefficient. The maximum required propagation time, the area setting information, and the correction gain coefficient are selected or generated by the host device 3 in response to the separation distance and the area to which the data transmitting unit 212 belongs, and then transmitted.

[0040] The maximum required propagation time is the time it takes for the sound wave from the farthest portable speaker device 2 to reach the origin 4. The farthest portable speaker device 2 refers to the portable speaker device 2 that is farthest from the origin 4 within the same belonging area as the portable speaker device itself.

[0041] The correction gain coefficient is a gain coefficient assigned to each portable speaker device 2 in order to equalize the outputs of each area. The output of one area is the sum of the outputs emitted by all portable speaker devices 2 arranged in that area. The correction gain coefficient is set in each portable speaker device 2 in order to equalize the sum of the outputs between the areas.

[0042] The acoustic signal receiving unit 216 includes a data communication unit 214 and a CPU 205. The acoustic signal receiving unit 216 receives a radio signal from the host device 3 and extracts an acoustic signal 304 from the radio signal. The acoustic signal receiving unit 216 converts the radio signal into an electrical signal and demodulates the acoustic signal 304 that has been modulated by, for example, Gaussian frequency shift keying, which is one of the frequency shift keying methods.

[0043] An individual ID indicating the destination of the data is added to the maximum required propagation time D, the correction gain coefficient G, and the acoustic signal 304. The portable speaker device 2 stores its own ID in advance, compares the individual ID with its own ID, and receives the maximum required propagation time, the correction gain coefficient, or the acoustic signal 304 if the individual ID matches its own ID.

[0044] The level adjustment unit 221 includes a CPU 205 and multiplies the acoustic signal 304 by a correction gain coefficient received from the host device 3. Consider a case where n portable speaker devices 2 exist in a first belonging area and m portable speaker devices 2 exist in a second belonging area. In this case, the correction gain coefficient for the first belonging area is 1 / n, and the correction gain coefficient for the second belonging area is 1 / m, and these correction gain coefficients equalize the output of the first belonging area and the output of the second belonging area.

[0045] Next, the delay correction filter calculation unit 211 is configured to include a CPU 205, and calculates a delay correction filter to impart a delay to the sound wave transmission of the device itself so that the sound waves of the device itself and the farthest portable speaker device 2 within the same belonging area arrive at the listener at the same time.

[0046] Specifically, as shown in the following formula (1), the separation distance d [m] estimated by the distance estimation unit 206 is divided by the sound speed c [m / s] to calculate the sound wave propagation time dt [s] to the origin 4. Next, as shown in the following formula (2), the maximum delay time ΔT [s], which is the difference between the maximum propagation time D [s] received by the data receiving unit 213 and the sound wave propagation time dt [s], is calculated. Then, a delay correction filter CorrD shown in the following formula (3) is generated from the maximum propagation time D and the maximum delay time ΔT. The delay correction filter CorrD is calculated by applying the maximum delay time ΔT to a phase linear filter H, whose phase does not increase or decrease uniformly in the frequency domain and is constant.

[0047]

number

[0048]

number

[0049]

number

[0050] The delay correction calculation unit 219 includes the CPU 205, and convolves the delay correction filter CorrD with the acoustic signal 304 to correct the delay time according to the position of the portable speaker device 2. The convolution calculation may be performed in the time domain or in the frequency domain after Fourier transformation.

[0051] This delay correction filter CorrD delays the sound emission to match that of the farthest portable speaker device 2. Therefore, this delay correction filter CorrD makes it possible for sound waves from all portable speaker devices 2 belonging to the same area to arrive at the listener at the same time, suppressing blurring of the sound image created by sound waves output from the same area and making it clearer, and preventing unintended cancellation due to interference between the sound waves of the portable speaker devices 2, thereby improving the accuracy of output adjustment between areas.

[0052] The delay correction calculation unit 219 and level adjustment unit 221 perform processing on the audio signal 304 that has undergone sound source separation and remixing by the channel selection unit 215, sound source separation processing unit 217, and audio signal selection unit 218.

[0053] The channel selection unit 215 includes a CPU 205 and a memory 204, and determines the channel that the portable speaker device 2 of the portable speaker device 2 will play. The portable speaker device 2 remixes the acoustic signal 304. The delay correction calculation unit 219 and the level adjustment unit 221 process the acoustic signal 304 of the channel generated by the remixing. Here, the data receiving unit 213 receives region setting information from the host device 3, and the channel selection unit 215 holds the region setting information received from the host device 3. The channel selection unit 215 determines the channel based on this region setting information and the belonging region determined by the region estimation unit 207. The region setting information may be stored in advance without being received from the host device 3.

[0054] The remixing is either upmixing or downmixing, and is determined by the arrangement of each portable speaker device 2. The area setting information is information set for each area to form a sound field for remixing by the portable speaker device 2. The area setting information has a channel type, a mixing function, and mixing setting information for each area of the area map. The channel type indicates which of the remixed channels to play. The mixing function is set with an acoustic processing function that serves as a means for performing remixing. The mixing setting information sets whether remixing is enabled or disabled.

[0055] The region setting information will be specifically described with reference to FIGS. 10 to 13. FIG. 10 is a schematic diagram showing region setting information corresponding to the region map shown in FIG. 8. In FIGS. 10 to 13, a space is divided into three vertical sections, three front-rear sections, and three left-right sections, centered on the origin 4, and portable speaker devices 2 are placed in these 27 divided regions. FIG. 10 shows a specific example of generating a monaural sound field by remixing a stereo audio signal 304. Before describing the sound field generation method, FIG. 10 will be used to explain how to view the region map and the region setting information for each region. In FIG. 10, the region setting information for each of the 27 divided regions is displayed as a single square on a plane. The names of the region setting information on the region map correspond to those shown in FIG. 8. Within each square, the upper row indicates the channel type, the lower row indicates the mixing function, and the color of the square indicates whether the mixing setting information is enabled or disabled. A blacked-out region setting information square indicates that the mixing setting information is disabled. Therefore, in FIG. 10 , the mixing setting information is set to enabled in all regions. For regions where mixing is not performed, the mixing setting information is set to disabled, and re-mixing is not performed. In this example, to generate a monaural sound field from a stereo signal, the sum (L+R) of the left and right channel acoustic signals 304 is calculated. In this case, the left and right channels of the acoustic signal 304 are set as the channel type. Also, an acoustic processing function that calculates the sum (L+R) of the left and right channel acoustic signals 304 is set as the mixing function. Furthermore, if it is desired to cut low frequencies for the portable speaker device 2 belonging to the upper region and output the signal, a high-pass filter (HRF) for cutting low frequencies is additionally set in the acoustic processing function for the upper regions U11 to U19. Also, if it is desired to cut high frequencies for the portable speaker device 2 belonging to the lower region and output the signal, a low-pass filter (LPF) for cutting high frequencies is additionally set in the acoustic processing function for the lower regions D21 to D29.

[0056] Next, Fig. 11 is a schematic diagram showing another example of region setting information corresponding to the region map shown in Fig. 8. Fig. 11 illustrates a specific example of configuring a sound field in which a stereo sound signal 304 is remixed and expanded to all 27 divided regions.

[0057] In this example, of the 27 divided regions, the left channel signal of the acoustic signal 304 is allocated to the left region, and the right channel signal of the acoustic signal 304 is allocated to the right region. For the left regions U11, U14, U17, M01, M04, M07, D21, D24, and D27 in the top, middle, and bottom rows of Fig. 11, the left channel of the stereo signal generated by remixing is set as the channel type, and the acoustic processing function for obtaining the left channel of the stereo signal is set as is as the mixing function.

[0058] For the right-hand areas U13, U16, U19, M03, M06, M09, D23, D26, and D29 of the top, middle, and bottom rows, the channel type is set to the right channel of the stereo signal after remixing, and the mixing function is set to an acoustic processing function for obtaining the right channel of the stereo signal. Furthermore, if it is desired to cut low frequencies for the portable speaker device 2 belonging to the upper area, a high-pass filter (HPF) for cutting low frequencies is additionally set to the acoustic processing function for the upper areas U11 to U19. Furthermore, if it is desired to cut high frequencies for the portable speaker device 2 belonging to the lower area, a low-pass filter (LPF) for cutting high frequencies is additionally set to the acoustic processing function for the lower areas D21 to D29.

[0059] In this example, the mixing setting information is invalid and is associated with the middle areas M02, M05, and M08, the upper areas U12, U15, and U18, and the lower areas D22, D25, and D28. Therefore, remixing is not performed for M02, M05, M08, U12, U15, U18, D22, D25, and D28.

[0060] Next, Fig. 12 is a schematic diagram showing yet another example of region setting information corresponding to the region map shown in Fig. 8. Fig. 12 illustrates a specific example of constructing a surround sound field by remixing the stereo audio signal 304. In this example, the number of regions is greater than the number of playback channels of the audio signal 304 received from the host device 3, and the surround sound field is constructed by remixing the stereo audio signal 304 so as to upmix it.

[0061] 12, the channel types are set for the remixed signals, with the front left channel set in the middle area M01, the center channel set in M02, the front right channel set in M03, the surround left channel set in M04, the surround right channel set in M06, the surround back left channel set in M07, the surround back center channel set in M08, and the surround back right channel set in M09. Similarly, the same channel types as in the middle area are set for the upper areas U11 to U19 and the lower areas D21 to D29.

[0062] The mixing functions are set in the middle region M01 to obtain the front left channel, M02 to obtain the center channel, M03 to obtain the front right channel, M04 to obtain the surround left channel, M06 to obtain the surround right channel, M07 to obtain the surround rear left channel, M08 to obtain the surround rear center channel, and M09 to obtain the surround rear right channel. The same mixing functions as those in the middle region are set in the upper regions U11 to U19 and the lower regions D21 to D29. Furthermore, if it is desired to cut low frequencies for the portable speaker devices 2 belonging to the upper region, a high-pass filter (HPF) for cutting low frequencies is additionally set in the acoustic processing function for the upper regions U11 to U19. Furthermore, if it is desired to cut high frequencies for the portable speaker devices 2 belonging to the lower region, a low-pass filter (LPF) for cutting high frequencies is additionally set in the acoustic processing function for the lower regions D21 to D29.

[0063] In this example, the mixing setting information for the central U15, M05, and D25 is set to invalid, and remixing is not performed.

[0064] Next, Fig. 13 is a schematic diagram showing yet another example of region setting information corresponding to the region map shown in Fig. 8. Fig. 13 explains a specific example of configuring a sound field in which a specific acoustic signal is extracted and remixed. In Fig. 13, only the channel type is shown in one box of the region-specific setting information, and the rest is omitted.

[0065] In this example, audio signal 304 contains a mixture of piano sounds, guitar sounds, and a female voice, each with its own direction, as in a binaural recording, and by remixing, a signal is generated in which the piano is extracted into channel 1 (CH1), the guitar into channel 2 (CH2), and the female voice into channel 3 (CH3). The channel types are assigned as follows: area U11 is assigned to channel 1 (CH1), i.e., the signal in which the piano sound has been extracted through remixing; area M06 is assigned to channel 2 (CH2), i.e., the signal in which the guitar sound has been extracted through remixing; and area M08 is assigned to channel 3 (CH3), i.e., the signal in which the female voice has been extracted through remixing.

[0066] The mixing function is set as follows: an acoustic processing function that extracts channel 1 (CH1), i.e., the sound of a piano, in region U11, an acoustic processing function that extracts channel 2 (CH2), i.e., the sound of a guitar, in region M06, and an acoustic processing function that extracts channel 3 (CH3), i.e., the sound of a female voice, in region M08. The mixing function is set with acoustic processing functions that have been trained using deep learning to extract the target sounds to be extracted, i.e., the sound of a piano in channel 1 (CH1), the sound of a guitar in channel 2 (CH2), and the sound of a female voice in channel 3 (CH3).

[0067] In this example, the mixing setting information is set to invalid for areas other than U11, M06, and M08, and remixing is not performed.

[0068] The sound source separation processing unit 217 includes the CPU 205, and generates an acoustic signal by remixing the acoustic signal 304 received by the acoustic signal receiving unit 216 in accordance with the set region setting information.

[0069] The acoustic processing function refers to a function set as a means for remixing an input acoustic signal and is set as a mixing function of the region setting information. One example of a function set as an acoustic processing function is a function that extracts components of the same amplitude and phase from two stereo acoustic signals 304 to generate a channel for an acoustic signal 304 that extracts a center vocal component that produces a perceived sound image in the center. Another example is a function that uses the correlation between the left and right channel signals of the stereo acoustic signal 304 to generate a channel for an acoustic signal 304 that extracts a surround reverberation component using an adaptive filter for components with low correlation. Another example is a function that is assigned a machine learning algorithm trained through deep learning and generates a channel for an acoustic signal 304 that extracts a specific musical instrument. Yet another example is a function that generates a channel for an acoustic signal 304 that has passed only a specific band through a high-pass filter, a low-pass filter, or the like.

[0070] The audio signal selection unit 218 is configured to include a CPU 205, and selects the audio signal 304 of the channel corresponding to the playback channel determined by the channel selection unit 215 from the audio signals 304 of each channel that have been remixed and generated by the sound source separation processing unit 217.

[0071] Attitude estimation unit 208 includes CPU 205, and estimates attitude information indicating the direction and inclination of portable speaker device 2 using the rotation angles (roll angle φ, pitch angle θ, yaw angle ψ) output by gyro sensor 202. The attitude information is made up of roll angle φ and pitch angle θ, assuming that the position of portable speaker device 2 when powered on is the origin of the three-dimensional space coordinate system (φ=0, θ=0, ψ=0).

[0072] The roll angle φ is the angle around an axis passing through the front and back surfaces of the portable speaker device 2. The front surface is the surface that is set to face the listener when the power is turned on, and has, for example, a logo mark 8 displayed on it. The pitch angle θ is the angle around an axis passing through the left and right surfaces. Each portable speaker device 2 is turned on with the roll angle φ and pitch angle θ set to the same.

[0073] The attitude information is initialized to a rotation angle (φ=0, θ=0, ψ=0) when the portable speaker device 2 is powered on. Then, the attitude estimation unit 208 periodically samples the rotation angle and accumulates it in the attitude information to determine the latest attitude information.

[0074] The correction filter memory unit 209 includes the memory 204 and stores a correction filter. The correction filter is a filter for canceling out the acoustic transfer characteristic convolved with the sound wave from the portable speaker device 2 in the direction and orientation indicated by the posture information to the origin 4, and for convolving the acoustic transfer characteristic from the portable speaker device 2 facing the origin 4 to the origin 4.

[0075] The portable speaker device 2 emits sound radially. At this time, the acoustic transfer characteristics until the sound wave reaches the origin 4 change depending on the orientation of the portable speaker device 2. The acoustic transfer characteristics change when the portable speaker device 2 is facing the origin 4 compared to when it is facing 90 degrees away. Therefore, regardless of the orientation and inclination of the portable speaker device 2, the acoustic transfer characteristics are corrected to those when the portable speaker device 2 is facing the origin 4.

[0076] The correction filter is composed of an inverse filter of a one-dimensional linear filter in a 360° direction from the center to the measurement point when measurement points are placed at equal distances and equal angular intervals around a single point in space. Figure 14 shows an example of a method for obtaining a correction filter. An omnidirectional speaker is placed in an anechoic chamber, and microphones 61 are placed at 1° intervals around the circumference of a sphere 62 with a radius of 1 meter centered at the position where the speaker is placed. The impulse responses from the speaker to all microphones 61 are measured, and a one-dimensional linear filter cut to an arbitrary filter length is obtained. The inverse matrix of the obtained one-dimensional linear filter is calculated and used as the correction filter.

[0077] Although it is desirable to have a high density of microphones 61 used for measurement by the compensation filter, i.e., measurement points, the data volume of the filter can be compressed and stored using mathematical techniques. For example, the measurement points can be thinned out every 5°, and the angles between them can be calculated by approximating the acoustic transfer characteristics at measurement points on the same sphere using spherical harmonics. Furthermore, the compensation filter can be generated from a one-dimensional linear filter of the acoustic transfer characteristics measured with a microphone that does not include the head, as in the example of obtaining the compensation filter described above, or it can be generated by measuring with a dummy head microphone. The compensation filter can be stored in the time domain, or it can be stored in the frequency domain after Fourier transform. Alternatively, the transfer characteristics stored in the frequency domain can be simulated and stored using an IIR filter.

[0078] The correction filter determination unit 210 is configured to include the CPU 205 and determines acoustic transfer characteristics corresponding to the tilt of the portable speaker device 2. Specifically, the correction filter is determined by referring to the correction filter memory unit 209 for information on the tilt and angle of the attitude information of the portable speaker device 2 estimated by the attitude estimation unit 208 and a correction filter corresponding to an angle specified by the tilt and angle of the attitude information. The correction filter convolution calculation unit 220 is configured to include the CPU 205 and convolves the acoustic signal 304 with the correction filter determined by the correction filter determination unit 210. The convolution calculation may be performed in the time domain or in the frequency domain after Fourier transform.

[0079] In this way, by correcting the levels by the level adjustment unit 221 so that the sum of the outputs of the portable speaker devices 2 in each area is equal, when a listener uses two or more portable speaker devices 2 in one area, a sound field can be generated without losing the balance of the volume between channels regardless of the number of portable speaker devices 2 used. Furthermore, by correcting the delay time by the delay correction calculation unit 219 so that all portable speaker devices 2 emit sound at the same timing, even if a plurality of portable speaker devices 2 are placed randomly in one area, the signals emitted from the portable speaker devices 2 do not interfere with each other and unintended cancellation occurs before they reach the listener, making it possible to perform accurate sound field control.

[0080] Next, the host device 3 will be described in more detail with reference to Fig. 15. The host device 3 is provided with a maximum required transmission time calculation unit 308, a correction gain calculation unit 309, a selection unit 310, a data receiving unit 312, and a data transmitting unit 311, which are configured by the CPU 302 executing a program in the memory 301.

[0081] The data receiving unit 312 includes the data communication unit 306 and the CPU 302, and receives the separation distance and the area to which each portable speaker device 2 belongs from each portable speaker device 2.

[0082] The maximum required transmission time calculation unit 308 includes the CPU 302, and identifies the distance of the farthest portable speaker device 2 in each belonging area by comparison from the distances received from each portable speaker device 2. The maximum required transmission time calculation unit 308 also converts the distance of the farthest portable speaker device 2 into a maximum required transmission time, which is the time required for a sound wave to travel from the farthest portable speaker device 2 to the origin 4. The maximum required transmission time is calculated for each belonging area.

[0083] The maximum required transmission time for each belonging area calculated by the maximum required transmission time calculation unit 308 is transmitted to each portable speaker device 2. Each portable speaker device 2 receives the maximum required transmission time for the belonging area to which the device itself belongs. The received maximum required transmission time is used for delaying the sound emission timing.

[0084] The correction gain calculation unit 309 includes the CPU 302 and calculates correction gain coefficients for all connected portable speaker devices 2. The correction gain coefficients are set individually for all portable speaker devices 2 in all regions so that the sum of the outputs of each region is the same between the regions.

[0085] Here, the correction gain coefficient is Gn. The separation distance between the portable speaker devices 2 for which the correction gain coefficient Gn is to be calculated is L1. The separation distance between the farthest portable speaker device 2 placed at the farthest position among all the connected portable speaker devices 2 is L2. The number of portable speaker devices 2 belonging to the area to which the portable speaker device 2 for which the correction gain coefficient Gn is to be calculated belongs is M. In this case, the correction gain coefficient Gn is calculated by multiplying the distance ratio L2 / L1 by the reciprocal of the number M of portable speaker devices 2 in the same area, as shown in the following formula (4).

[0086]

number

[0087] For example, there are two regions, region M01 and region M02, as shown in Fig. 16. In the case where region M01 contains portable speaker devices 51 and 52 which are portable speaker devices 2, and region M02 contains portable speaker device 53 which is also portable speaker device 2, a method for calculating the correction gain coefficient Gn in each portable speaker device 2 will be described.

[0088] 16, distance 54 indicates the distance from origin 4 to portable speaker device 51, distance 55 indicates the distance from origin 4 to portable speaker device 52, and distance 56 indicates the distance from origin 4 to portable speaker device 53. Furthermore, the distance from portable speaker device 2 placed at the farthest position is defined as L2.

[0089] The correction gain coefficient G1 of portable speaker device 51 is calculated by applying equation (4). In the example shown in FIG. 16, L1 corresponds to distance 54, and L2 corresponds to distance 56. The number M of portable speaker devices in area M01 is two, and when equation (4) is applied, the correction gain coefficient G1 of portable speaker device 51 can be calculated as (distance 56 / distance 54)×½. Similarly, the correction gain coefficient G2 of portable speaker device 52 is calculated as (distance 56 / distance 55)×½. Furthermore, the number M of portable speaker devices 2 in area M02 is one, and when equation (4) is applied, the correction gain coefficient G3 of portable speaker device 53 is calculated as (distance 56 / distance 56)×½=1.

[0090] The correction gain calculation unit 309 searches for the longest distance L1 from the separation distances transmitted from each portable speaker device 2. Furthermore, the correction gain calculation unit 309 sorts the area information transmitted from each portable speaker device 2, and counts the number M of portable speaker devices 2 belonging to the same area for each area.

[0091] The correction gain coefficient Gn may be set for each frequency using formula (4), or may be set for each band obtained by dividing the frequency band by any number, such as high, mid, and low bands, using formula (4). Alternatively, the coefficient may be set using formula (4) with all frequencies included.

[0092] The selection unit 310 is configured to include a CPU 302 and memory 301, and each parameter of the area map is set by the listener operating the GUI 303, and area setting information is generated using the area map. The method of setting the area map and the method of generating the area setting information will now be described. Fig. 17 is a schematic diagram of the GUI for setting area setting information from the area map. First, Fig. 17 shows an example in which the selection unit 310 sets area map parameters by operating the GUI.

[0093] As shown in Fig. 17, when U11 is selected from the setting GUI 95 for each region, a window opens in which the name of the selected region and the mixing function, region setting information, remixing function, and band equalizer settings are set using the components of the region map. For example, in the case of setting to generate the monaural sound field of Fig. 10, for region U11, the acoustic processing function for generating a monaural sound field, i.e., the acoustic processing function (L+R) that calculates the sum of the left and right channels of the audio signal, is selected as the mixing function, the region setting information is enabled, the acoustic processing function (L+R) is selected as the mixing function, and a high-pass filter is selected as the band equalizer setting.

[0094] The same settings are made for other channels. For unused areas, the area setting information is set to invalid.

[0095] Next, a method for generating the area setting information will be described. The area setting information has a channel type set to monaural, a mixing setting set to whether the area setting is enabled or disabled, and a mixing function set to the name or type of an acoustic processing function used in the mixing method and band equalizer setting. The acoustic processing function name set in the mixing function is called by a callback function or a corresponding function in a program deployed in the CPU 205 of the portable speaker device 2, and processing is performed by the called acoustic processing function.

[0096] The area-specific settings to be transmitted to the portable speaker device 2 may be transmitted only for areas where area setting information is valid, or only for areas where the portable speaker device 2 has been detected. Alternatively, the settings may be transmitted for all areas constituting the area map, regardless of whether the area setting information is valid or invalid.

[0097] Data transmission unit 311 is configured to include data communication unit 306 and CPU 302, and transmits the maximum required transmission time generated by maximum required transmission time calculation unit 308 of the program of host device 3, the area setting information generated by selection unit 310, and the correction gain coefficient generated by correction gain calculation unit 309 to portable speaker device 2. When transmitting, data transmission unit 311 modulates this information using, for example, a Gaussian frequency shift keying method, which is one of the frequency shift methods, and further converts it into a wireless signal to output to portable speaker device 2.

[0098] The operation of such an acoustic speaker system 1 will be described with reference to Fig. 18 to Fig. 21. Fig. 18 to Fig. 21 are flowcharts showing an example of the operation of this acoustic speaker system 1. The processing of each of the host device 3 and portable speaker device 2 that constitute the acoustic speaker system 1 will be described in chronological order. It is assumed that one host device 3 is connected to one or more portable speaker devices 2.

[0099] First, a flow from starting the program shown in FIG. 18 to setting up the portable speaker device 2 will be described.

[0100] The host device 3 starts a program from the CPU 302 (step S01). When the program is started, it transmits a connection request to the portable speaker device 2 to establish wireless communication with the portable speaker device 2 (step S02). When the portable speaker device 2 is powered on, the program is started from the CPU (step S03). After the portable speaker device 2 is powered on, it receives the connection request previously sent from the host device 3 and establishes connection (step S04). The connection wait state continues until the connection with the host device 3 is completed (step S04n), and when the connection is completed (step S04y), the portable speaker device 2 notifies the host device 3 of connection completion (step S05).

[0101] After the portable speaker device 2 is powered on, the acceleration sensor 201 is activated and the separation distance is initialized, i.e., the position of the portable speaker device 2 when the power was turned on is set as the origin 4 (x=0, y=0, z=0) (step S06). After the portable speaker device 2 is powered on, the gyro sensor 202 is activated and the attitude information is initialized, i.e., the tilt of the portable speaker device 2 when the power was turned on is set as the initial value (step S07). Note that each portable speaker device 2 is powered on at a common position, direction and tilt, and the initial positions are aligned.

[0102] The listener moves the portable speaker device 2 and determines the installation position (step S08). After the gyro sensor 202 is activated, the attitude estimation unit 208 generates attitude information (step S09). After the acceleration sensor 201 is activated, the distance estimation unit 206 generates a separation distance (step S10). Using the separation distance generated in step S10, the unit compares it with an area map, which is a space divided into an arbitrary number of parts and is stored in memory in advance, to determine the area to which the portable speaker device 2 belongs (step S11). The transmitter transmits the separation distance generated in step S10 and the area generated in step S11 to the host device 3 (step S12).

[0103] The host device 3 receives the separation distance and the area to which the portable speaker device 2 belongs transmitted in step S12 (step S14). When two or more portable speaker devices 2 are connected, this information is transmitted from each portable speaker device 2, and the host device 3 receives all of it. The host device 3 detects that all portable speaker devices 2 have been installed (step S15), continues to wait for installation until installation is complete (step S15n), detects that installation is complete (step S15y), and completes installation of all portable speaker devices 2 (step S16). Whether installation of the portable speaker devices 2 is complete may be determined by detecting that the received separation distance does not change for a certain period of time, or may be determined by the host device 3 detecting an installation completion trigger transmitted from the portable speaker device 2.

[0104] Next, a description will be given of the process performed mainly on the host device 3 side in the initial setting method for the acoustic setting information shown in Fig. 19. This process flow is performed after the portable speaker device 2 shown in Fig. 18 is installed.

[0105] First, the maximum required transmission time calculation unit 308 identifies the portable speaker device 2 that is farthest from the origin 4 based on the separation distance and the belonging area of each portable speaker device 2, and calculates the maximum required transmission time D, which is the delay time of the identified portable speaker device 2 (step S17). Next, the correction gain calculation unit 309 calculates, using the separation distance and the belonging area of each portable speaker device 2 received in step S14 and the maximum required transmission time D calculated in step S17, a correction gain coefficient G for each portable speaker device 2, which makes the sum of the outputs emitted by all portable speaker devices 2 belonging to the same area equal between the areas, using equation (4) (step S18).

[0106] Furthermore, the area setting information is generated by a GUI operation by the listener or a preset initial setting value (step S19). The data transmitting unit 311 transmits the maximum required transmission time D, the correction gain coefficient G, and the area setting information acquired in steps S17 to S20 to all the portable speaker devices 2 (step S20).

[0107] The data receiving unit 213 of the portable speaker device 2 receives the maximum required transmission time D, the correction gain coefficient G, and the area setting information transmitted from the host device 3 in step S20 (step S21).

[0108] Next, a description will be given mainly on the portable speaker device 2 side, continuing from the initial setting method for the acoustic setting information shown in Fig. 20. The correction filter determination unit 210 refers to the acoustic transfer function stored in the correction filter memory unit 209 corresponding to the posture information obtained in step S09, and determines a correction filter (step S22).

[0109] Next, the channel selection unit 215 extracts the area setting information corresponding to the belonging area set in step S20 from the area setting information received from the host device 3. The channel type is read from the area setting information, and the channel to be played by the portable speaker device 2 is determined (step S23).

[0110] The sound source separation processing unit 217 reads and sets the mixing setting information regarding whether mixing is enabled or disabled and the acoustic processing function set as a means for actually performing remixing from the area setting information received from the host device 3 in step S21 (step S24).

[0111] The delay correction filter calculation unit 211 calculates the sound wave propagation time dt [s] to the origin 4 by dividing the distance d [m] by the speed of sound c [m / s] using the maximum propagation time D received from the host device 3 in step S21 and the position of the portable speaker device 2, i.e., the distance d [m] from the origin 4, calculated in step S10, using equation (1). Next, the maximum delay time ΔT [s], which is the difference between the maximum propagation time D and the maximum delay time ΔT, is calculated using equation (3). The delay correction filter CorrD is calculated by applying a delay of the maximum delay time ΔT to a phase linear filter H, which has a constant phase and does not increase or decrease uniformly in the frequency domain (step S25).

[0112] Next, a flow of operations until the portable speaker device 2 shown in FIG. 21 reproduces an acoustic signal will be described.

[0113] The host device 3 transmits the acoustic signal 304 to the portable speaker device 2 (step S26), and the acoustic signal receiving unit 216 of the portable speaker device 2 receives the acoustic signal 304 (step S27).

[0114] The sound source separation processing unit 217 determines whether the remixing setting is valid or not using the mixing setting information of the region setting information set in step S24 (step S28). If the remixing setting is valid (step S28y), the sound source separation processing unit 217 applies the acoustic processing function set as the mixing function in step S24 to the acoustic signal 304 received in step S27 to generate another remixed signal (step S29y).

[0115] The flow up to the generation of another remixed signal by the sound source separation processing unit 217 will be specifically described using Fig. 12 as an example. In Fig. 12, a surround sound field is constructed by remixing a stereo audio signal 304. In this example, the number of regions is greater than the number of playback channels of the audio signal 304 received from the host device 3, and the surround sound field is constructed by remixing the stereo audio signal 304 in an upmixing manner.

[0116] The region setting information set in step S24 will be described with reference to FIG. 12. First, the channel types are set for the remixed signals, with the front left channel set to region M01, the center channel set to region M02, the front right channel set to region M03, the surround left channel set to region M04, the surround right channel set to region M06, the surround back left channel set to region M07, the surround back center channel set to region M08, and the surround back right channel set to region M09. Similarly, the same channel types as those in the middle region are set for regions U11 to U19 in the upper row and regions D21 to D29 in the lower row. Next, acoustic processing functions are set for the mixing functions to obtain the front left channel in region M01, the center channel set to region M02, the front right channel set to region M03, the surround left channel set to region M04, the surround right channel set to region M06, the surround right channel set to region M07, the surround back left channel set to region M08, the surround back center channel set to region M09. The same mixing function as the middle section is set for the upper sections U11 to U19 and the lower sections D21 to D29. Furthermore, a high-pass filter (HPF) to cut low frequencies is added to the acoustic processing function for the upper sections U11 to U19, and a low-pass filter (LPF) to cut high frequencies is added to the acoustic processing function for the lower sections D21 to D29. Finally, the mixing setting information is disabled for the central sections U15, M05, and D25, and enabled for the other sections.

[0117] The sound source separation processing unit 217 refers to the mixing setting information in the region setting information and determines whether the remixing setting is valid (step S28). For U15, M05, and D25 for which the remixing setting information is determined to be invalid, the sound source separation processing unit 217 passes the signals as they are without generating a mixed signal (step S28n). For regions for which the remixing setting is determined to be valid, the sound processing function set in the mixing function is applied to generate a different remixed signal (step S28y). A specific description will be given using region M02 as an example. If the mixing function for M02 is set to an sound processing function that obtains a center channel from the stereo sound signal 304, for example, an sound processing function that obtains the sum (L+R) of the left and right channels of the stereo sound signal 304, the sound source separation processing unit 217 generates the sum (L+R) of the left and right channels of the stereo sound signal 304 as the remixed signal. For other regions, remixed signals are generated in the same manner as above.

[0118] The acoustic signal selection unit 218 selects and extracts the acoustic signal corresponding to the channel selection unit 215 set in step S23 from the acoustic signals generated by applying the acoustic processing function and remixing in step S29y (step S30).

[0119] The delay correction calculation unit 219 calculates an acoustic signal 304 by convolving the acoustic signal of the specified channel after remixing extracted by the acoustic signal selection unit 218 in step S30 and the delay correction filter obtained in step S25, and corrects the delay time according to the position where the portable speaker device 2 is placed (step S31).

[0120] The correction filter convolution calculation unit 220 outputs the acoustic signal 304 obtained by convolving the correction filter determined by the correction filter determination unit 210 in step S22, i.e., the acoustic transfer function corresponding to the posture information, with the acoustic signal 304 whose delay time has been corrected by the delay correction calculation unit 219 in step S25, and corrects the acoustic signal 304 so that the portable speaker device 2 placed in any direction and in any inclination has a common acoustic transfer characteristic regardless of the direction and inclination (step S32). The convolution calculation may be performed in the time domain or the frequency domain.

[0121] Finally, the level adjustment unit 221 multiplies the acoustic signal 304 output by the convolution calculation performed by the correction filter convolution calculation unit 220 in step S32 by the correction gain coefficient G received from the host device 3 in step S21, and corrects the level so that the sum of the outputs emitted by the portable speaker devices 2 belonging to each area becomes equal (step S33).

[0122] The acoustic signal 304 whose level has been corrected in step S33 is D / A converted by the D / A converter 222 and amplified by the amplifier 223 (step S34), and the acoustic signal is output from the sound emitting unit 224 (step S35). As a result, the portable speaker device 2 reproduces the acoustic signal 304.

[0123] In this way, the acoustic speaker system 1 is configured so that the portable speaker device 2 includes the acceleration sensor 201, which is a position sensor, the distance estimation unit 206, and the area estimation unit 207. The host device 3 is also configured so that it includes the data receiving unit 312, which is a receiving unit, the correction gain calculation unit 309, and the data transmitting unit 311, which is a transmitting unit.

[0124] An acceleration sensor 201, which is a position sensor, detects changes in the position of the housing. A distance estimation unit 206 obtains position information and distance in three-dimensional space according to the position changes detected by the position sensor. An area estimation unit 207 compares pre-stored area information obtained by dividing the three-dimensional space with the position information estimated by the distance estimation unit 206 to determine the area to which the device belongs.

[0125] Data receiving unit 312 receives the distances and belonging areas of all portable speaker devices 2. Correction gain calculation unit 309 searches for the farthest distance among all portable speaker devices, and calculates a gain coefficient that equalizes the sum of the outputs emitted by portable speaker devices 2 in each belonging area, based on the number of portable speaker devices 2 belonging to the same belonging area. The gain coefficient is calculated by L2 / L1*1 / M, which is obtained by multiplying L2 / L1, the ratio of the distance L1 of the portable speaker device itself to the farthest distance L2, by 1 / M, which is the reciprocal of the number M of portable speaker devices 2 belonging to the belonging area. Data transmitting unit 311, which is a transmitting unit, transmits the correction gain coefficient to the portable speaker device 2.

[0126] The portable speaker device 2 further includes a level adjustment unit 221 that multiplies the acoustic signal by a correction gain coefficient, and a sound emission unit 224 that converts the acoustic signal that has passed through the level adjustment unit 221 into a sound wave and outputs it. This allows the volume levels of sounds emitted from each direction to be balanced, and a sound field can be generated regardless of the number of portable speaker devices 2 used.

[0127] Distance estimation unit 206 of portable speaker device 2 further calculates the distance from origin 4 to the position of the host device 3 based on the position change detected by acceleration sensor 201, which is a position sensor. Host device 3 further includes maximum required propagation time calculation unit 308 that calculates the time it takes for a sound wave output from the portable speaker device 2 that is farthest from origin 4, i.e., the farthest portable speaker device, to reach origin 4, i.e., the maximum required propagation time. Portable speaker device 2 further includes delay correction calculation unit 219 that delays sound emission based on the maximum required propagation time calculated by maximum required propagation time calculation unit 308 and the distance calculated by distance estimation unit 206, so that the sound wave from the farthest portable speaker device 2 and the sound wave output by the host device 3 arrive at origin 4 at the same time.

[0128] The portable speaker device transmits the separation distance calculated by distance estimation unit 206 to the host device. Data receiving unit 312 receives the separation distances of all portable speaker devices 2, and data transmitting unit 311, which is a transmitting unit, transmits the maximum required propagation time calculated by maximum required propagation time calculation unit 308 to portable speaker device 2. Based on the maximum required propagation time received from host device 3 and the separation distance calculated by distance estimation unit 206, delay correction calculation unit 219 delays sound emission so that the sound wave of the farthest portable speaker device and the sound wave output from the portable speaker device itself arrive at origin 4 at the same time. This aligns the time it takes for sound waves emitted from each direction to reach the listener, so that a sound field can be generated regardless of the position of the portable speaker device 2 used, even if multiple portable speaker devices 2 are placed randomly in one area.

[0129] (Second embodiment) An acoustic speaker system 1 according to the second embodiment will be described in detail with reference to the drawings. Note that the same functions and configurations as those in the first embodiment will be assigned the same reference numerals and detailed description thereof will be omitted.

[0130] In the first embodiment, an area map in which areas are divided into specified ranges and sizes is prepared in advance and this area map is used, but the area map may also be variable to match the size of the space in which the acoustic speaker system 1 generates the sound field.

[0131] The selection unit 310 of the host device 3 stores reference dimension information indicating the size of the room in which the portable speaker device 2 is to be played back. Furthermore, in order to generate area setting information in the GUI 303, the selection unit 310 displays an input screen for the size of the room in which the portable speaker device 2 is to be played back. The selection unit 310 calculates the ratio W1 / Wd between the room size W1 input via the GUI 303 and the size Wd indicated by the reference dimension information, and expands or reduces the area map in accordance with the ratio W1 / Wd. This changes the size of each divided area on the area map.

[0132] The data transmission unit 311 of the host device 3 transmits the area setting information, which has been changed to match the size of the room, to each portable speaker device 2, and the area estimation unit 207 of the portable speaker device 2 determines the area to which it belongs based on the area setting information, which has been changed to match the size of the room.

[0133] In this way, when determining the area to which each portable speaker device 2 belongs, the area information can be changed according to the size of the actual room. This makes it possible to determine the area to which each portable speaker device 2 belongs in accordance with the room, thereby generating a better sound field. For example, even if portable speaker devices 2 are distributed in a room smaller than the reference dimension information, it is possible to prevent the portable speaker devices 2 from belonging to the same area and only playing mono.

[0134] The size of the room may be estimated by displaying an input screen on the GUI 303 to accept manual input by the user, or by using a camera (not shown) included in the host device 3. The selection unit 310 calculates the size of the room by analyzing an image or video of the room captured by the camera.

[0135] (Third embodiment) An acoustic speaker system 1 according to the third embodiment will be described in detail with reference to the drawings. Note that the same functions and configurations as those in the first or second embodiment will be assigned the same reference numerals and detailed description thereof will be omitted.

[0136] The distance estimation unit 206 of the portable speaker device 2 stores upper distance limit information in advance. The upper distance limit information is information indicating a distance for determining whether the portable speaker device 2 needs to emit sound. The distance estimation unit 206 compares the separation distance of the portable speaker device 2 with the upper distance limit information. If the separation distance exceeds the upper distance limit information, the data transmission unit 212 stops transmitting the separation distance and affiliation information to the host device 3. If the separation distance of the portable speaker device 2 exceeds the upper distance limit information, the portable speaker device 2 may stop transmission by not establishing a connection with the host device 3.

[0137] A portable speaker device 2 that does not establish a connection with the host device 3 is not present in the host device 3 and cannot receive the acoustic signal 304. Furthermore, a portable speaker device 2 whose separation distance exceeds the upper distance limit information is excluded from the search and counting of the farthest portable speaker device 2 by the correction gain calculation unit 309 of the host device 3, and is also excluded from the comparison process for calculating the maximum required transmission time by the maximum required transmission time calculation unit 308.

[0138] Here, if the distance of a portable speaker device 2 that is located at a distance exceeding the distance indicated by the distance upper limit information is used to calculate the correction gain coefficient Gn in equation (4), L2 / L1 in equation (4) becomes too large, reducing the accuracy of volume control of the portable speaker device 2. In other words, by excluding a portable speaker device 2 that is located at a distance exceeding the distance indicated by the distance upper limit information from the search for and counting of the farthest portable speaker device 2, the total output of sound emitted by the portable speaker devices in each of the belonging areas can be accurately adjusted.

[0139] The comparison with the upper distance limit information may be performed on the host device 3 side, and the portable speaker device 2 at a distance exceeding the distance indicated by the upper distance limit information may be excluded from the calculation of the correction gain coefficient Gn in the correction gain calculation unit 309.

[0140] (Fourth embodiment) An acoustic speaker system 1 according to the fourth embodiment will be described in detail with reference to the drawings. Note that the same functions and configurations as those in the first to third embodiments will be assigned the same reference numerals and detailed description thereof will be omitted.

[0141] In the acoustic speaker systems 1 according to the first to third embodiments, the regions are divided according to the direction centered on the origin 4. As a method of dividing the regions, the regions may be further divided according to the structure of the room in which the portable speaker device 2 is played back. A region whose acoustic characteristics differ from other regions is identified from the structure of the room, and this region in which the acoustic characteristics are likely to change in a particular way is used as a division criterion for the regions, separate from the direction.

[0142] An area with acoustic characteristics different from other areas is, for example, an area where reflected sound is likely to occur, and has acoustic characteristics that include a large amount of reflected sound compared to other areas where reflected sound is relatively unlikely to occur. Areas with acoustic characteristics different from other areas can also be distinguished by the type of reflected sound; for example, an area close to a wall will have different reflected sounds than an area close to a window.

[0143] The selection unit 310 of the host device 3 analyzes the structure of the room from the image or video captured by the camera of the host device 3, and adds areas close to walls or windows to each area in the area map.

[0144] FIG. 22 shows an example of dividing a new area near a window. In FIG. 22, three portable speaker devices 2 (51 to 53) are placed in a room 101. In the area map, a boundary 102a is drawn from the origin 4 in the direction directly in front of the listener, dividing the room 100 into two areas, with area 103a on the left side and area 103b on the right side. First, the structure of the room 100, including the dimensions and shape, is acquired from an image or video captured by a camera or manually entered. By applying the room structure to the area map, it is determined that a window 101 is located on the left wall of the room 100 and that a portable speaker device 2B52 is located within 30 cm of the window. Based on this determination result, a boundary 102b is set at an equal distance of 30 cm from the window, and the listener manually divides the area by setting a new area 103c within that range in the area map. Alternatively, a process of automatically dividing the area by analyzing the area map and the room structure may be added.

[0145] (Other embodiments) Furthermore, the embodiments and examples of the present invention are presented as examples only and are not limited to the above embodiments and examples. The above embodiments and examples can be implemented in various other forms, and various omissions, substitutions, and modifications can be made without departing from the scope of the invention. The embodiments, examples, and their modifications are included within the scope of the present invention. [Explanation of symbols]

[0146] 1. Acoustic speaker system 2 Portable speaker device 201 Acceleration sensor 202 Gyro sensor 204 memory 205 CPU 206 Distance estimation unit 207 Area estimation part 208 Posture estimation section 209 Correction filter memory section 210 Correction filter determination unit 211 Delay correction filter calculation unit 212 Data transmission unit 213 Data receiving unit 214 Data Communications Department 215 Channel selection section 216 Acoustic signal receiving unit 217 Sound source separation processing unit 218 Acoustic signal selection unit 219 Delay Correction Calculation Unit 220 Correction filter convolution calculation unit 221 Level adjustment section 222 D / A converter 223 Amplifier 224 Sound emission section 3 Host Device 301 Memory 302 CPU 303 GUI 304 Acoustic Signals 306 Data Communications Department 308 Maximum transmission time calculation unit 309 Correction gain calculation unit 310 Selection Section 311 Data Transmission Unit 312 Data receiving unit U11 Upper left front area U12 Upper and anterior area U13 Upper right front area U14 Upper left area U15 Area directly above U16 Upper right area U17 Upper left rear area U18 Upper left rear area U19 Upper right rear area M01 Area in the left front direction M02 Forward area M03 Area in the right front direction M04 Leftward Area M05 Central Region M06 Rightward Area M07 Left rear area M08 Backward Region M09 Right rear area D21 Lower left front area D22 Lower front area D23 Lower right front area D24 Lower left area D25 Downward Area D26 Lower right area D27 Lower left rear area D28 Lower-posterior area D29 Lower right rear area 4. Origin 51, 52, 53 Portable speaker device 54, 55, 56 distance 61 Microphone 62 Spherical 63 center 7 Listeners 8. Logo 90 Area Map 91 Area map in the upward direction 92 Area map in the middle 93 Downward Area Map 94 Area maps for each area 95 GUI for setting each area 100 rooms 101 Window 102 Boundary 103 areas

Claims

1. An acoustic speaker system comprising a plurality of portable speaker devices and a host device, The portable speaker device a position sensor that detects a change in the position of the housing; a distance estimation unit that calculates position information and a distance in a three-dimensional space in response to the position change detected by the position sensor; a region estimation unit that determines a region to which the object belongs by comparing pre-stored information on regions obtained by dividing a three-dimensional space with the position information estimated by the distance estimation unit; a speaker-side transmitting unit that transmits the distance calculated by the distance estimating unit and the belonging area determined by the area estimating unit to the host device; Equipped with The host device a receiving unit that receives the distances and the areas to which all portable speaker devices belong; a correction gain calculation unit that calculates a correction gain coefficient that equalizes the sum of the outputs of the portable speaker devices in the respective belonging areas; a transmitter for transmitting the correction gain coefficient to the portable speaker device; Equipped with The correction gain calculation unit retrieving the farthest said distance among all said portable speaker devices; Counting the number M of the portable speaker devices that belong to the same belonging area; calculating a ratio L2 / L1 of the distance L2 of the portable speaker device to the distance L1 of the farthest distance; Calculating 1 / M, which is the reciprocal of the number M of the portable speaker devices belonging to the same belonging area; calculating a gain coefficient L2 / L1*1 / M obtained by multiplying the ratio L2 / L1 by the reciprocal 1 / M, and corresponding the gain coefficient to each of the portable speaker devices; The portable speaker device a level adjustment unit that multiplies the acoustic signal by the correction gain coefficient; a sound output unit that converts the acoustic signal that has passed through the level adjustment unit into a sound wave and outputs the sound wave; Further comprising: An acoustic speaker system comprising:

2. the distance estimation unit of the portable speaker device further calculates a separation distance from an origin to the position of the portable speaker device itself based on the position change detected by the position sensor; the speaker-side transmitter further transmits the separation distance to the host device; the host device further comprises a maximum required propagation time calculation unit that calculates a maximum required propagation time for a sound wave to reach the origin from a portable speaker device that is farthest from the origin within the same belonging area, based on the separation distances transmitted from each of the portable speaker devices; the transmitting unit of the host device further transmits the maximum required transmission time to the portable speaker device; the portable speaker device further comprises a delay correction calculation unit that delays sound emission based on the maximum required propagation time and the separation distance so that the sound wave of the farthest portable speaker device and the sound wave of the farthest portable speaker device arrive at the origin at the same time; 2. The acoustic speaker system according to claim 1, wherein:

3. the portable speaker device includes an acoustic signal receiving unit that receives the acoustic signal; the level adjustment unit multiplies the acoustic signal received by the acoustic signal receiving unit by the correction gain coefficient; 3. The acoustic speaker system according to claim 1 or 2, wherein:

4. the distance estimation unit pre-stores distance upper limit information indicating a predetermined distance, and compares the distance calculated in accordance with the position change with the predetermined distance indicated by the distance upper limit information; the speaker-side transmitter stops transmitting the distance and the belonging area when the distance exceeds the predetermined distance indicated by the distance upper limit information; 3. The acoustic speaker system according to claim 1 or 2, wherein:

Citation Information

Patent Citations

  • Method of detecting arrangement relation for speaker device in acoustic system, the acoustic system, server device, and speaker device

    JP2005198249A