Sound system with dynamically adjustable target listening point and elimination of environmental object interference
By combining sensor circuits and control circuits, the listening point of the audio system is dynamically adjusted and environmental interference is compensated, solving the problems of fixed position and interference from environmental objects in traditional audio systems, and achieving flexible optimization of the listening effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-18
- Publication Date
- 2026-04-10
AI Technical Summary
Traditional audio systems cannot dynamically adjust to the optimal listening point, requiring users to settle into fixed positions, and environmental interference can lead to poor listening quality.
The system employs a sound system that includes sensor circuitry, speakers, and a main unit. The sensor circuitry captures sound field environment information, identifies the user's position, and dynamically adjusts the listening point. The control circuitry performs channel basis compensation and object basis compensation to counteract environmental object interference and optimize audio output.
This system enables the audio system to dynamically adjust the listening point based on the user's location, eliminates interference from environmental objects, optimizes the listening experience, and enhances the flexibility and sound quality of the audio system.
Smart Images

Figure CN116261094B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to audio processing technology, which is actually a kind of sound system that can dynamically adjust the playing effect according to the changes in the sound field space. BACKGROUND
[0002] The existing sound system includes a plurality of speakers arranged around a target space to form a surround sound field environment. Each speaker can output corresponding channel audio. When configuring the surround sound field environment, the installer of the sound system usually assigns a central area of the target space as the best listening point as the basis for installing a plurality of speakers. When a plurality of speakers simultaneously play a plurality of channel audios, a user located at the best listening point can obtain an immersive listening effect.
[0003] However, in a real environment, the listening effect of the user is easily affected by various variables. For example, in a traditional sound system, the range of the best listening point is regionally limited. When the user moves to an area outside the best listening point, although the plurality of channel audios output by the sound system can still be heard, the listening effect of the plurality of channel audios at the user's location may have been greatly compromised or completely disabled. In addition, the room layout, furniture position and material in the target space are all environmental objects that can interfere with the listening effect. For example, sofas, windows, and curtains can absorb or reflect part of the sound energy, distorting the channel audio received at the best listening point.
[0004] In other words, the traditional sound system cannot dynamically adjust the position of the best listening point, and the user is forced to limit movement to accommodate the position of the best listening point, which is indeed inconvenient. On the other hand, the channel audio can be distorted by environmental objects, making the range of the best listening point more limited or even disappearing. In this way, the sound field environment built at high cost loses its meaning. SUMMARY
[0005] Therefore, how to make the sound system dynamically adjust the best listening point as the user moves and eliminate the interference of environmental objects in the target space is a problem to be solved.
[0006] The present specification provides an embodiment of a sound system, which can dynamically optimize the playing effect according to the user position, wherein the sound system comprises a sensor circuit, a first speaker and a second speaker, and a host device. The sensor circuit is configured to dynamically sense a target space to generate a sound field environment information. The first speaker and the second speaker are configured to play audio. The host device is coupled to the sensor circuit, the first speaker and the second speaker, and comprises an identification circuit, a control circuit and an audio transmission circuit. The identification circuit is configured to identify a user from the sound field environment information and determine a user position of the user in the target space. The control circuit is coupled to the identification circuit and configured to dynamically assign the user position as a target listening point. The audio transmission circuit is coupled to the control circuit, the first speaker and the second speaker, and configured to transmit audio. The sensor circuit comprises a camera configured to capture a sound field environment image of the target space. The identification circuit analyzes the sound field environment image to obtain spatial configuration information and acoustic attribute information of an environmental object in the target space. The control circuit performs channel base compensation operation according to the target listening point, and the spatial configuration information and the acoustic attribute information of the environmental object to generate a first channel audio and a second channel audio optimized for the target listening point. Finally, the control circuit outputs the first channel audio and the second channel audio to the corresponding first speaker and the second speaker through the audio transmission circuit.
[0007] The present disclosure provides an embodiment of a sound system, which can dynamically optimize the playing effect according to the user position, wherein the sound system comprises a sensor circuit, a first speaker and a second speaker, and a host device. The sensor circuit is configured to dynamically sense a target space to generate a sound field environment information. The first speaker and the second speaker are configured to play audio. The host device is coupled to the sensor circuit, the first speaker and the second speaker, and comprises an identification circuit, a control circuit and an audio transmission circuit. The identification circuit is configured to identify a user from the sound field environment information and determine a user position of the user in the target space. The control circuit is coupled to the identification circuit and configured to dynamically assign the user position as a target listening point. The audio transmission circuit is coupled to the control circuit, the first speaker and the second speaker, and configured to transmit audio. The sensor circuit comprises a camera configured to capture a sound field environment image of the target space. The identification circuit analyzes the sound field environment image to obtain spatial configuration information and acoustic attribute information of an environmental object in the target space. The control circuit corresponds the target space to an object base space, and correspondingly establishes a compensation sound source object in the object base space according to the environmental object. The relay data of the compensation sound source object comprises coordinate position, size, and reflectivity and absorptivity of the environmental object. The control circuit performs an object base compensation operation according to the target listening point and the relay data, to offset the interference of the environmental object to the target listening point and generate a first channel audio and a second channel audio optimized for the target listening point. The control circuit outputs the first channel audio and the second channel audio to the corresponding first speaker and second speaker through the audio transmission circuit.
[0008] The present specification provides an embodiment of a sound system, which can dynamically optimize the playing effect according to the user position, wherein the sound system comprises a sensor circuit, a first speaker and a second speaker, and a host device. The sensor circuit is configured to dynamically sense a target space to generate a sound field environment information. The first speaker and the second speaker are configured to play audio. The host device is coupled to the sensor circuit, the first speaker and the second speaker, and comprises an identification circuit, a control circuit, an audio transmission circuit and a human-computer interface circuit. The identification circuit is configured to identify a user from the sound field environment information and determine a user position of the user in the target space. The control circuit is coupled to the identification circuit and configured to dynamically assign the user position as a target listening point. The audio transmission circuit is coupled to the control circuit, the first speaker and the second speaker, and configured to transmit audio. The human-computer interface circuit is coupled to the control circuit and configured to run a configuration program to obtain spatial configuration information and acoustic attribute information of an environmental object in the target space. The control circuit performs a channel base compensation operation according to the target listening point, the spatial configuration information and the acoustic attribute information of the environmental object to generate a first channel audio and a second channel audio optimized for the target listening point. The control circuit outputs the first channel audio and the second channel audio to the corresponding first speaker and second speaker through the audio transmission circuit.
[0009] The present specification provides an embodiment of a sound system that dynamically optimizes playback based on a user's position. The sound system includes a sensor circuit, a first speaker and a second speaker, and a host device. The sensor circuit is configured to dynamically sense a target space to generate a sound field environment information. The first speaker and the second speaker are configured to play audio. The host device, coupled to the sensor circuit, the first speaker, and the second speaker, includes an identification circuit, a control circuit, an audio transmission circuit, and a human-machine interface circuit. The identification circuit is configured to identify a user from the sound field environment information and determine a user's position in the target space. The control circuit, coupled to the identification circuit, is configured to dynamically assign the user's position as a target listening point. The audio transmission circuit, coupled to the control circuit, the first speaker, and the second speaker, is configured to transmit audio. The human-machine interface circuit, coupled to the control circuit, is configured to run a configuration program to obtain spatial configuration information and acoustic property information of an environmental object in the target space. The control circuit maps the target space to an object base space and correspondingly establishes a compensation sound source object in the object base space based on the environmental object. A relay data of the compensation sound source object includes coordinate position, size, and reflectivity and absorptivity of sound of the environmental object. The control circuit performs an object base compensation operation based on the target listening point and the relay data to offset the interference of the environmental object to the target listening point and generate a first channel audio and a second channel audio optimized for the target listening point. The control circuit outputs the first channel audio and the second channel audio to the first speaker and the second speaker, respectively, through the audio transmission circuit.
[0010] One advantage of the above embodiment is that the sound system can dynamically track the user's position through the sensor and continuously optimize the playback for the user's position. The user does not need to compromise the fixed listening position to obtain the best experience.
[0011] Another advantage of the above embodiment is that the sound system can identify the environmental object in the target space and adjust the channel audio to offset the interference of the environmental object.
[0012] Other advantages of the present invention will be more fully understood from the following detailed description taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0013] Figure 1 A functional block diagram of a sound system according to an embodiment of the present invention.
[0014] Figure 2 A flowchart of a dynamic sound effect optimization method according to an embodiment of the present invention.
[0015] Figure 3A flowchart of a dynamic sound effect optimization method according to an embodiment of the present application.
[0016] Figure 4 A flowchart of a dynamic sound effect optimization method according to an embodiment of the present application.
[0017] Figure 5 A flowchart of a dynamic sound effect optimization method according to an embodiment of the present application.
[0018] Figure 6 A diagram of a target space according to an embodiment of the present application, for illustrating an embodiment of calculating an audio adjustment amount according to a position of a best listening point.
[0019] Figure 7 A diagram of a target space according to an embodiment of the present application, for illustrating an embodiment of calculating an audio adjustment amount according to an absorption rate of an environmental object.
[0020] Figure 8 A diagram of a target space according to an embodiment of the present application, for illustrating an embodiment of calculating an audio adjustment amount according to a reflectivity of an environmental object.
[0021] Figure 9 A flowchart of an object identification operation of a host device according to an embodiment of the present application.
[0022] Figure 10 A flowchart of an audio processing method according to an embodiment of the present application, for illustrating an embodiment of calculating an output compensation value according to a positional relationship of an environmental object.
[0023] Figure 11 A diagram of a target space according to an embodiment of the present application, for illustrating an embodiment of optimizing a sound field by an object base compensation operation.
[0024] Figure 12 A diagram of a target space according to an embodiment of the present application, for illustrating an embodiment of optimizing a sound field by an object base compensation operation.
[0025] Figure 13 A flowchart of an object base compensation operation according to an embodiment of the present application.
[0026] Symbol explanation
[0027] 100... sound system
[0028] 110... first speaker
[0029] 112... first channel audio
[0030] 120... second speaker
[0031] 122... second channel audio
[0032] 130... host device
[0033] 131... storage circuit
[0034] 132... control circuit
[0035] 133... human-machine interface circuit
[0036] 134... recognition circuit
[0037] 135... audio transmission circuit
[0038] 136... communication circuit
[0039] 140... sensor circuit
[0040] 150... user device
[0041] 160... remote database
[0042] 170... target space
[0043] 171... first position
[0044] 172... second position
[0045] 173... movement trajectory
[0046] 175... environmental object
[0047] 180... user
[0048] 202-218... processes
[0049] 312-316... processes
[0050] 410... process
[0051] 600... target space
[0052] 601... first position
[0053] 602... second position
[0054] 610... camera
[0055] 620... infrared sensor
[0056] 630... wireless detector
[0057] 700... target space
[0058] 800... target space
[0059] 902-910... processes
[0060] 1002-1012... processes
[0061] 1100... target space
[0062] 1103... object movement trajectory
[0063] 1105... virtual sound source object
[0064] 1110... first speaker
[0065] 1120... second speaker
[0066] 1130... third speaker
[0067] 1140... fourth speaker
[0068] P0... origin
[0069] P1... first position
[0070] P1 '... new first position
[0071] P2... second position
[0072] 1200... target space
[0073] 1201... target listening point
[0074] 1203... movement trajectory
[0075] 1210... first speaker
[0076] 1212... first channel output
[0077] 1220... second speaker
[0078] 1222... second channel output
[0079] 1230... third speaker
[0080] 1240... fourth speaker
[0081] 1250... fifth speaker
[0082] 1252... fifth channel output
[0083] 1260... sixth speaker
[0084] 1262... sixth channel output
[0085] 1304-1312... flow DETAILED DESCRIPTION
[0086] Embodiments of the present application will be described below with reference to the accompanying drawings. In the drawings, the same numbers represent the same or similar elements or method steps throughout.
[0087] Figure 1 A functional block diagram of a sound system 100 according to an embodiment of the present application.
[0088] The sound system 100 mainly comprises a host device 130 and a plurality of speakers. The host device 130 can control the plurality of speakers to play audio. The host device 130 can be a computer host, a stereo system, an embedded system, or a customized digital audio processing device. The host device 130 comprises a communication circuit 136, so that the host device 130 can be connected to a user device 150 via wired or wireless connection, and serve as an input channel for audio source signals or data.
[0089] The user device 150 can be a mobile phone, a computer, a TV stick, a game console, or other audio source providing devices, and provide music or sound stream to the host device 130 via the communication circuit 136. Further, the sound system 100 can use the communication circuit 136 to cooperate with the user device 150 or other multimedia devices, and form a home theater system with both video and audio functions. For example, the target space 170 can further comprise a projection screen, a screen, or a display (not shown), and display images under the control of the user device 150. For another example, the user device 150 can be a head-mounted virtual reality device. The user 180 can stand in the target space 170 and see images through the user device 150, and the host device 130 can be controlled by the user device 150 to play audio in synchronization with the images. The communication circuit 136 in the present embodiment can be (but not limited to) a High Definition Multimedia Interface (HDMI), a Sony / Philips Digital Interface Format (SPDIF), a wireless area network module, an Ethernet network module, a shortwave radio transceiver, or a Bluetooth Low Energy (BLE) version 4 or 5 evolution application, or a Universal Serial Bus (USB).
[0090] The host device 130 also includes an audio transmission circuit 135 for connecting multiple speakers and outputting multiple channels of audio to the speakers. The host device 130 can control the speakers through the audio transmission circuit 135 in a unidirectional digital or analog output, or a bidirectional synchronous communication protocol. The connection between the audio transmission circuit 135 and each speaker can be a wired interface, a wireless interface, or a combination of both. The wired interface can be, but is not limited to, a composite video and audio terminal, a digital transmission interface, or a high-definition multimedia interface. The wireless interface can be, but is not limited to, a wireless area network, a shortwave radio frequency transceiver, or a Bluetooth Low Energy version 4 or 5. In further embodiments, the audio transmission circuit 135 and the communication circuit 136 can be combined into a multifunctional bidirectional transmission interface module, as both circuits are interfaces for connecting external elements. The audio transmission circuit 135 and the communication circuit 136 can use various publicly available standard transmission technologies to connect and transmit between elements, which can increase the future expandability of the sound system 100 and reduce the replacement cost when an element is damaged.
[0091] Figure 1 The target space 170 in the sound system 100 can be understood as a three-dimensional space in which the user 180 can use the sound system 100. Each speaker can be configured at a different position in the target space 170 to play a channel of audio. The surround configuration of multiple speakers can create a surround sound environment in a target space 170. There are various standard specifications for the number and configuration of speakers. For example, in a 5.1 channel surround sound system, there are two front speakers, one center speaker, two surround channel speakers, and one subwoofer to create a surround sound space around a target listening point and play sound to the target listening point. In a 7.1 channel surround sound system, a pair of rear surround channel speakers are further configured behind the target listening point to provide a more three-dimensional sound field effect. In recent years, new specifications such as 5.1.2 channels and 7.2.2 channels have appeared, which include more speakers and specific directional channel configurations to achieve more realistic "panoramic sound", "sky sound effect", or "floor sound effect". For the convenience of explaining the technical features of the sound system 100 of the present embodiment, Figure 1Only the first speaker 110 and the second speaker 120 are shown for representation. The first speaker 110 receives and plays the first channel audio 112 provided by the host device 130, while the second speaker 120 receives and plays the second channel audio 122 provided by the host device 130. It must be understood that in practice, the sound system 100 of the present embodiment is not limited to only two speakers, but can be applied to 2.1 channel, 4.1 channel, 5.1 channel, 7.2 channel, or more channel configurations. Each speaker in the target space 170 can have different audio output specifications. For example, some speakers are good at outputting bass, while some speakers are good at outputting mid-high frequencies. The host device 130 can plan different characteristics of the sound field environment in the target space 170 according to different speaker specifications.
[0092] The term "channel" as referred to in the specification and claims refers to both physical channels and logical channels. A logical channel refers to an audio data stream transmitted within the system, while a physical channel refers to the source of the signal played by each speaker. In the present embodiment, the first channel audio 112 and the second channel audio 122 played by each speaker are physical channels, which can be the result of one or more logical channels being down-mixed. For example, a pair of headphones can have only two speakers, but can be able to hear the sound effects produced by multiple applications simultaneously. In other words, the sound effect data of multiple applications can be down-mixed by the system into two physical channels, and played as audible sound through the two speakers. Therefore, the first channel audio 112 and the second channel audio 122 in the present embodiment are not limited to audio signals containing only a single logical channel, but can also be audio signals mixed from multiple logical channels according to a predetermined ratio.
[0093] In Figure 1 The first speaker 110 and the second speaker 120 are configured on two sides of a target space 170 to play sound for a target listening point in the target space 170. The target listening point can be understood as a position in which the sound system 100 has the best playing effect. In some sound systems, the target listening point is also referred to as a listening sweet spot. In most cases, the target listening point is usually located in a specific area of the target space 170, such as a center point, an axis, a tangent plane, or an equivalent volume center of multiple speakers. In Figure 1In the target space 170, a first position 171 where the user 180 is located is used to represent a target listening point of the target space 170. When the user 180 moves from the first position 171 to a second position 172 along a moving trajectory 173, the listening effect received by the user 180 is deviated because the user 180 is away from the first speaker 110 and approaches the second speaker 120. The conventional sound system cannot track the movement of the user 180 and adjust the listening effect received at the second position 172 correspondingly. The solution proposed in the embodiment will be described later.
[0094] On the other hand, the target space 170 usually contains some environmental objects 175, such as sofas, tables, curtains, walls, ceilings, and floors. These environmental objects 175 will have different interference reactions to the sound played by the first speaker 110 and the second speaker 120 due to different materials, sizes, and positions. For example, a sofa or a curtain made of cloth will absorb sound, and a marble floor or wall will reflect sound. In other words, the presence of the environmental objects 175 will affect the first channel audio 112 and the second channel audio 122 received at the target listening point. The conventional sound system does not have the ability to identify the environmental objects 175 in the target space 170, nor does it have the function of compensating for the first channel audio 112 and the second channel audio 122 according to the size, material, and position of the environmental objects 175. The sound system 100 of the embodiment can calculate and eliminate the interference of all the environmental objects 175 in the target space 170 to the first channel audio 112 and the second channel audio 122. For the convenience of explanation, the sound system 100 of the embodiment is described below by taking only one environmental object 175 to explain the operation mode of the sound system 100. However, it must be understood that Figure 1 In the target space 170, a first position 171 where the user 180 is located is used to represent a target listening point of the target space 170. When the user 180 moves from the first position 171 to a second position 172 along a moving trajectory 173, the listening effect received by the user 180 is deviated because the user 180 is away from the first speaker 110 and approaches the second speaker 120. The conventional sound system cannot track the movement of the user 180 and adjust the listening effect received at the second position 172 correspondingly. The solution proposed in the embodiment will be described later. Figure 1 It must be understood that the target space 170 is not limited to contain only one environmental object 175. The solution to the interference of the environmental object 175 will be described later.
[0095] The host device 130 of the embodiment further includes a storage circuit 131. The storage circuit 131 can include a non-volatile memory for storing a related operating system, application software, or firmware required for the operation of the host device 130. The storage circuit 131 can also include a volatile memory for use as an operation memory of the control circuit 132. The host device 130 of the embodiment further includes a control circuit 132. The control circuit 132 can be a central processing unit, a digital signal processor, or a microcontroller. The control circuit 132 can read the pre-stored operating system, software, or firmware from the storage circuit 131 to control the host device 130, the first speaker 110, and the second speaker 120 to perform the audio playing operation. Further, the host device 130 of the embodiment uses the control circuit 132 to perform a series of sound field compensation operations to dynamically optimize the playing effect and solve the shortcomings that the conventional sound system cannot overcome.
[0096] To dynamically optimize playback at the target listening point, the audio system 100 of this embodiment includes a sensor circuit 140 configured to dynamically sense a target space 170 and generate sound field environment information. The sensor circuit 140 may be a component located outside the host device 130 and coupled to it. The sensor circuit 140 may be a combination of one or more of a camera 610, an infrared sensor 620, and a wireless detector 630. The form of the sound field environment information captured by the sensor circuit 140 may vary depending on how the sensor circuit 140 is implemented. For example, the sound field environment information may be a combination of one or more of images, pictures, thermal images, and radio wave imaging of the user and environmental objects. In one embodiment, the sensor circuit 140 is disposed around the target space 170. It is understood that although... Figure 1 Only one sensor circuit 140 is shown in the figure, but in practice, the audio system 100 may include multiple sensor circuits 140, which are respectively configured at different positions around the target space 170 to obtain more accurate sound field environment information.
[0097] In the host device 130 of this embodiment, an identification circuit 134 is included, coupled to the sensor circuit 140. The identification circuit 134 can identify key information affecting the sound field from the sound field environment information, enabling the control circuit 132 to dynamically adjust the first channel audio 112 and the second channel audio 122 played from the first speaker 110 and the second speaker 120. For example, the identification circuit 134 can identify a user from the sound field environment information and determine the user's position in the target space. Since the sound field environment information provided by the sensor circuit 140 can have various combinations, the identification circuit 134 can also implement different identification technology solutions accordingly. For example, when the sound field environment information is an image, the identification circuit 134 can use artificial intelligence recognition technology to distinguish the user in the image. Through the application of artificial intelligence, after analyzing the user in the image, the identification circuit 134 can further locate the user's head, face, and even ear positions. If the sensor circuit 140 can provide diverse information such as three-dimensional images with spatial depth, infrared thermal imaging, or wireless signals, it will help the recognition circuit 134 obtain more accurate recognition results.
[0098] To calculate the degree of interference caused by the environmental object 175 to the sound field environment, the host device 130 needs the spatial configuration information and the acoustic property information of the environmental object 175. The spatial configuration information can include the size, position, shape, and various external features of the environmental object 175. The acoustic property information can include the absorption rate, reflectivity, and resonance frequency of the material related features of the sound. In an embodiment, the identification circuit 134 can further identify the spatial configuration information of the environmental object 175 in the target space 170 from the sound field environment information when identifying the sound field environment information, and find the acoustic property information. To identify the environmental object, an object database is needed. In an embodiment, the storage circuit 131 in the host device 130 can also be used to store an object database. The object database can include various external feature information for identifying the environmental object, and various acoustic property information corresponding to each environmental object. For example, when the host device 130 needs to calculate the degree of interference caused by an environmental object 175 to the sound field environment, the identification circuit 134 can first analyze the object name of the environmental object 175, and then the host device 130 reads the storage circuit 131 to find the absorption rate and reflectivity corresponding to the environmental object 175.
[0099] In practice, the identification circuit 134 can be a custom processor chip that executes the artificial intelligence identification function in combination with the existing operating system, software or firmware in the storage circuit 131. The identification circuit 134 can also be one of the cores or execution thread circuits of the control circuit 132, which executes the existing artificial intelligence software product in the storage circuit 131 to achieve the identification function. The identification circuit 134 can also be a memory module of a specific artificial intelligence software product, which is executed by the control circuit 132 to complete the identification function.
[0100] The human interface circuit 133 in the host device 130 can be used by the user to control the operation of the host device 130. The human interface circuit 133 can include a display screen, buttons, a dial, or a touch screen, and can be used by the user to perform basic audio system 100 control functions, such as adjusting the volume, playing, and fast forwarding or rewinding. In one embodiment, the control circuit 132 can also execute a configuration program through the human interface circuit 133 to allow the user to set various sound field scenarios, or to inform the host device 130 of the spatial configuration information of the environmental objects 175 in the target space 170. For example, in the configuration program, the control circuit 132 receives object configuration data input by the user through the human interface circuit 133, such as the object name, type, size, and location of one or more environmental objects 175. After the control circuit 132 obtains the spatial configuration information, it looks up the corresponding absorption and reflection rates from the object database stored in the storage circuit 131, so as to perform subsequent sound field compensation operations. In a further embodiment, the human interface circuit 133 can also be provided by the user device 150. The user can operate the configuration program using the user device 150, and the user device 150 transmits the setting results to the control circuit 132 through the communication circuit 136.
[0101] The host device 130 can also be connected to a remote database 160 through the communication circuit 136. In a further embodiment, the object database originally stored in the storage circuit 131 can also be stored in the remote database 160. When the host device 130 needs to calculate the degree of disturbance caused by an environmental object 175 to the sound field environment, it can first analyze the sound field environment information through the recognition circuit 134 to obtain an object characteristic value, and then access the remote database 160 through the communication circuit 136 to find an environmental object 175 that matches the object characteristic value, and obtain the sound field attribute information of the environmental object 175. The remote database 160 can be a server located in the cloud or other system, and is connected to the host device 130 through wired or wireless bidirectional network communication technology. The remote database 160 can not only provide a lookup function, but also accept the uploading of update data to continuously expand the database content. For example, the host device 130 can communicate with the remote database 160 using structured query language (SQL).
[0102] Based on Figure 1Based on the system architecture, the audio system 100 proposed in this application can achieve at least the following technical effects. First, the audio system 100 can dynamically track the user's position as the target listening point. The audio system 100 can also dynamically acquire spatial configuration information of environmental objects as a basis for optimizing the sound field effect. Finally, the audio system 100 dynamically compensates the speaker output based on the user's position and the spatial configuration information of environmental objects to eliminate object interference and optimize the listening effect at the target listening point. The implementation of dynamically tracking the user's position can employ various technical solutions such as cameras, infrared sensors, or wireless positioning. The implementation of acquiring spatial configuration information of environmental objects can be automatic or manual. For example, the audio system 100 can use a camera to capture images and perform artificial intelligence recognition, or allow the user to manually input the environmental conditions through a configuration program. The implementation of compensating the speaker output can be based on several different algorithms. For example, this specification introduces the Channel Base algorithm and the Object Base algorithm.
[0103] The following is Figure 2 This describes an embodiment of an audio system 100 that dynamically tracks the user's position, uses a camera to acquire the sound field environment configuration, and compensates the speaker output using channel-based compensation.
[0104] Figure 2 This is a flowchart of a dynamic sound effect optimization method according to an embodiment of the present invention.
[0105] exist Figure 2 In the flowchart, the process located in the field belonging to a specific device represents the process performed by that specific device. For example, the part marked in the "Sensor Circuit" field is the process performed by sensor circuit 140; the part marked in the "Host Device" field is the process performed by host device 130; the part marked in the "Speaker" field is the process performed by first speaker 110 and / or second speaker 120; and so on. The aforementioned logic also applies to other subsequent flowcharts.
[0106] In block 202, the sound field environment information is dynamically sensed by the sensor circuit 140. In one embodiment, the sound field environment information can be optical, thermal, or electromagnetic wave information in the target space 170. For example, the sensor circuit 140 can include a video camera to continuously record video of the target space 170 or periodically capture still images of the target space 170. In another embodiment, the sensor circuit 140 can also include an infrared sensor configured to capture thermal imaging data of the target space. The thermal imaging data generated by the infrared sensor is extremely sensitive to temperature changes and includes information about the depth of the space, and thus is particularly suitable for tracking the location of a user. In another embodiment, the sensor circuit 140 can also include a wireless detector configured to detect wireless signals of an electronic device in the target space. When a user is holding an electronic device, the wireless detector can detect the beacon time difference or wireless signal strength of the electronic device as an auxiliary means to track the location of the user. The electronic device can be the user's own cell phone, a specially designed beacon generator, a head-mounted virtual reality device, a game controller, or a remote control of the sound system 100. It is understood that the number of sensor circuits 140 is not limited in the present embodiment, nor is the use of only one type of sensing scheme at a time. For example, the sound system 100 of the present embodiment can use multiple sensor circuits 140 to operate cooperatively from different locations, or use one or more video cameras, infrared sensors, and wireless detectors simultaneously. In this way, the host device 130 can obtain more complete sound field environment information and achieve more accurate recognition results in subsequent procedures.
[0107] In block 204, the sensor circuit 140 transmits the sensed sound field environment information to the host device 130. The sensor circuit 140 can transmit data continuously, such as video, or periodically return static data. The frequency of data transmission by the sensor circuit 140 can be adaptively determined according to the amount of information of the sound field environment information, the tracking accuracy requirement, and the computing power of the host device 130. The sensor circuit 140 and the host device 130 can be connected by a dedicated line or through the communication circuit 136. In a further derived embodiment, the sensor circuit 140 can share the audio transmission circuit 135 with the speakers, so as to transmit the sound field environment information to the host device 130 through the audio transmission circuit 135.
[0108] In procedure 206, the host device 130 determines the user position according to the sound field environment information received from the sensor circuit 140. The identification circuit 134 in the host device 130 can perform an identification process on the sound field environment information, for example, by applying artificial intelligence. The identification algorithm of the identification circuit 134 varies according to the sensing scheme of the sensor circuit 140. It is understood that the target space 170 and the user position can be represented in a two-dimensional space or a three-dimensional space. If only a single sensor circuit 140 is implemented in the sound system 100, at least the position information in a two-dimensional space can be sensed. If the number of sensor circuits 140 is increased or a multi-sensor sensing scheme is implemented in the sound system 100, the depth information in a three-dimensional space can be obtained to more accurately determine the user position or the user head position. In an embodiment, the identification circuit 134 can dynamically identify the user head position, the face orientation, or the ear position according to the sound field environment image captured by the camera. In another embodiment, the identification circuit 134 can analyze the moving track of the thermal imaging data generated by the infrared sensor to dynamically determine the position of the user 180. For example, the identification circuit 134 can dynamically locate a coordinate value of the electronic device in the target space 170 according to the characteristics of the wireless signal detected by the wireless detector. With the coordinate value, the control circuit 132 can further infer the user ear position.
[0109] In procedure 208, after the identification circuit 134 in the host device 130 analyzes the user position, the control circuit 132 in the host device 130 dynamically assigns the user position as the target listening point. For the convenience of describing the subsequent embodiments, the target space 170 is described as a two-dimensional coordinate space or a three-dimensional coordinate space, and the target listening point can be represented as a coordinate value in the target space 170. With different layouts of the plurality of speakers, the range of the target listening point can be more than a single point, can be a plane, or can be a three-dimensional region with length, width, and height. For example, after the identification circuit 134 analyzes the user head position or the ear position, the control circuit 132 can assign the user head position or the ear position as the target listening point. The control circuit 132 can compensate the target listening point to obtain a playback effect that is not affected by the user movement through subsequent compensation operations. In implementation, the control circuit 132 compensates the target listening point to obtain a listening effect by adjusting the first channel audio 112 and the second channel audio 122. It is understood that procedure 208 can be dynamically performed as the user position changes. Therefore, procedure 208 is not limited to be performed in the order shown. In other words, the target listening point can be updated in real time as the user position changes. The specific adjustment algorithm will be described later. Figure 2
[0110] In the process 210, the identification circuit 134 in the host device 130 further identifies the sound field environment information provided by the sensor circuit 140 to obtain the spatial configuration information of the environmental objects in the target space 170. In other words, the sound field environment information provided by the sensor circuit 140 can be used not only to determine the user's position, but also to determine various environmental objects 175 present in the target space 170. In an embodiment, after the camera in the sensor circuit 140 captures a sound field environment image of the target space 170, the identification circuit 134 analyzes the sound field environment image to identify one or more environmental objects 175 in the target space 170 and the spatial configuration information of the environmental objects 175. The spatial configuration information includes the size, position, shape, and appearance characteristics of the environmental objects 175. The identification circuit 134 can also determine the acoustic property information of each environmental object 175, such as the sound absorption and reflection rates, through artificial intelligence algorithms or database searches. In further derived embodiments, the identification circuit 134 can also determine the application scenario category of the target space 170 according to the sound field environment image. The application scenario category can include a theater, a living room, a bathroom, an outdoor space, etc. If the host device 130 knows the application scenario category of the target space 170, it can more quickly identify the environmental objects 175 in the target space 170 and reduce misjudgments. Related embodiments will be described in detail in the Figure 9 section.
[0111] In process 212, the control circuit 132 in the host device 130 can calculate the degree to which the playback effect of a speaker on a target listening point is affected by environmental objects. The playback effect of a speaker on a target listening point can be defined as the equal loudness or sound pressure level (SPL) received at the target listening point from the speaker. An equal loudness curve (Fletcher-Munson Curve) is defined in the ISO 226 standard to illustrate the equal loudness perceived by a user at different sub-bands, which actually corresponds to different sound pressure levels. In an embodiment, the control circuit 132 can use the equal loudness curve as a standard reference basis for the playback effect to calculate the sound pressure level received at the target listening point under various conditions. The control circuit 132 can use the spatial configuration information and acoustic property information of the environmental objects 175 to evaluate the interference caused by the environmental objects 175 to the target listening point, so as to further calculate the method for eliminating the interference. The influence of the spatial configuration information and property information of the environmental objects 175 includes many kinds of situations. For example, the larger the volume of the environmental objects 175, the greater the interference coefficient to the target listening point. Whether the position of the environmental objects 175 blocks the user 180 and the speaker also determines the degree of influence on the speaker. The environmental objects 175 can absorb sound or bounce sound as the material is different. Therefore, the control circuit 132 needs to select corresponding parameters or formulas to calculate the degree of influence on the speaker according to different acoustic properties.
[0112] In process 214, the control circuit 132 in the host device 130 uses a channel-based bass compensation operation to calculate the output compensation value required by each channel audio of each speaker. The channel-based bass compensation operation is calculated separately for each channel audio when judging the playback effect on the target listening point. Taking a first channel audio 112 played by a first speaker 110 in a plurality of speakers as an example, before the first channel audio 112 is transmitted to the target listening point through the air, the energy of the first channel audio 112 can be lost due to the interference of an environmental object 175. The position of the target listening point also affects the sound pressure level of the first channel audio 112 at the target listening point. Through the channel-based bass compensation operation, the control circuit 132 can calculate the change in the sound pressure level of the first channel audio 112 at the target listening point. The control circuit 132 in the present embodiment adds an output compensation value to the first channel audio 112 to offset the change in the sound pressure level, so that the first channel audio 112 received at the target listening point is restored to the state before being affected. In other words, the output compensation value has the same numerical value as the change in the sound pressure level, but has the opposite positive and negative polarity.
[0113] In process 216, control circuit 132 adjusts and outputs channel audio to the speakers based on the output compensation value. Since the adjusted channel audio has offset the effects of the user 180's displacement in the target space 170 and the interference caused by environmental objects 175, the listening effect perceived by the user 180 remains consistent. Taking the first speaker 110 and the second speaker 120 in the target space 170 as examples, control circuit 132 calculates and adjusts the sound pressure values of different sub-bands in the first channel audio 112 and the second channel audio 122, thereby offsetting the equivalent volume deviation perceived by the user 180 due to movement. On the other hand, the control circuit (132) compensates the first channel audio 112 and the second channel audio 122 accordingly based on the change in sound pressure value caused by the position, size, and acoustic properties of the environmental objects 175 at the target listening point.
[0114] In process 218, each speaker receives channel audio from the host device 130 via audio transmission circuit 135. Taking the first speaker 110 and the second speaker 120 in the target space 170 as an example, the control circuit 132 outputs the first channel audio 112 and the second channel audio 122 to the corresponding first speaker 110 and second speaker 120 via audio transmission circuit 135. Thus, the first speaker 110 and the second speaker 120 correspondingly play the adjusted first channel audio 112 and second channel audio 122, providing an optimized listening experience for the user 180's target listening point. For ease of explanation, Figure 1 In the embodiment of target space 170, only two speakers and one environmental object 175 are shown. However, it is understood that in practice, the host device 130 may contain more than two speakers, and the number of environmental objects 175 is not limited to one. In further derived embodiments, each speaker may be good at outputting different audio ranges. For example, some speakers are mid-high frequency speakers, and some are subwoofer speakers. When adjusting the channel audio, the control circuit 132 can further adjust the corresponding output first channel audio 112 and second channel audio 122 according to the characteristics of different speakers.
[0115] The following is Figure 3 This describes an embodiment of an audio system 100 that dynamically tracks the user's position, uses a camera to acquire the sound field environment configuration, and compensates the speaker output using object-based compensation.
[0116] Figure 3 This is a flowchart of a dynamic sound effect optimization method according to an embodiment of the present invention.
[0117] exist Figure 3In the flowchart, the flow marked in the column of a specific device means the flow performed by the specific device. For example, the part marked in the column of "sensor circuit" is the flow performed by the sensor circuit 140; the part marked in the column of "host device" is the flow performed by the host device 130; the part marked in the column of "speaker" is the flow performed by the first speaker 110 and / or the second speaker 120; and the rest is similar. The aforementioned logic is also applicable to the other flowcharts.
[0118] Figure 3 The flows 202, 204, 206, 208, and 210 in the flowchart are the same as those in the previous embodiment, and thus the description thereof is omitted.
[0119] When the sound system 100 of the present embodiment completes the flow 210, the control circuit 132 has tracked the position of the user 180 and assigned the target listening point, and also obtained the spatial configuration information of one or more environmental objects 175 in the target space 170. The object-based compensation operation is described later to adjust the channel audio of each speaker.
[0120] The object-based acoustic system is originated from the mixing technology of virtual reality, and can simulate the effect of moving sound source objects by using a limited number of physical speakers. Some existing software products, such as Dolby Atmos, Spatial Audio Workstation, or D-Spatial Reality, are object-based acoustic systems. A user can define the moving track of a sound source object in a virtual space through a human-machine interface. The object-based system can simulate the sound effect of the sound source object in the virtual space by using physical speakers. A user located at the target listening point can thus experience the moving sound source object in the space in reality.
[0121] The object-based acoustic system is based on the array operation of a large number of acoustic parameters. Each sound source object has a relay data for describing the type, position, size (length, width, and height), divergence, etc. of the sound source object. After the array operation of the object-based system, the sound represented by a sound source object is assigned to one or more speakers for playing together, and each speaker plays a part of the sound of the sound source object. In other words, the array operation of the object-based system can simulate the spatial effect of a sound source object by using multiple speakers. Figure 3 The embodiments of the present application propose an object-based compensation operation based on the object-based acoustic system to solve the problem of the conventional playing effect.
[0122] In process 312, the control circuit 132 in the host device 130 establishes a compensating sound object of the object basis in accordance with the environmental object 175. In implementation, the control circuit 132 first maps the target space 170 to an object basis space in virtual reality, and then establishes a compensating sound object in the object basis space in accordance with the environmental object 175, for generating a sound source effect that counteracts the environmental object 175. To the user 180 located at the target listening point, the environmental object 175 can also be simulated as a sound object. In actual application, the environmental object 175 can reflect the sound emitted by a loudspeaker to the target listening point. The environmental object 175 can also block or absorb a portion of the sound, so that the sound emitted by a loudspeaker to the target listening point is attenuated. In other words, after the control circuit 132 in the embodiment simulates the environmental object 175 as a sound object, it can correspondingly establish a negative sound object in the object basis space with an opposite sound source effect, as a means of counteracting the interference. In the sound source effect described in the embodiment, it can be the sound pressure value, the equivalent sound volume, or the gain value generated for the target listening point.
[0123] In process 314, the host device 130 substitutes the compensating sound object into the object basis compensation operation to generate the channel audio. The object basis compensation operation can use the object basis array operation module in the existing object basis acoustic product, and perform a large number of array operations related to acoustic interaction in accordance with the relay data of the sound object. For example, the relay data of the compensating sound object includes the coordinate position, size, and reflectivity and absorptivity of the environmental object 175. The control circuit 132 performs an object basis compensation operation in accordance with the target listening point and the relay data, to counteract the interference of the environmental object 175 to the target listening point and generate the first channel audio 112 and the second channel audio 122 optimized for the target listening point.
[0124] In an embodiment, the object basis compensation operation is performed on multiple sub-bands respectively. Due to the characteristics of sound transmission, the sound pressure value on each sub-band has different effects on the equivalent sound volume. Taking the influence of the first channel audio 112 generated by the first loudspeaker 110 on the environmental object 175 as an example, the control circuit 132 in the embodiment can calculate a sound source effect passively generated by the environmental object 175 affected by the first channel audio 112 on multiple sub-bands respectively in accordance with the coordinate position, size, and reflectivity and absorptivity of the environmental object 175. Then the control circuit 132 establishes the compensating sound object in accordance with the sound source effect. In the embodiment, the compensating sound object is established in accordance with the environmental object 175, and the relay data has the same coordinate position, size, and reflectivity and absorptivity of the environmental object 175, but the sign of the generated sound source effect is opposite to that of the environmental object 175.
[0125] It is known that the audible range of human ear is between 20 Hertz (Hz) and 20,000 Hz. The present embodiment can cut the audible range of human ear into multiple sub-band intervals and compensate them respectively. The interval size of each sub-band can be exponential interval. For example, exponential interval with base 10 can divide the sound frequency signal into 10 Hz to 100 Hz, 100 Hz to 1000 Hz, 1000 Hz to 10,000 Hz, and so on. In other embodiments, exponential interval can be divided with base 2 or base 4 according to the requirement of the precision of the playback quality. There already exists the processing technology of cutting multiple sub-bands in the equalizer in the field of audio processing, which will not be explained in depth here.
[0126] After the control circuit 132 obtains the negative sound source effect of the compensation sound source object, an object base compensation operation is performed to mix the negative sound source effect into the first channel audio 112 and the second channel audio 122 according to the proportion determined by the mixing operation result, so as to offset the interference of the environmental object 175 to the target listening point. Regarding the object base compensation operation, it will be described in detail in the embodiments of Figures 11 to 13 .
[0127] In the process 316, the host device 130 outputs the first channel audio 112 and the second channel audio 122 to the first speaker 110 and the second speaker 120 respectively according to the operation result of the process 314. Figure 3 The process 316 of the present embodiment is different from the process 216 of the Figure 2 embodiment. Figure 2 The compensation value is calculated for the existing channel audio to adjust the existing channel audio. When performing the object base compensation operation, the control circuit 132 directly calculates the corresponding channel audio of each speaker at one time according to all the relay data. The object base compensation operation mixes the interference component to be offset or compensated into the channel audio in the form of compensation sound source. In other words, because the channel audio contains the compensation sound source emitted by the compensation sound source object, the user 180 cannot feel the influence caused by the existence of the environmental object 175 at the target listening point.
[0128] As can be seen from the process 316, the object base compensation operation translates the target listening point and the environmental object into the relay data of the object base acoustic system and establishes the compensation sound source object, which simplifies the operation process of eliminating interference and optimizing the playback effect. It needs to be understood that the sound system 100 of the present embodiment can use the sensor circuit 140 to track the position of the user 180 in real time or periodically to dynamically update the target listening point. The object base compensation operation performed by the control circuit 132 can also be updated synchronously with the change of the target listening point and all the relay data related to the relative position of the target listening point in the target space 170.
[0129] Figure 3 The process 218 in this embodiment is the same as the previous embodiment, and will not be repeated here to save space.
[0130] The following is Figure 4 This describes an embodiment in which an audio system 100 dynamically tracks the user's position, runs a configuration program to obtain the sound field environment configuration, and compensates the speaker output using channel-based compensation.
[0131] Figure 4 This is a flowchart of a dynamic sound effect optimization method according to an embodiment of the present invention.
[0132] exist Figure 4 In the flowchart, the process located in the field belonging to a specific device represents the process performed by that specific device. For example, the part marked in the "Sensor Circuit" field is the process performed by sensor circuit 140; the part marked in the "Host Device" field is the process performed by host device 130; the part marked in the "Speaker" field is the process performed by first speaker 110 and / or second speaker 120; and so on. The aforementioned logic also applies to other subsequent flowcharts.
[0133] Figure 4 The processes 202, 204, 206 and 208 are the same as in the previous embodiment, and will not be repeated here to save space.
[0134] In this embodiment, when the audio system 100 completes process 210, the control circuit 132 has tracked the position of the user 180 and assigned it as the target listening point, and has also obtained the spatial configuration information of one or more environmental objects 175 in the target space 170. The next process is to use an object-based algorithm to adjust the channel audio of each speaker.
[0135] In order to eliminate interference in the sound field environment, the audio system 100 needs to obtain the spatial configuration information of various environmental objects 175 in the target space 170.
[0136] In process 410, the control circuit 132 in the host device 130 can run a configuration program to obtain the spatial configuration information of one or more environmental objects 175 in the target space 170. In the foregoing embodiments, the host device 130 automatically identifies the spatial configuration information of the environmental objects 175 using the sound field environment information captured by the sensor circuit 140. When running the configuration program, the host device 130 can interact with the user using a human interface circuit 133 to allow the user to manually input the spatial configuration information of the environmental objects 175. The human interface circuit 133 can provide a screen and an input method for the user to define the spatial configuration information of various objects in the target space 170 in a two-dimensional planar graph or a three-dimensional solid graph. The spatial configuration information of the environmental objects 175 can include the relative positions, sizes, names, and material types of the environmental objects 175 in the target space 170. In further derived embodiments, the user 180 can tell the host device 130 the application scenario category to which the target space 170 currently belongs through the human interface circuit 133. In different application scenarios, such as open outdoor spaces, theater spaces, or bathrooms, the common types of environmental objects 175 are also different, and the sound field atmosphere experienced by the user is also different. Optimizing the sound field for different application scenarios is one of the important functions of the sound system 100.
[0137] Different material types have different acoustic properties. When running the configuration program, the host device 130 further queries a database of objects according to the object names or material types input by the user to obtain the acoustic property information of the environmental objects 175, such as the sound absorption or reflection rate. In this way, the host device 130 can calculate the degree to which the playing effect of each speaker on the target listening point is affected by the environmental objects 175 in subsequent process 212 according to the foregoing spatial configuration information and acoustic property information. In further derived embodiments, the host device 130 can preferentially use the corresponding database of objects according to the application scenario category of the target space 170 to more quickly identify the environmental objects 175 in the target space 170. Related embodiments will be described in the following Figure 9 .
[0138] Figure 4 Processes 212, 214, 216, and 218 in the foregoing embodiments are the same as those in the foregoing embodiments, and will not be repeated for brevity.
[0139] Figure 4The embodiment illustrates that, in addition to dynamically tracking the user's position, the audio system 100 also allows the user 180 to configure the spatial configuration information of environmental objects 175 in the target space 170 via a configuration program. This configuration program provides a channel for manual input to compensate for any deficiencies in the recognition function. Besides actively inputting information to assist the host device 130 in making more accurate judgments, the user also has the opportunity to deliberately specify different application scenario categories or intentionally set imaginary virtual sound source objects to change the playback effect according to their preferences. The host device 130 will perform channel-based compensation operation, calculating the output compensation value corresponding to each speaker based on the spatial configuration information of the environmental objects 175 in the target space 170.
[0140] The following is Figure 5 This describes an embodiment in which an audio system 100 dynamically tracks the user's position, runs a configuration program to obtain the sound field environment configuration, and compensates the speaker output using object-based compensation.
[0141] Figure 5 This is a flowchart of a dynamic sound effect optimization method according to an embodiment of the present invention.
[0142] exist Figure 5 In the flowchart, the process located in the field belonging to a specific device represents the process performed by that specific device. For example, the part marked in the "Sensor Circuit" field is the process performed by sensor circuit 140; the part marked in the "Host Device" field is the process performed by host device 130; the part marked in the "Speaker" field is the process performed by first speaker 110 and / or second speaker 120; and so on. The aforementioned logic also applies to other subsequent flowcharts.
[0143] Figure 5 The processes 202, 204, 206, 208 and 210 are the same as in the previous embodiment, and will not be repeated here to save space.
[0144] and Figure 4 The implementation examples are similar, Figure 5 In order to eliminate interference in the sound field environment, the embodiment ran with Figure 4 The same process 410.
[0145] In process 410, the host device 130 runs a configuration program to obtain spatial configuration information of one or more environmental objects 175 in the target space 170. Figure 4In some embodiments, the host device 130 can receive spatial configuration information of the environmental objects 175 in the target space 170 from a user through the human interface circuit 133. In further embodiments, the host device 130 can also receive spatial configuration information from the user equipment 150 or other devices through the communication circuit 136. For example, the user equipment 150 can be a mobile phone running an application that provides similar functions as the human interface circuit 133. The application allows the user to define the size and shape of the target space 170, the locations of the speakers relative to the target space 170, the locations, sizes, names, and types of the environmental objects 175, and even the location of the user 180. The application can also communicate with the control circuit 132 through the communication circuit 136 to perform various playback operations, such as play, pause, fast forward, adjust volume, etc. In addition, the user can set the application scenario category of the target space 170 through the human interface circuit 133 to allow the host device 130 to produce diversified playback effects for the target space 170.
[0146] In further embodiments, the user equipment 150 connected to the host device 130 can be a virtual reality device or a game console. The user equipment 150 generates sound source signals for the host device 130 to play. The sound source signals can include virtual objects, such as a plane or a dragon, moving in a virtual reality space. The user equipment 150 can transmit relay data of the virtual objects to the host device 130 as part of the spatial configuration information of the environmental objects in the target space 170. In other words, the host device 130 can treat the virtual objects the same as the physical objects in the object-based acoustic system. Through the object-based compensation operation, the host device 130 can allow the user to feel the presence of the virtual objects in the target space 170, and can also allow the user to feel no disturbance from the physical objects in the target space 170. The implementation of the object-based compensation operation is further described in the embodiments of Figures 11 to 13
[0147] At the end of the process 410 of the acoustic system 100 in the present embodiment, the control circuit 132 has tracked the location of the user 180 and assigned it as the target listening point, and has also obtained the spatial configuration information of one or more environmental objects 175 in the target space 170. Then in the processes 312 to 316, the host device 130 adjusts the channel audio of each speaker using the object-based algorithm. Since the processes 312 to 316, and the process 218, are the same as in the previous embodiments, they are not repeated here for brevity.
[0148] Figure 5 The embodiments of the application illustrate that the sound system 100 not only dynamically tracks the position of the user 180, but also allows the user 180 to set the spatial configuration information of the environmental objects 175 in the target space 170 through a configuration program. The configuration program can be integrated with the existing virtual reality technology to receive the spatial configuration information of the virtual objects. The sound system 100 converts the physical environmental objects and the virtual objects into consistent relay data, and then applies all the relay data to the object basis array operation module of the existing object basis acoustic system to perform the object basis compensation operation. In this way, the control circuit 132 does not need to develop additional operation modules for different objects, which can reduce the cost and improve the execution efficiency.
[0149] The following describes Figure 6 several embodiments of the sensor circuit and the compensation algorithm of the channel basis.
[0150] Figure 6 A schematic diagram of a target space 600 of the application is shown to illustrate an embodiment of calculating the audio adjustment amount according to the position of the optimal listening point.
[0151] The sound system 100 of the present application uses the sensor circuit 140 to dynamically sense the target space 600 to generate the sound field environment information. The sound field environment information mainly includes the position of the user 180, and can also include the spatial configuration information of the environmental objects. There are multiple options for the technical solution of dynamic sensing. For example, the sensor circuit 140 can be a combination of one or more of the camera 610, the infrared sensor 620, and the wireless detector 630, which are respectively arranged at different positions around the target space 600 to provide sound field environment information with spatial depth to help the recognition circuit 134 and the control circuit 132 in the host device 130 track the position of the user 180 more efficiently. In this way, the recognition circuit 134 uses the sound field environment information provided by the sensor circuit 140 to not only identify the position of the user 180, but also identify the face direction, ear position, and even gestures or body posture. The control factors for adjusting the sound field can be applied, thus becoming more abundant. For example, focus detection, sleep detection, gesture control, etc.
[0152] In the target space 600 Figure 6 , a first loudspeaker 110 and a second loudspeaker 120 are arranged. The channel basis compensation operation can calculate the output compensation value for each loudspeaker respectively. In a preset condition, the target listening point is located at the center of the target space 600, i.e. Figure 6 the first position 601 in the target space 600. The distance between the first position 601 and the first loudspeaker 110 and the second loudspeaker 120 is also R1. At this time, the first loudspeaker 110 and the first channel audio 112 played by the first channel audio 112 and the second channel audio 122 are also in the preset state, and do not need any compensation processing for the position.
[0153] When the user 180 moves from the first position 601 to the second position 602 along the movement trajectory 173, the sensor circuit 140 detects the new position of the user 180 and assigns the target listening point of the sound system 100 as the second position 602. At this time, the distance between the user 180 and the first speaker 110 changes to R2, and the distance between the user 180 and the second speaker 120 changes to R2'. For the user 180, the first speaker 110 becomes farther, so the received first channel audio 112 is attenuated due to the distance. On the contrary, the second speaker 120 becomes closer, and the received second channel audio 122 is enhanced. In other words, the intensity of the first channel audio 112 and the second channel audio 122 received at the second position 602 has lost balance. The present embodiment uses the algorithm of the channel base to restore the listening effect received at the second position 602 to the same preset state as the first position 601. In other words, the control circuit 132 compensates the first channel audio 112 and the second channel audio 122 output by the first speaker 110 and the second speaker 120 to offset the listening effect deviation caused by the movement of the user 180. Figure 6 The target space 600 shown is not limited to only applicable to a horizontally arranged multi-speaker environment. In a three-dimensional sound field environment arranged with an upper speaker and a lower speaker, the distance deviation problem also occurs. For example, if the user changes from a standing position to a sitting position, the user will be farther away from the upper speaker and closer to the lower speaker.
[0154] To obtain a better compensation effect, the present embodiment uses equal loudness as the calculation standard. For example, the present embodiment can calculate the sound pressure value that needs to be compensated at the target listening point according to the equal loudness curve defined in the ISO226; 2003 protocol. Each channel audio is divided into multiple sub-bands for processing. In addition, different sound field formulas are used as the distance between the user 180 and the speaker changes. Since the equal loudness curve defines a linear relationship between the equal loudness and the sound pressure value, and the equal loudness and the gain value in "decibels" have a linear corresponding relationship. Therefore, in the present embodiment, it is not limited to adjusting in units of any one of equal loudness, sound pressure value, or gain value.
[0155] In an audio system 100, the space through which sound is transmitted due to air vibration is called the sound field. Due to the existence of reflection, sound in a closed room can be classified into several types: (1) Near Field: When the user 180 is located relatively close to the sound source, the physical effects of the sound source (such as pressure, displacement, vibration) will enhance the sound. (2) Reverberant Field: Sound is reflected by objects, resulting in wave superposition. (3) Free Field: A sound field that is not affected by the aforementioned near field and reverberant field. The above reverberant field and free field can be collectively referred to as the far field.
[0156] In many modern audio systems, the definitions of near and far sound fields differ. For example, assuming R is the distance (in meters) between the speaker and the user (180 degrees), L is the speaker's width (in meters), and λ is the representative wavelength (in meters) of a sub-band signal, then the conditions for satisfying the far sound field include the following types:
[0157] R>>λ / 2π (1)
[0158] R>>L (2)
[0159] R>>πL 2 / 2λ (3)
[0160] by Figure 1 Taking the first speaker 110 as an example. When the distance between the target listening point and the first speaker 110 is greater than a certain proportion of the wavelength of the sub-band signal or the size of the first speaker 110, the audio system 100 determines that the sound field type is a far sound field. When the distance between the target listening point and the first speaker 110 is less than the specific proportion of the wavelength of the sub-band signal or the size of the first speaker 110, the sound field type is determined to be a near sound field. In a simpler implementation, the audio system 100 can define twice the wavelength (2λ) corresponding to the center frequency of a sub-band signal as the boundary point between the far sound field and the near sound field of the sub-band signal.
[0161] In the far sound field, the relationship between the change in sound pressure level of a sub-band signal received by user 180 from the speaker and the change in distance is as follows:
[0162] SPL2 = SPL1 - 20 log 10 (R2 / R1) (4)
[0163] Wherein, SPL2 is the sound pressure level of the sub-band signal received at the new location, SPL1 is the sound pressure level of the sub-band signal received at the original location, R2 is the distance between the new location and the speaker, and R1 is the distance between the original location and the speaker.
[0164] From equation (4), the difference between SPL1 and SPL2 is the part that needs to be compensated back to the speaker.
[0165] SPL2' = SPL2 + 20 log 10 (R2 / R1) = SPL1 (5)
[0166] Where SPL2' is the sound pressure value of the sub-band signal received at the new position after compensation. From equation (5), the changed part is compensated back.
[0167] In the near sound field, the relationship between the sound pressure value change of the sub-band signal received by the user 180 from the speaker and the distance change is as follows:
[0168] SPL2 = SPL1 - 10 log 10 (R2 / R1) (6)
[0169] SPL2' = SPL2 + 20 log 10 (R2 / R1) = SPL1 (7)
[0170] From equations (6) and (7), the sound attenuation rate in the near sound field is slower than that in the far sound field, and other calculation logics are the same.
[0171] It can be understood that the above equations may have exceptions in some special cases. For example, when the user 180 moves from the first position 601 to the second position 602 and approaches the second speaker 120, the distance between the user 180 and the second speaker 120 changes from R1 to R2'. This may cause the calculation result of equation (7) to be negative. However, the sub-band signal output by the second speaker 120 cannot be negative, and the minimum can only be reduced to the lowest audible value of the human ear. For example, the sound pressure value of the sub-band signal output by the second speaker 120 is zero. On the other hand, when the user 180 moves from the first position 601 to the second position 602 and moves away from the first speaker 110, the distance between the user 180 and the first speaker 110 changes from R1 to R2. The maximum output limit of the first speaker 110 may not be able to meet equation (5). At this time, the sound system 100 can issue an over-limit prompt to the user 180.
[0172] Figure 6 The embodiments of the application highlight the following advantages. Through the compensation algorithm of the channel base, the optimal listening point of the user is not affected by movement. The calculation method of the channel base is simple and efficient, and can be applied in most target spaces 600.
[0173] Figure 6 The sound compensation method according to the movement of the user 180 has been described. The following describes a sound compensation method according to the movement of the target space 600.Figure 7 This describes the sound compensation method based on environmental object 175. The acoustic properties of environmental object 175 include its reflectivity and absorptivity. In this embodiment, an appropriate calculation method is used to calculate the acoustic impact of environmental object 175 based on its spatial configuration information.
[0174] Figure 7 This is a schematic diagram of a target space 700 of the present invention, used to illustrate an embodiment of calculating the audio adjustment amount based on the absorption rate of environmental objects.
[0175] Figure 7 The image shows an environmental object 175 located between a first speaker 110 and a user 180 in a target space 700. For example, the environmental object 175 could be a sofa or a pillar. In this case, the environmental object 175 may obstruct or attenuate the listening experience for the user 180. In other words, the sound pressure level received by the user 180 from the first speaker 110 may be blocked or absorbed. When the control circuit 132 interprets this layout using spatial configuration information, it uses the absorption rate of the environmental object 175 to calculate the degree to which the playback effect of the first speaker 110 at the target listening point (the user 180's position) is affected by the environmental object 175, in order to determine the equivalent volume, sound pressure level, or gain value that the first channel audio 112 needs to output.
[0176] In one embodiment, the sound loss absorbed by the environmental object 175 can be calculated based on the sound pressure level received by the environmental object 175 from the first speaker 110:
[0177] A t [n] = R[n] * SPL t (8)
[0178] Where n represents the sub-band number. That is, the first channel audio 112 output by the first speaker 110 can be divided into multiple sub-bands and calculated separately. A t [n] represents the gain value of the nth sub-band detected at time point t. R[n] represents the absorption rate of the nth sub-band. t This represents the sound pressure level from the first speaker 110 experienced by the environmental object 175 at time point t. Time point t can represent the time difference between the sound being transmitted from the first speaker 110 to the environmental object 175.
[0179] From formula (8), we can see that A t[n] represents a gain value of the first channel audio 112 absorbed by the environmental object 175 in the nth sub-band, and also represents an output compensation value required by the nth sub-band of the first channel audio 112. Therefore, the control circuit 132 increases the gain value of the nth sub-band of the first channel audio 112 by the gain value A t [n] represents a gain value of the first channel audio 112 absorbed by the environmental object 175 in the nth sub-band, and also represents an output compensation value required by the nth sub-band of the first channel audio 112. Therefore, the control circuit 132 increases the gain value of the nth sub-band of the first channel audio 112 by the gain value A
[0180] There are various situations in which the environmental object 175 is located between the first speaker 110 and the user 180. In this embodiment, the main criterion is whether the line of sight between the first speaker 110 and the user 180 is blocked, or further, the line of sight between the first speaker 110 and the ears of the user 180 is blocked. It can be understood that the SPLt itself is a function related to the distance and time between the environmental object 175 and the first speaker 110, and the influence degree of the calculated At[n] on the user 180 is a function related to the distance and time between the environmental object 175 and the user 180. After considering the arrangement and distance at different angles, various nonlinear correlations are involved. The present application does not limit the derivative changes of formula (8), such as adding other weight coefficients, parameters, and offset correction values according to the actual situation. For example, there may be a sofa between the user 180 and the first speaker 110. Although the sofa does not block the line of sight, it can still affect the sound pressure value received by the user 180 from the first speaker 110. The control circuit 132 can use interpolation or other correction formulas according to formula (8) to make the compensation result more in line with the requirements.
[0181] Figure 8 A schematic diagram of a target space 800 is shown for explaining an embodiment of calculating the audio adjustment amount according to the reflectivity of the environmental object.
[0182] Figure 8 It is shown that in a target space 800, a user 180 is located between a first speaker 110 and an environmental object 175. The environmental object 175 can be a wall, a ceiling, or a floor. In this case, the environmental object 175 will reflect the first channel audio 112 output by the first speaker 110 to the user 180. In other words, the sound pressure value received by the user 180 from the first speaker 110 will be superimposed or interfered. When the control circuit 132 interprets this layout through the spatial configuration information, the reflectivity of the environmental object 175 is used to calculate the degree to which the playback effect of the first speaker 110 at the target listening point (the position of the user 180) is affected by the environmental object 175, in order to determine the equivalent volume, sound pressure value, or gain value required by the first channel audio 112 to be output.
[0183] In this embodiment, the effect of the environmental object 175 can also be calculated according to equation (8), but R[n] is replaced by R[n] representing the reflectivity of the environmental object 175 in the nth sub-band.
[0184] The operation result A of equation (8) t [n] can represent the component of the first channel audio 112 reflected by the environmental object 175 to the user 180 in the nth sub-band. Therefore, the control circuit 132 can appropriately reduce the gain value of the first channel audio 112 when generating the first channel audio 112 through the first speaker 110, so that the total sound pressure value received by the user 180 from the first speaker 110 and the environmental object 175 is maintained at a preset level value.
[0185] Similar to the embodiments of Figure 7 , the case where the user 180 in Figure 8 is located in the middle of the first speaker 110 and the environmental object 175 can exist in various changing situations. This embodiment mainly takes whether the line of sight of the first speaker 110 and the environmental object 175 is blocked by the user 180 as the main basis. However, in practice, walls, ceilings, floors have a reflecting effect regardless of the angle at which they are located. Therefore, the operation formula of this embodiment is not limited to equation (8), and other nonlinear compensation calculation methods can also be derived according to the arrangement and distance relationship. For example, the target space 800 can be classified into different application scenarios, such as a living room, a study, a bathroom, a theater, or outdoors, due to the characteristics of the wall, ceiling, floor material, and room size and shape pattern. The host device 130 can first classify the application scenario to which the target space 800 belongs, and then use the corresponding parameters or formulas.
[0186] Figure 7 and the embodiments of Figure 8 highlight the following advantages. Through the channel basis compensation operation, the effect of the environmental object 175 on the listening effect of the user 180 is eliminated. The channel basis compensation operation can flexibly apply different object acoustic properties according to the configuration of the environmental object, and can effectively deal with the optimization problem of various complex environments.
[0187] In summary, the identification circuit 134 can receive the data of the sensor circuit 140 to identify the position of the user 180 in the target space 170, and assign the position of the user 180 as the target listening point dynamically by the control circuit 132. The compensation made by the control circuit 132 for the movement of the target listening point has been described in the embodiments of Figure 6 and equations (4) to (7). The compensation made by the control circuit 132 for the interference of the environmental object 175 has been described in the embodiments of Figures 7 to 8As explained in equation (8). These two compensation operations can be performed separately and applied to the sound channel audio. In other words, the optimized sound channel audio for final output contains the compensation values made for the target listening point movement, as well as the compensation made for the interference of environmental objects 175.
[0188] The identification circuit 134 identifies the position of the user 180 according to the sound field environment information captured by the sensor circuit 140. The identification process can also include the identification of the application scenario to facilitate the subsequent operations of the control circuit 132. The following describes the identification of the position of the user 180 according to the sound field environment information captured by the sensor circuit 140. Figure 9 The process of identifying objects by the host device 130 according to the application scenario category is described.
[0189] Figure 9 The flowchart of identifying objects by the host device 130 of an embodiment of the present application. The acoustic properties of environmental objects appearing in different application scenarios are usually significantly associated with the family, and the sound field reflection coefficients caused by the surrounding environment material or room size are different. Therefore, distinguishing the application scenario category in advance helps the sound system 100 to improve the efficiency of sound field optimization. It can be understood that, Figure 9 Each flow in the flowchart is performed by the host device 130, but is not limited to being performed by a single circuit or module therein, but can also be performed by the cooperative operation of multiple circuits.
[0190] In the flow 902, the host device 130 obtains the application scenario category of the target space 170. The host device 130 can obtain the application scenario category in several different ways. In an embodiment, the identification circuit 134 in the host device 130 can determine an application scenario category according to the sound field environment information provided by the sensor circuit 140 when identifying the sound field environment information. In another embodiment, the control circuit 132 in the host device 130 obtains the spatial configuration information of the environmental objects through the configuration program running by the human-computer interface circuit 133, and also obtains the application scenario category defined by the user 180 through the configuration program. In a further derived embodiment, the control circuit 132 in the host device 130 can obtain the relevant information of the application scenario category from a user device 150 through the communication circuit 136.
[0191] In procedure 904, to accelerate the query of the environmental object and improve the correctness, the host device 130 selects a relevant object database according to the application scenario category. The object database is usually a pre-established data set, which can be provided by various pipelines. For example, the storage circuit 131 in the host device 130 can pre-store one or more object databases corresponding to different application scenarios. In another embodiment, the host device 130 can connect to a remote database 160 by using the communication circuit 136. The remote database 160 can include a plurality of object databases corresponding to different application scenarios. Each object database includes the shape feature information of a plurality of environmental objects and the acoustic property information.
[0192] When the host device 130 obtains the application scenario category in procedure 902, it can preferentially select an object database related to the application scenario category from the storage circuit 131 or the remote database 160 for subsequent identification of the environmental object. In an embodiment, the identification circuit 134 analyzes the sound field environment information provided by the sensor circuit 140 to obtain one or more object shape feature information, and retrieves the object database according to the object shape feature information, so as to identify the environmental object that meets the object shape feature information, including the name, the absorption rate and the reflection rate. In another embodiment, the control circuit 132 executes a configuration program to obtain the name of an environmental object by using the human-computer interface circuit 133. The control circuit 132 looks up the object database according to the name of the environmental object 175 to obtain the absorption rate and the reflection rate corresponding to the environmental object.
[0193] In a further derived embodiment, the parameters used in the lookup process can be combined in multiple ways. For example, the identification circuit 134 can obtain the material, size, shape and other external characteristics of the environmental object 175 in the process of analyzing the sound field environment information. The identification circuit 134 transmits these external characteristic information to the object database for multi-condition cross comparison, and obtains a candidate object list sorted according to the matching score. If the information of the application scenario category is used as a lookup condition in the process of looking up the object database, it will help to narrow the possible range, accelerate the identification, and improve the correctness.
[0194] In procedure 906, the control circuit 132 looks up the absorption and reflection of the environmental object from the object database selected in procedure 904. In implementation, the acoustic property information of the environmental object stored in the object database is not limited to be stored in a plurality of independent object databases. The object database can be a relational database, which contains a plurality of fields connected together in a correlation coefficient. For example, the fields of the object database can include object name, application scenario category, material, absorption, reflection, and even external characteristics such as shape, color, and luster. The field values corresponding to each environmental object are not limited to be one-to-one, but can be one-to-many or many-to-one. The values stored in each field are not necessarily absolute values, but can be range values or probability values. In further implementation, the object database can be an adaptive database that can be iteratively corrected by machine learning. The user 180 can train the object database by feeding back the preferred set value through the human-computer interface circuit 133.
[0195] In procedure 908, the control circuit 132 adjusts the sound channel audio in a plurality of sub-bands according to the lookup result of the environmental object and the configuration condition. The acoustic properties of the environmental object 175 can be significantly different in different frequency bands. For example, a sofa can absorb a large amount of high-frequency signals, but does not affect the penetration of low-frequency signals. Therefore, the absorption or reflection found from the object database can be an array of values corresponding to a plurality of sub-bands, or a frequency response curve. The interval size or separation method of the sub-bands can be determined according to design requirements, which are not limited in the present embodiment. The control circuit 132 adjusts the gain value of the sound channel audio in a plurality of sub-bands, which can be simulated as an equalizer or a filter in implementation. In other words, the control circuit 132 can implement an equalizer for each speaker in the sound system 100, and customize the equalizer according to the output compensation value calculated in the foregoing embodiments, so that the corresponding sound channel audio is adjusted. Further embodiments of calculating the output compensation value will be described in the Figure 10 embodiment.
[0196] In procedure 910, the control circuit 132 outputs the adjusted sound channel audio to the corresponding speaker through the audio transmission circuit 135. The implementation of the audio transmission circuit 135 has been introduced in the Figure 1 embodiment, which will not be described here.
[0197] Figure 9 The embodiments highlight the following advantages. The operation of object recognition can refer to the application scenario category (automatic recognition or manual input) to increase the recognition efficiency. The object database adopts an extensible architecture, which continuously enhances the recognition ability in the long term under the feedback of cloud big data service and machine learning. The sound system 100 can apply the concept of equalizer to divide the sound channel audio into a plurality of sub-bands for separate processing, so that the final synthesized sound quality is effectively improved.
[0198] The following describes Figure 10 Further description is made on how the control circuit 132 calculates the output compensation value for each sound channel according to the spatial configuration information of the environmental objects 175.
[0199] Figure 10 The flowchart of the audio processing method for an embodiment of the present application describes an embodiment of calculating the output compensation value according to the positional relationship of the environmental objects. Figure 10 The flow is mainly executed by the control circuit 132 in the host device 130.
[0200] In the flow 1002, the control circuit 132 judges the relative positional relationship of the environmental objects, the target listening point, and the loudspeaker. The multiple loudspeakers and the multiple environmental objects 175 in the target space 170 can be arranged in combination with the target listening point to form multiple sets of positional relationship. Each set of positional relationship includes a loudspeaker, an environmental object 175, and the target listening point. The control circuit 132 checks and judges each set of positional relationship in the target space 170 and calculates the corresponding output compensation value. The following describes the compensation method made by the control circuit 132 for the interference caused by an environmental object 175 to a loudspeaker at the target listening point by taking one set of positional relationship in the sound system 100 as an example.
[0201] In the flow 1004, the control circuit 132 judges whether the environmental object 175 is between the target listening point and the loudspeaker. The position of the environmental object 175 in the target space 170 can also be obtained by the recognition circuit 134 or by the man-machine interface circuit 133 through a configuration program. The control circuit 132 can judge the relative positional relationship of each environmental object 175, the target listening point, and each loudspeaker after comprehensively considering the above information and make corresponding compensation operation for each loudspeaker. The case to be judged by the flow 1004 is the situation as shown in FIG. 10B. If the case is consistent, the flow 1008 is performed. If the case is not consistent, the flow 1006 is performed. Figure 7
[0202] In the flow 1006, the control circuit 132 judges whether the target listening point is between the environmental object 175 and the loudspeaker. The case to be judged by the flow 1006 is the situation as shown in FIG. 10C. If the case is consistent, the flow 1010 is performed. If the case is not consistent, the flow 1012 is performed. Figure 8
[0203] In the flow 1008, the control circuit 132 calculates the output compensation value of the sound channel audio using the absorption rate of the environmental object 175. In a preferred embodiment, the output compensation value of the sound channel audio of the loudspeaker is calculated in multiple sub-bands. The detailed calculation can refer to the description of the flow 1008 in FIG. 10A. Figure 7 the target space 700 and equation (8). The control circuit 132 can look up the absorption of the environmental object 175 from the object database and substitute into equation (8) to obtain the output compensation value.
[0204] In process 1010, the control circuit 132 calculates the output compensation value of the sound channel audio using the reflectivity of the environmental object 175. Referring to Figure 8 the target space 800 and equation (8), the control circuit 132 can look up the reflectivity of the environmental object 175 from the object database and substitute into equation (8) to obtain the output compensation value.
[0205] It can be appreciated that the output compensation value calculated based on the absorption of the environmental object 175 can cause the gain value, the sound pressure value, or the equivalent volume of the adjusted sound channel audio to be amplified to compensate for the energy absorbed. Conversely, the output compensation value calculated based on the reflectivity of the environmental object 175 can cause the gain value, the sound pressure value, or the equivalent volume of the adjusted sound channel audio to be reduced to balance the energy reflected back. In other words, the output compensation values calculated based on the absorption and the reflectivity are usually opposite in sign.
[0206] In process 1012, if the environmental object 175 does not meet the condition of process 1004 and does not meet the condition of process 1006, the control circuit 132 can determine that the environmental object 175 is located at a position that does not affect the playback of the target listening point by the speaker. In this case, the control circuit 132 can not calculate the effect of the environmental object 175 on the speaker and the target listening point for this set of position relationships. However, it needs to be appreciated that a target space 170 usually contains multiple speakers. The environmental object 175 can not affect the playback of the target listening point by one of the speakers, but can affect the playback of the target listening point by other speakers. In other words, the control circuit 132 needs to calculate the effect of the environmental object 175 on the target space 170 for each set of position relationships separately. Figure 10
[0207] In some specific cases, the presence of the environmental object 175 can be directly ignored. For example, if the reflectivity or the absorption of the environmental object 175 to sound is less than a specific threshold, it means that the presence of the environmental object 175 in the target space 170 can be ignored. On the other hand, if the control circuit 132 determines that the volume of the environmental object 175 is less than a specific size, the presence of the environmental object 175 can also be ignored.
[0208] In further derived embodiments, if more than one user is detected in the target space 170, the determination of the target listening point can be based on the position center point of the multiple users, or can be selectively based on the position of one of the users. As for the users who are not selected as the target listening point, the host device 130 can simulate them as environmental objects and calculate the effect of the environmental objects on the target listening point according toFigures 7 to 8 The example processing.
[0209] Figure 10 The embodiments highlight the following advantages. Figure 10 The embodiments continue Figure 7 and Figure 8 This approach simplifies complex environmental problems into multiple linear relationships, which are then solved separately. For specific environmental objects 175, the computational complexity can be further reduced by neglecting certain factors.
[0210] Figure 11 This is a schematic diagram of a target space 1100 of the present invention, used to illustrate an embodiment of optimizing the sound field by object substrate compensation operation.
[0211] The target space 1100 includes multiple speakers, such as a first speaker 1110, a second speaker 1120, a third speaker 1130, and a fourth speaker 1140. When the sound system 100 operates based on object-based compensation, the control circuit 132 logically treats the target space 1100 as a spatial coordinate system. This spatial coordinate system can be two-dimensional or three-dimensional planar coordinates. For ease of explanation, Figure 11 The illustration is presented in a two-dimensional planar coordinate system that includes an X-axis and a Y-axis.
[0212] In target space 1100, user 180 is located at origin P0. Control circuit 132 assigns user 180 as the target listening point. Figure 3 As described in the embodiments, the object-based acoustic system is built upon an array of acoustic parameters. Each sound source object has relay data describing its type, location, size (length, width, height), and divergence. After object-based computation, the sound represented by a sound source object is assigned to one or more speakers for simultaneous playback, with each speaker playing a portion of the sound from the sound source object. In other words, the object-based acoustic system can use multiple speakers to simulate the physical presence of a sound source object. For example, through object-based compensation, a user 180 at the target listening point can hear a virtual sound source object 1105 moving along a movement trajectory 1103 from a first position P1 to a new first position P1'.
[0213] The object basis compensation operation of the present embodiment can optimize the sound channel audio outputted by all the speakers for the target listening point. The object basis compensation operation utilizes the array operation module in the existing object basis acoustic system to parameterize various distance factors and sound field category parameters, and can perform operations similar to equations (4) to (7). For the sound system 100, the host device 130 only needs to apply the position information of the user 180 to the object basis compensation operation to optimize the sound channel audio outputted by all the speakers for the target listening point.
[0214] In an embodiment, the control circuit 132 can define the target listening point as the origin of the entire spatial coordinate system. When the user 180 moves, the entire spatial coordinate system moves with the origin. In other words, the relative position of the virtual sound source object 1105 remains unchanged relative to the origin. When the control circuit 132 plays the effect of the virtual sound source object 1105 through the object basis compensation operation, the user 180 experiences the relative position of the virtual sound source object 1105 does not change with the movement of the user 180.
[0215] In the target space 1100 of the present embodiment, there can be environmental objects 175 that can substantially affect the listening effect of the user 180. The control circuit 132 can obtain the spatial configuration information of the environmental objects in the target space 1100 through the flow 902 of Figure 9 When the user 180 moves, the origin of the entire spatial coordinate system changes with the user 180. The relative position of the environmental object 175 changes although it does not move. Therefore, it can be understood that in the spatial coordinate system after the movement, the coordinate value of the environmental object 175 moves in the opposite direction.
[0216] To offset the interference of the environmental object 175 to the user 180, the control circuit 132 of the present embodiment establishes a compensation sound source object of the object basis according to the environmental object 175. The relay data of the compensation sound source object includes the coordinate position, size, and reflectivity and absorptivity of sound of the environmental object 175. The reflectivity and absorptivity of sound of the environmental object 175 can be obtained through the flow 906 of Figure 9 The compensation sound source object is considered as a negative sound source object of the environmental object 175 and is applied to the object basis compensation operation to become a virtual sound source that can offset the environmental object 175.
[0217] It can be understood that the essence of the compensation sound source object is a negative sound source object corresponding to the environmental object 175, and its position overlaps with the environmental object 175, so in the Figure 11The target space 1100 is not shown in other configurations. In addition, the four-horn configuration of the target space 1100 is only an example. In actual applications of the sound system 100, there can be more horns, even a stereo configuration with upper and lower horns. The present description does not limit other possible configurations.
[0218] Figure 11 The embodiments of the object base compensation operation demonstrate its advantages. The control circuit 132 converts the information of the target space 1100 into the form of a spatial coordinate system, and simplifies the complex multi-object interaction operation into an array operation of relay data. The position of the moving user 180 is set as the origin of the spatial coordinate system, so that the processing of the virtual object is completely unaffected by the movement of the user 180, and the operation flow is simplified. The embodiments also propose the concept of compensating for the sound source object, which directly applies the object base compensation operation to offset the environmental object interference, and eliminates the complex multi-channel interaction operation.
[0219] The following describes the simplicity of the object base compensation operation and possible derivative applications. Figure 12 The following describes the simplicity of the object base compensation operation and possible derivative applications.
[0220] Figure 12 A target space 1200 is shown in the present description to demonstrate the embodiments of optimizing the sound field with the object base compensation operation.
[0221] The target space 1200 can include multiple horns, such as a first horn 1210, a second horn 1220, a third horn 1230, a fourth horn 1240, a fifth horn 1250, and a sixth horn 1260, arranged in a long strip-shaped sound field. Each horn corresponds to an ID. When the user 180 is at a first position P1, the relay data of a virtual sound source object (not shown) are mapped to the IDs of the first horn 1210 and the second horn 1220. After the control circuit 132 performs the object base compensation operation, the first horn 1210 and the second horn 1220 play the first channel output 1212 and the second channel output 1222, so that the user 180 feels the presence of the virtual sound source object. When the user 180 moves along the movement trajectory 1203 to a second position P2, the control circuit 132 recalculates the target listening point and maps the relay data of the virtual sound source object to the fifth horn 1250 and the sixth horn 1260. After the control circuit 132 performs the object base compensation operation, the fifth horn 1250 and the sixth horn 1260 play the fifth channel output 1252 and the sixth channel output 1262, so that the user 180 feels that the virtual sound source object still exists on the left and right of the user 180, and does not move away with the movement of the user 180.
[0222] This embodiment mainly explains the flexibility and simplicity of the object basis compensation operation. In many special cases, only a small amount of calculation is needed to optimize the sound field. For example, if the user 180 is located in a spherical sound field, the control circuit 132 only needs to perform a rotation coordinate calculation to make the user 180 feel consistent sound field effects in various directions.
[0223] The following describes the basic logic of the control circuit 132 when performing the object basis compensation operation. Figure 13 The basic logic of the control circuit 132 when performing the object basis compensation operation is summarized as follows.
[0224] Figure 13 The flowchart of the object basis compensation operation of an embodiment of the present application explains the concept of establishing a compensation sound source object.
[0225] In the flowchart 1304, the control circuit 132 establishes a corresponding compensation sound source object according to the environmental object 175. For the user 180 located at the target listening point, the environmental object 175 is an entity sound source. The environmental object 175 can reflect the sound emitted by a loudspeaker to the target listening point. The environmental object 175 can also block or absorb part of the sound, so that the sound emitted by a loudspeaker to the target listening point is attenuated. The compensation sound source object is a negative sound source object established for the environmental object 175. When the host device 130 substitutes the compensation sound source object into the object basis compensation operation to generate the channel audio, the presence of the environmental object 175 can be eliminated. The specific details of the object basis operation itself can be extended from the calculation method of existing object basis acoustic products, and a large number of related array operations are performed using the relay data of the sound source object. For example, the relay data of the compensation sound source object includes the coordinate position, size, and reflectivity and absorptivity of the sound of the environmental object 175.
[0226] In the flowchart 1306, the control circuit 132 calculates the sound source effect of the compensation sound source object. In this embodiment, the compensation sound source object is established according to the environmental object 175, and the relay data has the same coordinate position, size, and reflectivity and absorptivity of the sound as the environmental object 175, but the generated sound source effect is the inverse gain value of the environmental object 175.
[0227] Figure 13 The embodiments of the present application can also refer to similar calculations of the formula (8) and the formula (9). Figure 7 and Figure 8 The formula (8) can be derived into the formula (9) to calculate the gain value passively generated by the environmental object 175 according to the sound pressure value received by the environmental object 175 from the first loudspeaker 110:
[0228] A t [m][n]=R[n]*SPL t [m] (9)
[0229] wherein m represents the number of the loudspeaker, and n represents the number of the sub-band. A t [m][n] represents the gain value of the nthsub-band affected by the mthloudspeaker. R[n] represents the absorption rate of the nthsub-band. t [m] represents the sound pressure value of the mthloudspeaker received by the environmental object 175 at the tthtime point. The time point t can represent the time difference of the sound transmitted from the loudspeaker to the environmental object 175. If the time difference is greater than a non-negligible range, it indicates that there is a condition of echo in the target space 170.
[0230] According to the formula (9), the calculation result corresponding to each environmental object includes a gain value array of multiple loudspeakers and multiple sub-bands at a time point. The sound source effect of the compensation sound source object is the negative value of the gain value array. In other words, the object-based compensation operation based on the formula (9) includes array operations of multiple-dimensional parameter interaction permutation combinations. For the sake of convenience, the gain value corresponding to one loudspeaker and one sub-band at a time point is used for illustration.
[0231] Figure 13 The embodiments of the present application are similar to those of the Figure 7 and Figure 8 The embodiments of the present application are similar to those of the
[0232] For example, when an environmental object 175 absorbs the sound emitted by a loudspeaker, the volume effect received by the target listening point is reduced. At this time, the control circuit 132 establishes a virtual sound source object with a corresponding volume effect at the coordinate position of the environmental object 175 as compensation. Conversely, if an environmental object 175 reflects the sound of a loudspeaker, the target listening point receives excessive volume. At this time, the control circuit 132 establishes a virtual sound source object with a negative gain value at the coordinate position of the environmental object 175.
[0233] It is understood that a visual line is defined as a straight line connecting two objects in space. Since the objects have certain volume and area, the volume can be large, and the case of the visual line being blocked can include partial blocking and complete blocking. The embodiment can be based on formula (9), and different weight coefficients or different offset correction amounts can be multiplied according to various situations.
[0234] In process 1308, the control circuit 132 mixes the sound source effect of the compensation sound source object into the channel audio, which is played by the corresponding speaker. When performing the object basis compensation operation, the control circuit 132 can process complex object corresponding array operations, and mix the multiple sound source signals allocated to be played by each speaker into the corresponding channel audio. After applying the object basis compensation operation, the volume effect received at the target listening point includes the sound source effect generated by the compensation sound source object. In this way, the interference caused by the environmental object 175 can be effectively offset by the compensation sound source object.
[0235] In process 1310, the control circuit 132 determines whether the target listening point moves to a new position. As described in process 208, the sound system 100 can continuously track the movement of the user 180 to update the target listening point. If the target listening point moves, process 1312 is performed. Otherwise, the playing operation of process 1308 is continued.
[0236] In process 1312, the control circuit 132 updates the relay data of the compensation sound source object. In the embodiment, the control circuit 132 establishes an object basis space with the target listening point as a coordinate origin. If the target listening point moves to a new position, the control circuit 132 assigns the new position as a new coordinate origin of the object basis space. The difference between the new coordinate origin and the original coordinate origin position can be represented as a movement vector. The spatial coordinate values of the environmental object 175 relative to the target listening point are also inversely changed with the movement vector. The control circuit 132 then updates the relay data of the compensation sound source object corresponding to the environmental object 175 according to the movement vector. In a further embodiment, all the speakers in the object basis space can also be regarded as an object, which has corresponding ID, relay data and coordinate values.
[0237] In another embodiment, the sound system 100 does not limit the target listening point as the coordinate origin. The sound system 100 can also use a fixed reference point as the origin of the object basis space. When the relative position of the sound source object in the object basis space changes, the control circuit 132 correspondingly updates the coordinate values in the relay data of the sound source object.
[0238] After process 1312 is completed, the control circuit 132 repeats process 1308.
[0239] Figure 13The embodiments of the present application demonstrate the advantages of the object basis compensation operation. The control circuit 132 converts the information of the target space 1100 into the form of a spatial coordinate system, and simplifies the complex multi-object interaction operation into an array operation of relay data. The position of the moving user 180 is set as the origin of the spatial coordinate system, so that the processing of the virtual objects is completely independent of the movement of the user 180, and the operation flow is simplified. The present embodiment also proposes the concept of compensating for the sound source object, and directly applies the object basis compensation operation to offset the interference of the environmental objects, so that the complex multi-channel interaction operation is avoided.
[0240] In a further derived embodiment, if the host device 130 itself does not have the object basis mixing operation capability, the control circuit 132 can provide the function of channel mapping by executing software, so that the operation result of the object basis can be correctly mapped to each speaker.
[0241] In summary, the present application proposes an audio system 100, which can dynamically track the position of the user to optimize the sound field, and intelligently eliminate the interference caused by the environmental objects. The means for tracking the position of the user can be the individual use or combined use of various ways such as cameras, infrared rays, or wireless detectors. The spatial configuration information of the environmental objects 175 in the target space 170 can be obtained by recognizing the image captured by the camera, or can be manually input by the user. The way to optimize the sound field can be the channel basis compensation operation or the object basis compensation operation. When calculating the influence of the environmental objects 175 on the target listening point, the relative positional relationship between the environmental objects 175 and the speakers, and the target listening point can also be considered, and different calculation methods can be adopted. When using the object basis compensation operation, the control circuit 132 establishes a corresponding compensation sound source object for each environmental object 175, so that the sound channel audio generated by the final mixing eliminates the interference of the environmental objects 175 on the target listening point.
[0242] In the description and claims, certain terms are used to refer to particular elements. A person skilled in the art can use different names to refer to the same elements. The description and claims of the present application do not distinguish elements by name, but by function. In the description and claims, the term "comprising" is an open term, which should be interpreted as "comprising but not limited to". In addition, the term "coupled" herein includes any direct and indirect connection means. Therefore, if the first element is coupled to the second element in the description, it means that the first element can be directly connected to the second element by electrical connection or wireless transmission, optical transmission, or indirectly connected to the second element by other elements or connection means.
[0243] As used in the description, the expression "and / or" includes any combination of one or more items listed. In addition, unless the context clearly indicates otherwise, any singular grammatical form includes the plural.
[0244] The above merely describes the preferred embodiments of the present application, and any equivalent changes and modifications made according to the claims of the present application shall fall within the scope of the present application.
Claims
1. An audio system (100) capable of dynamically optimizing playback effect according to user position, comprising: a sensor circuit (140) comprising a camera (610) and configured to dynamically sense a target space (170) using the camera (610) to generate a soundfield environment information comprising a soundfield environment image, wherein the sound field environment information comprises a user position of a user in the target space (170), and spatial configuration information of an environmental object (175) in the target space (170); a first speaker (110) and a second speaker (120) configured to play audio; a host device (130) coupled to the sensor circuit (140), the first speaker (110) and the second speaker (120), comprising: a recognition circuit (134) configured to recognize the user position of the user in the target space (170) from the sound field environment information, and analyze the sound field environment information to obtain spatial configuration information and acoustic property information of the environmental object (175), wherein the acoustic property information of the environmental object (175) comprises at least one of sound absorption rate, sound reflection rate, and resonance frequency; a control circuit (132) coupled to the recognition circuit (134) and configured to dynamically assign the user position as a target listening point; and an audio transmission circuit (135) coupled to the control circuit (132), the first speaker (110) and the second speaker (120) and configured to transmit audio; wherein the control circuit (132) operates a channel basis compensation operation to generate a first channel audio (112) and a second channel audio (122) optimized for the target listening point according to the target listening point, and the spatial configuration information and the acoustic property information of the environmental object (175); wherein the control circuit (132) outputs the first channel audio (112) and the second channel audio (122) to the corresponding first speaker (110) and second speaker (120) through the audio transmission circuit (135); wherein the channel basis compensation operation comprises: splitting the first channel audio (112) emitted by the first speaker (110) into a plurality of sub-band signals; determining whether a sound field type generated by a sub-band signal in the plurality of sub-band signals at the target listening point belongs to a near sound field or a far sound field according to a wavelength of the sub-band signal and a distance between the target listening point and the first speaker (110); determining the sound field type as a far sound field when the distance between the target listening point and the first speaker (110) is greater than the wavelength of the sub-band signal or a certain proportion of the size of the first speaker (110); and determining the sound field type as a near sound field when the distance between the target listening point and the first speaker (110) is less than the wavelength of the sub-band signal or the certain proportion of the size of the first speaker (110); when the control circuit (132) determines that the position of the target listening point changes from a first distance (R1) to a second distance (R2) and the sound field type belongs to the far sound field, calculating the sound pressure value of the sub-band signal using a far sound field formula; wherein the far sound field formula comprises: SPL' = SPL + 20 log 10 (R2 / R1) wherein SPL is the sound pressure value of the sub-band signal before adjustment, SPL' is the sound pressure value of the sub-band signal after adjustment, R1 is the first distance, and R2 is the second distance; and wherein if the sound pressure value of the sub-band signal after adjustment is less than zero, the control circuit (132) sets the sound pressure value of the sub-band signal after adjustment to zero.
2. The sound system (100) of claim 1, wherein The channel base compensation operation further comprises, when the control circuit (132) determines that the distance between the target listening point and the first speaker (110) changes from the first distance (R1) to the second distance (R2), and the sound field type belongs to the near sound field, using a near sound field formula to calculate the sound pressure value of the sub-band signal; wherein the near sound field formula comprises: SPL' = SPL + 10 log 10 (R2 / R1) wherein SPL is the sound pressure value of the sub-band signal before adjustment, SPL' is the sound pressure value of the sub-band signal after adjustment, R1 is the first distance, and R2 is the second distance; and wherein if the sound pressure value of the sub-band signal after adjustment is less than zero, the control circuit (132) sets the sound pressure value of the sub-band signal after adjustment to zero.
3. The sound system (100) of claim 1, wherein, The spatial configuration information of the environmental object comprises the position, size and appearance characteristics of the environmental object, and the acoustic property information of the environmental object comprises the reflectivity and absorption rate of sound; wherein when the control circuit (132) generates the first channel audio, if the target listening point is located between the first speaker (110) and the visual line of the environmental object (175), the control circuit (132) calculates the degree to which the playing effect of the first speaker (110) at the target listening point is affected by the environmental object (175) according to the reflectivity of the environmental object, to determine the sound pressure value of the first channel audio (112); and wherein when the control circuit (132) generates the first channel audio, if the environmental object (175) is located between the target listening point and the visual line of the first speaker (110), the control circuit (132) calculates the degree to which the playing effect of the first speaker (110) at the target listening point is affected by the environmental object (175) according to the absorption rate of the environmental object, to determine the sound pressure value of the first channel audio (112).
4. The sound system (100) of claim 1, wherein, The identification circuit (134) dynamically identifies the head position, face direction or ear position of the user according to the sound field environment image captured by the camera (610), to determine the user position.
5. The sound system (100) of claim 1, wherein, The sensor circuit (140) further comprises an infrared sensor (620) arranged to capture a thermal imaging data in the target space; wherein the identification circuit (134) analyzes the moving track of the thermal imaging data to dynamically determine the user position.
6. The sound system (100) of claim 1, wherein, The sensor circuit (140) further comprises a wireless detector (630) arranged in the target space and detecting a wireless signal of an electronic device; wherein the identification circuit (134) dynamically locates the position of the electronic device according to the characteristics of the wireless signal detected by the wireless detector (630); and wherein the identification circuit (134) dynamically determines the user position according to the position of the electronic device.
7. The sound system (100) of claim 1, wherein, The host device (130) further comprises a storage circuit (131) coupled to the control circuit (132) and configured to store one or more object databases, each of which corresponds to an application scenario category and comprises shape feature information and acoustic property information of a plurality of environmental objects; wherein the identification circuit (134) identifies the sound field environment information to determine an applicable application scenario category when analyzing the sound field environment information; and wherein the control circuit (132) preferentially selects an object database related to the application scenario category from the storage circuit (131) to identify the environmental object and find the acoustic property information of the environmental object according to the application scenario category.
8. The sound system (100) as recited in claim 1, wherein, The host device (130) further comprises a communication circuit (136) coupled to the control circuit (132) and configured to be controlled by the control circuit (132) to connect to a remote database (160) corresponding to an application scenario category; the remote database (160) is configured to store one or more object databases, each of which corresponds to an application scenario category and comprises shape feature information and acoustic property information of a plurality of environmental objects; wherein the identification circuit (134) identifies the sound field environment information to determine an applicable application scenario category when analyzing the sound field environment information; and wherein the control circuit (132) preferentially selects an object database related to the application scenario category from the remote database (160) to identify the environmental object and find the acoustic property information of the environmental object according to the application scenario category.
Citation Information
Patent Citations
Playback device configuration based on proximity detection
CN106105271A
Audio processing method and electronic equipment
CN111050269A
Method And A System For Determining The Location Of An Object
US20150168542A1