Sound system with dynamically adjustable target listening point and elimination of environmental object interference
By combining sensor circuits and control circuits, the listening point of the audio system is dynamically adjusted and the audio output is compensated, solving the problems of fixed position and environmental interference in traditional audio systems and achieving a flexible listening experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-18
- Publication Date
- 2026-04-14
AI Technical Summary
Traditional audio systems cannot dynamically adjust to the optimal listening point, requiring users to settle into fixed positions, and environmental interference can lead to poor listening quality.
The system employs a sound system that includes sensor circuitry, speakers, and a main unit. The sensor circuitry captures sound field environment information, identifies the user's position, and dynamically adjusts the listening point. The control circuitry performs channel basis compensation and object basis compensation to counteract environmental object interference and optimize audio output.
It enables the audio system to dynamically optimize playback based on the user's location, eliminate interference from environmental objects, and provide a stable listening experience without requiring the user to be in a fixed position.
Smart Images

Figure CN116261093B_ABST
Abstract
Description
Technical Field
[0001] This application relates to audio processing technology, which is actually a sound system that can dynamically adjust the playback effect according to changes in the conditions in the sound field space. Background Technology
[0002] Existing audio systems consist of multiple speakers arranged around a target space to create a surround sound environment. Each speaker can output a corresponding channel of audio. When configuring a surround sound environment, the installer of the audio system usually designates a central area of the target space as the optimal listening point, which serves as the basis for installing multiple speakers. When multiple speakers simultaneously play multiple channels of audio, the user located at the optimal listening point can obtain an immersive listening experience.
[0003] However, in real-world environments, a user's listening experience is easily affected by various variables. For example, in traditional audio systems, the optimal listening point is geographically defined. When a user moves to an area outside this optimal listening point, although they can still hear the multiple channels of audio output from the system, the listening effect produced by these channels at the user's location may be significantly reduced or completely lost. Furthermore, the room layout, furniture placement, and materials within the target space are all environmental elements that can interfere with the listening experience. For instance, sofas, windows, and curtains can absorb or reflect some sound energy, distorting the audio received at the optimal listening point.
[0004] In other words, traditional audio systems cannot dynamically adjust the location of the optimal listening point, forcing users to restrict their movement to accommodate this location, which is inconvenient. Furthermore, the audio from each channel can be distorted by environmental interference, further limiting or even eliminating the range of the optimal listening point. This renders the costly soundstage construction meaningless. Summary of the Invention
[0005] Therefore, how to make the audio system dynamically adjust to the optimal listening point as the user moves, and how to eliminate interference from environmental objects in the target space, are problems that need to be solved.
[0006] This specification provides an embodiment of an audio system that dynamically optimizes playback based on the user's position. The audio system includes a sensor circuit, a first speaker, a second speaker, and a host device. The sensor circuit is configured to dynamically sense a target space and generate sound field environment information. The first and second speakers are configured to play audio. The host device, coupled to the sensor circuit, the first speaker, and the second speaker, includes an identification circuit, a control circuit, and an audio transmission circuit. The identification circuit is configured to identify a user from the sound field environment information and determine the user's position in the target space. The control circuit, coupled to the identification circuit, is configured to dynamically assign the user's position as a target listening point. The audio transmission circuit, coupled to the control circuit, the first speaker, and the second speaker, is configured to transmit audio. The sensor circuit includes a camera configured to capture an image of the sound field environment of the target space. The identification circuit analyzes the sound field environment image to obtain spatial configuration information and acoustic attribute information of environmental objects in the target space. The control circuit performs channel basis compensation based on the target listening point and the spatial configuration and acoustic properties of the surrounding objects to generate a first channel audio and a second channel audio optimized for the target listening point. Finally, the control circuit outputs the first channel audio and the second channel audio to the corresponding first speaker and second speaker respectively through the audio transmission circuit.
[0007] This specification provides an embodiment of an audio system that dynamically optimizes playback based on the user's position. The audio system includes a sensor circuit, a first speaker, a second speaker, and a host device. The sensor circuit is configured to dynamically sense a target space and generate sound field environment information. The first and second speakers are configured to play audio. The host device, coupled to the sensor circuit, the first speaker, and the second speaker, includes an identification circuit, a control circuit, and an audio transmission circuit. The identification circuit is configured to identify a user from the sound field environment information and determine the user's position in the target space. The control circuit, coupled to the identification circuit, is configured to dynamically assign the user's position as a target listening point. The audio transmission circuit, coupled to the control circuit, the first speaker, and the second speaker, is configured to transmit audio. The sensor circuit includes a camera configured to capture an image of the sound field environment of the target space. The identification circuit analyzes the sound field environment image to obtain spatial configuration information and acoustic attribute information of environmental objects in the target space. The control circuit maps the target space to an object base space and establishes a compensation sound source object in the object base space according to the environmental object. A relay data point for the compensation sound source object includes the coordinates, size, reflectivity, and absorptivity of the environmental object. Based on the target listening point and the relay data, the control circuit performs an object base compensation operation to counteract the interference of the environmental object on the target listening point, generating a first channel audio and a second channel audio optimized for the target listening point. The control circuit outputs the first channel audio and the second channel audio to the corresponding first and second speakers respectively through the audio transmission circuit.
[0008] This specification provides an embodiment of an audio system that dynamically optimizes playback based on the user's position. The audio system includes a sensor circuit, a first speaker and a second speaker, and a host device. The sensor circuit is configured to dynamically sense a target space and generate sound field environment information. The first speaker and the second speaker are configured to play audio. The host device, coupled to the sensor circuit, the first speaker, and the second speaker, includes an identification circuit, a control circuit, an audio transmission circuit, and a human-machine interface circuit. The identification circuit is configured to identify a user from the sound field environment information and determine the user's position in the target space. The control circuit, coupled to the identification circuit, is configured to dynamically assign the user's position as a target listening point. The audio transmission circuit, coupled to the control circuit, the first speaker, and the second speaker, is configured to transmit audio. The human-machine interface circuit, coupled to the control circuit, is configured to run a configuration program to obtain spatial configuration information and acoustic attribute information of environmental objects in the target space. The control circuit performs a channel basis compensation operation based on the target listening point and the spatial configuration and acoustic properties of the surrounding objects to generate a first channel audio and a second channel audio optimized for the target listening point. The control circuit then outputs the first channel audio and the second channel audio to the corresponding first speaker and second speaker respectively through the audio transmission circuit.
[0009] This specification provides an embodiment of an audio system that dynamically optimizes playback based on the user's position. The audio system includes a sensor circuit, a first speaker and a second speaker, and a host device. The sensor circuit is configured to dynamically sense a target space and generate sound field environment information. The first speaker and the second speaker are configured to play audio. The host device, coupled to the sensor circuit, the first speaker, and the second speaker, includes an identification circuit, a control circuit, an audio transmission circuit, and a human-machine interface circuit. The identification circuit is configured to identify a user from the sound field environment information and determine the user's position in the target space. The control circuit, coupled to the identification circuit, is configured to dynamically assign the user's position as a target listening point. The audio transmission circuit, coupled to the control circuit, the first speaker, and the second speaker, is configured to transmit audio. The human-machine interface circuit, coupled to the control circuit, is configured to run a configuration program to obtain spatial configuration information and acoustic attribute information of environmental objects in the target space. The control circuit maps the target space to an object base space and establishes a compensation sound source object in the object base space according to the environmental object. Relay data for the compensation sound source object includes the coordinates, size, reflectivity, and absorptivity of the environmental object. Based on the target listening point and the relay data, the control circuit performs an object base compensation operation to counteract the interference of the environmental object on the target listening point, generating a first channel audio and a second channel audio optimized for the target listening point. The control circuit outputs the first channel audio and the second channel audio to the corresponding first and second speakers respectively through the audio transmission circuit.
[0010] One advantage of the above embodiments is that the audio system can dynamically track the user's position through sensors and continuously optimize the playback effect according to the user's position. Users do not need to compromise a fixed listening position in order to obtain the best experience.
[0011] Another advantage of the above embodiments is that the audio system can identify environmental objects in the target space and adjust the channel audio accordingly to cancel out the interference of environmental objects.
[0012] Other advantages of the present invention will be explained in more detail with reference to the following description and drawings. Attached Figure Description
[0013] Figure 1 This is a functional block diagram of an audio system according to an embodiment of the present invention.
[0014] Figure 2 This is a flowchart of a dynamic sound effect optimization method according to an embodiment of the present invention.
[0015] Figure 3This is a flowchart of a dynamic sound effect optimization method according to an embodiment of the present invention.
[0016] Figure 4 This is a flowchart of a dynamic sound effect optimization method according to an embodiment of the present invention.
[0017] Figure 5 This is a flowchart of a dynamic sound effect optimization method according to an embodiment of the present invention.
[0018] Figure 6 This is a schematic diagram of a target space of the present invention, used to illustrate an embodiment of calculating audio adjustment based on the location of the optimal listening point.
[0019] Figure 7 This is a schematic diagram of a target space of the present invention, used to illustrate an embodiment of calculating audio adjustment based on the absorption rate of environmental objects.
[0020] Figure 8 This is a schematic diagram of a target space of the present invention, used to illustrate an embodiment of calculating audio adjustment based on the reflectivity of environmental objects.
[0021] Figure 9 This is a flowchart illustrating the identification of an object by a host device according to an embodiment of the present invention.
[0022] Figure 10 This is a flowchart of an audio processing method according to an embodiment of the present invention, illustrating an embodiment of calculating output compensation values based on the positional relationships of environmental objects.
[0023] Figure 11 This is a schematic diagram of the target space of the present invention, used to illustrate an embodiment of optimizing the sound field by object-based compensation operation.
[0024] Figure 12 This is a schematic diagram of the target space of the present invention, used to illustrate an embodiment of optimizing the sound field by object-based compensation operation.
[0025] Figure 13 This is a flowchart of an embodiment of the present invention for object substrate compensation operation.
[0026] Symbol Explanation
[0027] 100...sound system
[0028] 110...First loudspeaker
[0029] 112...First channel audio
[0030] 120...Second loudspeaker
[0031] 122...Second channel audio
[0032] 130...Main Unit
[0033] 131... Storage circuit
[0034] 132... Control Circuit
[0035] 133...Human-Machine Interface Circuit
[0036] 134...Identification Circuit
[0037] 135... Audio transmission circuit
[0038] 136... Communication circuit
[0039] 140... Sensor Circuit
[0040] 150... User Equipment
[0041] 160... Remote Database
[0042] 170...Target Space
[0043] 171...First position
[0044] 172...Second position
[0045] 173... Movement trajectory
[0046] 175...Environmental objects
[0047] 180... users
[0048] 202~218... Process
[0049] 312~316... Process
[0050] 410...process
[0051] 600... target space
[0052] 601...First position
[0053] 602...Second position
[0054] 610...camera
[0055] 620... Infrared sensor
[0056] 630... Wireless Detector
[0057] 700... target space
[0058] 800... target space
[0059] 902~910... Process
[0060] 1002~1012...process
[0061] 1100... target space
[0062] 1103...Object movement trajectory
[0063] 1105... Virtual audio source object
[0064] 1110...First loudspeaker
[0065] 1120...Second loudspeaker
[0066] 11:30... Third loudspeaker
[0067] 1140... Fourth loudspeaker
[0068] P0...Origin
[0069] P1...First Position
[0070] P1'...New First Position
[0071] P2...Second position
[0072] 1200... target space
[0073] 1201...Target Listening Point
[0074] 1203... Movement trajectory
[0075] 1210...First loudspeaker
[0076] 1212...First channel output
[0077] 1220...Second loudspeaker
[0078] 1222...Second channel output
[0079] 1230... Third loudspeaker
[0080] 1240... Fourth loudspeaker
[0081] 1250... Fifth loudspeaker
[0082] 1252... Fifth channel output
[0083] 1260...Sixth loudspeaker
[0084] 1262...Sixth channel output
[0085] 1304~1312...process Detailed Implementation
[0086] The embodiments of the present invention will be described below with reference to the accompanying drawings. In the drawings, the same reference numerals denote the same or similar elements or method flows.
[0087] Figure 1 This is a functional block diagram of an audio system 100 according to an embodiment of the present invention.
[0088] The audio system 100 mainly consists of a host device 130 and multiple speakers. The host device 130 can control the multiple speakers to play audio. The host device 130 can be a computer host, a barebone system, an embedded system, or a customized digital audio processing device. The host device 130 includes a communication circuit 136, enabling the host device 130 to connect to a user equipment 150 via wired or wireless means, serving as an input channel for audio signals or data.
[0089] User device 150 can be a mobile phone, computer, TV stick, game console, or other audio source providing device, providing music or sound streaming to host device 130 via communication circuit 136. Furthermore, the audio system 100 can utilize communication circuit 136 to operate in conjunction with user device 150 or other multimedia devices, forming a home theater system with both video and audio capabilities. For example, the target space 170 may also include a projection screen, display, or monitor (not shown), which is controlled by user device 150 to display images. As another example, user device 150 can be a head-mounted virtual reality device. User 180 can stand in target space 170 and see the images through user device 150, while host device 130 can be controlled by user device 150 to play audio synchronously with the images. The communication circuit 136 in this embodiment may be (but is not limited to) a High Definition Multimedia Interface (HDMI), a Sony / Philips Digital Interface Format (SPDIF), a wireless local area network module, an Ethernet module, a shortwave radio frequency transceiver, or an evolution of Bluetooth Low Energy (BLE) version 4 or 5, or a Universal Serial Bus.
[0090] The host device 130 also includes an audio transmission circuit 135 for connecting multiple speakers and outputting multiple channels of audio to the speakers for playback. The host device 130 controls the multiple speakers through the audio transmission circuit 135 in a manner that can be unidirectional digital or analog output, or a bidirectional synchronous communication protocol. The connection between the audio transmission circuit 135 and each speaker can be a wired interface, a wireless interface, or a hybrid of both. The wired interface can be (but is not limited to) a composite audio-visual terminal, a digital transmission interface, or a high-definition multimedia interface. The wireless interface can be (but is not limited to) a wireless local area network, a shortwave radio frequency transceiver, or an evolution of Bluetooth Low Energy version 4 or 5. In further derived embodiments, since both the audio transmission circuit 135 and the communication circuit 136 are functional interfaces for connecting to external components, they can be integrated into a multi-functional bidirectional transmission interface module. The audio transmission circuit 135 and the communication circuit 136 employ various publicly available standard transmission technologies to achieve connection and transmission between components, which can increase the future functional expandability of the audio system 100 and reduce the replacement cost when components are damaged.
[0091] Figure 1 The target space 170 can be understood as a three-dimensional space where the user 180 can use the audio system 100. Each speaker can be configured in a different position within the target space 170, corresponding to play one channel of audio. The surrounding configuration of multiple speakers can create a surround sound field environment within the target space 170. There are various standard specifications for the number and configuration of speakers. For example, a 5.1-channel surround sound system includes two front speakers, one center speaker, two surround channel speakers, and one subwoofer, creating a surround sound field space by surrounding a target listening point and playing sound towards that target listening point. In a 7.1-channel ambient sound system, further configuring a pair of rear surround channel speakers behind the target listening point can provide a more three-dimensional sound field effect. In recent years, new specifications such as 5.1.2-channel and 7.2.2-channel have also emerged, including more speakers and channel configurations in specific directions, achieving more realistic "Atmos," "sky effects," or "floor effects." To facilitate the explanation of the technical features of the audio system 100 in this embodiment, Figure 1Only the first speaker 110 and the second speaker 120 are shown as representative examples. The first speaker 110 receives and plays the first channel audio 112 provided by the host device 130, while the second speaker 120 receives and plays the second channel audio 122 provided by the host device 130. It must be understood that in practice, the audio system 100 of this embodiment is not limited to using only two speakers, but can be applied to configurations with 2.1 channels, 4.1 channels, 5.1 channels, 7.2 channels, or more. Each speaker in the target space 170 can have different audio output specifications. For example, some speakers excel at outputting deep bass, while others excel at outputting mid-high frequencies. The host device 130 can plan various sound field environments with different characteristics in the target space 170 according to the different speaker specifications.
[0092] The term "channel" as used in the specification and claims refers broadly to various physical channels and logical channels. A logical channel refers to the audio data stream transmitted within the system, while a physical channel refers to the signal source played by each speaker. In this embodiment, the first channel audio 112 and the second channel audio 122 played by each speaker are physical channels, which can be the result of down-mixing one or more logical channels. For example, a pair of headphones has only two speakers, but can simultaneously hear the sound effects generated by multiple applications. In other words, the sound effect data of multiple applications can be down-mixed by the system into two physical channels and played as audible sound through two speakers. Therefore, the first channel audio 112 and the second channel audio 122 in this embodiment are not limited to audio signals containing only a single logical channel, but can also be audio signals generated by mixing multiple logical channels according to a predetermined ratio.
[0093] exist Figure 1 In this system, a first speaker 110 and a second speaker 120 are positioned on either side of a target space 170 to play sound at a target listening point within that space. The target listening point can be understood as the location where the playback effect of the audio system 100 is optimized. In some audio systems, the target listening point is also called the listening sweet spot. In most cases, the target listening point is typically located in a specific area of the target space 170, such as the center point, on the axis, on a tangential plane, or at the equivalent volume center of multiple speakers. Figure 1In this embodiment, the first position 171 where the user 180 is located is used to represent the target listening point of the target space 170. When the user 180 moves from the first position 171 to the second position 172 along the movement trajectory 173, the listening effect received by the user 180 is deviated because the user 180 moves away from the first speaker 110 and closer to the second speaker 120. Conventional audio systems cannot track the movement of the user 180 and adjust the listening effect received at the second position 172 accordingly. The solution proposed in this embodiment will be described in detail later.
[0094] On the other hand, the target space 170 typically contains environmental objects 175, such as sofas, tables, curtains, walls, ceilings, and floors. These environmental objects 175, depending on their material, size, and location, will interfere with the sound played by the first speaker 110 and the second speaker 120 to varying degrees. For example, fabric sofas or curtains absorb sound, while marble floors or walls reflect sound. In other words, the presence of environmental objects 175 will affect the first channel audio 112 and the second channel audio 122 received by the target listening point. Traditional audio systems lack the ability to identify environmental objects 175 in the target space 170, nor do they have the function of compensating for the first channel audio 112 and the second channel audio 122 based on the size, material, and location of the environmental objects 175. The audio system 100 of this embodiment can calculate and eliminate the interference of all environmental objects 175 in the target space 170 on the first channel audio 112 and the second channel audio 122. For ease of explanation, this embodiment... Figure 1 Only one environmental object 175 is shown to explain the operation of the sound system 100. However, it must be understood that... Figure 1 The target space 170 is not intended to limit the number of environmental objects 175 to only one. The solution to the interference of environmental objects 175 will be detailed later.
[0095] The host device 130 in this embodiment also includes a storage circuit 131. The storage circuit 131 may include non-volatile memory for storing the relevant operating system, application software, or firmware required for the operation of the host device 130. The storage circuit 131 may also include volatile memory for use as the operational memory of the control circuit 132. The host device 130 in this embodiment also includes a control circuit 132. The control circuit 132 may be a central processing unit, a digital signal processor, or a microcontroller. The control circuit 132 can read the pre-stored operating system, software, or firmware from the storage circuit 131 to control the host device 130, the first speaker 110, and the second speaker 120 to perform audio playback operations. Furthermore, the host device 130 in this embodiment utilizes the control circuit 132 to perform a series of sound field compensation calculations to dynamically optimize the playback effect and overcome the shortcomings that traditional audio systems cannot overcome.
[0096] To dynamically optimize playback at the target listening point, the audio system 100 of this embodiment includes a sensor circuit 140 configured to dynamically sense a target space 170 and generate sound field environment information. The sensor circuit 140 may be a component located outside the host device 130 and coupled to it. The sensor circuit 140 may be a combination of one or more of a camera 610, an infrared sensor 620, and a wireless detector 630. The form of the sound field environment information captured by the sensor circuit 140 may vary depending on how the sensor circuit 140 is implemented. For example, the sound field environment information may be a combination of one or more of images, pictures, thermal images, and radio wave imaging of the user and environmental objects. In one embodiment, the sensor circuit 140 is disposed around the target space 170. It is understood that although... Figure 1 Only one sensor circuit 140 is shown in the figure, but in practice, the audio system 100 may include multiple sensor circuits 140, which are respectively configured at different positions around the target space 170 to obtain more accurate sound field environment information.
[0097] In the host device 130 of this embodiment, an identification circuit 134 is included, coupled to the sensor circuit 140. The identification circuit 134 can identify key information affecting the sound field from the sound field environment information, enabling the control circuit 132 to dynamically adjust the first channel audio 112 and the second channel audio 122 played from the first speaker 110 and the second speaker 120. For example, the identification circuit 134 can identify a user from the sound field environment information and determine the user's position in the target space. Since the sound field environment information provided by the sensor circuit 140 can have various combinations, the identification circuit 134 can also implement different identification technology solutions accordingly. For example, when the sound field environment information is an image, the identification circuit 134 can use artificial intelligence recognition technology to distinguish the user in the image. Through the application of artificial intelligence, after analyzing the user in the image, the identification circuit 134 can further locate the user's head, face, and even ear positions. If the sensor circuit 140 can provide diverse information such as three-dimensional images with spatial depth, infrared thermal imaging, or wireless signals, it will help the recognition circuit 134 obtain more accurate recognition results.
[0098] To calculate the degree of interference caused by environmental objects 175 to the sound field environment, the host device 130 requires spatial configuration information and acoustic attribute information of the environmental objects 175. The spatial configuration information may include the size, position, shape, and various external features of the environmental objects 175. The acoustic attribute information may include material-related characteristics such as sound absorption rate, reflectivity, and resonant frequency. In one embodiment, the identification circuit 134, while identifying sound field environment information, can further identify the spatial configuration information of the environmental objects 175 in the target space 170 from the sound field environment information and search for acoustic attribute information. An object database is needed to identify environmental objects. In one embodiment, the storage circuit 131 in the host device 130 can also be used to store an object database. The object database may contain various external feature information for identifying environmental objects, as well as various acoustic attribute information corresponding to each environmental object. For example, when the host device 130 needs to calculate the degree of interference caused by an environmental object 175 to the sound field environment, it can first analyze the object name of the environmental object 175 through the identification circuit 134, and then the host device 130 reads the storage circuit 131 to find the absorptivity and reflectivity corresponding to the environmental object 175.
[0099] In practice, the recognition circuit 134 can be a custom processor chip, working in conjunction with the existing operating system, software, or firmware in the storage circuit 131 to perform the artificial intelligence recognition function. The recognition circuit 134 can also be a core or thread circuit of the control circuit 132, executing the existing artificial intelligence software product in the storage circuit 131 to achieve the recognition function. Alternatively, the recognition circuit 134 can be a memory module of a specific artificial intelligence software product, executed by the control circuit 132 to complete the recognition function.
[0100] The human-machine interface circuit 133 in the host device 130 allows the user to control the operation of the host device 130. The human-machine interface circuit 133 may include a display screen, buttons, a dial, or a touchscreen, allowing the user to perform basic audio system 100 control functions, such as adjusting volume, playing, and fast-forwarding / rewinding. In one embodiment, the control circuit 132 can also execute a configuration program through the human-machine interface circuit 133 to allow the user to set various sound field scenarios or to inform the host device 130 of the spatial configuration information of environmental objects 175 in the target space 170. For example, in this configuration program, the control circuit 132 uses the human-machine interface circuit 133 to receive object configuration data input by the user, such as the object name, type, size, and position of one or more environmental objects 175. After obtaining this spatial configuration information, the control circuit 132 then searches for the corresponding absorptivity and reflectivity from the object database stored in the storage circuit 131 for subsequent sound field compensation operations. In further derived embodiments, the human-machine interface circuit 133 may also be provided by the user equipment 150. The user can operate the configuration program using the user equipment 150, and finally the user equipment 150 transmits the setting results to the control circuit 132 through the communication circuit 136.
[0101] The host device 130 can also be connected to a remote database 160 via a communication circuit 136. In a further embodiment, the object database originally stored using the storage circuit 131 can also be stored via the remote database 160. When the host device 130 needs to calculate the degree of interference caused by an environmental object 175 to the sound field environment, it can first analyze the sound field environment information through the identification circuit 134 to obtain the object's feature value, and then use the communication circuit 136 to access the remote database 160 to find an environmental object 175 that matches the object's feature value and obtain the sound field attribute information of the environmental object 175. The remote database 160 can be a server located in the cloud or other systems, connected to the host device 130 via wired or wireless bidirectional network communication technology. In addition to providing search functions, the remote database 160 can also accept the upload of updated data to continuously expand the database content. For example, the host device 130 can communicate with the remote database 160 using Structured Query Language (SQL).
[0102] based on Figure 1Based on the system architecture, the audio system 100 proposed in this application can achieve at least the following technical effects. First, the audio system 100 can dynamically track the user's position as the target listening point. The audio system 100 can also dynamically acquire spatial configuration information of environmental objects as a basis for optimizing the sound field effect. Finally, the audio system 100 dynamically compensates the speaker output based on the user's position and the spatial configuration information of environmental objects to eliminate object interference and optimize the listening effect at the target listening point. The implementation of dynamically tracking the user's position can employ various technical solutions such as cameras, infrared sensors, or wireless positioning. The implementation of acquiring spatial configuration information of environmental objects can be automatic or manual. For example, the audio system 100 can use a camera to capture images and perform artificial intelligence recognition, or allow the user to manually input the environmental conditions through a configuration program. The implementation of compensating the speaker output can be based on several different algorithms. For example, this specification introduces the Channel Base algorithm and the Object Base algorithm.
[0103] The following is Figure 2 This describes an embodiment of an audio system 100 that dynamically tracks the user's position, uses a camera to acquire the sound field environment configuration, and compensates the speaker output using channel-based compensation.
[0104] Figure 2 This is a flowchart of a dynamic sound effect optimization method according to an embodiment of the present invention.
[0105] exist Figure 2 In the flowchart, the process located in the field belonging to a specific device represents the process performed by that specific device. For example, the part marked in the "Sensor Circuit" field is the process performed by sensor circuit 140; the part marked in the "Host Device" field is the process performed by host device 130; the part marked in the "Speaker" field is the process performed by first speaker 110 and / or second speaker 120; and so on. The aforementioned logic also applies to other subsequent flowcharts.
[0106] In process 202, the sensor circuit 140 dynamically senses the target space to generate sound field environment information. In this embodiment, the sound field environment information can be optical, thermal, or electromagnetic wave information in the target space 170. For example, the sensor circuit 140 may include a camera to continuously record video of the target space 170 or periodically capture still photos of the target space 170. In another embodiment, the sensor circuit 140 may also include an infrared sensor configured to capture thermal imaging data in the target space. The thermal imaging data generated by the infrared sensor, in addition to containing spatial depth information, is also extremely sensitive to temperature changes, making it particularly suitable for tracking user location. In another embodiment, the sensor circuit 140 may also include a wireless detector located in the target space to detect the wireless signal of an electronic device. When a user holds an electronic device, the wireless detector can detect the beacon time difference or the strength of the wireless signal of the electronic device as an auxiliary means of tracking the user's location. The electronic device can be a user's own mobile phone, a specially designed beacon generator, a head-mounted virtual reality device, a game controller, or a remote control for the audio system 100. It is understood that this embodiment does not limit the number of sensor circuits 140, nor does it limit the use of only one sensing scheme at a time. For example, the audio system 100 of this embodiment can employ multiple sensor circuits 140 operating collaboratively from different locations, or simultaneously employ one or more cameras, infrared sensors, and wireless detectors. In this way, the host device 130 can obtain more complete sound field environment information and achieve more accurate recognition results in subsequent processes.
[0107] In process 204, the sensor circuit 140 transmits the sensed sound field environment information to the host device 130. The sensor circuit 140 may continuously transmit data, such as video, or periodically transmit static data. The frequency at which the sensor circuit 140 transmits data can be adaptively determined based on the amount of information in the sound field environment, the tracking accuracy requirements, and the computing power of the host device 130. The sensor circuit 140 and the host device 130 may be connected via a dedicated line or via a communication circuit 136. In a further derived embodiment, the sensor circuit 140 may share the audio transmission circuit 135 with the speaker to transmit the sound field environment information to the host device 130 via the audio transmission circuit 135.
[0108] In process 206, the host device 130 determines the user's position based on the sound field environment information received from the sensor circuit 140. The recognition circuit 134 in the host device 130 can perform a recognition program on the sound field environment information, such as applying artificial intelligence. The recognition algorithm of the recognition circuit 134 varies depending on the sensing scheme of the sensor circuit 140. It is understood that the target space 170 and the user's position can be represented in two-dimensional or three-dimensional space. If the audio system 100 implements only a single sensor circuit 140, it can at least perceive position information in two-dimensional space. If the audio system 100 increases the number of sensor circuits 140 or uses a multi-sensor scheme, it can obtain depth information in three-dimensional space to more accurately determine the user's position or the user's head position. In one embodiment, the recognition circuit 134 can dynamically identify the user's head position, face direction, or ear position based on sound field environment images captured by a camera. In another embodiment, the recognition circuit 134 can analyze the movement trajectory of thermal imaging data generated by an infrared sensor to dynamically determine the user's position 180. For example, the identification circuit 134 can dynamically locate the coordinates of the electronic device in the target space 170 based on the characteristics of the wireless signal detected by the wireless detector. Using this coordinate value, the control circuit 132 can further infer the position of the user's ear.
[0109] In process 208, after the identification circuit 134 in the host device 130 analyzes the user's position, the control circuit 132 in the host device 130 dynamically assigns the user's position as the target listening point. For ease of description of the following embodiments, the target space 170 is described here as a two-dimensional coordinate space or a three-dimensional coordinate space, and the target listening point can be represented as a coordinate value in the target space 170. Depending on the arrangement of the multiple speakers, the range of the target listening point can be more than a single point; it can also be a surface or a three-dimensional area with length, width, and height. For example, after the identification circuit 134 analyzes the user's head position or ear position, the control circuit 132 can assign the user's head position or ear position as the target listening point. The control circuit 132 will then perform subsequent compensation operations to ensure that the playback effect obtained at the target listening point is not affected by the user's movement. In practice, the control circuit 132 compensates for the listening effect obtained at the target listening point by adjusting the first channel audio 112 and the second channel audio 122. It is understood that process 208 may be executed dynamically as the user's position changes. Therefore, process 208 is not limited to following... Figure 2 The execution sequence is shown. In other words, the target listening point can be updated in real time as the user's position changes. The specific adjustment calculations will be described later.
[0110] In process 210, the recognition circuit 134 in the host device 130 further recognizes the sound field environment information provided by the sensor circuit 140 to obtain spatial configuration information of environmental objects in the target space 170. In other words, the sound field environment information provided by the sensor circuit 140 can not only be used to determine the user's position, but also to determine the various environmental objects 175 present in the target space 170. In one embodiment, after the camera in the sensor circuit 140 captures a sound field environment image of the target space 170, the recognition circuit 134 analyzes the sound field environment image to identify one or more environmental objects 175 in the target space 170, as well as the spatial configuration information of these environmental objects 175. The spatial configuration information includes the size, position, shape, and appearance features of the environmental objects 175. The recognition circuit 134 can also determine the acoustic attribute information of each environmental object 175, such as its sound absorption rate and reflectivity, through artificial intelligence calculations or database retrieval. In a further derived embodiment, the recognition circuit 134 can also determine the application scenario category of the target space 170 based on the sound field environment image. The application scenario category can include theater, living room, bathroom, outdoors, etc. If the host device 130 knows the application scenario category of the target space 170, it can more quickly identify environmental objects 175 in the target space 170 and reduce false positives. Related embodiments will be discussed later. Figure 9 The explanation is as follows.
[0111] In process 212, the control circuit 132 in the host device 130 can calculate the degree to which the playback effect of a speaker at the target listening point is affected by environmental objects. The playback effect of a speaker at the target listening point can be defined as the equivalent loudness or sound pressure level (SPL) received from the speaker at that target listening point. The ISO 226 standard defines an equal loudness curve (Fletcher-Munson Curve), illustrating that the equivalent loudness perceived by a user in different sub-bands actually corresponds to different sound pressure levels. In one embodiment, the control circuit 132 can use the equal loudness curve as a standard reference for the playback effect to calculate the sound pressure level received at the target listening point under various conditions. The control circuit 132 can utilize the spatial configuration information and acoustic attribute information of the environmental object 175 to evaluate the interference caused by the environmental object 175 to the target listening point, in order to further calculate methods to eliminate the interference. The influence of the spatial configuration information and attribute information of the environmental object 175 includes many scenarios. For example, the larger the volume of the environmental object 175, the greater the interference coefficient it may have on the target listening point. Whether the position of the environmental object 175 obstructs the user 180 and the speaker also determines the degree to which the speaker is affected. Depending on the material, the environmental object 175 may absorb or reflect sound. Therefore, the control circuit 132 needs to select corresponding parameters or formulas for different acoustic properties to calculate the degree to which the speaker is affected.
[0112] In process 214, the control circuit 132 in the host device 130 employs channel-based compensation to calculate the required output compensation value for each channel audio of each speaker. The channel-based compensation operation calculates the playback effect on the target listening point separately for each channel audio. For example, the first channel audio 112 played by the first speaker 110 may lose energy due to interference from an environmental object 175 before being transmitted through the air to the target listening point. Changes in the target listening point's location also affect the sound pressure level generated by the first channel audio 112 at that point. Through channel-based compensation, the control circuit 132 can calculate the change in sound pressure level of the first channel audio 112 at the target listening point. In this embodiment, the control circuit 132 adds an output compensation value to the first channel audio 112 to offset the change in sound pressure level, restoring the first channel audio 112 received by the target listening point to its state before being affected. In other words, the output compensation value has the same numerical value as the change in sound pressure level, but with the opposite positive or negative polarity.
[0113] In process 216, control circuit 132 adjusts and outputs channel audio to the speakers based on the output compensation value. Since the adjusted channel audio has offset the effects of the user 180's displacement in the target space 170 and the interference caused by environmental objects 175, the listening effect perceived by the user 180 remains consistent. Taking the first speaker 110 and the second speaker 120 in the target space 170 as examples, control circuit 132 calculates and adjusts the sound pressure values of different sub-bands in the first channel audio 112 and the second channel audio 122, thereby offsetting the equivalent volume deviation perceived by the user 180 due to movement. On the other hand, the control circuit (132) compensates the first channel audio 112 and the second channel audio 122 accordingly based on the change in sound pressure value caused by the position, size, and acoustic properties of the environmental objects 175 at the target listening point.
[0114] In process 218, each speaker receives channel audio from the host device 130 via audio transmission circuit 135. Taking the first speaker 110 and the second speaker 120 in the target space 170 as an example, the control circuit 132 outputs the first channel audio 112 and the second channel audio 122 to the corresponding first speaker 110 and second speaker 120 via audio transmission circuit 135. Thus, the first speaker 110 and the second speaker 120 correspondingly play the adjusted first channel audio 112 and second channel audio 122, providing an optimized listening experience for the user 180's target listening point. For ease of explanation, Figure 1 In the embodiment of target space 170, only two speakers and one environmental object 175 are shown. However, it is understood that in practice, the host device 130 may contain more than two speakers, and the number of environmental objects 175 is not limited to one. In further derived embodiments, each speaker may be good at outputting different audio ranges. For example, some speakers are mid-high frequency speakers, and some are subwoofer speakers. When adjusting the channel audio, the control circuit 132 can further adjust the corresponding output first channel audio 112 and second channel audio 122 according to the characteristics of different speakers.
[0115] The following is Figure 3 This describes an embodiment of an audio system 100 that dynamically tracks the user's position, uses a camera to acquire the sound field environment configuration, and compensates the speaker output using object-based compensation.
[0116] Figure 3 This is a flowchart of a dynamic sound effect optimization method according to an embodiment of the present invention.
[0117] exist Figure 3In the flowchart, the process located in the field belonging to a specific device represents the process performed by that specific device. For example, the part marked in the "Sensor Circuit" field is the process performed by sensor circuit 140; the part marked in the "Host Device" field is the process performed by host device 130; the part marked in the "Speaker" field is the process performed by first speaker 110 and / or second speaker 120; and so on. The aforementioned logic also applies to other subsequent flowcharts.
[0118] Figure 3 The processes 202, 204, 206, 208 and 210 are the same as in the previous embodiment, and will not be repeated here to save space.
[0119] When the audio system 100 in this embodiment completes process 210, the control circuit 132 has tracked the position of the user 180 and assigned it as the target listening point, and has also obtained the spatial configuration information of one or more environmental objects 175 in the target space 170. The object-based compensation operation will then be described in the subsequent process to adjust the channel audio of each speaker.
[0120] Object-based acoustic systems originated from virtual reality mixing technology and can simulate the movement of sound source objects using a limited number of physical speakers. Existing software products, such as Dolby Atmos, Spatial Audio Workstations, and Digital Spatial Reality, all fall under the category of object-based acoustic systems. Users can define the movement trajectory of sound source objects in a virtual space through a human-computer interface. The object-based system then uses physical speakers to simulate the sound effects of these objects within the virtual space. Users at the target listening point can thus realistically perceive the movement of sound source objects in space.
[0121] An object-based acoustic system is built upon an array of acoustic parameters. Each sound source object has relay data describing its type, location, size (length, width, height), and divergence. After the object-based array calculation, the sound represented by a sound source object is assigned to one or more speakers for simultaneous playback, with each speaker playing a portion of the sound from the sound source object. In other words, the object-based array calculation can utilize multiple speakers to simulate the spatial effect of a single sound source object. Figure 3 The embodiments propose an object substrate compensation operation based on an object substrate acoustic system to solve the traditional playback effect problem.
[0122] In process 312, the control circuit 132 in the host device 130 establishes a compensating sound source object for the object base based on the environmental object 175. In practice, the control circuit 132 first maps the target space 170 to an object base space in virtual reality, and then establishes a compensating sound source object in the object base space corresponding to the environmental object 175 to generate a sound source effect that cancels out the environmental object 175. For the user 180 located at the target listening point, the presence of the environmental object 175 can also be simulated as a sound source object. In practical applications, the environmental object 175 may reflect the sound emitted by a speaker to the target listening point. The environmental object 175 may also block or absorb some sound, causing the sound emitted by a speaker to the target listening point to be attenuated. In other words, after the control circuit 132 of this embodiment simulates the environmental object 175 as a sound source object, it can correspondingly establish a negative sound source object with the opposite sound source effect in the object base space as a means of canceling interference. The sound source effect described in this embodiment can be the sound pressure level, equivalent volume, or gain value generated for the target listening point.
[0123] In process 314, the host device 130 substitutes the compensation sound source object into the object-based compensation operation to generate channel audio. The object-based compensation operation can utilize the object-based array calculation module in existing object-based acoustic products to perform a large number of array calculations related to acoustic interaction based on the relay data of the sound source object. For example, a relay data of the compensation sound source object includes: the coordinate position, size, reflectivity, and absorptivity of the environmental object 175. The control circuit 132 performs an object-based compensation operation based on the target listening point and the relay data to cancel the interference of the environmental object 175 on the target listening point and generate a first channel audio 112 and a second channel audio 122 optimized for the target listening point.
[0124] In one embodiment, the object basis compensation operation is performed on multiple sub-bands. Due to the characteristics of sound transmission, the sound pressure level on each sub-band has a different effect on the equivalent volume. Taking the effect of the first channel audio 112 generated by the first speaker 110 on the environmental object 175 as an example, the control circuit 132 in this embodiment can calculate the passive sound source effect generated by the environmental object 175 under the influence of the first channel audio 112 on multiple sub-bands based on the coordinate position, size, reflectivity, and absorptivity of the environmental object 175. Then, the control circuit 132 establishes the compensated sound source object based on the sound source effect. In this embodiment, the compensated sound source object is established correspondingly based on the environmental object 175, wherein the relay data has the same coordinate position, size, reflectivity, and absorptivity as the environmental object 175, but the sign of the generated sound source effect is opposite to that of the environmental object 175.
[0125] It is known that the audible range of the human ear is between 20 Hz and 20000 Hz. This embodiment can divide the audible range into multiple sub-band intervals and compensate for them separately. The size of each sub-band interval can be an exponential interval. For example, an exponential interval with a base of 10 can divide the audio signal into multiple sub-band ranges such as 10Hz to 100Hz, 100Hz to 1000Hz, and 1000Hz to 10000Hz. In other embodiments, the exponential intervals can also be divided with a base of 2 or a base of 4, depending on the required level of detail in playback quality. Equalizers in the field of audio processing already have techniques for dividing multiple sub-bands, which will not be explained in detail here.
[0126] After the control circuit 132 obtains the negative sound source effect of the compensated sound source object, it runs an object-based compensation operation, mixing the negative sound source effect into the first channel audio 112 and the second channel audio 122 according to the proportion determined by the mixing operation, thereby canceling the interference of the environmental object 175 on the target listening point. Regarding the object-based compensation operation, it will be discussed later. Figures 11 to 13 The embodiments are described in detail.
[0127] In process 316, the host device 130 outputs the first channel audio 112 and the second channel audio 122 to the first speaker 110 and the second speaker 120 respectively according to the calculation result of process 314. Figure 3 Process 316 and Figure 2 The process 216 in the embodiment is different. Figure 2 The compensation value is calculated based on the existing channel audio, and the existing channel audio is adjusted accordingly. During object-based compensation, the control circuit 132 directly calculates the channel audio for each speaker based on all relay data. The object-based compensation operation mixes the interference components that need to be canceled or compensated into the channel audio in the form of a compensation sound source. In other words, because the channel audio contains the compensation sound source emitted by the compensation sound source object, the user 180 does not perceive the influence of the environmental object 175 at the target listening point.
[0128] As shown in process 316, the object-based compensation operation translates the target listening point and environmental objects into relay data of the object-based acoustic system and establishes compensated sound source objects, simplifying the calculation process for eliminating interference and optimizing playback effects. It should be understood that the audio system 100 in this embodiment can dynamically update the target listening point by using the sensor circuit 140 to track the position of the user 180 in real time or periodically. The object-based compensation operation performed by the control circuit 132 can also synchronously update all relay data in the target space 170 related to the relative position of the target listening point as the target listening point changes.
[0129] Figure 3 The process 218 in this embodiment is the same as the previous embodiment, and will not be repeated here to save space.
[0130] The following is Figure 4 This describes an embodiment in which an audio system 100 dynamically tracks the user's position, runs a configuration program to obtain the sound field environment configuration, and compensates the speaker output using channel-based compensation.
[0131] Figure 4 This is a flowchart of a dynamic sound effect optimization method according to an embodiment of the present invention.
[0132] exist Figure 4 In the flowchart, the process located in the field belonging to a specific device represents the process performed by that specific device. For example, the part marked in the "Sensor Circuit" field is the process performed by sensor circuit 140; the part marked in the "Host Device" field is the process performed by host device 130; the part marked in the "Speaker" field is the process performed by first speaker 110 and / or second speaker 120; and so on. The aforementioned logic also applies to other subsequent flowcharts.
[0133] Figure 4 The processes 202, 204, 206 and 208 are the same as in the previous embodiment, and will not be repeated here to save space.
[0134] In this embodiment, when the audio system 100 completes process 210, the control circuit 132 has tracked the position of the user 180 and assigned it as the target listening point, and has also obtained the spatial configuration information of one or more environmental objects 175 in the target space 170. The next process is to use an object-based algorithm to adjust the channel audio of each speaker.
[0135] In order to eliminate interference in the sound field environment, the audio system 100 needs to obtain the spatial configuration information of various environmental objects 175 in the target space 170.
[0136] In process 410, the control circuit 132 in the host device 130 can run a configuration program to obtain spatial configuration information of one or more environmental objects 175 in the target space 170. In the previous embodiment, the host device 130 automatically identifies the spatial configuration information of the environmental objects 175 using sound field environment information captured by the sensor circuit 140. When running the configuration program, the host device 130 can interact with the user using a human-machine interface circuit 133, allowing the user to manually input the spatial configuration information of the environmental objects 175. The human-machine interface circuit 133 can provide a screen and an input method, allowing the user to define the spatial configuration information of various objects in the target space 170 in a two-dimensional plan view or a three-dimensional solid view. The spatial configuration information of the environmental objects 175 can include the relative position, size, name, and material type of the environmental objects 175 in the target space 170. In a further derived embodiment, the user 180 can tell the host device 130 the application scenario category to which the current target space 170 belongs through the human-machine interface circuit 133. In different application scenarios, such as open outdoor spaces, theater spaces, or bathrooms, the types of common environmental objects 175 are not the same, and the sound field atmosphere perceived by users is also different. Optimizing the sound field for different application scenarios is also one of the important functions of the audio system 100.
[0137] Different materials possess different acoustic properties. When the host device 130 runs the configuration program, it further queries an object database based on the object name or material type input by the user to obtain acoustic property information of the environmental object 175, such as its sound absorption or reflection rate. Therefore, in subsequent process 212, the host device 130 can calculate the degree to which the playback effect of each speaker at the target listening point is affected by the environmental object 175, based on the aforementioned spatial configuration information and acoustic property information. In further derived embodiments, the host device 130 can prioritize the use of the corresponding object database based on the application scenario category of the target space 170 to more quickly identify the environmental objects 175 in the target space 170. Related embodiments will be discussed later. Figure 9 The explanation is as follows.
[0138] Figure 4 The processes 212, 214, 216 and 218 are the same as in the previous embodiment, and will not be repeated here to save space.
[0139] Figure 4The embodiment illustrates that, in addition to dynamically tracking the user's position, the audio system 100 also allows the user 180 to configure the spatial configuration information of environmental objects 175 in the target space 170 via a configuration program. This configuration program provides a channel for manual input to compensate for any deficiencies in the recognition function. Besides actively inputting information to assist the host device 130 in making more accurate judgments, the user also has the opportunity to deliberately specify different application scenario categories or intentionally set imaginary virtual sound source objects to change the playback effect according to their preferences. The host device 130 will perform channel-based compensation operation, calculating the output compensation value corresponding to each speaker based on the spatial configuration information of the environmental objects 175 in the target space 170.
[0140] The following is Figure 5 This describes an embodiment in which an audio system 100 dynamically tracks the user's position, runs a configuration program to obtain the sound field environment configuration, and compensates the speaker output using object-based compensation.
[0141] Figure 5 This is a flowchart of a dynamic sound effect optimization method according to an embodiment of the present invention.
[0142] exist Figure 5 In the flowchart, the process located in the field belonging to a specific device represents the process performed by that specific device. For example, the part marked in the "Sensor Circuit" field is the process performed by sensor circuit 140; the part marked in the "Host Device" field is the process performed by host device 130; the part marked in the "Speaker" field is the process performed by first speaker 110 and / or second speaker 120; and so on. The aforementioned logic also applies to other subsequent flowcharts.
[0143] Figure 5 The processes 202, 204, 206, 208 and 210 are the same as in the previous embodiment, and will not be repeated here to save space.
[0144] and Figure 4 The implementation examples are similar, Figure 5 In order to eliminate interference in the sound field environment, the embodiment ran with Figure 4 The same process 410.
[0145] In process 410, the host device 130 runs a configuration program to obtain spatial configuration information of one or more environmental objects 175 in the target space 170. Figure 4In the embodiments described, the host device 130 can receive spatial configuration information of environmental objects 175 manually input by the user via a human-machine interface circuit 133. In further derivative embodiments, the host device 130 can also receive spatial configuration information transmitted from the user device 150 or other devices via a communication circuit 136. For example, the user device 150 may be a mobile phone running an application to provide functions similar to the human-machine interface circuit 133. This application allows the user to define the range and size of the target space 170, the position of each speaker relative to the target space 170, the position, size, name, and type of various environmental objects 175, and even the location of the user 180 itself. The application can also communicate with the control circuit 132 via the communication circuit 136 to perform various playback operations, such as play, pause, fast forward, and adjust the volume. In addition, the user can set the application scenario category of the target space 170 through the human-machine interface circuit 133, enabling the host device 130 to produce diversified playback effects on the target space 170.
[0146] In a further derived embodiment, the user device 150 connected to the host device 130 may be a virtual reality device or a game console. The user device 150 generates an audio signal, which the host device 130 then plays. This audio signal may contain virtual objects moving in a virtual reality space, such as an airplane or a fire-breathing dragon. The user device 150 can relay the data of these virtual objects to the host device 130, making it part of the environmental object space configuration information of the target space 170. In other words, the host device 130 can use an object-based acoustic system to treat virtual and physical objects equally. Through object-based compensation, the host device 130 can make the user perceive the presence of a virtual object in the target space 170, and also prevent the user from perceiving the interference of a physical object in the target space 170. Regarding the implementation of the object-based compensation operation, in Figures 11 to 13 Further details are provided in the embodiments.
[0147] In this embodiment, when the audio system 100 completes process 410, the control circuit 132 has tracked the position of the user 180 and assigned it as the target listening point, and has also obtained spatial configuration information of one or more environmental objects 175 in the target space 170. Then, in processes 312 to 316, the host device 130 uses an object-based algorithm to adjust the channel audio of each speaker. Since processes 312 to 316, and process 218, are the same as in the previous embodiment, they will not be described again for brevity.
[0148] Figure 5The embodiment illustrates that, in addition to dynamically tracking the user's position, the audio system 100 also allows the user 180 to set the spatial configuration information of environmental objects 175 in the target space 170 through a configuration program. This configuration program can be integrated with existing virtual reality technology to receive the spatial configuration information of virtual objects. The audio system 100 converts physical environmental objects and virtual objects into relay data of a consistent format, and then applies all relay data to the object substrate array computing module of the existing object substrate acoustic system to perform object substrate compensation operations. Therefore, the control circuit 132 does not need to develop additional computing modules for different objects, reducing costs and improving execution efficiency.
[0149] The following is Figure 6 This paper describes several implementation methods of the sensor circuit and explains the compensation algorithm for the channel substrate.
[0150] Figure 6 This is a schematic diagram of a target space 600 of the present invention, used to illustrate an embodiment of calculating audio adjustment based on the location of the optimal listening point.
[0151] The audio system 100 of this application uses a sensor circuit 140 to dynamically sense the target space 600 to generate sound field environment information. The sound field environment information mainly includes the position of the user 180, and may also include spatial configuration information of environmental objects. Various options are available for dynamic sensing techniques. For example, the sensor circuit 140 can be a combination of one or more of a camera 610, an infrared sensor 620, and a wireless detector 630, respectively configured at different locations around the target space 600, providing sound field environment information with spatial depth to help the recognition circuit 134 and control circuit 132 in the host device 130 more efficiently track the user 180's position. Thus, the recognition circuit 134, using the sound field environment information provided by the sensor circuit 140, can not only identify the user 180's position, but also the direction the face is facing, the position of the ears, and even gestures or body postures. This allows for richer control factors that can be applied to adjust the sound field, such as focus detection, sleep detection, and gesture control.
[0152] exist Figure 6 Within the target space 600, a first speaker 110 and a second speaker 120 are configured. Channel basis compensation operation can calculate the output compensation value for each speaker separately. Under default conditions, the target listening point is located at the center of the target space 600, i.e. Figure 6 The first position 601 is located at the same distance R1 from the first speaker 110 and the second speaker 120. At this time, the first speaker 110 and the first channel audio 112 played by the first channel audio 112 are also in the preset state, and no position compensation processing is required.
[0153] When user 180 moves from first position 601 to second position 602 along movement trajectory 173, sensor circuit 140 detects the new position of user 180 and assigns the target listening point of audio system 100 to second position 602. At this time, the distance between user 180 and first speaker 110 changes to R2, and the distance between user 180 and second speaker 120 changes to R2'. For user 180, first speaker 110 is farther away, so the received first channel audio 112 is attenuated due to distance. Conversely, second speaker 120 is closer, and the received second channel audio 122 is enhanced. In other words, the intensity of first channel audio 112 and second channel audio 122 received at second position 602 has become unbalanced. This embodiment uses a channel-based algorithm to restore the listening effect received at second position 602 to the same preset state as first position 601. In other words, the control circuit 132 compensates for the first channel audio 112 and the second channel audio 122 output by the first speaker 110 and the second speaker 120 to offset the listening effect deviation caused by the user 180's movement. Figure 6 The displayed target space of 600 is not limited to horizontally configured multi-speaker environments. Distance deviation issues also arise in three-dimensional sound field environments with upper and lower speakers. For example, if a user changes from a standing to a sitting position, they will move away from the upper speaker and closer to the lower speaker.
[0154] To achieve better compensation results, this embodiment uses equal loudness as the calculation standard. For example, this embodiment can calculate the sound pressure level that needs to be compensated at the target listening point based on the equal loudness curve defined by the ISO 226:2003 protocol. Each audio channel is divided into multiple sub-bands for separate processing. Furthermore, the sound field formula used varies depending on the distance between the user (180°) and the speaker. Since the equal loudness curve defines a linear relationship between equal loudness and sound pressure level, and the sum of the equal loudness and the gain value (in decibels) also has a linear relationship, this embodiment does not limit the adjustment to using only equal loudness, sound pressure level, or gain value as the unit.
[0155] In an audio system 100, the space through which sound is transmitted due to air vibration is called the sound field. Due to the existence of reflection, sound in a closed room can be classified into several types: (1) Near Field: When the user 180 is located relatively close to the sound source, the physical effects of the sound source (such as pressure, displacement, vibration) will enhance the sound. (2) Reverberant Field: Sound is reflected by objects, resulting in wave superposition. (3) Free Field: A sound field that is not affected by the aforementioned near field and reverberant field. The above reverberant field and free field can be collectively referred to as the far field.
[0156] In many modern audio systems, the definitions of near and far sound fields differ. For example, assuming R is the distance (in meters) between the speaker and the user (180 degrees), L is the speaker's width (in meters), and λ is the representative wavelength (in meters) of a sub-band signal, then the conditions for satisfying the far sound field include the following types:
[0157] R>>λ / 2π (1)
[0158] R>>L (2)
[0159] R>>πL 2 / 2λ (3)
[0160] by Figure 1 Taking the first speaker 110 as an example. When the distance between the target listening point and the first speaker 110 is greater than a certain proportion of the wavelength of the sub-band signal or the size of the first speaker 110, the audio system 100 determines that the sound field type is a far sound field. When the distance between the target listening point and the first speaker 110 is less than the specific proportion of the wavelength of the sub-band signal or the size of the first speaker 110, the sound field type is determined to be a near sound field. In a simpler implementation, the audio system 100 can define twice the wavelength (2λ) corresponding to the center frequency of a sub-band signal as the boundary point between the far sound field and the near sound field of the sub-band signal.
[0161] In the far sound field, the relationship between the change in sound pressure level of a sub-band signal received by user 180 from the speaker and the change in distance is as follows:
[0162] SPL2 = SPL1 - 20 log 10 (R2 / R1) (4)
[0163] Wherein, SPL2 is the sound pressure level of the sub-band signal received at the new location, SPL1 is the sound pressure level of the sub-band signal received at the original location, R2 is the distance between the new location and the speaker, and R1 is the distance between the original location and the speaker.
[0164] As can be seen from formula (4), the difference between SPL1 and SPL2 is the part of the speaker that needs to be compensated back.
[0165] SPL2' = SPL2 + 20 log 10 (R2 / R1) = SPL1 (5)
[0166] Where SPL2' is the sound pressure level of the sub-band signal received at the new compensated position. As can be seen from formula (5), this embodiment compensates back the changed part.
[0167] In the near-field sound field, the relationship between the change in sound pressure level of the sub-band signal received by user 180 from the speaker and the change in distance is as follows:
[0168] SPL2 = SPL1 - 10 log 10 (R2 / R1) (6)
[0169] SPL2' = SPL2 + 20 log 10 (R2 / R1) = SPL1 (7)
[0170] As can be seen from formulas (6) and (7), the rate of change of sound attenuation in the near sound field is more moderate than that in the far sound field, while the other calculation logic is the same.
[0171] It is understandable that the above formula may have exceptions in some special cases. For example, when user 180 moves from the first position 601 to the second position 602 and gets closer to the second speaker 120, the distance between user 180 and the second speaker 120 decreases from R1 to R2', which may cause the calculation result of formula (7) to become negative. However, the sub-band signal output by the second speaker 120 cannot be negative; it can only be reduced to the lowest audible value for the human ear. For example, the sound pressure level of the sub-band signal output by the second speaker 120 may be zero. On the other hand, when user 180 moves from the first position 601 to the second position 602 and gets away from the first speaker 110, the distance between user 180 and the first speaker 110 increases from R1 to R2. The maximum output limit of the first speaker 110 may not be able to satisfy formula (5). In this case, the sound system 100 can issue an over-limit warning to user 180.
[0172] Figure 6 The embodiments highlight the following advantages. Through the channel-based compensation algorithm, the user's optimal listening point is unaffected by movement. The channel-based calculation method is simple and efficient, and applicable to most target spaces.
[0173] Figure 6 The sound compensation method based on the user's 180° movement has already been explained. The following will use... Figure 7 The sound compensation method based on environmental object 175 is explained. The acoustic attribute information of environmental object 175 includes its reflectivity and absorptivity. In this embodiment, an appropriate calculation method is used to calculate the acoustic impact of environmental object 175 based on its spatial configuration information.
[0174] Figure 7 This is a schematic diagram of a target space 700 of the present invention, used to illustrate an embodiment of calculating the audio adjustment amount based on the absorption rate of environmental objects.
[0175] Figure 7 The image shows an environmental object 175 located between a first speaker 110 and a user 180 in a target space 700. For example, the environmental object 175 could be a sofa or a pillar. In this case, the environmental object 175 may obstruct or attenuate the listening experience for the user 180. In other words, the sound pressure level received by the user 180 from the first speaker 110 may be blocked or absorbed. When the control circuit 132 interprets this layout using spatial configuration information, it uses the absorption rate of the environmental object 175 to calculate the degree to which the playback effect of the first speaker 110 at the target listening point (the user 180's position) is affected by the environmental object 175, in order to determine the equivalent volume, sound pressure level, or gain value that the first channel audio 112 needs to output.
[0176] In one embodiment, the sound loss absorbed by the environmental object 175 can be calculated based on the sound pressure level received by the environmental object 175 from the first speaker 110:
[0177] A t [n] = R[n] * SPL t (8)
[0178] Where n represents the sub-band number. That is, the first channel audio 112 output by the first speaker 110 can be divided into multiple sub-bands and calculated separately. A t [n] represents the gain value of the nth sub-band detected at time point t. R[n] represents the absorption rate of the nth sub-band. t This represents the sound pressure level from the first speaker 110 experienced by the environmental object 175 at time point t. Time point t can represent the time difference between the sound being transmitted from the first speaker 110 to the environmental object 175.
[0179] From formula (8), we can see that A t[n] represents the gain value of the first channel audio 112 that is absorbed by the environmental object 175 in the nth sub-band, and also represents the output compensation value required for the nth sub-band of the first channel audio 112. Therefore, when the control circuit 132 generates the first channel audio 112 through the first speaker 110, it increases the gain value of the nth sub-band of the first channel audio 112 by the gain value A. t [n].
[0180] There may be various scenarios where the environmental object 175 is located between the first speaker 110 and the user 180. This embodiment primarily uses whether the line of sight between the first speaker 110 and the user 180 is obstructed, or further uses the line of sight between the first speaker 110 and the user 180's ears as the criterion. It is understood that SPLt itself is a function related to the distance and time between the environmental object 175 and the first speaker 110, and the degree of influence of the calculated At[n] on the user 180 is also a function related to the distance and time between the environmental object 175 and the user 180. After considering different angles of arrangement and distance relationships, various nonlinear correlations are involved. This application does not limit the derivative changes of formula (8), such as adding other weight coefficients, parameters, and offset correction values depending on the actual situation. For example, a sofa may be placed between the user 180 and the first speaker 110. Although the sofa does not obstruct the line of sight, it may still affect the sound pressure value received by the user 180 from the first speaker 110. The control circuit 132 can use formula (8) in combination with interpolation or other modified formulas to make the compensation result more in line with the requirements.
[0181] Figure 8 This is a schematic diagram of a target space 800 of the present invention, used to illustrate an embodiment of calculating audio adjustment based on the reflectivity of environmental objects.
[0182] Figure 8 The image shows a target space 800 where a user 180 is positioned between a first speaker 110 and an ambient object 175. The ambient object 175 could be a wall, ceiling, or floor. In this case, the ambient object 175 will reflect the first channel audio 112 output from the first speaker 110 back to the user 180. In other words, the sound pressure level received by the user 180 from the first speaker 110 will be superimposed or interfered with. When the control circuit 132 interprets this layout using spatial configuration information, it uses the reflectivity of the ambient object 175 to calculate the degree to which the playback effect of the first speaker 110 at the target listening point (the user 180's position) is affected by the ambient object 175, in order to determine the equivalent volume, sound pressure level, or gain value that the first channel audio 112 needs to output.
[0183] In this embodiment, the influence caused by the environmental object 175 can also be calculated according to formula (8), but R[n] is changed to represent the reflectivity of the environmental object 175 in the nth sub-band.
[0184] The result A of formula (8) t [n] can represent the component of the first channel audio 112 reflected to the user 180 by the environmental object 175 in the nth sub-band. Therefore, when the control circuit 132 generates the first channel audio 112 through the first speaker 110, it can appropriately reduce the gain value of the first channel audio 112 so that the total sound pressure level received by the user 180 from the first speaker 110 and the environmental object 175 is maintained at a preset level.
[0185] and Figure 7 The implementation examples are similar, Figure 8 The user 180 may be positioned between the first speaker 110 and the environmental object 175 in various possible scenarios. This embodiment primarily relies on whether the visible lines of sight of the first speaker 110 and the environmental object 175 are blocked by the user 180. However, in practice, walls, ceilings, and floors reflect light regardless of their angle. Therefore, the calculation formula in this embodiment is not limited to formula (8), and other nonlinear compensation calculation methods may be derived based on the arrangement and distance relationships. For example, the target space 800 can be classified into different application scenarios due to the characteristics of the wall, ceiling, and floor materials, as well as the size, shape, and layout of the room, such as a living room, study, bathroom, theater, or outdoors. The host device 130 can first classify the application scenario to which the target space 800 belongs, and then use the corresponding parameters or formulas respectively.
[0186] Figure 7 and Figure 8 The embodiments highlight the following advantages. Through channel-floor compensation, the influence of environmental objects 175 on the listening experience of the user 180 is eliminated. Channel-floor compensation can flexibly apply different acoustic properties of environmental objects based on their configuration, effectively addressing optimization problems in various complex environments.
[0187] In summary, the identification circuit 134 can receive data from the sensor circuit 140 to identify the position of the user 180 in the target space 170, so that the control circuit 132 can dynamically assign the position of the user 180 as the target listening point. The compensation made by the control circuit 132 for movement of the target listening point has been... Figure 6 The embodiments and formulas (4) to (7) are described. The compensation made by the control circuit 132 for interference from the environmental object 175 has been described in the embodiments and formulas (4) to (7). Figures 7 to 8As explained in Formula (8). These two compensation operations can be performed separately and applied to the channel audio. In other words, the final output optimized channel audio includes compensation values for the movement of the target listening point, as well as compensation for interference from environmental objects 175.
[0188] The recognition circuit 134 identifies the user 180's position based on the sound field environment information captured by the sensor circuit 140. The recognition process may also include the identification of the application scenario to help accelerate subsequent calculations by the control circuit 132. The following uses... Figure 9 This describes the process by which the host device 130 identifies objects based on the application scenario category.
[0189] Figure 9 This is a flowchart illustrating the object recognition process of a host device 130 according to an embodiment of the present invention. Environmental objects appearing in different application scenarios typically exhibit significant group-related acoustic properties, and the sound field reflection coefficients caused by surrounding environmental materials or room size also differ. Therefore, pre-distinguishing application scenario categories helps the audio system 100 improve the efficiency of sound field optimization. It is understood that... Figure 9 Each process in the process is executed by the host device 130, but it is not limited to being executed by a single circuit or module; it can also be the coordinated operation of multiple circuits.
[0190] In process 902, the host device 130 acquires the application scenario category of the target space 170. The host device 130 can acquire the application scenario category in several different ways. In one embodiment, the recognition circuit 134 in the host device 130 can determine an applicable application scenario category based on the sound field environment information provided by the recognition sensor circuit 140. In another embodiment, the control circuit 132 in the host device 130, while obtaining spatial configuration information of environmental objects through a configuration program run by the human-machine interface circuit 133, also simultaneously obtains the application scenario category defined by the user 180 through the configuration program. In a further derived embodiment, the control circuit 132 in the host device 130 can obtain relevant information about the application scenario category from a user device 150 through the communication circuit 136.
[0191] In process 904, to accelerate the query of environmental objects and improve accuracy, the host device 130 prioritizes relevant object databases based on application scenario categories. Object databases are typically pre-established data sets that can be provided by various pipelines. For example, the storage circuit 131 in the host device 130 can pre-store one or more object databases corresponding to different application scenarios. In another embodiment, the host device 130 can connect to a remote database 160 using the communication circuit 136. The remote database 160 may contain multiple object databases corresponding to different application scenarios. Each object database contains the external feature information and acoustic attribute information of multiple environmental objects.
[0192] After the host device 130 obtains the application scenario category in process 902, it can preferentially select an object database related to that application scenario category from the storage circuit 131 or the remote database 160 for subsequent identification of environmental objects. In one embodiment, the identification circuit 134 analyzes the sound field environment information provided by the sensor circuit 140 to obtain one or more object shape feature information, and retrieves the object database based on the object shape feature information to identify environmental objects that match the object shape feature information, including name, absorptivity, and reflectivity. In another embodiment, the control circuit 132 executes a configuration program and obtains the name of an environmental object using the human-machine interface circuit 133. The control circuit 132 searches the object database based on the name of the environmental object 175 to obtain the absorptivity and reflectivity corresponding to the environmental object.
[0193] In further derived embodiments, the parameters used in the search process can be combined in multiple ways. For example, during the analysis of sound field environment information, the recognition circuit 134 can obtain external features such as the material, size, and shape of the environmental object 175. The recognition circuit 134 transmits this external feature information to the object database for multi-condition cross-comparison to obtain a list of candidate objects sorted according to matching scores. If application scenario category information is used as a search condition during the search of the object database, it will help narrow down the possible range, accelerate the recognition, and improve accuracy.
[0194] In process 906, control circuit 132 retrieves the absorptivity and reflectivity of environmental objects from the object database selected in process 904. In practice, the acoustic attribute information of environmental objects stored in the object database is not limited to being stored in multiple independent object databases. The object database can be a relational database, containing multiple fields connected together by correlation coefficients. For example, the fields of the object database can include object name, application scenario category, material, absorptivity, reflectivity, and even external features such as shape, color, and gloss. The field values corresponding to each environmental object are not limited to a one-to-one relationship, but can be one-to-many or many-to-one. The values stored in each field are not necessarily absolute values, but range values or probability values. In further derivative implementations, the object database can be an adaptive database that can be continuously iteratively corrected by machine learning. User 180 can train the object database by providing feedback on preferred settings through human-machine interface circuit 133.
[0195] In process 908, control circuit 132 adjusts the channel audio in multiple sub-bands based on the search results and configuration of environmental objects. The acoustic properties of environmental objects 175 may vary significantly across different frequency bands. For example, a sofa may absorb a large amount of high-frequency signals but not affect the penetration of low-frequency signals. Therefore, the absorption rate or reflectivity retrieved from the object database can be an array value corresponding to multiple sub-bands or a frequency response curve. The size or separation method of the sub-bands can be determined according to design requirements and is not limited in this embodiment. Control circuit 132 adjusts the gain value of the channel audio in multiple sub-bands, which can be simulated as an equalizer or filter concept in practice. In other words, control circuit 132 can implement an equalizer for each speaker in the audio system 100 and customize the equalizer according to the output compensation value calculated in the aforementioned embodiment, so that the corresponding channel audio is adjusted. Further embodiments regarding the calculation of output compensation values will be discussed later. Figure 10 The explanation is as follows.
[0196] In process 910, the control circuit 132 outputs the adjusted channel audio to the corresponding speaker through the audio transmission circuit 135. The implementation of the audio transmission circuit 135 has been described in [the following text is missing from the original] Figure 1 As previously explained, I will not repeat myself here.
[0197] Figure 9 The embodiments highlight the following advantages: Object recognition operations can be performed according to application scenario categories (automatic recognition or manual input) to increase recognition efficiency. The object database employs a scalable architecture, continuously enhancing recognition capabilities over the long term with feedback from cloud-based big data services and machine learning. The audio system 100 can apply the concept of an equalizer to divide channel audio into multiple sub-bands for separate processing, effectively improving the final synthesized sound quality.
[0198] The following Figure 10 This further explains how the control circuit 132 calculates the output compensation value for each channel based on the spatial configuration information of the environmental objects 175.
[0199] Figure 10 This is a flowchart of an audio processing method according to an embodiment of the present invention, illustrating an embodiment of calculating output compensation values based on the positional relationships of environmental objects. Figure 10 The process is mainly executed by the control circuit 132 in the host device 130.
[0200] In process 1002, control circuit 132 determines the relative positional relationship between environmental objects, the target listening point, and the speaker. Multiple speakers and multiple environmental objects 175 in the target space 170 can be arranged and combined with the target listening point to form multiple sets of positional relationships. Each set of positional relationships includes one speaker, one environmental object 175, and the target listening point. Control circuit 132 checks and judges each combination of positional relationships in the target space 170 and calculates the corresponding output compensation value. The following uses one set of positional relationships in the audio system 100 as an example to illustrate the compensation method taken by control circuit 132 for interference caused by an environmental object 175 to a speaker at the target listening point.
[0201] In process 1004, control circuit 132 determines whether environmental object 175 is between the target listening point and the speaker. The position of environmental object 175 in the target space 170 can also be obtained by recognition circuit 134, or by human-machine interface circuit 133 through a configuration program. After integrating the above information, control circuit 132 can determine the relative positional relationship between each environmental object 175, the target listening point, and each speaker, and perform corresponding compensation calculations for each speaker. The situation to be determined in process 1004 is as follows: Figure 7 The situation is as shown. If the situation is met, proceed to process 1008. If the situation is not met, proceed to process 1006.
[0202] In process 1006, control circuit 132 determines whether the target listening point is located between environmental object 175 and the speaker. The condition to be determined in process 1006 is as follows: Figure 8 The situation is shown. If the situation is met, proceed to process 1010. If the situation is not met, proceed to process 1012.
[0203] In process 1008, control circuit 132 uses the absorption rate of environmental object 175 to calculate the output compensation value of the channel audio. In a preferred embodiment, the output compensation value of the speaker's channel audio is calculated separately for multiple sub-bands. Detailed calculations can be found in [reference needed]. Figure 7The target space 700 and formula (8). The control circuit 132 can find the absorption rate of the environmental object 175 from the object database and substitute it into formula (8) to obtain the output compensation value.
[0204] In process 1010, control circuit 132 uses the reflectivity of environmental object 175 to calculate the output compensation value for the channel audio. (Reference) Figure 8 Given the target space 800 and formula (8), the control circuit 132 can search for the reflectivity of the environmental object 175 from the object database and substitute it into formula (8) to obtain the output compensation value.
[0205] Understandably, the output compensation value calculated based on the absorptivity of the ambient object 175 may amplify the gain, sound pressure level, or equivalent volume of the adjusted channel audio to compensate for the absorbed energy. Conversely, the output compensation value calculated based on the reflectivity of the ambient object 175 may reduce the gain, sound pressure level, or equivalent volume of the adjusted channel audio to balance the reflected energy. In other words, the output compensation values calculated based on absorptivity and reflectivity are usually opposite in sign.
[0206] In process 1012, if environmental object 175 does not meet the conditions of process 1004 or process 1006, then control circuit 132 can determine that environmental object 175 is located in a position that will not affect the speaker's playback to the target listening point. In this case, control circuit 132 may not calculate the impact of environmental object 175 on the speaker and the target listening point for this set of positional relationships. However, it should be understood that a target space 170 typically contains multiple speakers. Environmental object 175 may not affect the playback of one speaker to the target listening point, but it may still affect the playback of other speakers to the target listening point. In other words, control circuit 132 needs to perform separate calculations for each set of positional relationships in the target space 170. Figure 10 The process.
[0207] In certain specific cases, the presence of environmental object 175 can be directly ignored. For example, if the reflectivity or absorption rate of sound by environmental object 175 is less than a certain threshold, its presence in the target space 170 can be ignored. On the other hand, if the control circuit 132 determines that the volume of environmental object 175 is less than a certain size, the presence of environmental object 175 can also be ignored.
[0208] In a further derived embodiment, if more than one user is detected in the target space 170, the determination of the target listening point can be based on the center points of multiple users' locations, or selectively based on the location of one user. As for users not selected as target listening points, the host device 130 can simulate them as environmental objects, according to... Figures 7 to 8 The example processing.
[0209] Figure 10 The embodiments highlight the following advantages. Figure 10 The embodiments continue Figure 7 and Figure 8 This approach simplifies complex environmental problems into multiple linear relationships, which are then solved separately. For specific environmental objects 175, the computational complexity can be further reduced by neglecting certain factors.
[0210] Figure 11 This is a schematic diagram of a target space 1100 of the present invention, used to illustrate an embodiment of optimizing the sound field by object substrate compensation operation.
[0211] The target space 1100 includes multiple speakers, such as a first speaker 1110, a second speaker 1120, a third speaker 1130, and a fourth speaker 1140. When the sound system 100 operates based on object-based compensation, the control circuit 132 logically treats the target space 1100 as a spatial coordinate system. This spatial coordinate system can be two-dimensional or three-dimensional planar coordinates. For ease of explanation, Figure 11 The illustration is presented in a two-dimensional planar coordinate system that includes an X-axis and a Y-axis.
[0212] In target space 1100, user 180 is located at origin P0. Control circuit 132 assigns user 180 as the target listening point. Figure 3 As described in the embodiments, the object-based acoustic system is built upon an array of acoustic parameters. Each sound source object has relay data describing its type, location, size (length, width, height), and divergence. After object-based computation, the sound represented by a sound source object is assigned to one or more speakers for simultaneous playback, with each speaker playing a portion of the sound from the sound source object. In other words, the object-based acoustic system can use multiple speakers to simulate the physical presence of a sound source object. For example, through object-based compensation, a user 180 at the target listening point can hear a virtual sound source object 1105 moving along a movement trajectory 1103 from a first position P1 to a new first position P1'.
[0213] The object-based compensation operation in this embodiment optimizes the audio output of all speakers for the target listening point. The object-based compensation operation utilizes the array calculation module in the existing object-based acoustic system to parameterize various distance factors and sound field categories, and can perform calculations similar to formulas (4) to (7). For the audio system 100, the host device 130 only needs to apply the user 180's position information to the object-based compensation operation to optimize the audio output of all speakers for the target listening point.
[0214] In one embodiment, the control circuit 132 can define the target listening point as the origin of the entire spatial coordinate system. When the user 180 moves, the entire spatial coordinate system moves with the origin. In other words, the position of the virtual sound source object 1105 relative to the origin remains unchanged. When the control circuit 132 plays the effect of the virtual sound source object 1105 through object-based compensation operation, the relative position of the virtual sound source object 1105 perceived by the user 180 will not change with the movement of the user 180.
[0215] In the target space 1100 of this embodiment, there may be environmental objects 175 that could substantially affect the listening experience of the user 180. The control circuit 132 can... Figure 9 The process 902 obtains the spatial configuration information of environmental objects in the target space 1100, revealing that environmental object 175 is located at the second position P2. When user 180 moves, the origin of the entire spatial coordinate system changes accordingly. Although environmental object 175 does not move, its relative position to the origin changes. Therefore, it can be understood that in the spatial coordinate system after the movement, the coordinate values of environmental object 175 move in the opposite direction.
[0216] To counteract the interference caused by the environmental object 175 to the user 180, the control circuit 132 of this embodiment establishes a compensating sound source object based on the environmental object 175. The relay data of this compensating sound source object includes: the coordinate position and size of the environmental object 175, as well as its reflectivity and absorptivity. The reflectivity and absorptivity of the environmental object 175 can be determined by… Figure 9 The compensation sound source object is obtained through process 906. The compensation sound source object is regarded as the negative sound source object of the environment object 175 and is applied to the object base compensation operation, becoming a virtual sound source that can cancel out the environment object 175.
[0217] It is understandable that the essence of the compensating sound source object is the negative sound source object corresponding to the environmental object 175, and its position overlaps with the environmental object 175. Therefore, in Figure 11Unless otherwise specified, the four-speaker configuration of the target space 1100 is merely an example. In actual applications of the audio system 100, the number of speakers can be greater, even including a stereo configuration of upper and lower speakers. This description does not limit other possible configurations.
[0218] Figure 11 The embodiments illustrate the advantages of object-based compensation operations. Control circuit 132 converts information from the target space 1100 into a spatial coordinate system, simplifying complex multi-object interaction calculations into array operations of relay data. Setting the position of the moving user 180 as the origin of the spatial coordinate system completely eliminates the impact of user 180's movement on the processing of virtual objects, thus simplifying the computational process. This embodiment also proposes the concept of compensating for sound source objects, directly applying object-based compensation operations to counteract environmental object interference, eliminating the need for complex multi-channel interactive calculations.
[0219] The following is Figure 12 This section explains the ease of object base compensation operations and their possible derivative applications.
[0220] Figure 12 This is a schematic diagram of a target space 1200 of the present invention, used to illustrate an embodiment of optimizing the sound field by object substrate compensation operation.
[0221] The target space 1200 may contain multiple speakers, such as the first speaker 1210, the second speaker 1220, the third speaker 1230, the fourth speaker 1240, the fifth speaker 1250, and the sixth speaker 1260, arranged in a long strip sound field. Each speaker corresponds to an ID. When the user 180 is in the first position P1, the relay data of a virtual sound source object (not shown) is mapped to the IDs of the first speaker 1210 and the second speaker 1220. After the control circuit 132 performs object basis compensation, it causes the first speaker 1210 and the second speaker 1220 to play the first channel output 1212 and the second channel output 1222, allowing the user 180 to perceive the presence of the virtual sound source object. When the user 180 moves along the movement trajectory 1203 to the second position P2, the control circuit 132 recalculates the target listening point and maps the relay data of the virtual sound source object to the fifth speaker 1250 and the sixth speaker 1260. After the control circuit 132 performs object base compensation, it will cause the fifth speaker 1250 to play the fifth channel output 1252 and the sixth channel output 1262, so that the user 180 can feel that the virtual sound source object still exists to the left and right of the user 180 and does not leave with the movement of the user 180.
[0222] This embodiment primarily illustrates the flexible application and simplicity of object substrate compensation operations. In many special cases, sound field optimization can be achieved with only a small amount of computation. For example, if user 180 is located in a spherical sound field, the control circuit 132 only needs to perform rotation coordinate calculations to ensure that user 180 experiences a consistent sound field effect regardless of the direction they are facing.
[0223] The following is Figure 13 The basic logic of control circuit 132 when performing object base compensation operation is summarized.
[0224] Figure 13 This is a flowchart of an embodiment of the present invention for object base compensation operation, illustrating the concept of establishing a compensation sound source object.
[0225] In process 1304, control circuit 132 establishes a corresponding compensation sound source object based on environmental object 175. For user 180 located at the target listening point, the presence of environmental object 175 constitutes a physical sound source. Environmental object 175 may reflect sound emitted by a speaker to the target listening point. Environmental object 175 may also block or absorb some sound, causing attenuation of the sound emitted by a speaker to the target listening point. The compensation sound source object is a negative sound source object established for environmental object 175. When host device 130 substitutes the compensation sound source object into the object-based compensation operation to generate channel audio, the presence of environmental object 175 can be eliminated. The specific details of the object-based operation itself can utilize the calculation methods of existing object-based acoustic products, using relay data of the sound source object to perform a large number of related array operations. For example, a relay data of the compensation sound source object includes: the coordinate position and size of the environmental object 175, as well as its reflectivity and absorptivity.
[0226] In process 1306, control circuit 132 calculates the sound source effect of the compensation sound source object. In this embodiment, the compensation sound source object is established based on the environmental object 175, wherein the relay data has the same coordinate position, size, and reflectivity and absorptivity of sound as the environmental object 175, but the resulting sound source effect is the inverse gain value of the environmental object 175.
[0227] Figure 13 The embodiments can also be referred to in similar ways. Figure 7 and Figure 8 The calculation. Formula (8) can be derived into formula (9), which calculates the passively generated gain value of the environmental object 175 based on the sound pressure value received by the environmental object 175 from the first speaker 110:
[0228] A t [m][n]=R[n]*SPL t [m] (9)
[0229] Where m represents the speaker number and n represents the sub-band number. A t [m][n] represents the gain value of the nth sub-band due to the influence of the mth speaker. R[n] represents the absorption rate of the nth sub-band. SPL t [m] represents the sound pressure level of the environmental object 175 at time point t, which is received by the m-th speaker. Time point t represents the time difference between the sound transmission from the speaker to the environmental object 175. If the time difference is greater than a non-negligible range, it indicates that there is an echo in the target space 170.
[0230] As can be seen from formula (9), the calculation result for each environmental object includes an array of gain values for multiple speakers and multiple sub-bands at a given time point. The sound source effect of the compensated sound source object is the negative value of this gain value array. In other words, the object-based compensation operation based on formula (9) involves array operations involving the interactive arrangement and combination of parameters in multiple dimensions. For ease of explanation, the following explanation will use the gain value of one speaker and one sub-band at a given time point as an example.
[0231] Figure 13 Implementation examples and Figure 7 and Figure 8 Similar to other embodiments, this embodiment can use appropriate calculation methods to calculate the acoustic effects of environmental objects 175 based on their spatial configuration information. For example, if the target listening point is located between a speaker and the visible line of sight of the environmental object 175, the control circuit 132 calculates the sound source effect of the compensated sound source object based on the reflectivity of the environmental object 175. Conversely, if the environmental object 175 is located between the target listening point and the visible line of sight of the speaker, the control circuit 132 calculates the sound source effect of the compensated sound source object based on the absorptivity of the environmental object 175.
[0232] For example, when an environmental object 175 absorbs sound emitted by a speaker, reducing the volume received by the target listening point, the control circuit 132 creates a virtual sound source object at the coordinates of the environmental object 175 that produces the corresponding volume effect as compensation. Conversely, if an environmental object 175 reflects sound from a speaker, causing the target listening point to receive too much volume, the control circuit 132 creates a virtual sound source object with a negative gain value at the coordinates of the environmental object 175.
[0233] It is understandable that a visible line of sight is defined as a straight line connecting two objects in space. Since objects have a certain volume and area, the volume may be very large, and the obstruction of the visible line of sight may include partial obstruction and complete obstruction. This embodiment can be based on formula (9), and then multiplied by different weighting coefficients or added with different offset corrections depending on various situations.
[0234] In process 1308, control circuit 132 mixes the sound source effect of the compensated sound source object into the channel audio, causing the corresponding speaker to play it. When performing the object-based compensation operation, control circuit 132 can handle complex object correspondence array operations, mixing the multiple sound source signals assigned to each speaker into a corresponding channel audio. After applying the object-based compensation operation, the volume effect received at the target listening point will include the sound source effect generated by the compensated sound source object. In this way, interference caused by environmental objects 175 can be effectively canceled out by the compensated sound source object.
[0235] In process 1310, control circuit 132 determines whether the target listening point has moved to a new position. As described in process 208, the audio system 100 can continuously track the movement of user 180 and update the target listening point accordingly. If the target listening point has moved, process 1312 is performed. Otherwise, the playback operation of process 1308 continues.
[0236] In process 1312, control circuit 132 updates the relay data of the compensation sound source object. In this embodiment, control circuit 132 establishes an object base space with the target listening point as a coordinate origin. If the target listening point moves to a new position, control circuit 132 assigns this new position as the new coordinate origin of the object base space. The difference between the new coordinate origin and the original coordinate origin can be represented as a movement vector. The spatial coordinates of the environmental object 175 relative to the target listening point also change in the opposite direction with this movement vector. Control circuit 132 then updates the relay data of the compensation sound source object corresponding to the environmental object 175 based on this movement vector. In a further embodiment, all speakers in the object base space can also be regarded as an object, having corresponding IDs, relay data, and coordinate values.
[0237] In another embodiment, the audio system 100 is not limited to using the target listening point as the origin of the coordinate system. The audio system 100 may also use a fixed reference point as the origin of the object base space. When the relative position of the sound source object in the object base space changes, the control circuit 132 updates the coordinate values in the relay data of the sound source object accordingly.
[0238] Once process 1312 is completed, control circuit 132 repeats process 1308.
[0239] Figure 13The embodiments illustrate the advantages of object-based compensation operations. Control circuit 132 converts information from the target space 1100 into a spatial coordinate system, simplifying complex multi-object interaction calculations into array operations of relay data. Setting the position of the moving user 180 as the origin of the spatial coordinate system completely eliminates the impact of user 180's movement on the processing of virtual objects, thus simplifying the computational process. This embodiment also proposes the concept of compensating for sound source objects, directly applying object-based compensation operations to counteract environmental object interference, eliminating the need for complex multi-channel interactive calculations.
[0240] In a further derived embodiment, if the host device 130 itself does not have the ability to perform mixing operations on the object substrate, the control circuit 132 can provide the function of channel mapping by executing software, so that the calculation results of the object substrate can be correctly mapped to each speaker.
[0241] In summary, this application proposes a sound system 100 that can dynamically track the user's position to optimize the sound field and intelligently eliminate interference caused by environmental objects. The means of tracking the user's position can be the individual or combined use of various methods such as cameras, infrared sensors, or wireless detectors. The spatial configuration information of environmental objects 175 in the target space 170 can be obtained by recognizing images captured by a camera or by manual input by the user. The sound field optimization method can be channel-based compensation or object-based compensation. When calculating the impact of environmental objects 175 on the target listening point, the relative positional relationship between the environmental objects 175 and the speaker, and between the environmental objects 175 and the target listening point, can be considered, and different calculation methods can be used. When using object-based compensation, the control circuit 132 establishes a corresponding compensation sound source object for each environmental object 175, so that the final mixed channel audio eliminates the interference caused by the environmental objects 175 on the target listening point.
[0242] Certain terms are used in the specification and claims to refer to specific elements, and those skilled in the art may use different names to refer to the same element. This specification and claims do not distinguish elements by differences in name, but rather by differences in function. The term "comprising" as used in the specification and claims is an open-ended term and should be interpreted as "comprising but not limited to". Furthermore, the term "coupled" herein includes any direct and indirect connection means. Therefore, if the text describes a first element coupled to a second element, it means that the first element can be directly connected to the second element through electrical connection or signal connection methods such as wireless transmission or optical transmission, or indirectly electrically or signal-connected to the second element through other elements or connection means.
[0243] The use of "and / or" in this specification includes any combination of one or more of the listed items. Furthermore, unless otherwise specified in this specification, any singular term also includes the meaning of the plural form.
[0244] The above are merely preferred embodiments of the present invention. All equivalent changes and modifications made in accordance with the claims of the present invention shall fall within the scope of the present invention.
Claims
1. A sound system (100) capable of dynamically optimizing playback based on the user's position, comprising: A sensor circuit (140) is configured to dynamically sense a target space (170) and generate sound field environment information, wherein, The sound field environment information includes the location of a user in the target space (170); A first speaker (110) and a second speaker (120) are configured to play audio; A host device (130), coupled to the sensor circuit (140), the first speaker (110), and the second speaker (120), includes: An identification circuit (134) is configured to identify the user's location in the target space from the sound field environment information; A control circuit (132), coupled to the identification circuit (134), is configured to dynamically assign the user's location as a target listening point; An audio transmission circuit (135), coupled to the control circuit (132), the first speaker (110), and the second speaker (120), is configured to transmit audio; and A human-machine interface circuit (133), coupled to the control circuit (132), is configured to be controlled by the control circuit (132) to run a configuration program to obtain spatial configuration information and acoustic attribute information of an environmental object in the target space (170), wherein the acoustic attribute information of the environmental object includes at least one of the following: sound absorption rate, sound reflectivity, and resonant frequency. The control circuit (132) performs a channel basis compensation operation based on the target listening point and the spatial configuration information and acoustic attribute information of the environmental objects to generate a first channel audio (112) and a second channel audio (122) optimized for the target listening point. The control circuit (132) outputs the first channel audio (112) and the second channel audio (122) to the corresponding first speaker (110) and second speaker (120) respectively through the audio transmission circuit (135); The channel base compensation operation includes: The first channel audio (112) emitted by the first speaker (110) is split into multiple sub-band signals; Based on the wavelength of one of the multiple sub-band signals and the distance between the target listening point and the first speaker (110), it is determined whether the sound field type generated by the sub-band signal at the target listening point is a near sound field or a far sound field. When the distance between the target listening point and the first speaker (110) is greater than a certain proportion of the wavelength of the sub-band signal or the size of the first speaker (110), the sound field type is determined to be a far sound field; and When the distance between the target listening point and the first speaker (110) is less than the wavelength of the sub-band signal or a specific proportion of the size of the first speaker (110), the sound field type is determined to be a near sound field. When the control circuit (132) determines that the distance between the target listening point and the first speaker (110) changes from a first distance (R1) to a second distance (R2), and the sound field type belongs to the far sound field, it uses a far sound field formula to calculate the sound pressure value of the sub-band signal. The formula for the far sound field includes: SPL’=SPL+20log 10 (R2 / R1) Wherein, SPL is the sound pressure level of the sub-band signal before adjustment, SPL' is the sound pressure level of the sub-band signal after adjustment, R1 is the first distance, and R2 is the second distance; and If the sound pressure value of the adjusted sub-band signal is less than zero, the control circuit (132) makes the sound pressure value of the adjusted sub-band signal zero.
2. The audio system (100) as claimed in claim 1, wherein, The channel base compensation operation also includes, when the control circuit (132) determines that the distance between the position of the target listening point and the first speaker (110) changes from the first distance (R1) to the second distance (R2), and the sound field type belongs to the near sound field, using a near sound field formula to calculate the sound pressure value of the sub-band signal; The near-field formula includes: SPL’=SPL+10log 10 (R2 / R1) Wherein, SPL is the sound pressure level of the sub-band signal before adjustment, SPL' is the sound pressure level of the sub-band signal after adjustment, R1 is the first distance, and R2 is the second distance; and If the sound pressure value of the adjusted sub-band signal is less than zero, the control circuit (132) makes the sound pressure value of the adjusted sub-band signal zero.
3. The audio system (100) as claimed in claim 1, wherein, The spatial configuration information of the environmental object includes its location, size, and appearance features, while the acoustic attribute information of the environmental object includes its reflectivity and absorption rate of sound. When the control circuit (132) generates the first channel audio, if the target listening point is located between the first speaker (110) and the visible line of sight of the environmental object, the control circuit (132) calculates the degree to which the playback effect of the first speaker (110) at the target listening point is affected by the environmental object based on the reflectivity of the environmental object, in order to determine the sound pressure level of the first channel audio; and When the control circuit (132) generates the first channel audio, if the environmental object is located between the target listening point and the line of sight of the first speaker (110), the control circuit (132) calculates the degree to which the playback effect of the first speaker (110) at the target listening point is affected by the environmental object based on the absorption rate of the environmental object, so as to determine the sound pressure value of the first channel audio.
4. The audio system (100) as claimed in claim 1, wherein, The sensor circuit (140) includes a camera (610) configured to capture an image of the sound field environment of the target space (170); The recognition circuit (134) dynamically identifies the user's head position, face direction, or ear position based on the sound field environment image captured by the camera (610) to determine the user's position.
5. The audio system (100) as claimed in claim 1, wherein, The sensor circuit (140) also includes an infrared sensor (620) configured to capture thermal imaging data in the target space; The identification circuit (134) analyzes the movement trajectory of the thermal imaging data to dynamically determine the user's location.
6. The audio system (100) as claimed in claim 1, wherein, The sensor circuit (140) also includes a wireless detector (630) disposed in the target space and detects the wireless signal of an electronic device; The identification circuit (134) dynamically locates the position of the electronic device based on the characteristics of the wireless signal detected by the wireless detector (630); and The identification circuit (134) dynamically determines the user's location based on the location of the electronic device.
7. The audio system (100) as claimed in claim 1, wherein, The host device (130) also includes a storage circuit (131) coupled to the control circuit (132) and configured to store one or more object databases, wherein each object database corresponds to an application scenario category and contains the shape feature information and acoustic attribute information of multiple environmental objects. When the human-machine interface circuit (133) is controlled by the control circuit (132) to run the configuration program, it also acquires an application scenario category of the target space (170); and The control circuit (132) selects an object database related to the application scenario category from the storage circuit (131) according to the application scenario category to identify the environmental object and find the acoustic attribute information of the environmental object.
8. The audio system (100) as claimed in claim 1, wherein, The host device (130) also includes a communication circuit (136) coupled to the control circuit (132) and configured to be controlled by the control circuit (132) to connect to a remote database (160) corresponding to an application scenario category; the remote database (160) is configured to store one or more object databases, wherein each object database corresponds to an application scenario category and contains the shape feature information and acoustic attribute information of multiple environmental objects; When the human-machine interface circuit (133) is controlled by the control circuit (132) to run the configuration program, it also acquires an application scenario category of the target space (170); and The control circuit (132) selects an object database related to the application scenario category from the remote database (160) based on the application scenario category to identify the environmental object and search for the acoustic attribute information of the environmental object.
Citation Information
Patent Citations
Playback device configuration based on proximity detection
CN106105271A
Audio processing method and electronic equipment
CN111050269A
Distributed wireless speaker system with automatic configuration determination when new speakers are added
US20150208188A1