Sound system with dynamically adjustable target listening point and elimination of environmental object interference

By combining sensor circuits and control circuits, the listening point of the audio system is dynamically adjusted and the audio output is compensated, solving the problems of fixed position and interference from environmental objects in traditional audio systems, and achieving a flexible listening experience and optimized sound quality.

CN116261096BActive Publication Date: 2025-12-30REALTEK SEMICON CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210992774.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-03-20
Filing Date
2022-08-18
Publication Date
2025-12-30
Estimated Expiration
2042-08-18

AI Technical Summary

Technical Problem

Traditional audio systems cannot dynamically adjust to the optimal listening point, requiring users to settle into fixed positions, and environmental interference can lead to poor listening quality.

Method used

The system employs a sound system that includes sensor circuitry, speakers, and a main unit. The sensor circuitry captures sound field environment information, identifies the user's position, and dynamically adjusts the listening point. The control circuitry performs channel basis compensation and object basis compensation to counteract environmental object interference and optimize audio output.

Benefits of technology

This system enables the audio system to dynamically adjust the listening point based on the user's location, eliminates interference from environmental objects, optimizes the listening experience, and enhances the flexibility and sound quality of the audio system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116261096B_ABST
    Figure CN116261096B_ABST
Patent Text Reader

Abstract

An audio system dynamically optimizes playback based on a user position. A sensor circuit dynamically senses a target space to generate an acoustic environment information. A first speaker and a second speaker play audio. A host device identifies a user from the acoustic environment information and determines a user position of the user in the target space, and dynamically assigns the user position as a target listening point. The sensor circuit includes a camera that captures an acoustic environment image of the target space. A control circuit uses a human interface circuit to run a configuration program to obtain spatial configuration information and acoustic property information of an environmental object in the target space, and causes the control circuit to generate a first channel audio and a second channel audio that are optimized for the target listening point using object-based compensation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to audio processing technology, which is actually a kind of sound system that can dynamically adjust the playing effect according to the changes in the sound field space. BACKGROUND

[0002] The existing sound system includes a plurality of speakers arranged around a target space to form a surround sound field environment. Each speaker can output corresponding channel audio. When configuring the surround sound field environment, the installer of the sound system usually assigns a central area of the target space as the best listening point as the basis for installing a plurality of speakers. When a plurality of speakers simultaneously play a plurality of channel audios, a user located at the best listening point can obtain an immersive listening effect.

[0003] However, in a real environment, the listening effect of the user is easily affected by various variables. For example, in a traditional sound system, the range of the best listening point is regionally limited. When the user moves to an area outside the best listening point, although the plurality of channel audios output by the sound system can still be heard, the listening effect of the plurality of channel audios at the user's location may have been greatly compromised or completely disabled. In addition, the room layout, furniture position and material in the target space are all environmental objects that can interfere with the listening effect. For example, sofas, windows, and curtains can absorb or reflect part of the sound energy, distorting the channel audio received at the best listening point.

[0004] In other words, the traditional sound system cannot dynamically adjust the position of the best listening point, and the user is forced to limit movement to accommodate the position of the best listening point, which is indeed inconvenient. On the other hand, the channel audio can be distorted by environmental objects, making the range of the best listening point more limited or even disappearing. In this way, the sound field environment built at high cost loses its meaning. SUMMARY

[0005] Therefore, how to make the sound system dynamically adjust the best listening point as the user moves and eliminate the interference of environmental objects in the target space is a problem to be solved.

[0006] The present specification provides an embodiment of a sound system, which can dynamically optimize the playing effect according to the user position, wherein the sound system comprises a sensor circuit, a first speaker and a second speaker, and a host device. The sensor circuit is configured to dynamically sense a target space to generate a sound field environment information. The first speaker and the second speaker are configured to play audio. The host device is coupled to the sensor circuit, the first speaker and the second speaker, and comprises an identification circuit, a control circuit and an audio transmission circuit. The identification circuit is configured to identify a user from the sound field environment information and determine a user position of the user in the target space. The control circuit is coupled to the identification circuit and configured to dynamically assign the user position as a target listening point. The audio transmission circuit is coupled to the control circuit, the first speaker and the second speaker, and configured to transmit audio. The sensor circuit comprises a camera configured to capture a sound field environment image of the target space. The identification circuit analyzes the sound field environment image to obtain spatial configuration information and acoustic attribute information of an environmental object in the target space. The control circuit performs channel base compensation operation according to the target listening point, and the spatial configuration information and the acoustic attribute information of the environmental object to generate a first channel audio and a second channel audio optimized for the target listening point. Finally, the control circuit outputs the first channel audio and the second channel audio to the corresponding first speaker and the second speaker through the audio transmission circuit.

[0007] The present disclosure provides an embodiment of a sound system, which can dynamically optimize the playing effect according to the user position, wherein the sound system comprises a sensor circuit, a first speaker and a second speaker, and a host device. The sensor circuit is configured to dynamically sense a target space to generate a sound field environment information. The first speaker and the second speaker are configured to play audio. The host device is coupled to the sensor circuit, the first speaker and the second speaker, and comprises an identification circuit, a control circuit and an audio transmission circuit. The identification circuit is configured to identify a user from the sound field environment information and determine a user position of the user in the target space. The control circuit is coupled to the identification circuit and configured to dynamically assign the user position as a target listening point. The audio transmission circuit is coupled to the control circuit, the first speaker and the second speaker, and configured to transmit audio. The sensor circuit comprises a camera configured to capture a sound field environment image of the target space. The identification circuit analyzes the sound field environment image to obtain spatial configuration information and acoustic attribute information of an environmental object in the target space. The control circuit corresponds the target space to an object base space, and correspondingly establishes a compensation sound source object in the object base space according to the environmental object. The relay data of the compensation sound source object comprises coordinate position, size, and reflectivity and absorptivity of the environmental object. The control circuit performs an object base compensation operation according to the target listening point and the relay data, to offset the interference of the environmental object to the target listening point and generate a first channel audio and a second channel audio optimized for the target listening point. The control circuit outputs the first channel audio and the second channel audio to the corresponding first speaker and second speaker through the audio transmission circuit.

[0008] The present disclosure provides an embodiment of a sound system that dynamically optimizes playback based on a user position. The sound system includes a sensor circuit, a first speaker and a second speaker, and a host device. The sensor circuit is configured to dynamically sense a target space to generate a sound field environment information. The first speaker and the second speaker are configured to play audio. The host device is coupled to the sensor circuit, the first speaker and the second speaker, and includes an identification circuit, a control circuit, an audio transmission circuit, and a human-machine interface circuit. The identification circuit is configured to identify a user from the sound field environment information and determine a user position of the user in the target space. The control circuit is coupled to the identification circuit and configured to dynamically assign the user position as a target listening point. The audio transmission circuit is coupled to the control circuit, the first speaker and the second speaker, and configured to transmit audio. The human-machine interface circuit is coupled to the control circuit and configured to run a configuration program to obtain spatial configuration information and acoustic property information of an environmental object in the target space. The control circuit performs a channel basis compensation operation based on the target listening point, the spatial configuration information and the acoustic property information of the environmental object to generate a first channel audio and a second channel audio that are optimized for the target listening point. The control circuit outputs the first channel audio and the second channel audio to the first speaker and the second speaker, respectively, through the audio transmission circuit.

[0009] The present specification provides an embodiment of a sound system that dynamically optimizes playback based on a user's position. The sound system includes a sensor circuit, a first speaker and a second speaker, and a host device. The sensor circuit is configured to dynamically sense a target space to generate a sound field environment information. The first speaker and the second speaker are configured to play audio. The host device, coupled to the sensor circuit, the first speaker, and the second speaker, includes an identification circuit, a control circuit, an audio transmission circuit, and a human-machine interface circuit. The identification circuit is configured to identify a user from the sound field environment information and determine a user's position in the target space. The control circuit, coupled to the identification circuit, is configured to dynamically assign the user's position as a target listening point. The audio transmission circuit, coupled to the control circuit, the first speaker, and the second speaker, is configured to transmit audio. The human-machine interface circuit, coupled to the control circuit, is configured to run a configuration program to obtain spatial configuration information and acoustic property information of an environmental object in the target space. The control circuit maps the target space to an object base space and correspondingly establishes a compensation sound source object in the object base space based on the environmental object. A relay data of the compensation sound source object includes coordinate position, size, and reflectivity and absorptivity of sound of the environmental object. The control circuit performs an object base compensation operation based on the target listening point and the relay data to offset the interference of the environmental object to the target listening point and generate a first channel audio and a second channel audio optimized for the target listening point. The control circuit outputs the first channel audio and the second channel audio to the first speaker and the second speaker, respectively, through the audio transmission circuit.

[0010] One advantage of the above embodiment is that the sound system can dynamically track the user's position through the sensor and continuously optimize the playback for the user's position. The user does not need to compromise the fixed listening position to obtain the best experience.

[0011] Another advantage of the above embodiment is that the sound system can identify the environmental object in the target space and adjust the channel audio to offset the interference of the environmental object.

[0012] Other advantages of the present invention will be more fully understood from the following detailed description taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0013] Figure 1 A functional block diagram of a sound system according to an embodiment of the present invention.

[0014] Figure 2 A flowchart of a dynamic sound effect optimization method according to an embodiment of the present invention.

[0015] Figure 3A flowchart of a dynamic sound effect optimization method according to an embodiment of the present application.

[0016] Figure 4 A flowchart of a dynamic sound effect optimization method according to an embodiment of the present application.

[0017] Figure 5 A flowchart of a dynamic sound effect optimization method according to an embodiment of the present application.

[0018] Figure 6 A diagram of a target space according to an embodiment of the present application, for illustrating an embodiment of calculating an audio adjustment amount according to a position of a best listening point.

[0019] Figure 7 A diagram of a target space according to an embodiment of the present application, for illustrating an embodiment of calculating an audio adjustment amount according to an absorption rate of an environmental object.

[0020] Figure 8 A diagram of a target space according to an embodiment of the present application, for illustrating an embodiment of calculating an audio adjustment amount according to a reflectivity of an environmental object.

[0021] Figure 9 A flowchart of an object identification operation of a host device according to an embodiment of the present application.

[0022] Figure 10 A flowchart of an audio processing method according to an embodiment of the present application, for illustrating an embodiment of calculating an output compensation value according to a positional relationship of an environmental object.

[0023] Figure 11 A diagram of a target space according to an embodiment of the present application, for illustrating an embodiment of optimizing a sound field by an object base compensation operation.

[0024] Figure 12 A diagram of a target space according to an embodiment of the present application, for illustrating an embodiment of optimizing a sound field by an object base compensation operation.

[0025] Figure 13 A flowchart of an object base compensation operation according to an embodiment of the present application.

[0026] Symbol explanation

[0027] 100... sound system

[0028] 110... first speaker

[0029] 112... first channel audio

[0030] 120... second speaker

[0031] 122... second channel audio

[0032] 130... host device

[0033] 131... storage circuit

[0034] 132... control circuit

[0035] 133... human-machine interface circuit

[0036] 134... recognition circuit

[0037] 135... audio transmission circuit

[0038] 136... communication circuit

[0039] 140... sensor circuit

[0040] 150... user device

[0041] 160... remote database

[0042] 170... target space

[0043] 171... first position

[0044] 172... second position

[0045] 173... movement trajectory

[0046] 175... environmental object

[0047] 180... user

[0048] 202-218... processes

[0049] 312-316... processes

[0050] 410... process

[0051] 600... target space

[0052] 601... first position

[0053] 602... second position

[0054] 610... camera

[0055] 620... infrared sensor

[0056] 630... wireless detector

[0057] 700... target space

[0058] 800... target space

[0059] 902-910... processes

[0060] 1002-1012... processes

[0061] 1100... target space

[0062] 1103... object movement trajectory

[0063] 1105... virtual sound source object

[0064] 1110... first speaker

[0065] 1120... second speaker

[0066] 1130... third speaker

[0067] 1140... fourth speaker

[0068] P0... origin

[0069] P1... first position

[0070] P1 '... new first position

[0071] P2... second position

[0072] 1200... target space

[0073] 1201... target listening point

[0074] 1203... movement trajectory

[0075] 1210... first speaker

[0076] 1212... first channel output

[0077] 1220... second speaker

[0078] 1222... second channel output

[0079] 1230... third speaker

[0080] 1240... fourth speaker

[0081] 1250... fifth speaker

[0082] 1252... fifth channel output

[0083] 1260... sixth speaker

[0084] 1262... sixth channel output

[0085] 1304-1312... flow DETAILED DESCRIPTION

[0086] Embodiments of the present application will be described below with reference to the accompanying drawings. In the drawings, the same numbers represent the same or similar elements or method steps throughout.

[0087] Figure 1 A functional block diagram of a sound system 100 according to an embodiment of the present application.

[0088] The sound system 100 mainly comprises a host device 130 and a plurality of speakers. The host device 130 can control the plurality of speakers to play audio. The host device 130 can be a computer host, a stereo system, an embedded system, or a customized digital audio processing device. The host device 130 comprises a communication circuit 136, so that the host device 130 can be connected to a user device 150 via wired or wireless connection, and serve as an input channel for audio source signals or data.

[0089] The user device 150 can be a mobile phone, a computer, a TV stick, a game console, or other audio source providing devices, and provide music or sound stream to the host device 130 via the communication circuit 136. Further, the sound system 100 can use the communication circuit 136 to cooperate with the user device 150 or other multimedia devices, and form a home theater system with both video and audio functions. For example, the target space 170 can further comprise a projection screen, a screen, or a display (not shown), and display images under the control of the user device 150. For another example, the user device 150 can be a head-mounted virtual reality device. The user 180 can stand in the target space 170 and see images through the user device 150, and the host device 130 can be controlled by the user device 150 to play audio in synchronization with the images. The communication circuit 136 in the present embodiment can be (but not limited to) a High Definition Multimedia Interface (HDMI), a Sony / Philips Digital Interface Format (SPDIF), a wireless area network module, an Ethernet network module, a shortwave radio transceiver, or a Bluetooth Low Energy (BLE) version 4 or 5 evolution application, or a Universal Serial Bus (USB).

[0090] The host device 130 also includes an audio transmission circuit 135 for connecting multiple speakers and outputting multiple channels of audio to the speakers. The host device 130 can control the speakers through the audio transmission circuit 135 in a unidirectional digital or analog output, or a bidirectional synchronous communication protocol. The connection between the audio transmission circuit 135 and each speaker can be a wired interface, a wireless interface, or a combination of both. The wired interface can be, but is not limited to, a composite video and audio terminal, a digital transmission interface, or a high-definition multimedia interface. The wireless interface can be, but is not limited to, a wireless area network, a shortwave radio frequency transceiver, or a Bluetooth Low Energy version 4 or 5. In further embodiments, the audio transmission circuit 135 and the communication circuit 136 can be combined into a multifunctional bidirectional transmission interface module, as both circuits are interfaces for connecting external elements. The audio transmission circuit 135 and the communication circuit 136 can use various publicly available standard transmission technologies to connect and transmit between elements, which can increase the future expandability of the sound system 100 and reduce the replacement cost when an element is damaged.

[0091] Figure 1 The target space 170 in the sound system 100 can be understood as a three-dimensional space in which the user 180 can use the sound system 100. Each speaker can be configured at a different position in the target space 170 to play a channel of audio. The surround configuration of multiple speakers can create a surround sound environment in a target space 170. There are various standard specifications for the number and configuration of speakers. For example, in a 5.1 channel surround sound system, there are two front speakers, one center speaker, two surround channel speakers, and one subwoofer to create a surround sound space around a target listening point and play sound to the target listening point. In a 7.1 channel surround sound system, a pair of rear surround channel speakers are further configured behind the target listening point to provide a more three-dimensional sound field effect. In recent years, new specifications such as 5.1.2 channels and 7.2.2 channels have appeared, which include more speakers and specific directional channel configurations to achieve more realistic "panoramic sound", "sky sound effect", or "floor sound effect". For the convenience of explaining the technical features of the sound system 100 of the present embodiment, Figure 1Only the first speaker 110 and the second speaker 120 are shown for representation. The first speaker 110 receives and plays the first channel audio 112 provided by the host device 130, while the second speaker 120 receives and plays the second channel audio 122 provided by the host device 130. It must be understood that in practice, the sound system 100 of the present embodiment is not limited to only two speakers, but can be applied to 2.1 channel, 4.1 channel, 5.1 channel, 7.2 channel, or more channel configurations. Each speaker in the target space 170 can have different audio output specifications. For example, some speakers are good at outputting bass, while some speakers are good at outputting mid-high frequencies. The host device 130 can plan different characteristics of the sound field environment in the target space 170 according to different speaker specifications.

[0092] The term "channel" as referred to in the specification and claims refers to both physical channels and logical channels. A logical channel refers to an audio data stream transmitted within the system, while a physical channel refers to the source of the signal played by each speaker. In the present embodiment, the first channel audio 112 and the second channel audio 122 played by each speaker are physical channels, which can be the result of one or more logical channels being down-mixed. For example, a pair of headphones can have only two speakers, but can be able to hear the sound effects produced by multiple applications simultaneously. In other words, the sound effect data of multiple applications can be down-mixed by the system into two physical channels, and played as audible sound through the two speakers. Therefore, the first channel audio 112 and the second channel audio 122 in the present embodiment are not limited to audio signals containing only a single logical channel, but can also be audio signals mixed from multiple logical channels according to a predetermined ratio.

[0093] In Figure 1 The first speaker 110 and the second speaker 120 are configured on two sides of a target space 170 to play sound for a target listening point in the target space 170. The target listening point can be understood as a position in which the sound system 100 has the best playing effect. In some sound systems, the target listening point is also referred to as a listening sweet spot. In most cases, the target listening point is usually located in a specific area of the target space 170, such as a center point, an axis, a tangent plane, or an equivalent volume center of multiple speakers. In Figure 1In the target space 170, a first position 171 where the user 180 is located is used to represent a target listening point of the target space 170. When the user 180 moves from the first position 171 to a second position 172 along a moving trajectory 173, the listening effect received by the user 180 is deviated because the user 180 is away from the first speaker 110 and approaches the second speaker 120. The conventional sound system cannot track the movement of the user 180 and adjust the listening effect received at the second position 172 correspondingly. The solution proposed in the embodiment will be described later.

[0094] On the other hand, the target space 170 usually contains some environmental objects 175, such as sofas, tables, curtains, walls, ceilings, and floors. These environmental objects 175 will have different interference reactions to the sound played by the first speaker 110 and the second speaker 120 due to different materials, sizes, and positions. For example, a sofa or a curtain made of cloth will absorb sound, and a marble floor or wall will reflect sound. In other words, the presence of the environmental objects 175 will affect the first channel audio 112 and the second channel audio 122 received at the target listening point. The conventional sound system does not have the ability to identify the environmental objects 175 in the target space 170, nor does it have the function of compensating for the first channel audio 112 and the second channel audio 122 according to the size, material, and position of the environmental objects 175. The sound system 100 of the embodiment can calculate and eliminate the interference of all the environmental objects 175 in the target space 170 to the first channel audio 112 and the second channel audio 122. For the convenience of explanation, the sound system 100 of the embodiment is described below by taking only one environmental object 175 to explain the operation mode of the sound system 100. However, it must be understood that Figure 1 In the target space 170, a first position 171 where the user 180 is located is used to represent a target listening point of the target space 170. When the user 180 moves from the first position 171 to a second position 172 along a moving trajectory 173, the listening effect received by the user 180 is deviated because the user 180 is away from the first speaker 110 and approaches the second speaker 120. The conventional sound system cannot track the movement of the user 180 and adjust the listening effect received at the second position 172 correspondingly. The solution proposed in the embodiment will be described later. Figure 1 It must be understood that the target space 170 is not limited to contain only one environmental object 175. The solution to the interference of the environmental object 175 will be described later.

[0095] The host device 130 of the embodiment further includes a storage circuit 131. The storage circuit 131 can include a non-volatile memory for storing a related operating system, application software, or firmware required for the operation of the host device 130. The storage circuit 131 can also include a volatile memory for use as an operation memory of the control circuit 132. The host device 130 of the embodiment further includes a control circuit 132. The control circuit 132 can be a central processing unit, a digital signal processor, or a microcontroller. The control circuit 132 can read the pre-stored operating system, software, or firmware from the storage circuit 131 to control the host device 130, the first speaker 110, and the second speaker 120 to perform the audio playing operation. Further, the host device 130 of the embodiment uses the control circuit 132 to perform a series of sound field compensation operations to dynamically optimize the playing effect and solve the shortcomings that the conventional sound system cannot overcome.

[0096] To dynamically optimize playback at the target listening point, the audio system 100 of this embodiment includes a sensor circuit 140 configured to dynamically sense a target space 170 and generate sound field environment information. The sensor circuit 140 may be a component located outside the host device 130 and coupled to it. The sensor circuit 140 may be a combination of one or more of a camera 610, an infrared sensor 620, and a wireless detector 630. The form of the sound field environment information captured by the sensor circuit 140 may vary depending on how the sensor circuit 140 is implemented. For example, the sound field environment information may be a combination of one or more of images, pictures, thermal images, and radio wave imaging of the user and environmental objects. In one embodiment, the sensor circuit 140 is disposed around the target space 170. It is understood that although... Figure 1 Only one sensor circuit 140 is shown in the figure, but in practice, the audio system 100 may include multiple sensor circuits 140, which are respectively configured at different positions around the target space 170 to obtain more accurate sound field environment information.

[0097] In the host device 130 of this embodiment, an identification circuit 134 is included, coupled to the sensor circuit 140. The identification circuit 134 can identify key information affecting the sound field from the sound field environment information, enabling the control circuit 132 to dynamically adjust the first channel audio 112 and the second channel audio 122 played from the first speaker 110 and the second speaker 120. For example, the identification circuit 134 can identify a user from the sound field environment information and determine the user's position in the target space. Since the sound field environment information provided by the sensor circuit 140 can have various combinations, the identification circuit 134 can also implement different identification technology solutions accordingly. For example, when the sound field environment information is an image, the identification circuit 134 can use artificial intelligence recognition technology to distinguish the user in the image. Through the application of artificial intelligence, after analyzing the user in the image, the identification circuit 134 can further locate the user's head, face, and even ear positions. If the sensor circuit 140 can provide diverse information such as three-dimensional images with spatial depth, infrared thermal imaging, or wireless signals, it will help the recognition circuit 134 obtain more accurate recognition results.

[0098] To calculate the degree of interference caused by environmental objects 175 to the sound field environment, the host device 130 requires spatial configuration information and acoustic attribute information of the environmental objects 175. The spatial configuration information may include the size, position, shape, and various external features of the environmental objects 175. The acoustic attribute information may include material-related characteristics such as sound absorption rate, reflectivity, and resonant frequency. In one embodiment, the identification circuit 134, while identifying sound field environment information, can further identify the spatial configuration information of the environmental objects 175 in the target space 170 from the sound field environment information and search for acoustic attribute information. An object database is needed to identify environmental objects. In one embodiment, the storage circuit 131 in the host device 130 can also be used to store an object database. The object database may contain various external feature information for identifying environmental objects, as well as various acoustic attribute information corresponding to each environmental object. For example, when the host device 130 needs to calculate the degree of interference caused by an environmental object 175 to the sound field environment, it can first analyze the object name of the environmental object 175 through the identification circuit 134, and then the host device 130 reads the storage circuit 131 to find the absorptivity and reflectivity corresponding to the environmental object 175.

[0099] In practice, the recognition circuit 134 can be a custom processor chip, working in conjunction with the existing operating system, software, or firmware in the storage circuit 131 to perform the artificial intelligence recognition function. The recognition circuit 134 can also be a core or thread circuit of the control circuit 132, executing the existing artificial intelligence software product in the storage circuit 131 to achieve the recognition function. Alternatively, the recognition circuit 134 can be a memory module of a specific artificial intelligence software product, executed by the control circuit 132 to complete the recognition function.

[0100] The human-machine interface circuit 133 in the host device 130 allows the user to control the operation of the host device 130. The human-machine interface circuit 133 may include a display screen, buttons, a dial, or a touchscreen, allowing the user to perform basic audio system 100 control functions, such as adjusting volume, playing, and fast-forwarding / rewinding. In one embodiment, the control circuit 132 can also execute a configuration program through the human-machine interface circuit 133 to allow the user to set various sound field scenarios or to inform the host device 130 of the spatial configuration information of environmental objects 175 in the target space 170. For example, in this configuration program, the control circuit 132 uses the human-machine interface circuit 133 to receive object configuration data input by the user, such as the object name, type, size, and position of one or more environmental objects 175. After obtaining this spatial configuration information, the control circuit 132 then searches for the corresponding absorptivity and reflectivity from the object database stored in the storage circuit 131 for subsequent sound field compensation operations. In further derived embodiments, the human-machine interface circuit 133 may also be provided by the user equipment 150. The user can operate the configuration program using the user equipment 150, and finally the user equipment 150 transmits the setting results to the control circuit 132 through the communication circuit 136.

[0101] The host device 130 can also be connected to a remote database 160 via a communication circuit 136. In a further embodiment, the object database originally stored using the storage circuit 131 can also be stored via the remote database 160. When the host device 130 needs to calculate the degree of interference caused by an environmental object 175 to the sound field environment, it can first analyze the sound field environment information through the identification circuit 134 to obtain the object's feature value, and then use the communication circuit 136 to access the remote database 160 to find an environmental object 175 that matches the object's feature value and obtain the sound field attribute information of the environmental object 175. The remote database 160 can be a server located in the cloud or other systems, connected to the host device 130 via wired or wireless bidirectional network communication technology. In addition to providing search functions, the remote database 160 can also accept the upload of updated data to continuously expand the database content. For example, the host device 130 can communicate with the remote database 160 using Structured Query Language (SQL).

[0102] based on Figure 1Based on the system architecture, the audio system 100 proposed in this application can achieve at least the following technical effects. First, the audio system 100 can dynamically track the user's position as the target listening point. The audio system 100 can also dynamically acquire spatial configuration information of environmental objects as a basis for optimizing the sound field effect. Finally, the audio system 100 dynamically compensates the speaker output based on the user's position and the spatial configuration information of environmental objects to eliminate object interference and optimize the listening effect at the target listening point. The implementation of dynamically tracking the user's position can employ various technical solutions such as cameras, infrared sensors, or wireless positioning. The implementation of acquiring spatial configuration information of environmental objects can be automatic or manual. For example, the audio system 100 can use a camera to capture images and perform artificial intelligence recognition, or allow the user to manually input the environmental conditions through a configuration program. The implementation of compensating the speaker output can be based on several different algorithms. For example, this specification introduces the Channel Base algorithm and the Object Base algorithm.

[0103] The following is Figure 2 This describes an embodiment of an audio system 100 that dynamically tracks the user's position, uses a camera to acquire the sound field environment configuration, and compensates the speaker output using channel-based compensation.

[0104] Figure 2 This is a flowchart of a dynamic sound effect optimization method according to an embodiment of the present invention.

[0105] exist Figure 2 In the flowchart, the process located in the field belonging to a specific device represents the process performed by that specific device. For example, the part marked in the "Sensor Circuit" field is the process performed by sensor circuit 140; the part marked in the "Host Device" field is the process performed by host device 130; the part marked in the "Speaker" field is the process performed by first speaker 110 and / or second speaker 120; and so on. The aforementioned logic also applies to other subsequent flowcharts.

[0106] In process 202, the sensor circuit 140 dynamically senses the target space to generate sound field environment information. In this embodiment, the sound field environment information can be optical, thermal, or electromagnetic wave information in the target space 170. For example, the sensor circuit 140 may include a camera to continuously record video of the target space 170 or periodically capture still photos of the target space 170. In another embodiment, the sensor circuit 140 may also include an infrared sensor configured to capture thermal imaging data in the target space. The thermal imaging data generated by the infrared sensor, in addition to containing spatial depth information, is also extremely sensitive to temperature changes, making it particularly suitable for tracking user location. In another embodiment, the sensor circuit 140 may also include a wireless detector located in the target space to detect the wireless signal of an electronic device. When a user holds an electronic device, the wireless detector can detect the beacon time difference or the strength of the wireless signal of the electronic device as an auxiliary means of tracking the user's location. The electronic device can be a user's own mobile phone, a specially designed beacon generator, a head-mounted virtual reality device, a game controller, or a remote control for the audio system 100. It is understood that this embodiment does not limit the number of sensor circuits 140, nor does it limit the use of only one sensing scheme at a time. For example, the audio system 100 of this embodiment can employ multiple sensor circuits 140 operating collaboratively from different locations, or simultaneously employ one or more cameras, infrared sensors, and wireless detectors. In this way, the host device 130 can obtain more complete sound field environment information and achieve more accurate recognition results in subsequent processes.

[0107] In process 204, the sensor circuit 140 transmits the sensed sound field environment information to the host device 130. The sensor circuit 140 may continuously transmit data, such as video, or periodically transmit static data. The frequency at which the sensor circuit 140 transmits data can be adaptively determined based on the amount of information in the sound field environment, the tracking accuracy requirements, and the computing power of the host device 130. The sensor circuit 140 and the host device 130 may be connected via a dedicated line or via a communication circuit 136. In a further derived embodiment, the sensor circuit 140 may share the audio transmission circuit 135 with the speaker to transmit the sound field environment information to the host device 130 via the audio transmission circuit 135.

[0108] In process 206, the host device 130 determines the user's position based on the sound field environment information received from the sensor circuit 140. The recognition circuit 134 in the host device 130 can perform a recognition program on the sound field environment information, such as applying artificial intelligence. The recognition algorithm of the recognition circuit 134 varies depending on the sensing scheme of the sensor circuit 140. It is understood that the target space 170 and the user's position can be represented in two-dimensional or three-dimensional space. If the audio system 100 implements only a single sensor circuit 140, it can at least perceive position information in two-dimensional space. If the audio system 100 increases the number of sensor circuits 140 or uses a multi-sensor scheme, it can obtain depth information in three-dimensional space to more accurately determine the user's position or the user's head position. In one embodiment, the recognition circuit 134 can dynamically identify the user's head position, face direction, or ear position based on sound field environment images captured by a camera. In another embodiment, the recognition circuit 134 can analyze the movement trajectory of thermal imaging data generated by an infrared sensor to dynamically determine the user's position 180. For example, the identification circuit 134 can dynamically locate the coordinates of the electronic device in the target space 170 based on the characteristics of the wireless signal detected by the wireless detector. Using this coordinate value, the control circuit 132 can further infer the position of the user's ear.

[0109] In process 208, after the identification circuit 134 in the host device 130 analyzes the user's position, the control circuit 132 in the host device 130 dynamically assigns the user's position as the target listening point. For ease of description of the following embodiments, the target space 170 is described here as a two-dimensional coordinate space or a three-dimensional coordinate space, and the target listening point can be represented as a coordinate value in the target space 170. Depending on the arrangement of the multiple speakers, the range of the target listening point can be more than a single point; it can also be a surface or a three-dimensional area with length, width, and height. For example, after the identification circuit 134 analyzes the user's head position or ear position, the control circuit 132 can assign the user's head position or ear position as the target listening point. The control circuit 132 will then perform subsequent compensation operations to ensure that the playback effect obtained at the target listening point is not affected by the user's movement. In practice, the control circuit 132 compensates for the listening effect obtained at the target listening point by adjusting the first channel audio 112 and the second channel audio 122. It is understood that process 208 may be executed dynamically as the user's position changes. Therefore, process 208 is not limited to following... Figure 2 The execution sequence is shown. In other words, the target listening point can be updated in real time as the user's position changes. The specific adjustment calculations will be described later.

[0110] In process 210, the recognition circuit 134 in the host device 130 further recognizes the sound field environment information provided by the sensor circuit 140 to obtain spatial configuration information of environmental objects in the target space 170. In other words, the sound field environment information provided by the sensor circuit 140 can not only be used to determine the user's position, but also to determine the various environmental objects 175 present in the target space 170. In one embodiment, after the camera in the sensor circuit 140 captures a sound field environment image of the target space 170, the recognition circuit 134 analyzes the sound field environment image to identify one or more environmental objects 175 in the target space 170, as well as the spatial configuration information of these environmental objects 175. The spatial configuration information includes the size, position, shape, and appearance features of the environmental objects 175. The recognition circuit 134 can also determine the acoustic attribute information of each environmental object 175, such as its sound absorption rate and reflectivity, through artificial intelligence calculations or database retrieval. In a further derived embodiment, the recognition circuit 134 can also determine the application scenario category of the target space 170 based on the sound field environment image. The application scenario category can include theater, living room, bathroom, outdoors, etc. If the host device 130 knows the application scenario category of the target space 170, it can more quickly identify environmental objects 175 in the target space 170 and reduce false positives. Related embodiments will be discussed later. Figure 9 The explanation is as follows.

[0111] In process 212, the control circuit 132 in the host device 130 can calculate the degree to which the playback effect of a speaker at the target listening point is affected by environmental objects. The playback effect of a speaker at the target listening point can be defined as the equivalent loudness or sound pressure level (SPL) received from the speaker at that target listening point. The ISO 226 standard defines an equal loudness curve (Fletcher-Munson Curve), illustrating that the equivalent loudness perceived by a user in different sub-bands actually corresponds to different sound pressure levels. In one embodiment, the control circuit 132 can use the equal loudness curve as a standard reference for the playback effect to calculate the sound pressure level received at the target listening point under various conditions. The control circuit 132 can utilize the spatial configuration information and acoustic attribute information of the environmental object 175 to evaluate the interference caused by the environmental object 175 to the target listening point, in order to further calculate methods to eliminate the interference. The influence of the spatial configuration information and attribute information of the environmental object 175 includes many scenarios. For example, the larger the volume of the environmental object 175, the greater the interference coefficient it may have on the target listening point. Whether the position of the environmental object 175 obstructs the user 180 and the speaker also determines the degree to which the speaker is affected. Depending on the material, the environmental object 175 may absorb or reflect sound. Therefore, the control circuit 132 needs to select corresponding parameters or formulas for different acoustic properties to calculate the degree to which the speaker is affected.

[0112] In process 214, the control circuit 132 in the host device 130 employs channel-based compensation to calculate the required output compensation value for each channel audio of each speaker. The channel-based compensation operation calculates the playback effect on the target listening point separately for each channel audio. For example, the first channel audio 112 played by the first speaker 110 may lose energy due to interference from an environmental object 175 before being transmitted through the air to the target listening point. Changes in the target listening point's location also affect the sound pressure level generated by the first channel audio 112 at that point. Through channel-based compensation, the control circuit 132 can calculate the change in sound pressure level of the first channel audio 112 at the target listening point. In this embodiment, the control circuit 132 adds an output compensation value to the first channel audio 112 to offset the change in sound pressure level, restoring the first channel audio 112 received by the target listening point to its state before being affected. In other words, the output compensation value has the same numerical value as the change in sound pressure level, but with the opposite positive or negative polarity.

[0113] In process 216, control circuit 132 adjusts and outputs channel audio to the speakers based on the output compensation value. Since the adjusted channel audio has offset the effects of the user 180's displacement in the target space 170 and the interference caused by environmental objects 175, the listening effect perceived by the user 180 remains consistent. Taking the first speaker 110 and the second speaker 120 in the target space 170 as examples, control circuit 132 calculates and adjusts the sound pressure values ​​of different sub-bands in the first channel audio 112 and the second channel audio 122, thereby offsetting the equivalent volume deviation perceived by the user 180 due to movement. On the other hand, the control circuit (132) compensates the first channel audio 112 and the second channel audio 122 accordingly based on the change in sound pressure value caused by the position, size, and acoustic properties of the environmental objects 175 at the target listening point.

[0114] In process 218, each speaker receives channel audio from the host device 130 via audio transmission circuit 135. Taking the first speaker 110 and the second speaker 120 in the target space 170 as an example, the control circuit 132 outputs the first channel audio 112 and the second channel audio 122 to the corresponding first speaker 110 and second speaker 120 via audio transmission circuit 135. Thus, the first speaker 110 and the second speaker 120 correspondingly play the adjusted first channel audio 112 and second channel audio 122, providing an optimized listening experience for the user 180's target listening point. For ease of explanation, Figure 1 In the embodiment of target space 170, only two speakers and one environmental object 175 are shown. However, it is understood that in practice, the host device 130 may contain more than two speakers, and the number of environmental objects 175 is not limited to one. In further derived embodiments, each speaker may be good at outputting different audio ranges. For example, some speakers are mid-high frequency speakers, and some are subwoofer speakers. When adjusting the channel audio, the control circuit 132 can further adjust the corresponding output first channel audio 112 and second channel audio 122 according to the characteristics of different speakers.

[0115] The following is Figure 3 This describes an embodiment of an audio system 100 that dynamically tracks the user's position, uses a camera to acquire the sound field environment configuration, and compensates the speaker output using object-based compensation.

[0116] Figure 3 This is a flowchart of a dynamic sound effect optimization method according to an embodiment of the present invention.

[0117] exist Figure 3In the flowchart, the process located in the field belonging to a specific device represents the process performed by that specific device. For example, the part marked in the "Sensor Circuit" field is the process performed by sensor circuit 140; the part marked in the "Host Device" field is the process performed by host device 130; the part marked in the "Speaker" field is the process performed by first speaker 110 and / or second speaker 120; and so on. The aforementioned logic also applies to other subsequent flowcharts.

[0118] Figure 3 The processes 202, 204, 206, 208 and 210 are the same as in the previous embodiment, and will not be repeated here to save space.

[0119] When the audio system 100 in this embodiment completes process 210, the control circuit 132 has tracked the position of the user 180 and assigned it as the target listening point, and has also obtained the spatial configuration information of one or more environmental objects 175 in the target space 170. The object-based compensation operation will then be described in the subsequent process to adjust the channel audio of each speaker.

[0120] Object-based acoustic systems originated from virtual reality mixing technology and can simulate the movement of sound source objects using a limited number of physical speakers. Existing software products, such as Dolby Atmos, Spatial Audio Workstations, and Digital Spatial Reality, all fall under the category of object-based acoustic systems. Users can define the movement trajectory of sound source objects in a virtual space through a human-computer interface. The object-based system then uses physical speakers to simulate the sound effects of these objects within the virtual space. Users at the target listening point can thus realistically perceive the movement of sound source objects in space.

[0121] An object-based acoustic system is built upon an array of acoustic parameters. Each sound source object has relay data describing its type, location, size (length, width, height), and divergence. After the object-based array calculation, the sound represented by a sound source object is assigned to one or more speakers for simultaneous playback, with each speaker playing a portion of the sound from the sound source object. In other words, the object-based array calculation can utilize multiple speakers to simulate the spatial effect of a single sound source object. Figure 3 The embodiments propose an object substrate compensation operation based on an object substrate acoustic system to solve the traditional playback effect problem.

[0122] In process 312, the control circuit 132 in the host device 130 establishes a compensating sound source object for the object base based on the environmental object 175. In practice, the control circuit 132 first maps the target space 170 to an object base space in virtual reality, and then establishes a compensating sound source object in the object base space corresponding to the environmental object 175 to generate a sound source effect that cancels out the environmental object 175. For the user 180 located at the target listening point, the presence of the environmental object 175 can also be simulated as a sound source object. In practical applications, the environmental object 175 may reflect the sound emitted by a speaker to the target listening point. The environmental object 175 may also block or absorb some sound, causing the sound emitted by a speaker to the target listening point to be attenuated. In other words, after the control circuit 132 of this embodiment simulates the environmental object 175 as a sound source object, it can correspondingly establish a negative sound source object with the opposite sound source effect in the object base space as a means of canceling interference. The sound source effect described in this embodiment can be the sound pressure level, equivalent volume, or gain value generated for the target listening point.

[0123] In process 314, the host device 130 substitutes the compensation sound source object into the object-based compensation operation to generate channel audio. The object-based compensation operation can utilize the object-based array calculation module in existing object-based acoustic products to perform a large number of array calculations related to acoustic interaction based on the relay data of the sound source object. For example, a relay data of the compensation sound source object includes: the coordinate position, size, reflectivity, and absorptivity of the environmental object 175. The control circuit 132 performs an object-based compensation operation based on the target listening point and the relay data to cancel the interference of the environmental object 175 on the target listening point and generate a first channel audio 112 and a second channel audio 122 optimized for the target listening point.

[0124] In one embodiment, the object basis compensation operation is performed on multiple sub-bands. Due to the characteristics of sound transmission, the sound pressure level on each sub-band has a different effect on the equivalent volume. Taking the effect of the first channel audio 112 generated by the first speaker 110 on the environmental object 175 as an example, the control circuit 132 in this embodiment can calculate the passive sound source effect generated by the environmental object 175 under the influence of the first channel audio 112 on multiple sub-bands based on the coordinate position, size, reflectivity, and absorptivity of the environmental object 175. Then, the control circuit 132 establishes the compensated sound source object based on the sound source effect. In this embodiment, the compensated sound source object is established correspondingly based on the environmental object 175, wherein the relay data has the same coordinate position, size, reflectivity, and absorptivity as the environmental object 175, but the sign of the generated sound source effect is opposite to that of the environmental object 175.

[0125] It is known that the audible range of the human ear is between 20 Hz and 20000 Hz. This embodiment can divide the audible range into multiple sub-band intervals and compensate for them separately. The size of each sub-band interval can be an exponential interval. For example, an exponential interval with a base of 10 can divide the audio signal into multiple sub-band ranges such as 10Hz to 100Hz, 100Hz to 1000Hz, and 1000Hz to 10000Hz. In other embodiments, the exponential intervals can also be divided with a base of 2 or a base of 4, depending on the required level of detail in playback quality. Equalizers in the field of audio processing already have techniques for dividing multiple sub-bands, which will not be explained in detail here.

[0126] After the control circuit 132 obtains the negative sound source effect of the compensated sound source object, it runs an object-based compensation operation, mixing the negative sound source effect into the first channel audio 112 and the second channel audio 122 according to the proportion determined by the mixing operation, thereby canceling the interference of the environmental object 175 on the target listening point. Regarding the object-based compensation operation, it will be discussed later. Figures 11 to 13 The embodiments are described in detail.

[0127] In process 316, the host device 130 outputs the first channel audio 112 and the second channel audio 122 to the first speaker 110 and the second speaker 120 respectively according to the calculation result of process 314. Figure 3 Process 316 and Figure 2 The process 216 in the embodiment is different. Figure 2 The compensation value is calculated based on the existing channel audio, and the existing channel audio is adjusted accordingly. During object-based compensation, the control circuit 132 directly calculates the channel audio for each speaker based on all relay data. The object-based compensation operation mixes the interference components that need to be canceled or compensated into the channel audio in the form of a compensation sound source. In other words, because the channel audio contains the compensation sound source emitted by the compensation sound source object, the user 180 does not perceive the influence of the environmental object 175 at the target listening point.

[0128] As shown in process 316, the object-based compensation operation translates the target listening point and environmental objects into relay data of the object-based acoustic system and establishes compensated sound source objects, simplifying the calculation process for eliminating interference and optimizing playback effects. It should be understood that the audio system 100 in this embodiment can dynamically update the target listening point by using the sensor circuit 140 to track the position of the user 180 in real time or periodically. The object-based compensation operation performed by the control circuit 132 can also synchronously update all relay data in the target space 170 related to the relative position of the target listening point as the target listening point changes.

[0129] Figure 3 The process 218 in this embodiment is the same as the previous embodiment, and will not be repeated here to save space.

[0130] The following is Figure 4 This describes an embodiment in which an audio system 100 dynamically tracks the user's position, runs a configuration program to obtain the sound field environment configuration, and compensates the speaker output using channel-based compensation.

[0131] Figure 4 This is a flowchart of a dynamic sound effect optimization method according to an embodiment of the present invention.

[0132] exist Figure 4 In the flowchart, the process located in the field belonging to a specific device represents the process performed by that specific device. For example, the part marked in the "Sensor Circuit" field is the process performed by sensor circuit 140; the part marked in the "Host Device" field is the process performed by host device 130; the part marked in the "Speaker" field is the process performed by first speaker 110 and / or second speaker 120; and so on. The aforementioned logic also applies to other subsequent flowcharts.

[0133] Figure 4 The processes 202, 204, 206 and 208 are the same as in the previous embodiment, and will not be repeated here to save space.

[0134] In this embodiment, when the audio system 100 completes process 210, the control circuit 132 has tracked the position of the user 180 and assigned it as the target listening point, and has also obtained the spatial configuration information of one or more environmental objects 175 in the target space 170. The next process is to use an object-based algorithm to adjust the channel audio of each speaker.

[0135] In order to eliminate interference in the sound field environment, the audio system 100 needs to obtain the spatial configuration information of various environmental objects 175 in the target space 170.

[0136] In process 410, the control circuit 132 in the host device 130 can run a configuration program to obtain spatial configuration information of one or more environmental objects 175 in the target space 170. In the previous embodiment, the host device 130 automatically identifies the spatial configuration information of the environmental objects 175 using sound field environment information captured by the sensor circuit 140. When running the configuration program, the host device 130 can interact with the user using a human-machine interface circuit 133, allowing the user to manually input the spatial configuration information of the environmental objects 175. The human-machine interface circuit 133 can provide a screen and an input method, allowing the user to define the spatial configuration information of various objects in the target space 170 in a two-dimensional plan view or a three-dimensional solid view. The spatial configuration information of the environmental objects 175 can include the relative position, size, name, and material type of the environmental objects 175 in the target space 170. In a further derived embodiment, the user 180 can tell the host device 130 the application scenario category to which the current target space 170 belongs through the human-machine interface circuit 133. In different application scenarios, such as open outdoor spaces, theater spaces, or bathrooms, the types of common environmental objects 175 are not the same, and the sound field atmosphere perceived by users is also different. Optimizing the sound field for different application scenarios is also one of the important functions of the audio system 100.

[0137] Different materials possess different acoustic properties. When the host device 130 runs the configuration program, it further queries an object database based on the object name or material type input by the user to obtain acoustic property information of the environmental object 175, such as its sound absorption or reflection rate. Therefore, in subsequent process 212, the host device 130 can calculate the degree to which the playback effect of each speaker at the target listening point is affected by the environmental object 175, based on the aforementioned spatial configuration information and acoustic property information. In further derived embodiments, the host device 130 can prioritize the use of the corresponding object database based on the application scenario category of the target space 170 to more quickly identify the environmental objects 175 in the target space 170. Related embodiments will be discussed later. Figure 9 The explanation is as follows.

[0138] Figure 4 The processes 212, 214, 216 and 218 are the same as in the previous embodiment, and will not be repeated here to save space.

[0139] Figure 4The embodiment illustrates that, in addition to dynamically tracking the user's position, the audio system 100 also allows the user 180 to configure the spatial configuration information of environmental objects 175 in the target space 170 via a configuration program. This configuration program provides a channel for manual input to compensate for any deficiencies in the recognition function. Besides actively inputting information to assist the host device 130 in making more accurate judgments, the user also has the opportunity to deliberately specify different application scenario categories or intentionally set imaginary virtual sound source objects to change the playback effect according to their preferences. The host device 130 will perform channel-based compensation operation, calculating the output compensation value corresponding to each speaker based on the spatial configuration information of the environmental objects 175 in the target space 170.

[0140] The following is Figure 5 This describes an embodiment in which an audio system 100 dynamically tracks the user's position, runs a configuration program to obtain the sound field environment configuration, and compensates the speaker output using object-based compensation.

[0141] Figure 5 This is a flowchart of a dynamic sound effect optimization method according to an embodiment of the present invention.

[0142] exist Figure 5 In the flowchart, the process located in the field belonging to a specific device represents the process performed by that specific device. For example, the part marked in the "Sensor Circuit" field is the process performed by sensor circuit 140; the part marked in the "Host Device" field is the process performed by host device 130; the part marked in the "Speaker" field is the process performed by first speaker 110 and / or second speaker 120; and so on. The aforementioned logic also applies to other subsequent flowcharts.

[0143] Figure 5 The processes 202, 204, 206, 208 and 210 are the same as in the previous embodiment, and will not be repeated here to save space.

[0144] and Figure 4 The implementation examples are similar, Figure 5 In order to eliminate interference in the sound field environment, the embodiment ran with Figure 4 The same process 410.

[0145] In process 410, the host device 130 runs a configuration program to obtain spatial configuration information of one or more environmental objects 175 in the target space 170. Figure 4In the embodiments described, the host device 130 can receive spatial configuration information of environmental objects 175 manually input by the user via a human-machine interface circuit 133. In further derivative embodiments, the host device 130 can also receive spatial configuration information transmitted from the user device 150 or other devices via a communication circuit 136. For example, the user device 150 may be a mobile phone running an application to provide functions similar to the human-machine interface circuit 133. This application allows the user to define the range and size of the target space 170, the position of each speaker relative to the target space 170, the position, size, name, and type of various environmental objects 175, and even the location of the user 180 itself. The application can also communicate with the control circuit 132 via the communication circuit 136 to perform various playback operations, such as play, pause, fast forward, and adjust the volume. In addition, the user can set the application scenario category of the target space 170 through the human-machine interface circuit 133, enabling the host device 130 to produce diversified playback effects on the target space 170.

[0146] In a further derived embodiment, the user device 150 connected to the host device 130 may be a virtual reality device or a game console. The user device 150 generates an audio signal, which the host device 130 then plays. This audio signal may contain virtual objects moving in a virtual reality space, such as an airplane or a fire-breathing dragon. The user device 150 can relay the data of these virtual objects to the host device 130, making it part of the environmental object space configuration information of the target space 170. In other words, the host device 130 can use an object-based acoustic system to treat virtual and physical objects equally. Through object-based compensation, the host device 130 can make the user perceive the presence of a virtual object in the target space 170, and also prevent the user from perceiving the interference of a physical object in the target space 170. Regarding the implementation of the object-based compensation operation, in Figures 11 to 13 Further details are provided in the embodiments.

[0147] In this embodiment, when the audio system 100 completes process 410, the control circuit 132 has tracked the position of the user 180 and assigned it as the target listening point, and has also obtained spatial configuration information of one or more environmental objects 175 in the target space 170. Then, in processes 312 to 316, the host device 130 uses an object-based algorithm to adjust the channel audio of each speaker. Since processes 312 to 316, and process 218, are the same as in the previous embodiment, they will not be described again for brevity.

[0148] Figure 5The embodiment illustrates that, in addition to dynamically tracking the user's position, the audio system 100 also allows the user 180 to set the spatial configuration information of environmental objects 175 in the target space 170 through a configuration program. This configuration program can be integrated with existing virtual reality technology to receive the spatial configuration information of virtual objects. The audio system 100 converts physical environmental objects and virtual objects into relay data of a consistent format, and then applies all relay data to the object substrate array computing module of the existing object substrate acoustic system to perform object substrate compensation operations. Therefore, the control circuit 132 does not need to develop additional computing modules for different objects, reducing costs and improving execution efficiency.

[0149] The following is Figure 6 This paper describes several implementation methods of the sensor circuit and explains the compensation algorithm for the channel substrate.

[0150] Figure 6 This is a schematic diagram of a target space 600 of the present invention, used to illustrate an embodiment of calculating audio adjustment based on the location of the optimal listening point.

[0151] The audio system 100 of this application uses a sensor circuit 140 to dynamically sense the target space 600 to generate sound field environment information. The sound field environment information mainly includes the position of the user 180, and may also include spatial configuration information of environmental objects. Various options are available for dynamic sensing techniques. For example, the sensor circuit 140 can be a combination of one or more of a camera 610, an infrared sensor 620, and a wireless detector 630, respectively configured at different locations around the target space 600, providing sound field environment information with spatial depth to help the recognition circuit 134 and control circuit 132 in the host device 130 more efficiently track the user 180's position. Thus, the recognition circuit 134, using the sound field environment information provided by the sensor circuit 140, can not only identify the user 180's position, but also the direction the face is facing, the position of the ears, and even gestures or body postures. This allows for richer control factors that can be applied to adjust the sound field, such as focus detection, sleep detection, and gesture control.

[0152] exist Figure 6 Within the target space 600, a first speaker 110 and a second speaker 120 are configured. Channel basis compensation operation can calculate the output compensation value for each speaker separately. Under default conditions, the target listening point is located at the center of the target space 600, i.e. Figure 6 The first position 601 is located at the same distance R1 from the first speaker 110 and the second speaker 120. At this time, the first speaker 110 and the first channel audio 112 played by the first channel audio 112 are also in the preset state, and no position compensation processing is required.

[0153] When user 180 moves from first position 601 to second position 602 along movement trajectory 173, sensor circuit 140 detects the new position of user 180 and assigns the target listening point of audio system 100 to second position 602. At this time, the distance between user 180 and first speaker 110 changes to R2, and the distance between user 180 and second speaker 120 changes to R2'. For user 180, first speaker 110 is farther away, so the received first channel audio 112 is attenuated due to distance. Conversely, second speaker 120 is closer, and the received second channel audio 122 is enhanced. In other words, the intensity of first channel audio 112 and second channel audio 122 received at second position 602 has become unbalanced. This embodiment uses a channel-based algorithm to restore the listening effect received at second position 602 to the same preset state as first position 601. In other words, the control circuit 132 compensates for the first channel audio 112 and the second channel audio 122 output by the first speaker 110 and the second speaker 120 to offset the listening effect deviation caused by the user 180's movement. Figure 6 The displayed target space of 600 is not limited to horizontally configured multi-speaker environments. Distance deviation issues also arise in three-dimensional sound field environments with upper and lower speakers. For example, if a user changes from a standing to a sitting position, they will move away from the upper speaker and closer to the lower speaker.

[0154] To achieve better compensation results, this embodiment uses equal loudness as the calculation standard. For example, this embodiment can calculate the sound pressure level that needs to be compensated at the target listening point based on the equal loudness curve defined by the ISO 226:2003 protocol. Each audio channel is divided into multiple sub-bands for separate processing. Furthermore, the sound field formula used varies depending on the distance between the user (180°) and the speaker. Since the equal loudness curve defines a linear relationship between equal loudness and sound pressure level, and the sum of the equal loudness and the gain value (in decibels) also has a linear relationship, this embodiment does not limit the adjustment to using only equal loudness, sound pressure level, or gain value as the unit.

[0155] In an audio system 100, the space through which sound is transmitted due to air vibration is called the sound field. Due to the existence of reflection, sound in a closed room can be classified into several types: (1) Near Field: When the user 180 is located relatively close to the sound source, the physical effects of the sound source (such as pressure, displacement, vibration) will enhance the sound. (2) Reverberant Field: Sound is reflected by objects, resulting in wave superposition. (3) Free Field: A sound field that is not affected by the aforementioned near field and reverberant field. The above reverberant field and free field can be collectively referred to as the far field.

[0156] In many modern audio systems, the definitions of near and far sound fields differ. For example, assuming R is the distance (in meters) between the speaker and the user (180 degrees), L is the speaker's width (in meters), and λ is the representative wavelength (in meters) of a sub-band signal, then the conditions for satisfying the far sound field include the following types:

[0157] R>>λ / 2π (1)

[0158] R>>L (2)

[0159] R>>πL 2 / 2λ (3)

[0160] by Figure 1 Taking the first speaker 110 as an example. When the distance between the target listening point and the first speaker 110 is greater than a certain proportion of the wavelength of the sub-band signal or the size of the first speaker 110, the audio system 100 determines that the sound field type is a far sound field. When the distance between the target listening point and the first speaker 110 is less than the specific proportion of the wavelength of the sub-band signal or the size of the first speaker 110, the sound field type is determined to be a near sound field. In a simpler implementation, the audio system 100 can define twice the wavelength (2λ) corresponding to the center frequency of a sub-band signal as the boundary point between the far sound field and the near sound field of the sub-band signal.

[0161] In the far sound field, the relationship between the change in sound pressure level of a sub-band signal received by user 180 from the speaker and the change in distance is as follows:

[0162] SPL2 = SPL1 - 20 log 10 (R2 / R1) (4)

[0163] Wherein, SPL2 is the sound pressure level of the sub-band signal received at the new location, SPL1 is the sound pressure level of the sub-band signal received at the original location, R2 is the distance between the new location and the speaker, and R1 is the distance between the original location and the speaker.

[0164] As can be seen from formula (4), the difference between SPL1 and SPL2 is the part of the speaker that needs to be compensated back.

[0165] SPL2' = SPL2 + 20 log 10 (R2 / R1) = SPL1 (5)

[0166] Where SPL2' is the sound pressure level of the sub-band signal received at the new compensated position. As can be seen from formula (5), this embodiment compensates back the changed part.

[0167] In the near-field sound field, the relationship between the change in sound pressure level of the sub-band signal received by user 180 from the speaker and the change in distance is as follows:

[0168] SPL2 = SPL1 - 10 log 10 (R2 / R1) (6)

[0169] SPL2' = SPL2 + 20 log 10 (R2 / R1) = SPL1 (7)

[0170] As can be seen from formulas (6) and (7), the rate of change of sound attenuation in the near sound field is more moderate than that in the far sound field, while the other calculation logic is the same.

[0171] It is understandable that the above formula may have exceptions in some special cases. For example, when user 180 moves from the first position 601 to the second position 602 and gets closer to the second speaker 120, the distance between user 180 and the second speaker 120 decreases from R1 to R2', which may cause the calculation result of formula (7) to become negative. However, the sub-band signal output by the second speaker 120 cannot be negative; it can only be reduced to the lowest audible value for the human ear. For example, the sound pressure level of the sub-band signal output by the second speaker 120 may be zero. On the other hand, when user 180 moves from the first position 601 to the second position 602 and gets away from the first speaker 110, the distance between user 180 and the first speaker 110 increases from R1 to R2. The maximum output limit of the first speaker 110 may not be able to satisfy formula (5). In this case, the sound system 100 can issue an over-limit warning to user 180.

[0172] Figure 6 The embodiments highlight the following advantages. Through the channel-based compensation algorithm, the user's optimal listening point is unaffected by movement. The channel-based calculation method is simple and efficient, and applicable to most target spaces.

[0173] Figure 6 The sound compensation method based on the user's 180° movement has already been explained. The following will use...Figure 7 This describes the sound compensation method based on environmental object 175. The acoustic properties of environmental object 175 include its reflectivity and absorptivity. In this embodiment, an appropriate calculation method is used to calculate the acoustic impact of environmental object 175 based on its spatial configuration information.

[0174] Figure 7 This is a schematic diagram of a target space 700 of the present invention, used to illustrate an embodiment of calculating the audio adjustment amount based on the absorption rate of environmental objects.

[0175] Figure 7 The image shows an environmental object 175 located between a first speaker 110 and a user 180 in a target space 700. For example, the environmental object 175 could be a sofa or a pillar. In this case, the environmental object 175 may obstruct or attenuate the listening experience for the user 180. In other words, the sound pressure level received by the user 180 from the first speaker 110 may be blocked or absorbed. When the control circuit 132 interprets this layout using spatial configuration information, it uses the absorption rate of the environmental object 175 to calculate the degree to which the playback effect of the first speaker 110 at the target listening point (the user 180's position) is affected by the environmental object 175, in order to determine the equivalent volume, sound pressure level, or gain value that the first channel audio 112 needs to output.

[0176] In one embodiment, the sound loss absorbed by the environmental object 175 can be calculated based on the sound pressure level received by the environmental object 175 from the first speaker 110:

[0177] A t [n] = R[n] * SPL t (8)

[0178] Where n represents the sub-band number. That is, the first channel audio 112 output by the first speaker 110 can be divided into multiple sub-bands and calculated separately. A t [n] represents the gain value of the nth sub-band detected at time point t. R[n] represents the absorption rate of the nth sub-band. t This represents the sound pressure level from the first speaker 110 experienced by the environmental object 175 at time point t. Time point t can represent the time difference between the sound being transmitted from the first speaker 110 to the environmental object 175.

[0179] From formula (8), we can see that A t[n] represents the gain value of the first channel audio 112 that is absorbed by the environmental object 175 in the nth sub-band, and also represents the output compensation value required for the nth sub-band of the first channel audio 112. Therefore, when the control circuit 132 generates the first channel audio 112 through the first speaker 110, it increases the gain value of the nth sub-band of the first channel audio 112 by the gain value A. t [n].

[0180] There may be various scenarios where the environmental object 175 is located between the first speaker 110 and the user 180. This embodiment primarily uses whether the line of sight between the first speaker 110 and the user 180 is obstructed, or further uses the line of sight between the first speaker 110 and the user 180's ears as the criterion. It is understood that SPLt itself is a function related to the distance and time between the environmental object 175 and the first speaker 110, and the degree of influence of the calculated At[n] on the user 180 is also a function related to the distance and time between the environmental object 175 and the user 180. After considering different angles of arrangement and distance relationships, various nonlinear correlations are involved. This application does not limit the derivative changes of formula (8), such as adding other weight coefficients, parameters, and offset correction values ​​depending on the actual situation. For example, a sofa may be placed between the user 180 and the first speaker 110. Although the sofa does not obstruct the line of sight, it may still affect the sound pressure value received by the user 180 from the first speaker 110. The control circuit 132 can use formula (8) in combination with interpolation or other modified formulas to make the compensation result more in line with the requirements.

[0181] Figure 8 This is a schematic diagram of a target space 800 of the present invention, used to illustrate an embodiment of calculating audio adjustment based on the reflectivity of environmental objects.

[0182] Figure 8 The image shows a target space 800 where a user 180 is positioned between a first speaker 110 and an ambient object 175. The ambient object 175 could be a wall, ceiling, or floor. In this case, the ambient object 175 will reflect the first channel audio 112 output from the first speaker 110 back to the user 180. In other words, the sound pressure level received by the user 180 from the first speaker 110 will be superimposed or interfered with. When the control circuit 132 interprets this layout using spatial configuration information, it uses the reflectivity of the ambient object 175 to calculate the degree to which the playback effect of the first speaker 110 at the target listening point (the user 180's position) is affected by the ambient object 175, in order to determine the equivalent volume, sound pressure level, or gain value that the first channel audio 112 needs to output.

[0183] In this embodiment, the influence caused by the environmental object 175 can also be calculated according to formula (8), but R[n] is changed to represent the reflectivity of the environmental object 175 in the nth sub-band.

[0184] The result A of formula (8) t [n] can represent the component of the first channel audio 112 reflected to the user 180 by the environmental object 175 in the nth sub-band. Therefore, when the control circuit 132 generates the first channel audio 112 through the first speaker 110, it can appropriately reduce the gain value of the first channel audio 112 so that the total sound pressure level received by the user 180 from the first speaker 110 and the environmental object 175 is maintained at a preset level.

[0185] and Figure 7 The implementation examples are similar, Figure 8 The user 180 may be positioned between the first speaker 110 and the environmental object 175 in various possible scenarios. This embodiment primarily relies on whether the visible lines of sight of the first speaker 110 and the environmental object 175 are blocked by the user 180. However, in practice, walls, ceilings, and floors reflect light regardless of their angle. Therefore, the calculation formula in this embodiment is not limited to formula (8), and other nonlinear compensation calculation methods may be derived based on the arrangement and distance relationships. For example, the target space 800 can be classified into different application scenarios due to the characteristics of the wall, ceiling, and floor materials, as well as the size, shape, and layout of the room, such as a living room, study, bathroom, theater, or outdoors. The host device 130 can first classify the application scenario to which the target space 800 belongs, and then use the corresponding parameters or formulas respectively.

[0186] Figure 7 and Figure 8 The embodiments highlight the following advantages. Through channel-based compensation, the influence of environmental objects 175 on the listening experience of the user 180 is eliminated. Channel-based compensation can flexibly apply different acoustic properties of environmental objects based on their configuration, effectively addressing optimization problems in various complex environments.

[0187] In summary, the identification circuit 134 can receive data from the sensor circuit 140 to identify the position of the user 180 in the target space 170, so that the control circuit 132 can dynamically assign the position of the user 180 as the target listening point. The compensation made by the control circuit 132 for movement of the target listening point has been... Figure 6 The embodiments and formulas (4) to (7) are described. The compensation made by the control circuit 132 for interference from the environmental object 175 has been described in the embodiments and formulas (4) to (7). Figures 7 to 8As explained in Formula (8). These two compensation operations can be performed separately and applied to the channel audio. In other words, the final output optimized channel audio includes compensation values ​​for the movement of the target listening point, as well as compensation for interference from environmental objects 175.

[0188] The recognition circuit 134 identifies the user 180's position based on the sound field environment information captured by the sensor circuit 140. The recognition process may also include the identification of the application scenario to help accelerate subsequent calculations by the control circuit 132. The following uses... Figure 9 This describes the process by which the host device 130 identifies objects based on the application scenario category.

[0189] Figure 9 This is a flowchart illustrating the object recognition process of a host device 130 according to an embodiment of the present invention. Environmental objects appearing in different application scenarios typically exhibit significant group-related acoustic properties, and the sound field reflection coefficients caused by surrounding environmental materials or room size also differ. Therefore, pre-distinguishing application scenario categories helps the audio system 100 improve the efficiency of sound field optimization. It is understood that... Figure 9 Each process in the process is executed by the host device 130, but it is not limited to being executed by a single circuit or module; it can also be the coordinated operation of multiple circuits.

[0190] In process 902, the host device 130 acquires the application scenario category of the target space 170. The host device 130 can acquire the application scenario category in several different ways. In one embodiment, the recognition circuit 134 in the host device 130 can determine an applicable application scenario category based on the sound field environment information provided by the recognition sensor circuit 140. In another embodiment, the control circuit 132 in the host device 130, while obtaining spatial configuration information of environmental objects through a configuration program run by the human-machine interface circuit 133, also simultaneously obtains the application scenario category defined by the user 180 through the configuration program. In a further derived embodiment, the control circuit 132 in the host device 130 can obtain relevant information about the application scenario category from a user device 150 through the communication circuit 136.

[0191] In process 904, to accelerate the query of environmental objects and improve accuracy, the host device 130 prioritizes relevant object databases based on application scenario categories. Object databases are typically pre-established data sets that can be provided by various pipelines. For example, the storage circuit 131 in the host device 130 can pre-store one or more object databases corresponding to different application scenarios. In another embodiment, the host device 130 can connect to a remote database 160 using the communication circuit 136. The remote database 160 may contain multiple object databases corresponding to different application scenarios. Each object database contains the external feature information and acoustic attribute information of multiple environmental objects.

[0192] After the host device 130 obtains the application scenario category in process 902, it can preferentially select an object database related to that application scenario category from the storage circuit 131 or the remote database 160 for subsequent identification of environmental objects. In one embodiment, the identification circuit 134 analyzes the sound field environment information provided by the sensor circuit 140 to obtain one or more object shape feature information, and retrieves the object database based on the object shape feature information to identify environmental objects that match the object shape feature information, including name, absorptivity, and reflectivity. In another embodiment, the control circuit 132 executes a configuration program and obtains the name of an environmental object using the human-machine interface circuit 133. The control circuit 132 searches the object database based on the name of the environmental object 175 to obtain the absorptivity and reflectivity corresponding to the environmental object.

[0193] In further derived embodiments, the parameters used in the search process can be combined in multiple ways. For example, during the analysis of sound field environment information, the recognition circuit 134 can obtain external features such as the material, size, and shape of the environmental object 175. The recognition circuit 134 transmits this external feature information to the object database for multi-condition cross-comparison to obtain a list of candidate objects sorted according to matching scores. If application scenario category information is used as a search condition during the search of the object database, it will help narrow down the possible range, accelerate the recognition, and improve accuracy.

[0194] In process 906, control circuit 132 retrieves the absorptivity and reflectivity of environmental objects from the object database selected in process 904. In practice, the acoustic attribute information of environmental objects stored in the object database is not limited to being stored in multiple independent object databases. The object database can be a relational database, containing multiple fields connected together by correlation coefficients. For example, the fields of the object database can include object name, application scenario category, material, absorptivity, reflectivity, and even external features such as shape, color, and gloss. The field values ​​corresponding to each environmental object are not limited to a one-to-one relationship, but can be one-to-many or many-to-one. The values ​​stored in each field are not necessarily absolute values, but range values ​​or probability values. In further derivative implementations, the object database can be an adaptive database that can be continuously iteratively corrected by machine learning. User 180 can train the object database by providing feedback on preferred settings through human-machine interface circuit 133.

[0195] In process 908, control circuit 132 adjusts the channel audio in multiple sub-bands based on the search results and configuration of environmental objects. The acoustic properties of environmental objects 175 may vary significantly across different frequency bands. For example, a sofa may absorb a large amount of high-frequency signals but not affect the penetration of low-frequency signals. Therefore, the absorption rate or reflectivity retrieved from the object database can be an array value corresponding to multiple sub-bands or a frequency response curve. The size or separation method of the sub-bands can be determined according to design requirements and is not limited in this embodiment. Control circuit 132 adjusts the gain value of the channel audio in multiple sub-bands, which can be simulated as an equalizer or filter concept in practice. In other words, control circuit 132 can implement an equalizer for each speaker in the audio system 100 and customize the equalizer according to the output compensation value calculated in the aforementioned embodiment, so that the corresponding channel audio is adjusted. Further embodiments regarding the calculation of output compensation values ​​will be discussed later. Figure 10 The explanation is as follows.

[0196] In process 910, the control circuit 132 outputs the adjusted channel audio to the corresponding speaker through the audio transmission circuit 135. The implementation of the audio transmission circuit 135 has been described in [the following text is missing from the original] Figure 1 As previously explained, I will not repeat myself here.

[0197] Figure 9 The embodiments highlight the following advantages: Object recognition operations can be performed according to application scenario categories (automatic recognition or manual input) to increase recognition efficiency. The object database employs a scalable architecture, continuously enhancing recognition capabilities over the long term with feedback from cloud-based big data services and machine learning. The audio system 100 can apply the concept of an equalizer to divide channel audio into multiple sub-bands for separate processing, effectively improving the final synthesized sound quality.

[0198] The following Figure 10 This further explains how the control circuit 132 calculates the output compensation value for each channel based on the spatial configuration information of the environmental objects 175.

[0199] Figure 10 This is a flowchart of an audio processing method according to an embodiment of the present invention, illustrating an embodiment of calculating output compensation values ​​based on the positional relationships of environmental objects. Figure 10 The process is mainly executed by the control circuit 132 in the host device 130.

[0200] In process 1002, control circuit 132 determines the relative positional relationship between environmental objects, the target listening point, and the speaker. Multiple speakers and multiple environmental objects 175 in the target space 170 can be arranged and combined with the target listening point to form multiple sets of positional relationships. Each set of positional relationships includes one speaker, one environmental object 175, and the target listening point. Control circuit 132 checks and judges each combination of positional relationships in the target space 170 and calculates the corresponding output compensation value. The following uses one set of positional relationships in the audio system 100 as an example to illustrate the compensation method taken by control circuit 132 for interference caused by an environmental object 175 to a speaker at the target listening point.

[0201] In process 1004, control circuit 132 determines whether environmental object 175 is between the target listening point and the speaker. The position of environmental object 175 in the target space 170 can also be obtained by recognition circuit 134, or by human-machine interface circuit 133 through a configuration program. After integrating the above information, control circuit 132 can determine the relative positional relationship between each environmental object 175, the target listening point, and each speaker, and perform corresponding compensation calculations for each speaker. The situation to be determined in process 1004 is as follows: Figure 7 The situation is as shown. If the situation is met, proceed to process 1008. If the situation is not met, proceed to process 1006.

[0202] In process 1006, control circuit 132 determines whether the target listening point is located between environmental object 175 and the speaker. The condition to be determined in process 1006 is as follows: Figure 8 The situation is shown. If the situation is met, proceed to process 1010. If the situation is not met, proceed to process 1012.

[0203] In process 1008, control circuit 132 uses the absorption rate of environmental object 175 to calculate the output compensation value of the channel audio. In a preferred embodiment, the output compensation value of the speaker's channel audio is calculated separately for multiple sub-bands. Detailed calculations can be found in [reference needed]. Figure 7The target space 700 and formula (8). The control circuit 132 can find the absorption rate of the environmental object 175 from the object database and substitute it into formula (8) to obtain the output compensation value.

[0204] In process 1010, control circuit 132 uses the reflectivity of environmental object 175 to calculate the output compensation value for the channel audio. (Reference) Figure 8 Given the target space 800 and formula (8), the control circuit 132 can search for the reflectivity of the environmental object 175 from the object database and substitute it into formula (8) to obtain the output compensation value.

[0205] Understandably, the output compensation value calculated based on the absorptivity of the ambient object 175 may amplify the gain, sound pressure level, or equivalent volume of the adjusted channel audio to compensate for the absorbed energy. Conversely, the output compensation value calculated based on the reflectivity of the ambient object 175 may reduce the gain, sound pressure level, or equivalent volume of the adjusted channel audio to balance the reflected energy. In other words, the output compensation values ​​calculated based on absorptivity and reflectivity are usually opposite in sign.

[0206] In process 1012, if environmental object 175 does not meet the conditions of process 1004 or process 1006, then control circuit 132 can determine that environmental object 175 is located in a position that will not affect the speaker's playback to the target listening point. In this case, control circuit 132 may not calculate the impact of environmental object 175 on the speaker and the target listening point for this set of positional relationships. However, it should be understood that a target space 170 typically contains multiple speakers. Environmental object 175 may not affect the playback of one speaker to the target listening point, but it may still affect the playback of other speakers to the target listening point. In other words, control circuit 132 needs to perform separate calculations for each set of positional relationships in the target space 170. Figure 10 The process.

[0207] In certain specific cases, the presence of environmental object 175 can be directly ignored. For example, if the reflectivity or absorption rate of sound by environmental object 175 is less than a certain threshold, its presence in the target space 170 can be ignored. On the other hand, if the control circuit 132 determines that the volume of environmental object 175 is less than a certain size, the presence of environmental object 175 can also be ignored.

[0208] In a further derived embodiment, if more than one user is detected in the target space 170, the determination of the target listening point can be based on the center points of multiple users' locations, or selectively based on the location of one user. As for users not selected as target listening points, the host device 130 can simulate them as environmental objects, according to...Figures 7 to 8 The example processing.

[0209] Figure 10 The embodiments highlight the following advantages. Figure 10 The embodiments continue Figure 7 and Figure 8 This approach simplifies complex environmental problems into multiple linear relationships, which are then solved separately. For specific environmental objects 175, the computational complexity can be further reduced by neglecting certain factors.

[0210] Figure 11 This is a schematic diagram of a target space 1100 of the present invention, used to illustrate an embodiment of optimizing the sound field by object substrate compensation operation.

[0211] The target space 1100 includes multiple speakers, such as a first speaker 1110, a second speaker 1120, a third speaker 1130, and a fourth speaker 1140. When the sound system 100 operates based on object-based compensation, the control circuit 132 logically treats the target space 1100 as a spatial coordinate system. This spatial coordinate system can be two-dimensional or three-dimensional planar coordinates. For ease of explanation, Figure 11 The illustration is presented in a two-dimensional planar coordinate system that includes an X-axis and a Y-axis.

[0212] In target space 1100, user 180 is located at origin P0. Control circuit 132 assigns user 180 as the target listening point. Figure 3 As described in the embodiments, the object-based acoustic system is built upon an array of acoustic parameters. Each sound source object has relay data describing its type, location, size (length, width, height), and divergence. After object-based computation, the sound represented by a sound source object is assigned to one or more speakers for simultaneous playback, with each speaker playing a portion of the sound from the sound source object. In other words, the object-based acoustic system can use multiple speakers to simulate the physical presence of a sound source object. For example, through object-based compensation, a user 180 at the target listening point can hear a virtual sound source object 1105 moving along a movement trajectory 1103 from a first position P1 to a new first position P1'.

[0213] The object-based compensation operation in this embodiment optimizes the audio output of all speakers for the target listening point. The object-based compensation operation utilizes the array calculation module in the existing object-based acoustic system to parameterize various distance factors and sound field categories, and can perform calculations similar to formulas (4) to (7). For the audio system 100, the host device 130 only needs to apply the user 180's position information to the object-based compensation operation to optimize the audio output of all speakers for the target listening point.

[0214] In one embodiment, the control circuit 132 can define the target listening point as the origin of the entire spatial coordinate system. When the user 180 moves, the entire spatial coordinate system moves with the origin. In other words, the position of the virtual sound source object 1105 relative to the origin remains unchanged. When the control circuit 132 plays the effect of the virtual sound source object 1105 through object-based compensation operation, the relative position of the virtual sound source object 1105 perceived by the user 180 will not change with the movement of the user 180.

[0215] In the target space 1100 of this embodiment, there may be environmental objects 175 that could substantially affect the listening experience of the user 180. The control circuit 132 can... Figure 9 The process 902 obtains the spatial configuration information of environmental objects in the target space 1100, revealing that environmental object 175 is located at the second position P2. When user 180 moves, the origin of the entire spatial coordinate system changes accordingly. Although environmental object 175 does not move, its relative position to the origin changes. Therefore, it can be understood that in the spatial coordinate system after the movement, the coordinate values ​​of environmental object 175 move in the opposite direction.

[0216] To counteract the interference caused by the environmental object 175 to the user 180, the control circuit 132 of this embodiment establishes a compensating sound source object based on the environmental object 175. The relay data of this compensating sound source object includes: the coordinate position and size of the environmental object 175, as well as its reflectivity and absorptivity. The reflectivity and absorptivity of the environmental object 175 can be determined by… Figure 9 The compensation sound source object is obtained through process 906. The compensation sound source object is regarded as the negative sound source object of the environment object 175 and is applied to the object base compensation operation, becoming a virtual sound source that can cancel out the environment object 175.

[0217] It is understandable that the essence of the compensating sound source object is the negative sound source object corresponding to the environmental object 175, and its position overlaps with the environmental object 175. Therefore, in Figure 11Unless otherwise specified, the four-speaker configuration of the target space 1100 is merely an example. In actual applications of the audio system 100, the number of speakers can be greater, even including a stereo configuration of upper and lower speakers. This description does not limit other possible configurations.

[0218] Figure 11 The embodiments illustrate the advantages of object-based compensation operations. Control circuit 132 converts information from the target space 1100 into a spatial coordinate system, simplifying complex multi-object interaction calculations into array operations of relay data. Setting the position of the moving user 180 as the origin of the spatial coordinate system completely eliminates the impact of user 180's movement on the processing of virtual objects, thus simplifying the computational process. This embodiment also proposes the concept of compensating for sound source objects, directly applying object-based compensation operations to counteract environmental object interference, eliminating the need for complex multi-channel interactive calculations.

[0219] The following is Figure 12 This section explains the ease of object base compensation operations and their possible derivative applications.

[0220] Figure 12 This is a schematic diagram of a target space 1200 of the present invention, used to illustrate an embodiment of optimizing the sound field by object substrate compensation operation.

[0221] The target space 1200 may contain multiple speakers, such as the first speaker 1210, the second speaker 1220, the third speaker 1230, the fourth speaker 1240, the fifth speaker 1250, and the sixth speaker 1260, arranged in a long strip sound field. Each speaker corresponds to an ID. When the user 180 is in the first position P1, the relay data of a virtual sound source object (not shown) is mapped to the IDs of the first speaker 1210 and the second speaker 1220. After the control circuit 132 performs object basis compensation, it causes the first speaker 1210 and the second speaker 1220 to play the first channel output 1212 and the second channel output 1222, allowing the user 180 to perceive the presence of the virtual sound source object. When the user 180 moves along the movement trajectory 1203 to the second position P2, the control circuit 132 recalculates the target listening point and maps the relay data of the virtual sound source object to the fifth speaker 1250 and the sixth speaker 1260. After the control circuit 132 performs object base compensation, it will cause the fifth speaker 1250 to play the fifth channel output 1252 and the sixth channel output 1262, so that the user 180 feels that the virtual sound source object still exists to the left and right of the user 180 and does not leave with the movement of the user 180.

[0222] This embodiment primarily illustrates the flexible application and simplicity of object substrate compensation operations. In many special cases, sound field optimization can be achieved with only a small amount of computation. For example, if user 180 is located in a spherical sound field, the control circuit 132 only needs to perform rotation coordinate calculations to ensure that user 180 experiences a consistent sound field effect regardless of the direction they are facing.

[0223] The following is Figure 13 The basic logic of control circuit 132 when performing object base compensation operation is summarized.

[0224] Figure 13 This is a flowchart of an embodiment of the present invention for object base compensation operation, illustrating the concept of establishing a compensation sound source object.

[0225] In process 1304, control circuit 132 establishes a corresponding compensation sound source object based on environmental object 175. For user 180 located at the target listening point, the presence of environmental object 175 constitutes a physical sound source. Environmental object 175 may reflect sound emitted by a speaker to the target listening point. Environmental object 175 may also block or absorb some sound, causing attenuation of the sound emitted by a speaker to the target listening point. The compensation sound source object is a negative sound source object established for environmental object 175. When host device 130 substitutes the compensation sound source object into the object-based compensation operation to generate channel audio, the presence of environmental object 175 can be eliminated. The specific details of the object-based operation itself can utilize the calculation methods of existing object-based acoustic products, using relay data of the sound source object to perform a large number of related array operations. For example, a relay data of the compensation sound source object includes: the coordinate position and size of the environmental object 175, as well as its reflectivity and absorptivity.

[0226] In process 1306, control circuit 132 calculates the sound source effect of the compensation sound source object. In this embodiment, the compensation sound source object is established based on the environmental object 175, wherein the relay data has the same coordinate position, size, and reflectivity and absorptivity of sound as the environmental object 175, but the resulting sound source effect is the inverse gain value of the environmental object 175.

[0227] Figure 13 The embodiments can also be referred to in similar ways. Figure 7 and Figure 8 The calculation. Formula (8) can be derived into formula (9), which calculates the passively generated gain value of the environmental object 175 based on the sound pressure value received by the environmental object 175 from the first speaker 110:

[0228] A t [m][n]=R[n]*SPL t [m] (9)

[0229] Where m represents the speaker number and n represents the sub-band number. A t [m][n] represents the gain value of the nth sub-band due to the influence of the mth speaker. R[n] represents the absorption rate of the nth sub-band. SPL t [m] represents the sound pressure level of the environmental object 175 at time point t, which is received by the m-th speaker. Time point t represents the time difference between the sound transmission from the speaker to the environmental object 175. If the time difference is greater than a non-negligible range, it indicates that there is an echo in the target space 170.

[0230] As can be seen from formula (9), the calculation result for each environmental object includes an array of gain values ​​for multiple speakers and multiple sub-bands at a given time point. The sound source effect of the compensated sound source object is the negative value of this gain value array. In other words, the object-based compensation operation based on formula (9) involves array operations involving the interactive arrangement and combination of parameters in multiple dimensions. For ease of explanation, the following explanation will use the gain value of one speaker and one sub-band at a given time point as an example.

[0231] Figure 13 Implementation examples and Figure 7 and Figure 8 Similar to other embodiments, this embodiment can use appropriate calculation methods to calculate the acoustic effects of environmental objects 175 based on their spatial configuration information. For example, if the target listening point is located between a speaker and the visible line of sight of the environmental object 175, the control circuit 132 calculates the sound source effect of the compensated sound source object based on the reflectivity of the environmental object 175. Conversely, if the environmental object 175 is located between the target listening point and the visible line of sight of the speaker, the control circuit 132 calculates the sound source effect of the compensated sound source object based on the absorptivity of the environmental object 175.

[0232] For example, when an environmental object 175 absorbs sound emitted by a speaker, reducing the volume received by the target listening point, the control circuit 132 creates a virtual sound source object at the coordinates of the environmental object 175 that produces the corresponding volume effect as compensation. Conversely, if an environmental object 175 reflects sound from a speaker, causing the target listening point to receive too much volume, the control circuit 132 creates a virtual sound source object with a negative gain value at the coordinates of the environmental object 175.

[0233] It is understandable that a visible line of sight is defined as a straight line connecting two objects in space. Since objects have a certain volume and area, the volume may be very large, and the obstruction of the visible line of sight may include partial obstruction and complete obstruction. This embodiment can be based on formula (9), and then multiplied by different weighting coefficients or added with different offset corrections depending on various situations.

[0234] In process 1308, control circuit 132 mixes the sound source effect of the compensated sound source object into the channel audio, causing the corresponding speaker to play it. When performing the object-based compensation operation, control circuit 132 can handle complex object correspondence array operations, mixing the multiple sound source signals assigned to each speaker into a corresponding channel audio. After applying the object-based compensation operation, the volume effect received at the target listening point will include the sound source effect generated by the compensated sound source object. In this way, interference caused by environmental objects 175 can be effectively canceled out by the compensated sound source object.

[0235] In process 1310, control circuit 132 determines whether the target listening point has moved to a new position. As described in process 208, the audio system 100 can continuously track the movement of user 180 and update the target listening point accordingly. If the target listening point has moved, process 1312 is performed. Otherwise, the playback operation of process 1308 continues.

[0236] In process 1312, control circuit 132 updates the relay data of the compensation sound source object. In this embodiment, control circuit 132 establishes an object base space with the target listening point as a coordinate origin. If the target listening point moves to a new position, control circuit 132 assigns this new position as the new coordinate origin of the object base space. The difference between the new coordinate origin and the original coordinate origin can be represented as a movement vector. The spatial coordinates of the environmental object 175 relative to the target listening point also change in the opposite direction with this movement vector. Control circuit 132 then updates the relay data of the compensation sound source object corresponding to the environmental object 175 based on this movement vector. In a further embodiment, all speakers in the object base space can also be regarded as an object, having corresponding IDs, relay data, and coordinate values.

[0237] In another embodiment, the audio system 100 is not limited to using the target listening point as the origin of the coordinate system. The audio system 100 may also use a fixed reference point as the origin of the object base space. When the relative position of the sound source object in the object base space changes, the control circuit 132 updates the coordinate values ​​in the relay data of the sound source object accordingly.

[0238] Once process 1312 is completed, control circuit 132 repeats process 1308.

[0239] Figure 13The embodiments illustrate the advantages of object-based compensation operations. Control circuit 132 converts information from the target space 1100 into a spatial coordinate system, simplifying complex multi-object interaction calculations into array operations of relay data. Setting the position of the moving user 180 as the origin of the spatial coordinate system completely eliminates the impact of user 180's movement on the processing of virtual objects, thus simplifying the computational process. This embodiment also proposes the concept of compensating for sound source objects, directly applying object-based compensation operations to counteract environmental object interference, eliminating the need for complex multi-channel interactive calculations.

[0240] In a further derived embodiment, if the host device 130 itself does not have the ability to perform mixing operations on the object substrate, the control circuit 132 can provide the function of channel mapping by executing software, so that the calculation results of the object substrate can be correctly mapped to each speaker.

[0241] In summary, this application proposes a sound system 100 that can dynamically track the user's position to optimize the sound field and intelligently eliminate interference caused by environmental objects. The means of tracking the user's position can be the individual or combined use of various methods such as cameras, infrared sensors, or wireless detectors. The spatial configuration information of environmental objects 175 in the target space 170 can be obtained by recognizing images captured by a camera or by manual input by the user. The sound field optimization method can be channel-based compensation or object-based compensation. When calculating the impact of environmental objects 175 on the target listening point, the relative positional relationship between the environmental objects 175 and the speaker, and between the environmental objects 175 and the target listening point, can be considered, and different calculation methods can be used. When using object-based compensation, the control circuit 132 establishes a corresponding compensation sound source object for each environmental object 175, so that the final mixed channel audio eliminates the interference caused by the environmental objects 175 on the target listening point.

[0242] Certain terms are used in the specification and claims to refer to specific elements, and those skilled in the art may use different names to refer to the same element. This specification and claims do not distinguish elements by differences in name, but rather by differences in function. The term "comprising" as used in the specification and claims is an open-ended term and should be interpreted as "comprising but not limited to". Furthermore, the term "coupled" herein includes any direct and indirect connection means. Therefore, if the text describes a first element coupled to a second element, it means that the first element can be directly connected to the second element through electrical connection or signal connection methods such as wireless transmission or optical transmission, or indirectly electrically or signal-connected to the second element through other elements or connection means.

[0243] The use of "and / or" in this specification includes any combination of one or more of the listed items. Furthermore, unless otherwise specified in this specification, any singular term also includes the meaning of the plural form.

[0244] The above are merely preferred embodiments of the present invention. All equivalent changes and modifications made in accordance with the claims of the present invention shall fall within the scope of the present invention.

Claims

1. An audio system (100) capable of dynamically optimizing a playback effect according to a user position, comprising: a sensor circuit (140) configured to dynamically sense a target space (170) to generate an acoustic field environment information, wherein a sound field environment information comprising a user position of a user in a target space (170); a first speaker (110) and a second speaker (120) configured to play audio; a host device (130) coupled to the sensor circuit (140), the first speaker (110) and the second speaker (120), comprising: a recognition circuit (134) configured to recognize the user position of the user in the target space from the sound field environment information; a control circuit (132) coupled to the recognition circuit (134) and configured to dynamically assign the user position as a target listening point; an audio transmission circuit (135) coupled to the control circuit (132), the first speaker (110) and the second speaker (120) and configured to transmit audio; and a human-machine interface circuit (133) coupled to the control circuit (132) and configured to be controlled by the control circuit (132) to run a configuration program to obtain spatial configuration information and acoustic property information of an environmental object in the target space (170), wherein the acoustic property information of the environmental object comprises at least one of sound absorption, sound reflection and resonance frequency; wherein the control circuit (132) corresponds the target space to an object base space and correspondingly establishes a compensation sound source object in the object base space according to the environmental object; wherein a relay data of the compensation sound source object comprises the spatial configuration information and the acoustic property information of the environmental object; wherein the control circuit (132) performs an object base compensation operation according to the target listening point and the relay data to offset the interference of the environmental object to the target listening point to generate a first channel audio (112) and a second channel audio (122) optimized for the target listening point; wherein the control circuit (132) outputs the first channel audio (112) and the second channel audio (122) to the corresponding first speaker (110) and second speaker (120) through the audio transmission circuit (135); wherein the spatial configuration information of the environmental object comprises the position, size and appearance characteristics of the environmental object, and the acoustic property information of the environmental object comprises sound reflection and absorption; wherein the object base compensation operation comprises: calculating an acoustic source effect of the environmental object passively generated under the influence of the first channel audio on a plurality of sub-bands according to the coordinate position, size and sound reflection or absorption of the environmental object; establishing the compensation sound source object according to the acoustic source effect, so that the compensation sound source object has a negative acoustic source effect opposite in polarity to the acoustic source effect; and mixing the negative acoustic source effect of the compensation sound source object into the first channel audio (112) to offset the interference of the environmental object to the target listening point; wherein the control circuit corresponds the target space to the object base space by establishing the object base space with the target listening point as a coordinate origin. wherein, when the identification circuit (134) judges that the user position moves, the control circuit (132) reassigns the moved user position as a new target listening point, and reconstructs the object base space with the new target listening point as a new coordinate origin; and The control circuit (132) correspondingly updates the coordinate position of the environmental object according to the new coordinate origin and a moving vector of the coordinate origin.

2. The sound system (100) of claim 1, wherein The object base compensation operation further includes: If the target listening point is located between the first loudspeaker (110) and the visual line of the environmental object, the control circuit (132) calculates the sound source effect according to the reflectivity of the environmental object.

3. The sound system (100) of claim 1, wherein, The object base compensation operation further includes: If the environmental object is located between the target listening point and the visual line of the first loudspeaker (110), the control circuit (132) calculates the sound source effect according to the absorption of the environmental object.

4. The sound system (100) of claim 1, wherein, The sensor circuit (140) includes a camera (610) configured to capture an audio field environment image of the target space (170); Wherein, the identification circuit (134) dynamically identifies the head position, face direction, or ear position of the user according to the audio field environment image captured by the camera (610) to judge the user position.

5. The sound system (100) of claim 1, wherein, The sensor circuit (140) further includes an infrared sensor (620) configured to capture a thermal imaging data in the target space; wherein the identification circuit (134) analyzes the moving track of the thermal imaging data to dynamically judge the user position.

6. The sound system (100) of claim 1, wherein, The sensor circuit (140) further includes a wireless detector (630) arranged in the target space and detecting a wireless signal of an electronic device; Wherein, the identification circuit (134) dynamically locates the position of the electronic device according to the characteristics of the wireless signal detected by the wireless detector (630); and Wherein, the identification circuit (134) dynamically judges the user position according to the position of the electronic device.

7. The sound system (100) of claim 1, wherein, The host device (130) further includes a storage circuit (131) coupled to the control circuit (132) and configured to store one or more object databases, wherein each object database corresponds to an application scenario category and includes the external feature information and acoustic property information of a plurality of environmental objects; Wherein, when the control circuit (132) controls the human-computer interface circuit (133) to run the configuration program, the human-computer interface circuit (133) further acquires an application scenario category of the target space (170); and Wherein, the control circuit (132) preferentially selects an object database related to the application scenario category from the storage circuit (131) to identify the environmental object and find the acoustic property information of the environmental object according to the application scenario category.

8. The sound system (100) as recited in claim 1, wherein, The host device (130) further includes a communication circuit (136) coupled to the control circuit (132) and configured to be controlled by the control circuit (132) to connect to a remote database (160) corresponding to an application scenario category; the remote database (160) is configured to store one or more object databases, wherein each object database corresponds to an application scenario category and includes shape feature information and acoustic attribute information of a plurality of environmental objects; Wherein, when the human-computer interface circuit (133) is controlled by the control circuit (132) to run the configuration program, an application scenario category of the target space (170) is further obtained; and Wherein, the control circuit (132) selects an object database related to the application scenario category from the remote database (160) to identify the environmental object and find the acoustic attribute information of the environmental object according to the application scenario category.

Citation Information

Patent Citations

  • Playback device configuration based on proximity detection

    CN106105271A

  • Audio processing method and electronic equipment

    CN111050269A

  • Distributed wireless speaker system with automatic configuration determination when new speakers are added

    US20150208188A1