A method and system for speaker sound field adaptation

By automatically adjusting audio parameters through real-time sensing and machine learning models of the speaker system, the problem of insufficient acoustic adaptability of traditional speaker systems after location changes is solved. This achieves intelligent sound quality optimization and multi-speaker collaborative sound field correction, thus improving the user experience.

CN122160678APending Publication Date: 2026-06-05HANSONG NANJING TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANSONG NANJING TECH LTD
Filing Date
2026-03-10
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Traditional speaker systems cannot detect and automatically adjust the acoustic environment in real time after the location changes, resulting in sound image shift and frequency response imbalance. This is especially true in multi-speaker scenarios where coordinated correction is difficult, increasing the user's operational burden.

Method used

By acquiring the speaker's evaluation parameters, the machine learning model is used to perceive the speaker's placement and environmental acoustic characteristics in real time, and automatically adjusts the parameters of the audio processing unit to optimize the sound field performance. This includes the combination of inertial measurement unit and environmental perception unit, and constructs an acoustic environment model for adaptive audio adjustment.

Benefits of technology

It enables automatic acoustic environment adaptation of the speaker system after position changes, reduces the need for manual calibration by users, and improves the level of intelligence in sound quality optimization and listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122160678A_ABST
    Figure CN122160678A_ABST
Patent Text Reader

Abstract

Embodiments of the present specification provide a sound box sound field self-adaptive method and system. The method comprises: obtaining an evaluation parameter of a sound box; determining a current placement state of the sound box based on the evaluation parameter; determining whether the placement state of the sound box is changed based on a historical placement state and the current placement state; in response to the placement state of the sound box being changed: determining an acoustic characteristic identifier in the current placement state and an audio adjustment scheme corresponding to the acoustic characteristic identifier based on the evaluation parameter through an acoustic environment model; the acoustic environment model is a machine learning model; and adjusting a parameter of an audio processing unit of the sound box based on the audio adjustment scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This manual relates to the field of speakers, and in particular to a speaker sound field adaptive method and system. Background Technology

[0002] In modern home and office environments, speaker systems often experience sound quality degradation issues such as sound image shift and frequency response imbalance due to changes in placement (e.g., cleaning, layout adjustments). Traditional solutions often rely on manual calibration by the user, which is cumbersome and cannot adapt to changes in the acoustic environment in real time. While some existing technologies support automatic acoustic calibration via test tones and microphones, this is usually only performed once during installation or after manual triggering, failing to continuously detect changes in speaker position and orientation, and making it difficult to address collaborative sound field misalignment issues in multi-speaker scenarios. Users often need to repeat complex calibrations after moving speakers, increasing their operational burden.

[0003] Therefore, there is an urgent need for a technology that can sense the placement of speakers in real time, automatically analyze the acoustic characteristics of the environment, and make adaptive audio adjustments to achieve continuous sound quality optimization without human intervention and support collaborative sound field correction of multi-speaker systems, thereby improving the intelligence level and listening experience of audio equipment. Summary of the Invention

[0004] One embodiment of this specification provides a speaker sound field adaptive method, the method comprising: acquiring evaluation parameters of the speaker, the evaluation parameters including speaker state parameters and / or environmental distance data; determining the current placement state of the speaker based on the evaluation parameters; determining whether the placement state of the speaker has changed based on historical placement states and the current placement state; responding to a change in the placement state of the speaker: determining acoustic feature identifiers and corresponding audio adjustment schemes for the current placement state based on the evaluation parameters and an acoustic environment model; the acoustic environment model being a machine learning model; and adjusting the parameters of the speaker's audio processing unit based on the audio adjustment scheme.

[0005] One embodiment of this specification provides a speaker sound field adaptive system. The system includes a control device configured to: acquire evaluation parameters of the speaker; determine the current placement state of the speaker based on the evaluation parameters; determine whether the speaker's placement state has changed based on historical placement states and the current placement state; and, in response to a change in the speaker's placement state: determine acoustic feature identifiers and corresponding audio adjustment schemes for the current placement state based on the evaluation parameters and an acoustic environment model; the acoustic environment model is a machine learning model; and adjust the parameters of the speaker's audio processing unit based on the audio adjustment scheme. Attached Figure Description

[0006] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein:

[0007] Figure 1 This is a schematic diagram illustrating the application scenario of the speaker sound field adaptive system according to some embodiments of this specification; Figure 2 This is an exemplary block diagram of a speaker sound field adaptive system according to some embodiments of this specification; Figure 3 This is an exemplary flowchart of a speaker sound field adaptive method according to some embodiments of this specification; Figure 4 This is an exemplary flowchart illustrating the determination of collaborative compensation parameters according to some embodiments of this specification; Figure 5 This is an exemplary schematic diagram illustrating, according to some embodiments of this specification, whether the placement of the speaker has changed. Detailed Implementation

[0008] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of this specification. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.

[0009] It should be understood that the terms “system,” “device,” “unit,” and / or “module” used herein are one way to distinguish different components, elements, parts, sections, or assemblies at different levels. However, if other terms can achieve the same purpose, they may be replaced by other expressions.

[0010] As indicated in this specification and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of expressly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.

[0011] Flowcharts are used in this specification to illustrate the operations performed by the system according to embodiments of this specification. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the steps can be processed in reverse order or simultaneously. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.

[0012] Figure 1 This is a schematic diagram illustrating the application scenario of a speaker sound field adaptive system according to some embodiments of this specification.

[0013] like Figure 1 As shown, the application scenario 100 of the speaker sound field adaptive system includes the speaker sound field adaptive system 110, network 120, audio equipment 130 and storage device 140. It can be applied to audio playback scenarios such as smart speakers in home audio-visual systems, multi-speaker layouts in conference rooms, car audio systems, and other audio equipment that needs to automatically optimize the sound field effect according to the placement environment.

[0014] The speaker sound field adaptive system 110 refers to a system capable of sensing the speaker's placement and surrounding acoustic environment in real time, and automatically adjusting audio output parameters accordingly to continuously optimize sound field performance. In some embodiments, the speaker sound field adaptive system 110 includes a control device 210. More information about the speaker sound field adaptive system 110 can be found at [link to relevant documentation]. Figure 2 Related descriptions.

[0015] Network 120 is used to connect the speaker sound field adaptive system 110, the audio equipment 130, and the storage device 140. Network 120 enables communication between the various components of the speaker sound field adaptive system application scenario, facilitating the exchange of data and / or information. In some embodiments, network 120 can be any one or more of a wired network or a wireless network. For example, network 120 includes a cable network, a fiber optic network, a telecommunications network, or any combination thereof.

[0016] Audio device 130 refers to a device component used to play audio. In some embodiments, audio device 130 may include a speaker (also referred to as a main speaker) 1301.

[0017] Speaker 1301 refers to an independent audio device that serves as the core of adaptive adjustment in an audio system.

[0018] In some embodiments, the speaker 1301 includes an audio processing unit 13011, a reference microphone 13012, and a loudspeaker 13013.

[0019] In some embodiments, the audio processing unit 13011 may be connected to the reference microphone 13012 and the speaker 13013 respectively to receive the ambient noise signal collected by the reference microphone 13012 and adjust the parameters of the audio processing unit 13011 based on the ambient noise signal, and then generate an audio signal based on the adjusted parameters (such as volume parameters and / or frequency response parameters) and output it to the speaker for playback.

[0020] The audio processing unit 13011 refers to the functional module or circuit inside the speaker 1301 that performs various processing operations (such as equalization, delay, gain control, etc.) on the audio signal. For example, the audio processing unit 13011 can be a digital signal processor (DSP) or a system on chip (SoC).

[0021] Reference microphone 13012 refers to a microphone used to acquire ambient sound signals for system reference or analysis. For example, reference microphone 13012 can be used to pick up ambient noise or test signals for acoustic calibration.

[0022] In some embodiments, the reference microphone 13012 may be built into the speaker 1301. In some embodiments, the reference microphone 13012 may be located outside the speaker 1301 and communicate with the speaker 1301.

[0023] The loudspeaker 13013 refers to the acoustic transducer inside the speaker enclosure 1301 used to convert processed electrical signals into audible sound. In some embodiments, the loudspeaker 13013 may be built into the speaker enclosure 1301.

[0024] In some embodiments, the audio device 130 may further include a co-speaker 1302. The speaker sound field adaptive system 110 can control the speaker 1301 and the co-speaker 1302 to automatically sense their own and each other's placement, orientation and acoustic environment, and coordinately adjust the audio output parameters of the main speaker and the co-speaker to optimize the overall sound field performance.

[0025] Collaborating speaker 1302 refers to other speaker devices that work in conjunction with speaker 1301 to jointly form a unified sound field. For example, in a stereo or surround sound system, other speakers besides the main speaker can act as collaborating speakers.

[0026] Storage device 140 refers to a device used to store data, instructions, and / or any other information. In some embodiments, the storage device stores data and / or instructions related to the speaker adaptive system. For example, the storage device stores data related to speaker evaluation parameters, status parameters, environmental distance parameters, audio adjustment schemes, and parameters of the audio processing unit 13011.

[0027] The above description is illustrative and does not limit the scope of this disclosure. Many alternatives, modifications, and variations will be apparent to those skilled in the art. The features, structures, methods, and other characteristics of the exemplary embodiments described herein can be combined in various ways to obtain other and / or alternative exemplary embodiments. For example, the configuration and / or functionality of the speaker sound field adaptive system may be varied or modified depending on the specific implementation scenario. However, these variations and modifications do not depart from the scope of this disclosure.

[0028] Figure 2 This is an exemplary block diagram of a speaker sound field adaptive system according to some embodiments of this specification. In some embodiments, the speaker sound field adaptive system (referred to as the system) 200 may include a control device 210.

[0029] Control device 210 refers to a device used to execute specific control logic and algorithms to manage and operate one or more devices or systems. For example, control device 210 can be integrated inside speaker 1301 or operate as a standalone central controller. Its core may include one or more processors, which may integrate data processing unit 2101, logic judgment unit 2102, state control unit 2103, etc. Through the coordinated work of the above units, control device 210 can achieve centralized management and execution of the sound field adaptive process.

[0030] The data processing unit 2101 refers to a unit used for data processing. For example, a microcontroller.

[0031] In some embodiments, the data processing unit 2101 may acquire evaluation parameters, wherein the evaluation parameters include the speaker's state parameters and / or environmental distance data.

[0032] In some embodiments, the data processing unit 2101 may include an attitude sensing unit 21011 and an environment sensing unit 21012.

[0033] The attitude sensing unit 21011 refers to a sensing device or module used to detect and acquire spatial attitude information of the speaker 1301. For example, the attitude sensing unit 21011 may include an inertial measurement unit (IMU) and a magnetic field sensor.

[0034] In some embodiments, the attitude sensing unit 21011 can acquire the state parameters of the speaker 1301.

[0035] An inertial measurement unit (IMU) is a sensor module used to measure the attitude and acceleration of an object along three axes (such as X, Y, and Z axes). In some embodiments, an IMU includes a three-axis gyroscope and a three-axis accelerometer. A magnetic field sensor (such as a magnetometer) is a sensor used to detect the strength and direction of a magnetic field.

[0036] In some embodiments, the inertial measurement unit can be used to detect the tilt angle, rotation angle, and horizontal state of the speaker 1301.

[0037] In some embodiments, a magnetic field sensor can be used to detect the orientation of the speaker 1301, and an auxiliary logic judgment unit 2102 can determine the positional relationship between the speaker 1301 and surrounding magnetic objects or structures.

[0038] The environmental sensing unit 21012 refers to a sensing device or module used to acquire the spatial relationship between the speaker 1301 and its surrounding environment. In some embodiments, the environmental sensing unit 21012 can acquire environmental distance data of the speaker 1301.

[0039] In some embodiments, the environmental sensing unit 21012 may include an infrared sensor array.

[0040] An infrared sensor array is a collection of multiple infrared sensors used for multi-directional detection. For example, an infrared sensor array can consist of multiple infrared ranging probes (such as Time-of-Flight (ToF) probes) distributed around the periphery of the speaker 1301 (e.g., arranged at equal intervals around the central axis of the speaker 1301). Each infrared ranging probe is configured to independently emit an infrared beam and receive reflected signals, thereby measuring the distance to objects in that specific direction. This array-style multi-point deployment ensures that the system can achieve 360-degree coverage of the environment surrounding the speaker 1301, avoiding blind spots.

[0041] For more information regarding the status parameters and ambient distance data of speaker 1301, please refer to [link / reference]. Figure 3 Related descriptions.

[0042] The logic judgment unit 2102 refers to the unit used to determine whether the placement state of the speaker 1301 has changed. For example, a logic processor.

[0043] In some embodiments, the logic judgment unit 2102 may determine the current placement state of the speaker 1301 based on evaluation parameters; and determine whether the placement state of the speaker 1301 has changed based on the historical placement state and the current placement state.

[0044] The status control unit 2103 refers to the unit that controls the parameters of the audio processing unit 13011 of the speaker 1301. For example, it is an audio processing parameter controller, a cooperative system coordinator, etc.

[0045] In some embodiments, the state control unit 2103 may, in response to a change in the placement state of the speaker 1301, determine, based on evaluation parameters and an acoustic environment model, the acoustic feature identifiers and corresponding audio adjustment schemes for the current placement state; and adjust the parameters of the audio processing unit 13011 of the speaker 1301 based on the audio adjustment scheme. For more information on the acoustic environment model, acoustic feature identifiers, audio adjustment schemes, and parameters of the audio processing unit 13011, please refer to [link to relevant documentation]. Figure 3 Related descriptions.

[0046] It should be noted that the above description of the speaker sound field adaptive system and its modules is for convenience only and should not be construed as limiting this specification to the embodiments described. It is understood that those skilled in the art, after understanding the principles of this system, may arbitrarily combine the various modules or construct subsystems connected to other modules without departing from these principles. In some embodiments, Figure 2 The control devices and speakers disclosed herein can be different modules within a system, or a single module can perform the functions of the two modules mentioned above. For example, modules can share a single storage module, or each module can have its own independent storage module. Such variations are all within the scope of protection of this specification.

[0047] Figure 3 This is an exemplary flowchart illustrating a speaker sound field adaptive method according to some embodiments of this specification. Figure 3 As shown, process 300 includes steps 310 to 350. In some embodiments, process 300 may be executed by control device 210.

[0048] Step 310: Obtain the speaker's evaluation parameters.

[0049] A speaker is a device used to convert electrical signals into sound signals. For example, a speaker can be a smart speaker, a wireless speaker, or a sound-producing unit in a stereo system.

[0050] Evaluation parameters refer to data used to assess the physical state and environmental characteristics of a speaker. In some embodiments, evaluation parameters may include speaker status parameters and / or environmental distance data.

[0051] State parameters refer to data that characterize the speaker's physical posture and spatial orientation. For example, state parameters may include one or more of the speaker's three-dimensional posture angles, absolute orientation, relative / absolute height (the height of the speaker in vertical space), and horizontal reference plane deviation (the angle between the bottom surface of the speaker and the horizontal plane in the direction of gravity).

[0052] Environmental distance data refers to data characterizing the spatial distance between a speaker and surrounding objects or boundaries. For example, environmental distance data can be a set of distance values ​​detected by the speaker in different directions relative to obstacles such as walls and furniture. In some embodiments, environmental distance data can be represented as distance measurements in multiple directions centered on the speaker. For example, environmental distance data can be represented as a distance vector containing eight dimensions, representing the distances to the nearest obstacles in the front, back, left, right, and four diagonal directions of the speaker (e.g., [0.5m, 0.7m, 0.2m, ...]).

[0053] For example, state parameters can be obtained through the speaker's built-in attitude sensing unit. The inertial measurement unit can integrate devices such as gyroscopes and accelerometers to measure and output the speaker's three-dimensional attitude angles in real time.

[0054] For example, environmental distance data can be acquired through an environmental sensing unit mounted on the speaker. In some embodiments, the environmental sensing unit may integrate multiple ultrasonic sensors, which can emit ultrasonic waves in different directions and receive echoes; the control device 210 can determine the distance to surrounding walls or obstacles by calculating the flight time from the emission to the reception of the ultrasonic waves. In some embodiments, the environmental sensing unit may integrate a LiDAR (Light Detection and Ranging) or Time of Flight (ToF) camera to obtain more accurate and comprehensive 360-degree environmental distance information.

[0055] In some embodiments, the control device 210 may also acquire evaluation parameters in other ways. For example, it may capture images of the surrounding environment using the speaker's camera and use computer vision algorithms to identify corners, walls, etc., thereby estimating environmental distance data.

[0056] In some embodiments, state parameters are acquired through an attitude sensing unit, and the state parameters include the speaker's three-dimensional attitude angle and absolute orientation; the attitude sensing unit includes an inertial measurement unit and a magnetic field sensor, and the inertial measurement unit includes a three-axis gyroscope and a three-axis accelerometer.

[0057] In some embodiments, environmental distance data is acquired through an environmental sensing unit, which includes an infrared sensor array distributed around the periphery of the speaker enclosure. The infrared sensor array is configured to construct a polar coordinate distance map centered on the speaker, and the environmental distance data is determined based on the polar coordinate distance map.

[0058] For more information on the attitude sensing unit, environment sensing unit, and infrared sensor array, please refer to [link to relevant documentation]. Figure 2 Related descriptions.

[0059] In some embodiments, when the speaker is moved by an external force, a three-axis accelerometer can capture displacement acceleration, a three-axis gyroscope can capture instantaneous angular velocity during the flipping process, and a magnetic field sensor can re-lock the absolute position of the speaker after the movement stops.

[0060] Three-dimensional attitude angles refer to angles that describe the rotational attitude of a speaker in three-dimensional space. In some embodiments, three-dimensional attitude angles may include pitch, roll, and yaw.

[0061] Absolute heading refers to the angle of a speaker on a horizontal plane relative to a geographic or geomagnetic reference direction. For example, absolute heading can be the angle between the front of the speaker (such as the center line of the front panel) and the magnetic north pole.

[0062] The inertial measurement unit (IMU) includes a three-axis gyroscope that measures the angular velocity of the speaker around three axes in real time, and a three-axis accelerometer that measures the acceleration of the gravitational component. In some embodiments, the control device 210 can acquire the raw physical signals collected by the three-axis gyroscope and the three-axis accelerometer, eliminate sensor drift through an internally integrated attitude fusion algorithm (such as complementary filtering or Kalman filtering), and calculate a high-precision three-dimensional attitude angle.

[0063] In some embodiments, the control device 210 can measure the local geomagnetic field vector using a magnetic field sensor. By projecting this vector onto a horizontal plane, the yaw angle of the speaker relative to the geomagnetic north pole, i.e., the absolute orientation, can be calculated. In some embodiments, the control device 210 can also employ data fusion algorithms such as Kalman filtering and complementary filtering to fuse data from the gyroscope, accelerometer, and magnetic field sensor, ultimately obtaining a stable and accurate three-dimensional attitude angle and absolute orientation.

[0064] In some embodiments, state parameters can also be obtained in other ways. For example, the attitude sensing unit can also integrate a visual sensor, and the control device 210 can use visual SLAM (simultaneous localization and mapping) technology to analyze changes in the features of the surrounding environment image to help determine the speaker's attitude and orientation.

[0065] A polar coordinate distance map is a map that represents the distribution of surrounding objects using angles and radial distances, centered on a specific point (e.g., the center of a speaker). In some embodiments, multiple infrared sensors (e.g., eight ToF infrared sensors) can be evenly distributed along the outer perimeter of the speaker enclosure (e.g., a horizontal circle). Each sensor is responsible for detecting the distance to obstacles at a specific angular direction (e.g., one every 45 degrees). The control device 210 can construct a polar coordinate distance map consisting of multiple discrete points, using the geometric center of the speaker as the pole, the installation angle of each sensor as the polar angle, and the distance measured by the corresponding sensor as the polar radius. In some embodiments, the polar coordinate distance map can also be constructed using a single rotatable infrared sensor. For example, a laser rangefinder (e.g., LiDAR) can be placed on top of the speaker and continuously acquired by scanning 360 degrees at a constant speed to obtain distance information at various angles, thereby constructing a polar coordinate distance map.

[0066] For example, a polar coordinate distance map can be represented as {(r1, α1), (r2, α3), ..., (rn, αn)}. Here, α represents the azimuth angle (polar angle) of the infrared sensors arranged around the speaker, r represents the real-time distance (polar radius) between the obstacle and the speaker measured by the azimuth sensor, and n is the total number of infrared sensors. By sequentially connecting these discrete radial values ​​in a polar coordinate system, a "horizontal cross-sectional profile" of the physical space where the speaker is located can be drawn.

[0067] In some embodiments, the control device 210 determines environmental distance data based on various angles and corresponding distance values ​​in a polar coordinate distance map. For example, the environmental distance data can be represented as a multi-dimensional environmental distance vector, where each element of the vector corresponds to a distance measurement value in an angular direction. For example, assuming that eight infrared ranging probes (spaced 45 degrees apart) are arranged at equal intervals around the speaker's outer perimeter with the central axis of the speaker as a reference, the environmental distance data can be represented as (d1, d2, d3, d4, d5, d6, d7, d8), where d1, d2, d3, ... represent the distances between the speaker and obstacles in the 0-degree direction (speaker's positive direction), 45-degree direction, 90-degree direction, ... directions, respectively.

[0068] For example, if the back of the speaker is flush against the wall, the elements in the environmental distance data corresponding to the back of the speaker (such as d4, d5, and d6) will have characteristic values ​​close to 0 or less than a preset threshold (such as 10cm), while the values ​​of elements in other directions will be relatively large.

[0069] In some embodiments of this specification, by fusing an inertial measurement unit and a magnetic field sensor, the speaker's three-dimensional attitude angles and absolute orientation are accurately obtained, effectively suppressing the cumulative drift of attitude calculation. Simultaneously, a polar coordinate distance map constructed using an infrared sensor array enables comprehensive perception of the speaker's surrounding environment. This method of combining precise self-attitude data with distance data from the surrounding environment provides a high-dimensional and accurate data foundation for adaptive sound field adjustment, making sound field optimization more targeted and accurate, and significantly improving the listening experience in complex spatial layouts.

[0070] Step 320: Determine the current placement of the speaker based on the evaluation parameters.

[0071] Placement status refers to the physical posture and orientation of a speaker in space. For example, placement status can be represented by a set of data describing its posture, orientation, and relationship to its surroundings. Placement status can include the current placement status and historical placement status.

[0072] The current placement state refers to the placement state of the speaker at the current moment. In some embodiments, the current placement state of the speaker can be represented by a combination of the speaker's current state parameters and environmental distance data.

[0073] For example, state parameters and environmental distance data can be fused into a multidimensional feature vector, which represents the current placement state of the speaker. For instance, if the acquired state parameters are three-dimensional attitude angles [15°, -3°, 90°] and absolute orientation [120°], and the acquired environmental distance data is a vector representing distance values ​​in eight directions [0.5, 0.7, 0.6, 0.9, 0.8, 1.2, 1.1, 0.6], the control device 210 can concatenate the acquired three-dimensional attitude angles, absolute orientation, and environmental distance data into a current placement state vector [15, -3, 90, 120, 0.5, 0.7, 0.6, 0.9, 0.8, 1.2, 1.1, 0.6] to characterize the speaker's placement state in the current three-dimensional coordinate system.

[0074] In some embodiments, the determination of the current placement state may also employ other feature fusion techniques, such as weighted fusion or feature extraction and fusion through a small neural network, to generate a more representative state vector.

[0075] Step 330: Based on the historical placement status and the current placement status, determine whether the placement status of the speaker has changed.

[0076] Historical placement state refers to the placement state of the speaker recorded at a certain point in the past. For example, the historical placement state could be the placement state recorded when the speaker last completed sound field adaptive adjustment. In some embodiments, the control device 210 can acquire historical state parameters (including historical three-dimensional attitude angles and historical absolute orientation) and historical environmental distance data collected when the sound field adaptive adjustment was last completed, and concatenate the historical three-dimensional attitude angles, historical absolute orientation, and historical environmental distance data to form a historical placement state vector to represent the historical placement state.

[0077] In some embodiments, the control device 210 can determine whether the placement state has changed by calculating the difference between the current placement state vector and the historical placement state vector and comparing the difference with a preset change threshold (such as a preset threshold).

[0078] For example, control device 210 can calculate the Euclidean distance between the current placement state vector and the historical placement state vector as the difference. If the calculated Euclidean distance is greater than a preset distance threshold (e.g., 0.1), it is determined that the placement state has changed. As another example, control device 210 can calculate the cosine similarity between two state vectors. If the calculated cosine similarity is less than a preset similarity threshold (e.g., 0.98), it is determined that the difference between the two vectors is significant, i.e., the placement state has changed.

[0079] In some embodiments, in response to a change in the placement of the speaker, the control device 210 continues to execute steps 340 to 350 below.

[0080] Step 340: Based on the evaluation parameters, determine the acoustic feature identifiers and corresponding audio adjustment schemes for the current placement state through the acoustic environment model.

[0081] The acoustic environment model is a machine learning model.

[0082] Acoustic signature refers to a qualitative description of the acoustic properties of the physical environment in which a speaker is currently located. In some embodiments, acoustic signature may include signature type, location of occurrence, frequency band of influence, and degree of influence. For more information on acoustic signature, please refer to [link to relevant documentation]. Figure 5 And its related descriptions.

[0083] An audio adjustment scheme refers to a set of audio processing parameters set to improve the sound quality of speakers. In some embodiments, an audio adjustment scheme may include adjustment values ​​for parameters such as volume, frequency response, channel balance, and speaker delay. For more information on audio adjustment schemes, please refer to [link to relevant documentation]. Figure 5 And its related descriptions.

[0084] In some embodiments, each audio adjustment scheme corresponds to a specific acoustic feature identifier, and this correspondence is reflected in customized compensation logic for different acoustic environment characteristics. For example, if the feature type of the acoustic feature identifier is "sound reflection enhancement", the corresponding audio adjustment scheme adopts suppression equalization (Cut EQ) processing; if the feature type of the acoustic feature identifier is "sound absorption attenuation", the corresponding audio adjustment scheme adopts compensatory gain (Boost EQ) processing, and so on.

[0085] In some embodiments, the control device 210 can input evaluation parameters into a pre-trained acoustic environment model, obtain the model's output results, and determine the acoustic feature identifiers and corresponding audio adjustment schemes under the current placement state.

[0086] An acoustic environment model is a model used to establish a mapping relationship between the placement of speakers and their acoustic characteristics. In some embodiments, the acoustic environment model can be a machine learning model, such as a convolutional neural network (CNN) or a deep neural network (DNN).

[0087] In some embodiments, the inputs to the acoustic environment model include evaluation parameters (state parameters, environmental distance data), historical state parameters and historical environmental distance data from the last time the sound field adaptive adjustment was completed; the outputs include acoustic feature identifiers and audio adjustment schemes corresponding to the acoustic feature identifiers.

[0088] In some embodiments, the acoustic environment model can be obtained by training an initial acoustic environment model using multiple sets of first training samples with first labels. In some embodiments, the first training samples may include sample evaluation parameters, sample historical state parameters, and sample historical environmental distance data. In some embodiments, the first training samples and their corresponding first labels may be obtained based on preferred adjustment records from historical adjustment records. Historical adjustment records include records of multiple acoustic tests and parameter adjustments performed in the past, such as recording the evaluation parameters at each adjustment, the acoustic feature identifier at the time of adjustment, the state parameters before adjustment, the environmental distance data before adjustment, the audio adjustment scheme used, and the results of objective acoustic index tests (such as frequency response flatness, transient response index) or subjective listening evaluation results after adjustment. Preferred adjustment records refer to records in historical adjustment records that, after parameter optimization, have reached the preset standard through objective acoustic index tests or have been confirmed as having excellent sound quality through subjective listening evaluation.

[0089] In some embodiments, the control device 210 may use the evaluation parameters during adjustment, the state parameters before adjustment, and the environmental distance data before adjustment from a set of preferred adjustment records as sample evaluation parameters, sample historical state parameters, and sample historical environmental distance data, respectively, to generate a set of first training samples; and use the acoustic feature identifiers during adjustment and the audio adjustment scheme used from the set of preferred adjustment records as the first training label corresponding to the first training sample.

[0090] In some embodiments, the control device 210 can input multiple first training samples with first labels into an initial acoustic environment model, construct a loss function using the first labels and the results of the initial acoustic environment model, and iteratively update the parameters of the initial acoustic environment model based on the loss function through gradient descent or other methods. When preset conditions are met, the model training is complete, and a trained acoustic environment model is obtained. The preset conditions may include loss function convergence, the number of iterations reaching a threshold, etc.

[0091] Step 350: Based on the audio adjustment scheme, adjust the parameters of the speaker's audio processing unit.

[0092] For more information about the audio processing unit, please refer to [link / reference]. Figure 1 Related descriptions.

[0093] In some embodiments, the control device 210 generates tuning instructions based on an audio adjustment scheme; the tuning instructions are then sent to an audio processing unit, which adjusts the speaker parameters based on the tuning instructions. For example, if the audio adjustment scheme indicates a -3dB attenuation of the 80Hz frequency band, the control device 210 converts this indication into a tuning instruction that the audio processing unit (such as a DSP) can recognize and sends it to the audio processing unit. The audio processing unit then adjusts the filter coefficients of the parametric equalizer (PEQ) according to the tuning instructions, thereby changing the speaker's frequency response in real time and suppressing excessive low frequencies.

[0094] In some embodiments, the control device 210 is further configured to: acquire recorded data of the speaker in a specific placement state, the recorded data including state parameters, environmental distance data, and scheme adjustment data corresponding to the specific placement state; based on the recorded data, a preference prediction model is generated; the preference prediction model is a machine learning model; in response to the matching degree between the current placement state and the specific placement state meeting a preset condition, based on the recorded data, the preferred prediction model is used to determine a corrected audio adjustment scheme.

[0095] A specific placement state refers to a particular placement state that the speaker has historically been in, with relevant recorded data. For example, a specific placement state could be "placed in the upper left corner of a wooden desk in a study, 20 centimeters away from the back wall." In this case, the control device 210 would record a set of state parameter vectors (such as a pitch angle of 0°) and an environmental distance data vector at that moment. In some embodiments, when the control device detects that the user has manually intervened in the speaker's parameters through a companion app or physical knob, and the parameters are not modified again within a preset time period (such as 5 minutes), a "snapshot" mechanism will be triggered. The control device 210 will then mark and store the current placement state as a specific placement state.

[0096] Recorded data refers to archived information associated with a speaker's specific placement. For example, recorded data may include state parameters for that specific placement, environmental distance data, audio adjustment schemes, and adjustment data of the schemes used (such as adjustment data of schemes manually adjusted by the user).

[0097] Scheme adjustment data refers to the modification or setting data made to the audio adjustment scheme. For example, scheme adjustment data can be a record of data on the user's modification of the audio adjustment scheme generated by the control device 210 based on subjective listening preferences. For instance, assuming that the low-frequency gain indicated in the audio adjustment scheme generated by the control device 210 is -3dB, but the user manually adjusts it to +2dB, then the scheme adjustment data can be recorded as the incremental value "+5dB".

[0098] In some embodiments, when the control device 210 marks and stores a specific placement state of the speaker (such as when it detects that the user has manually intervened in the audio adjustment scheme), it will automatically associate the intervention behavior in this specific placement state with the speaker state at that time and archive it. The archived content is a record data, which includes the state parameters when the intervention occurred, environmental distance data, audio adjustment scheme, and specific data of the user's manual adjustment (i.e., scheme adjustment data).

[0099] For example, when a user manually increases the low-frequency gain by 3dB in the audio adjustment settings of a speaker placed in a corner via a mobile app, the control device 210 will capture this "low-frequency gain +3dB" adjustment data and simultaneously record the audio adjustment settings generated by the control device 210, the speaker's status parameters, and the ambient distance data. This data will be packaged together into a single record and stored locally or on a cloud server.

[0100] In some embodiments, the control device 210 may also acquire the recorded data in other ways. For example, the control device may periodically poll the audio parameter settings, and trigger the recording process when it finds that the settings are inconsistent with those automatically configured by the control device.

[0101] A preference prediction model is a model used to learn and predict a user's personalized preferences for audio effects. In some embodiments, the preference prediction model can be a machine learning model, such as a convolutional neural network (CNN), a multilayer perceptron (MLP), or a gradient boosting decision tree (GBDT).

[0102] In some embodiments, the input to the preference prediction model may include the speaker's current state parameters, environmental distance data, and audio adjustment scheme; the output may include user preference parameters. The user preference parameters refer to the predicted possible adjustments the user might make to the upcoming audio adjustment scheme. For example, if the audio adjustment scheme generated by the control device 210 has a low-frequency gain of -2dB, but the preference prediction model predicts that the user might manually increase the low-frequency gain by 5dB (i.e., adjust it to +3dB), then the user preference parameter for low-frequency gain would be an increase of 5dB.

[0103] In some embodiments, the prediction preference model can be obtained by training an initial prediction preference model with multiple sets of second training samples with second labels.

[0104] In some embodiments, the second training sample and its corresponding second label can be obtained based on the aforementioned recorded data. For example, the control device 210 can use the state parameters, environmental distance data, and audio adjustment scheme in a certain recorded data as the sample state parameters, sample environmental distance data, and sample audio adjustment scheme in the second training sample; and use the scheme adjustment data in the recorded data as the second label corresponding to the second training sample.

[0105] The method of training a preference prediction model based on the second training sample and its corresponding second label is similar to the aforementioned method of training an acoustic environment model based on the first training sample and its corresponding first label, and will not be repeated here.

[0106] Matching degree refers to the degree of similarity between two states or sets of data. In some embodiments, matching degree can be the degree of similarity between the current placement state and a specific placement state in the historical record in terms of physical feature space.

[0107] In some embodiments, the control device 210 can combine the state parameters corresponding to the current placement state and the environmental distance data to form a multi-dimensional feature vector (current placement state vector); combine the state parameters and environmental distance data of each specific placement state stored in the historical records to form each multi-dimensional feature vector (specific placement state vector); and calculate the similarity between the current placement state vector and each specific placement state vector, which is the matching degree between the current placement state and each specific placement state.

[0108] In some embodiments, the matching degree can be calculated using Euclidean distance; the smaller the distance, the higher the matching degree. In some embodiments, the matching degree can also be calculated using other vector similarity measures such as cosine similarity and Mahalanobis distance.

[0109] The revised audio adjustment scheme refers to the final adjustment scheme after adding user preference parameters to the generated audio adjustment scheme.

[0110] In some embodiments, when the matching degree between a specific placement state and the current placement state meets a preset condition, the control device 210 inputs the current state parameters, environmental distance data, and audio adjustment scheme into the preference prediction model. The preference prediction model then outputs the aforementioned user preference parameters. The control device 210 applies the user preference parameters to the audio adjustment scheme to obtain a corrected audio adjustment scheme. For example, when the determined user preference parameter is "low-frequency gain +2.8dB", the control device 210 can add an additional 2.8dB of low-frequency gain to the low-frequency gain value indicated by the audio adjustment scheme to form the final corrected audio adjustment scheme for playback.

[0111] The preset condition can be that the matching degree is greater than the matching degree threshold. The matching threshold can be a value between 0 and 1 preset based on historical experience or actual needs, such as 0.95.

[0112] In some embodiments of this specification, a model capable of predicting user preferences is trained by acquiring and learning manual sound tuning data from users in specific placement states. When the speaker is in a similar state again, the control device can automatically incorporate the user's personalized preferences based on the generated audio adjustment scheme, such as automatically boosting the bass that the user is accustomed to adding in that scenario. This achieves an upgrade from "environmental adaptation" to "user preference adaptation," making sound effect adjustment both scientific and personalized, significantly improving the user experience, and reducing repetitive manual operations by the user.

[0113] In some embodiments of this specification, the control device can automatically acquire the speaker's evaluation parameters and intelligently detect whether its placement has changed. When a change is detected, a pre-trained acoustic environment model can accurately map the current complex physical placement state to specific acoustic problems and corresponding optimization solutions. This avoids the complex process of manual adjustment by the user, achieving a "place and use" intelligent experience. Compared with traditional fixed audio parameters or simple scene modes, it can more effectively solve problems such as low-frequency booming and blurred sound field caused by improper placement (such as near a corner), significantly improving sound quality and listening experience, enabling the speaker to exhibit optimal acoustic performance in various home environments.

[0114] Figure 4 This is an exemplary flowchart illustrating the determination of collaborative compensation parameters according to some embodiments of this specification. Figure 4 As shown, process 400 includes steps 410 to 430. In some embodiments, process 400 may be executed by control device 210.

[0115] Step 410: Discover and identify the collaborative speaker through the wireless communication protocol.

[0116] A wireless communication protocol is a communication standard or specification used to establish data transmission links and achieve identity handshakes and information exchange between different electronic devices. In some embodiments, wireless communication protocols may include Bluetooth, Wi-Fi, ZigBee, etc.

[0117] Collaborating speakers refer to other speakers that work in conjunction with the current speaker (i.e., the speaker whose collaboration parameters are to be determined) to form an audio system. It's important to note that the roles of "current speaker" and "collaborating speaker" are relative, not physically fixed master-slave relationships. For example, in a stereo system consisting of speaker A and speaker B, when speaker A performs environmental awareness and parameter adjustment, speaker A is the current speaker, and speaker B is speaker A's collaborating speaker; conversely, when speaker B performs environmental awareness and parameter adjustment, speaker B is considered the current speaker, and speaker A becomes speaker B's collaborating speaker.

[0118] In some embodiments, after a speaker connects to a local Wi-Fi network, it broadcasts a specific discovery protocol data packet. Upon receiving this data packet, other compatible collaborating speakers within the same network respond with information such as their device identifier, model, and supported collaborating capabilities. The current speaker's control device 210 can then identify available collaborating speakers based on this response data and establish a communication link with them.

[0119] In some embodiments, the process of discovering and identifying the collaborating speaker can also be implemented via the Bluetooth protocol. For example, the current speaker can enable Bluetooth Low Energy (BLE) scanning mode and continuously listen to surrounding Bluetooth broadcasts. The collaborating speaker will then send data packets containing its identity information in broadcast mode; after the current speaker's control device 210 scans these broadcasts, it can identify the nearby collaborating speaker, make a preliminary judgment on its approximate distance based on the signal strength, and then establish a Bluetooth pairing connection.

[0120] Step 420: Determine the relative positional relationship between the speaker and the cooperating speaker through acoustic positioning to form a speaker topology map.

[0121] Acoustic positioning refers to the technical methods used to measure and determine the location of objects by utilizing acoustic characteristics. For example, acoustic positioning methods can include positioning techniques based on signal time of flight (ToF) or time difference of arrival (TDOA), such as ultrasonic ranging or broadband chirp positioning.

[0122] Relative positional relationship refers to the spatial position information of one speaker relative to another speaker. In some embodiments, the relative positional relationship may include the straight-line distance, horizontal angle, and height difference between the two speakers. For example, the relative positional relationship can be described as: "The collaborating speaker is located at a 45-degree angle to the right front of the current speaker, with a straight-line distance of 2.5 meters and the same height."

[0123] In some embodiments, each speaker takes turns playing a high-frequency coded acoustic signal (such as ultrasound above 18kHz) that is imperceptible to humans, and the microphone arrays on the other speakers receive the signal; the control device 210 can accurately calculate the direction (i.e., azimuth and pitch) of the sound source relative to the receiving speaker by analyzing the minute time differences in the arrival of the signal at different microphones in the array; and by combining the direction data measured by all speakers, the relative positional relationship between each speaker is calculated using a triangulation algorithm.

[0124] In some embodiments, the relative positional relationship can also be determined based on Time-of-Flight (ToF) technology. For example, all speakers in the scene first achieve precise time synchronization via Network Time Protocol (NTP); one speaker is controlled to emit a positioning signal at a precise time t1, and the receiving speaker records the arrival time t2 of the signal; the time difference (t2-t1) is calculated and multiplied by the speed of sound to obtain the straight-line distance between the two speakers. By measuring the distances between all pairs of speakers, the control device 210 can use a polygonal positioning algorithm to construct their relative positional relationship in space.

[0125] In some embodiments, in addition to acoustic positioning, other positioning technologies can be combined to assist or independently determine the relative positional relationship. For example, high-precision ranging can be performed using ultra-wideband (UWB) radio signals, or the position can be determined by visual positioning technology (such as using a camera to identify specific markers).

[0126] A speaker topology map is a data image describing the geometric distribution of all speakers in a scene under a unified coordinate system. In some embodiments, a speaker topology map can be a spatial structure diagram with each speaker as a node and their relative positional relationships as connection information.

[0127] In some embodiments, after obtaining the relative positional relationships between all speakers, the control device 210 establishes a local three-dimensional coordinate system with one of the speakers (such as the user-preset or system-default main speaker) as the origin. Then, it integrates the positional information of all speakers (such as (x, y, z) coordinates) into this coordinate system to form the final speaker topology map. For example, the control device 210 uses a multilateration algorithm or a graph optimization algorithm to map the relative positional relationships between the speakers to a unified Cartesian coordinate system, calculates the global relative coordinates of each speaker, and finally generates a complete speaker topology map.

[0128] Step 430: Determine the collaborative compensation parameters based on the speaker topology map.

[0129] Co-compensation parameters refer to a set of adjustment parameters set to optimize the overall sound field of a system composed of multiple speakers. For example, co-compensation parameters may include group delay parameters, crossfeed parameters, and / or co-equalization parameters.

[0130] Group delay parameters refer to the time delay value applied to each channel to ensure that the sound emitted by each speaker arrives at the listening position synchronously. The purpose of setting group delay parameters is to ensure that the sound waves emitted by each speaker arrive at the target listening position simultaneously (this can be preset by the user). For example, if the speaker topology map shows that speaker A is 3 meters away from the target listening position and the cooperating speaker B is 2 meters away, the control device will set a delay parameter of approximately 2.9 ms for speaker B (calculated based on the speed of sound 343 m / s), thereby achieving precise sound positioning and avoiding sound dispersion.

[0131] Crossfeed parameters are control parameters used to adjust the degree of mixing between the left and right channel signals, thereby adjusting the width and depth of the stereo sound field. For example, in a stereo pair, the control device 210 can use crossfeed parameters to delay the left channel signal by about 300μs and attenuate the high frequencies before mixing it into the right channel, simulating the "inter-auditory time difference (ITD)" and "inter-auditory level difference (ILD)" of human hearing, so that the user can feel a wide sound field even if they are not in the center.

[0132] Co-equalization parameters refer to equalization adjustment parameters set to correct frequency response imbalances caused by sound wave interference or reflection in a multi-speaker system. For example, if speaker A is in an open area and co-speaker B is in a corner, the control device 210 will generate a special low-frequency suppression equalization parameter for speaker B, while increasing the gain for speaker A, so that the two devices achieve a balance in energy output and spectrum performance.

[0133] In some embodiments, the process of determining the collaborative compensation parameters is performed on a round-robin basis for each speaker in the system (e.g., distributed or polling processing). For example, assuming there are speakers A and B in the scenario, speaker A is first designated as the current speaker (while speaker B is the collaborating speaker), and the group delay parameters and cross-feed parameters for speaker A are calculated. Subsequently, or in parallel, speaker B is switched to the current speaker (while speaker A is the collaborating speaker), and the group delay parameters and cross-feed parameters for speaker B are calculated.

[0134] In some embodiments, the control device 210 may determine the collaborative compensation parameters based on the speaker topology map using various methods. In some embodiments, the control device 210 may determine the collaborative compensation parameters based on the speaker topology map using a sound field optimization algorithm. For example, the control device 210 can extract geometric distribution features reflecting the spatial distribution of each speaker from the speaker topology map through graph theory analysis or geometric coordinate analysis (such as the relative radial distance of each speaker node relative to the target listening position, the angle between adjacent speakers, and the spatial symmetry of each speaker relative to the main axis). The extracted geometric distribution features are quantitatively compared with the preset ideal sound field model in the memory to identify the output sound field deviation. For example, the current geometric distribution features are mapped to the ideal model space to identify the deviation between the actual physical position and the ideal acoustic position. For the identified sound field deviation, the control device 210 can use a preset acoustic optimization algorithm to calculate the cooperative compensation parameters that can minimize the sound field deviation. For example, the control device 210 can use the Least Mean Square (LMS) algorithm or the Least Squares Optimization algorithm to construct a cost function with the goal of minimizing the sound field deviation and perform optimization to obtain the cooperative compensation parameters of each speaker.

[0135] The ideal sound field model is an optimal placement reference defined by standard acoustic specifications (such as ITU-R BS.775). For example, if the ideal sound field model is a standard stereo triangle placement, but the actual geometric distribution characteristics show that the left speaker is far from the target listening position and the azimuth angle deviates from the preset value, then the control device 210 can identify two types of sound field deviations: "sound center offset" and "phase distortion".

[0136] In some embodiments, the control device encapsulates the calculated collaborative compensation parameters into control commands and sends them to each collaborative speaker via a wireless communication link. The audio processing unit of each collaborative speaker receives and applies these control commands (adjusting its own parameters according to the control commands), thereby coupling the originally scattered physical devices into a unified sound field system that meets the requirements of an ideal model in terms of acoustic performance.

[0137] In some embodiments, the group delay parameter is determined based on the geometric distance difference between each speaker and a preset virtual listening point in the speaker topology map; the cross-feed coefficient is determined based on the horizontal angle and straight-line distance between the left and right speakers in the speaker topology map; the cooperative equalization parameter is determined based on the abnormal response, which is caused by the indirect acoustic path. The indirect acoustic path is the path of sound emitted by any speaker after being reflected by the room boundary to the location of other speakers, which is identified based on the speaker topology map.

[0138] A preset virtual listening point refers to a reference target location defined in the sound field space for sound field optimization calculations. For example, the preset virtual listening point can be the center point of the area where the user's seat is located, or it can be a point preset by the user. For instance, in a two-speaker stereo mode, the preset virtual listening point can be set on the extension line of the central axis of the two speakers, and at the vertex of an equilateral triangle formed by the two speakers.

[0139] The geometric distance difference refers to the difference in the straight-line distance between different speakers and the same target point. For example, if speaker A is 3.2 meters away from the preset virtual listening point and speaker B is 2.8 meters away from the preset virtual listening point, then the geometric distance difference between the two is 0.4 meters.

[0140] In some embodiments, the control device 210 can calculate the straight-line distance (di) from each speaker to a preset virtual listening point based on the speaker topology map; determine the maximum value among all straight-line distances (di) as the reference distance (dref); for any speaker, the corresponding geometric distance difference is (dref - di); divide the geometric distance difference by the speed of sound in air (e.g., about 343 m / s at room temperature), and the result is the group delay parameter that the speaker needs to compensate for.

[0141] In some embodiments, to improve the accuracy of the acquired group delay parameters, the control device 210 can dynamically adjust the sound propagation speed based on information such as the current ambient temperature and humidity. For example, a temperature and humidity sensor can be integrated, and a table showing the correspondence between sound speed and environmental parameters can be queried based on real-time monitored data to obtain a more accurate sound speed value, thereby calculating a more accurate group delay parameter.

[0142] In a multi-speaker cooperative system, left and right speakers refer to audio output devices located in the left and right spatial regions, respectively, relative to the speaker (current speaker) whose cooperative compensation parameters need to be determined. In some embodiments, if the speaker topology map shows that multiple speakers exist in the left or right spatial region of the current speaker, the left and right speakers can be determined based on the straight-line distances between these speakers and the current speaker. For example, the straight-line distances between all speakers located in the left spatial region and the current speaker are calculated, and the speaker with the smallest straight-line distance (i.e., closest) is selected as the "left speaker" of the current speaker; similarly, the speaker with the smallest straight-line distance to the current speaker in the right spatial region is selected as the "right speaker" of the current speaker.

[0143] The horizontal angle refers to the angle between the frontal orientations of two speakers. For example, the horizontal angle between the left and right speakers is the angle between the rays connecting the geometric centers of the left and right speakers with a preset virtual listening point as the vertex. In some embodiments, the control device can calculate the horizontal angle between the left and right speakers based on the coordinates of three points (left speaker, right speaker, and preset virtual listening point) in the speaker topology map, using the cosine theorem or the vector angle formula.

[0144] The straight-line distance refers to the shortest distance between two speakers in space. In some embodiments, the control device 210 can obtain the straight-line distance by extracting the coordinates of the nodes corresponding to the left and right speakers in the speaker topology map and directly calculating the magnitude of their spatial vectors.

[0145] In some embodiments, the control device 210 can consult a preset first preset table based on the horizontal angle and straight-line distance between the left and right speakers to determine the crossfeed parameters. The first preset table records the optimal crossfeed coefficient values ​​corresponding to different combinations of "horizontal angle" and "straight-line distance". The optimal crossfeed coefficient values ​​can be obtained through extensive experiments in a standard acoustic environment. For example, multiple different speaker layouts (different sets of horizontal angles and straight-line distances) can be set up, and professional listeners can subjectively evaluate the listening experience under different crossfeed parameters, selecting the crossfeed parameters with the best sound quality and their corresponding speaker layouts (horizontal angle and straight-line distance) and storing them in the table. In some embodiments, the crossfeed parameters are negatively correlated with the straight-line distance; that is, the crossfeed parameters increase as the straight-line distance or the horizontal angle decreases.

[0146] An anomalous response refers to a non-ideal peak or trough in the frequency response curve caused by the propagation, reflection, and interference of sound waves in the environment. For example, an anomalous response could be an energy surge (standing wave) at 200 Hz or a signal drop at 1 kHz due to reflection cancellation.

[0147] An indirect acoustic path refers to the propagation path of sound from its source, after reflection at the room boundaries, to reach another speaker (or listening point). For example, the path of sound emitted by speaker A after reflection through the side walls into the area where speaker B is located.

[0148] As an example only, an indirect acoustic path refers to a path in which sound does not travel in a straight line, but is reflected by the boundaries of a room, such as the walls, floor, or ceiling, and travels from one speaker to another (or the listening point).

[0149] A room boundary refers to the physical interface of an interior space. For example, a room boundary can include walls, ceilings, and floors.

[0150] In some embodiments, the method for identifying indirect acoustic paths can employ the mirror source method. For example, given a sound source (speaker) and receiver (other speaker locations) in a speaker topology map, and a known room boundary (such as a wall), an indirect acoustic path with a single reflection can be determined by calculating the mirror point of the sound source with respect to that boundary, then connecting the mirror point to the receiver; the intersection of this line with the boundary is the reflection point. By performing this operation on all boundaries, the main indirect acoustic paths can be identified.

[0151] In some embodiments, ray tracing can also be used to identify indirect acoustic paths. A large number of virtual acoustic "rays" are emitted from the sound source in all directions, and the propagation and reflection of each ray between the room boundaries are traced. When a ray reaches the target speaker location after one or more reflections, its trajectory constitutes an indirect acoustic path.

[0152] In some embodiments, after identifying the indirect acoustic path, the control device 210 can calculate the length difference between the direct acoustic path (from the sound source directly to the receiving point) and the identified indirect acoustic path; calculate the phase difference between the direct sound and the reflected sound at different frequencies when they reach the receiving point based on the length difference; and use a comb filtering algorithm to predict the interference effect caused by this phase difference, thereby obtaining a frequency response anomaly curve (containing peaks and troughs) as an anomaly response.

[0153] In some embodiments, after acquiring an abnormal response, the control device 210 can determine the specific characteristics (frequency, bandwidth, gain value) of the abnormal response; based on the specific characteristics of the abnormal response, a set of corresponding inverse filtering parameters (inverse EQ) is generated, which are the collaborative equalization parameters. For example, if the abnormal response shows a +6dB reflection enhancement in a specific frequency band, the collaborative equalization parameters are configured to suppress reflections by -6dB in that frequency band. The collaborative equalization parameters are applied to the source speaker that generates the indirect path or the affected target speaker to counteract the acoustic coloration caused by boundary reflections.

[0154] In some embodiments, the control device 210 may also design an inverse filter whose frequency response characteristics are exactly the opposite of the abnormal curve (i.e., attenuation at the peak and boost at the trough), and the coefficients of the inverse filter are determined as the cooperative equalization parameters.

[0155] In some embodiments of this specification, group delay, crossfeed, and co-equalization parameters are determined through a speaker topology map, precisely correcting the time delay of each speaker and achieving accurate sound image localization. Simultaneously, crossfeed is optimized based on the actual speaker placement, enhancing the immersive stereo sound. More importantly, by identifying and compensating for frequency response anomalies caused by specific indirect acoustic paths, sound coloration caused by room reflections is effectively eliminated, significantly improving the frequency response flatness and sound fidelity at the listening position, and achieving automated, high-precision sound field calibration.

[0156] In some embodiments of this specification, automated sound field calibration of a multi-speaker system is implemented. By automatically discovering and locating the cooperating speakers, constructing a speaker topology map, and calculating the cooperating compensation parameters accordingly, it can accurately compensate for problems such as sound image shift and frequency response imbalance caused by poor speaker placement. This eliminates the need for manual configuration by the user, significantly improving the fidelity of the audio system and the convenience of the listening experience.

[0157] Figure 5 This is a schematic diagram illustrating how to determine whether the placement of the speaker has changed, based on some embodiments of this specification.

[0158] In some embodiments, the state parameters include the three-dimensional attitude angle 5501 and absolute orientation 5502 of the speaker 1301; determining whether the speaker's placement state has changed based on the historical placement state 510 and the current placement state 530 includes, for example... Figure 5 As shown: Based on the historical placement state 510, a historical reference sequence 520 is determined; based on the current placement state 530, a current state sequence 540 is determined; comparing the current state sequence 540 and the historical reference sequence 520, in response to the change 550 of at least one dimension of the three-dimensional attitude angle 5501, absolute orientation 5502, or environmental distance data 5503 exceeding its corresponding preset tolerance threshold, it is determined that the placement state of the speaker has changed.

[0159] For more information on state parameters, environmental distance data, historical placement status, current placement status, 3D attitude angles, and absolute orientation, please refer to [link to relevant documentation]. Figure 3 Related descriptions.

[0160] Historical reference sequence 520 refers to a sequence of state information recorded at a previous point in time, used as a comparison reference. For example, the historical reference sequence could be three-dimensional attitude angle, absolute orientation, and environmental distance data saved after the last sound field adjustment was completed.

[0161] In some embodiments, after the speaker completes a sound field adaptive adjustment, the control device structures and encapsulates the recorded state parameters (e.g., three-dimensional attitude angles and absolute orientation) and the environmental distance data measured by the environmental sensing unit to form a historical reference sequence. This sequence can be timestamped and stored in a storage device as a reference for subsequent state comparisons.

[0162] For example, a historical baseline sequence can be a data object containing multiple key-value pairs, such as {pitch: 1.2°, roll: -0.5°, yaw: 93.4°, heading: 15°_NE, env_vector: [d1, d2, ..., dn]}. Here, “env_vector” is a vector composed of environmental distance data.

[0163] The current state sequence 540 refers to the real-time acquired information sequence describing the speaker's current state. For example, the current state sequence can be a vector composed of currently measured three-dimensional attitude angles, absolute orientation, and environmental distance data.

[0164] In some embodiments, the process of determining the current state sequence is performed in real time or periodically. The control device acquires three-dimensional attitude angles in real time through a built-in attitude sensing unit, obtains the absolute orientation through a magnetometer, and acquires current environmental distance data through an environmental sensing unit. The control device integrates these real-time data into a current state sequence with the same data structure as the historical reference sequence.

[0165] For example, the control device can be set to collect data from each sensor every second and combine it into a current state sequence. This periodic acquisition method ensures that the control device can promptly capture changes in the speaker's state.

[0166] In some embodiments, the current state sequence can also be determined based on event triggering. For example, the control device initiates a complete process of acquiring and determining the current state sequence only when the speaker's built-in motion sensor detects an acceleration change exceeding a certain threshold. This approach can effectively reduce the power consumption of the control device when it is stationary.

[0167] The preset tolerance threshold refers to the pre-defined numerical limits used to determine whether the changes in the three-dimensional attitude angle, absolute orientation, or environmental distance data are significant. For example, different preset tolerance thresholds can be set for the three-dimensional attitude angle, absolute orientation, and environmental distance data. For instance, a preset tolerance threshold T1 can be set for the three-dimensional attitude angle, a preset tolerance threshold T2 can be set for the absolute orientation, and a preset tolerance threshold T3 can be set for the environmental distance data.

[0168] In some embodiments, the preset tolerance threshold can be set by a technician based on experience.

[0169] For example, the preset tolerance threshold T1 for the three-dimensional attitude angle can be set to 5 degrees, because even slight tilting can affect the vertical diffusion of sound; the preset tolerance threshold T2 for absolute orientation can be set to 10 degrees to tolerate slight rotation; and the change in environmental distance data can be measured by calculating the Euclidean distance or cosine similarity between two distance vectors, and a normalized preset tolerance threshold T3 can be set for it. If the change in at least one dimension exceeds its corresponding preset tolerance threshold, the control device determines that the speaker's placement has changed.

[0170] In some embodiments, if a change in the speaker's placement is detected, the control device will automatically trigger a sound field adaptive adjustment process to adjust the parameters of the speaker's audio processing unit. For example, based on the current placement, an audio adjustment scheme is generated using an acoustic environment model and executed in real time to optimize sound output. More information on adjusting the parameters of the audio processing unit can be found in [link to relevant documentation]. Figure 3 Related descriptions.

[0171] Some embodiments in this specification combine three-dimensional attitude angles, absolute orientation, and environmental distance data across multiple dimensions, and set independent preset tolerance thresholds for each dimension, enabling precise identification of positional and attitude changes that substantially affect the sound field. This method avoids unnecessary sound field recalibration triggered by minute vibrations or meaningless slight movements, significantly improving the accuracy and robustness of state change detection compared to schemes relying solely on a single motion sensor. Ultimately, while ensuring adaptive sound quality, it effectively reduces system power consumption and computational resource consumption, optimizing the user experience.

[0172] In some embodiments, the acoustic environment model is a neural network model, which includes a detection layer and a compensation layer connected in sequence; the detection layer is configured to determine the acoustic feature identifiers in the current placement state based on the current state sequence and the historical reference sequence; the compensation layer is configured to determine the audio adjustment scheme based on the current state sequence, the historical reference sequence and the acoustic feature identifiers in the current placement state.

[0173] For more information on acoustic feature identification and audio adjustment schemes, please refer to [link / reference]. Figure 3 Related descriptions.

[0174] In some embodiments, the acoustic feature identifier in the current placement state can be represented by an acoustic feature vector.

[0175] In some embodiments, the control device can construct an acoustic feature vector based on the characteristic type, location, frequency band, and degree of impact of the acoustic anomaly in the current placement state.

[0176] Feature type refers to the classification of the physical properties of acoustic anomalies. Examples include enhanced boundary reflections, ultra-low frequency accumulation, directivity deviation, and phase overlap interference.

[0177] The location of occurrence refers to the location where acoustic anomalies are concentrated, such as the left, right, rear, above, or front.

[0178] The affected frequency band refers to the frequency range in which acoustic anomalies primarily affect the sound. For example, the affected frequency bands include ultra-low frequency (<60Hz), low frequency (60-250Hz), mid frequency (250Hz-2kHz), mid-high frequency (2kHz-6kHz), and high frequency (>6kHz).

[0179] Impact level refers to a quantitative indicator of the severity of acoustic anomalies. In some embodiments, the control device can normalize the impact level to a value between 0 and 1 to characterize the impact level; the larger the value, the greater the impact.

[0180] For example, if the acoustic feature identifier of the current placement state is: the speaker is moved to the right, which reduces the distance between the right side of the speaker and the obstacle, thereby increasing the boundary reflection on the right side of the speaker, and causing the low-frequency boom and the overall sound to shift to the left, then the acoustic feature vector is [feature type: enhanced boundary reflection, location of occurrence: right side of the speaker, affected frequency band: low frequency, degree of influence: 0.5].

[0181] For example, if the acoustic characteristics of the current placement state are: there are obstacles at close range behind and to the side rear of the speaker, causing ultra-low frequency accumulation at the rear boundary of the speaker, resulting in muffled and unclear sound, then the acoustic feature vector is [feature type: standing wave accumulation, location of occurrence: rear, affected frequency band: ultra-low frequency, degree of influence: 0.8].

[0182] In some embodiments, audio adjustment schemes can also be represented as vectors. For example, the control device can construct an audio adjustment vector based on adjustments to volume parameters, frequency response parameters, channel balance parameters, and speaker delay.

[0183] The volume parameter refers to the global volume gain adjustment value of the speaker sound field adaptive system, and the unit is decibel (dB).

[0184] Frequency response parameters refer to the set of gain control parameters for a specific frequency band, including center frequency, gain value, and Q value.

[0185] The center frequency refers to the specific frequency point in a filter where the gain or attenuation effect is most pronounced.

[0186] Gain refers to the amount by which the volume parameter is increased or decreased at the frequency point corresponding to the center frequency.

[0187] The Q value is the quality factor of a filter, which can be used to adjust the bandwidth. The Q value is negatively correlated with bandwidth; the higher the Q value, the narrower the bandwidth.

[0188] Channel balance parameters refer to the gain offset of a specific channel relative to a reference level. For example, +1dB for the left channel, or +2dB for the tweeter. The reference level is the original output level used as a standard reference when adjusting channel balance. In some embodiments, the reference level may be set by a technician based on experience or by a control device based on historical data.

[0189] Audio delay refers to the time delay applied to a specific audio output channel (such as the subwoofer channel, mid-frequency channel, etc.), and the unit can be milliseconds (ms).

[0190] As an example, if the acoustic feature vector corresponding to the acoustic feature identifier in the current placement state is [Feature Type: Enhanced Boundary Reflection, Occurrence Location: Right Side, Affected Frequency Band: Low Frequency, Influence Level: 0.5] (i.e., "Speaker shifted to the right"), the control device can determine an attenuation curve (e.g., center frequency: 100Hz, gain: -3dB, Q value: 1.5) in the low-frequency band (e.g., 80Hz-120Hz) on the right side of the speaker to suppress the newly added boundary reflection on the right side of the speaker. The frequency response corresponding to this attenuation curve is used as the frequency response parameter in the acoustic feature vector. The control device can increase the gain of the mid / high frequency unit on the left side of the speaker to correct the leftward shift of the sound image caused by the enhanced reflection on the right side of the speaker, and use the sound gain on the left side of the speaker as the channel balance parameter in the acoustic feature vector. The corresponding audio adjustment scheme is [Volume parameter: 0dB, Frequency response parameter: {100Hz, -3dB, 1.5}, Channel balance parameter: Left high frequency +1.0dB, Audio delay: 0ms].

[0191] In some embodiments, the acoustic environment model is a neural network model, such as a convolutional neural network (CNN) or a recurrent neural network (RNN). In some embodiments, the neural network model includes a detection layer and a compensation layer connected in sequence.

[0192] The sequential connection between the detection layer and the compensation layer means that the output of the detection layer serves as one of the inputs to the compensation layer, thereby decoupling and processing the diagnosis of acoustic problems (detection layer) and the generation of solutions (compensation layer).

[0193] A detection layer is a functional layer in an acoustic environment model used to identify or detect specific features or patterns in the input data. For example, a detection layer can be configured to identify corresponding acoustic feature identifiers based on an input state sequence.

[0194] In some embodiments, the input to the detection layer is the current placement state sequence and the historical baseline sequence, and the output is the acoustic feature identifier of the current placement state.

[0195] In some embodiments, the detection layer can be trained using a third training sample and a third label. In some embodiments, the control device can acquire the third training sample and the third label based on historical data. For example, the historical data includes multiple acoustic test records performed by the control device over a historical period. The control device can generate the third training sample and its corresponding third label based on multiple preferred diagnostic samples identified in the historical data.

[0196] Preferred diagnostic samples refer to historical data records that accurately correspond to physical states and acoustic defects, verified by acoustic measurement instruments or manually marked by experts.

[0197] In some embodiments, the control device may use the current placement state sequence and historical benchmark sequence corresponding to the preferred diagnostic sample as the third training sample, and use the acoustic feature identifier confirmed in the preferred diagnostic sample as the corresponding third label.

[0198] A compensation layer is a functional layer in an acoustic environment model used to generate corresponding compensation or adjustment schemes based on detected features or sound anomalies. For example, a compensation layer can be configured to determine an audio adjustment scheme for correcting acoustic problems based on acoustic feature identifiers.

[0199] In some embodiments, the input to the compensation layer is the current placement state sequence, the historical reference sequence, and the acoustic feature identifiers under the current placement state, and the output is the audio adjustment scheme.

[0200] In some embodiments, the compensation layer may be obtained by training with a fourth training sample and a fourth label. In some embodiments, the control device may obtain the fourth training sample and the fourth label based on historical data. For example, the historical data includes records of multiple parameter adjustments performed by the control device over a historical period. The control device may determine the fourth training sample and the fourth label based on multiple preferred configuration samples obtained from the historical data.

[0201] The preferred configuration sample refers to the configuration records in historical data that, after adjusting the speaker parameters according to the configuration in the sample, have a high flatness of frequency response curve or a high user satisfaction score.

[0202] In some embodiments, the control device may use the placement state sequence, historical reference sequence, and acoustic feature identifiers in the current placement state corresponding to the preferred configuration sample as the fourth training sample, and the audio adjustment scheme vector finally adopted in the preferred configuration sample as the corresponding fourth label.

[0203] In some embodiments, the acoustic environment model can be obtained by training the detection layer and the compensation layer separately. For the training methods of the detection layer and the compensation layer, please refer to the relevant description of the training methods of the acoustic environment model above.

[0204] In some embodiments, the acoustic environment model can be obtained by jointly training the detection layer and the compensation layer.

[0205] In some embodiments, the control device can input a third training sample into the detection layer to obtain the acoustic feature identifier in the current placement state; then input the third training sample and the acoustic feature identifier as a fourth training sample into the compensation layer, and construct a loss function based on the audio adjustment scheme output by the compensation layer and the fourth label, and iteratively update the parameters of the detection layer and the compensation layer based on the loss function until the preset conditions are met, the training is completed, and the trained detection layer and compensation layer are obtained. The preset conditions can be that the loss function is less than a threshold, convergence, or the training period reaches a threshold.

[0206] In some embodiments of this specification, a neural network model consisting of a detection layer and a compensation layer connected in sequence is employed. This two-stage model structure decouples "problem diagnosis" from "solution generation." First, the detection layer accurately identifies structured acoustic feature identifiers, and then the compensation layer formulates a targeted audio adjustment scheme based on these explicit acoustic feature identifiers. This method not only improves the accuracy and robustness of acoustic environment model adjustment but also makes the intermediate diagnostic results (acoustic feature identifiers) of the acoustic environment model physically interpretable, facilitating analysis and optimization. Compared to end-to-end black-box models, acoustic environment models including detection and compensation layers can achieve more refined and reliable automatic acoustic compensation, significantly improving the user's listening experience in complex acoustic environments.

[0207] In some embodiments, after determining the audio adjustment scheme, the control device can also acquire the ambient noise signal through the reference microphone built into the speaker and calculate the background noise sound pressure level; based on the background noise sound pressure level, adjust the volume parameters and / or frequency response parameters in the audio adjustment scheme.

[0208] For more information about the reference microphone, please see [link / reference]. Figure 2 Related descriptions.

[0209] Environmental noise signals refer to sound signals generated in the environment by non-target sound sources. For example, environmental noise signals can include the sound of air conditioners running, people talking, or traffic noise outside the window.

[0210] The target sound source refers to the sound played based on the audio data built into the speaker.

[0211] In some embodiments, the control device may activate the built-in reference microphone of the speaker to continuously collect sound signals in the physical environment in which the speaker is located, and convert the collected analog sound signals into a digital audio data stream, i.e., an ambient noise signal.

[0212] Background noise sound pressure level is a physical measure that quantifies the intensity of environmental noise signals. For example, background noise sound pressure level can be expressed in decibels (dB).

[0213] In some embodiments, the control device may periodically sample the ambient noise signal acquired by a reference microphone, for example, every 500 milliseconds (ms). Then, the average power of the signal within that sampling period is calculated using the Root Mean Square (RMS) method. Finally, the calculated average power is converted to background noise sound pressure level in decibels (dB) using a standard sound pressure level conversion formula.

[0214] In other embodiments, the control device can also perform a Fast Fourier Transform (FFT) on the acquired ambient noise signal, converting it from the time domain to the frequency domain. Then, the signal energy is calculated in the frequency domain; for example, an A-weighting network can be applied to weight the energy at different frequencies to simulate the auditory characteristics of the human ear. Finally, the weighted total energy is converted into the background noise sound pressure level. This method can more accurately reflect the noise loudness actually perceived by the human ear.

[0215] Those skilled in the art will understand that the background noise sound pressure level can also be calculated in other ways, such as by using a pre-trained acoustic model to directly estimate the sound pressure level value from the environmental noise signal. This specification does not impose any specific limitations on this.

[0216] In some embodiments, the control device can compare the calculated current background noise sound pressure level with a preset reference quiet environment threshold (e.g., 35 dB). Based on the difference between the two, a dynamic gain compensation amount is calculated using a preset mapping relationship or gain algorithm (such as the sigmoid function, logarithmic function, etc.). For example, for every 10 dB increase in background noise, the volume parameter is increased by 3 dB from its original value, and a gain cap is set to prevent excessive volume. Finally, this dynamic gain compensation amount is superimposed on the volume parameter of the original audio adjustment scheme to form the updated volume parameter.

[0217] In some embodiments, the control device can perform spectral analysis on the ambient noise signal to determine the frequency bands where the noise energy is mainly concentrated. For example, if the noise is detected to be mainly concentrated in the low-frequency band of 100Hz-300Hz, a dynamic equalization (EQ) compensation amount can be calculated. This compensation amount is used to moderately boost the response of key frequency bands in the audio (such as 1kHz-4kHz) to ensure that this part of the audio content is not masked by low-frequency noise. Then, this dynamic equalization compensation amount is applied to the frequency response parameters (such as equalizer settings) of the original audio adjustment scheme to form updated frequency response parameters.

[0218] In some embodiments, the control device can also simultaneously adjust volume parameters and frequency response parameters to achieve a better listening experience. Furthermore, the adjustment can also be achieved through other technical means. For example, a pre-trained machine learning model (such as a neural network) can be used, taking background noise sound pressure level and noise spectrum characteristics as input, and the model can directly output optimized volume and frequency response parameters. Alternatively, a preset set of volume and frequency response parameter adjustment values ​​can be directly matched according to different noise level ranges by querying a preset look-up table.

[0219] In some embodiments, the volume parameter is adjusted based on the background noise sound pressure level and determined by a mapping function; the frequency response parameter is determined based on the noise concentration band, which is one or more frequency bands whose noise energy meets preset conditions by performing spectral analysis on the environmental noise signal.

[0220] A mapping function is a mathematical function pre-defined in the system that describes the correspondence between "background noise sound pressure level" and "volume gain value". For example, a mapping function can be a mathematical formula or lookup table that maps a background noise sound pressure level value to a desired volume gain adjustment.

[0221] In some embodiments, the control device acquires a real-time measured background noise sound pressure level, then uses this sound pressure level value as input to calculate or query a corresponding volume gain value through a preset mapping function. This volume gain value is then used as a volume parameter in the audio adjustment scheme.

[0222] For example, the mapping function can be preset as a lookup table. This lookup table stores the mapping relationship between different background noise sound pressure level ranges and their corresponding volume gain values. For instance, when a noise sound pressure level of 65 dB SPL is detected, the control device obtains the corresponding volume gain value of +6 dB by querying the lookup table. If the original volume is set to 80 dB, the adjusted final output volume will be 86 dB.

[0223] In some embodiments, the lookup table may be constructed by a technician based on experience or experimental data.

[0224] For example, the mapping function can be a pre-defined mathematical formula, such as a non-linear polynomial function. This function takes the input background noise sound pressure level as the independent variable and directly calculates the output volume gain value. This function can be obtained by fitting and training a large amount of historical data (including sound pressure levels under different noise environments and volume settings with good user feedback).

[0225] In some embodiments, the volume parameter can be determined in various ways. For example, in addition to using lookup tables or mathematical formulas, rule-based logical judgment systems can be employed, or the mapping relationship can be fitted to a large amount of historical data through machine learning models (such as neural networks) to achieve a more adaptive adjustment.

[0226] The noise concentration band refers to one or more frequency ranges in the noise spectrum where the noise energy meets preset conditions. For example, the noise energy of an air conditioner fan is significantly concentrated in the 200Hz to 400Hz frequency band.

[0227] Preset conditions refer to the quantitative criteria defined in the process of identifying noise concentration bands to determine whether a certain frequency band belongs to "noise concentration". For example, if the preset condition is that the noise energy of a certain frequency band exceeds the energy threshold or energy peak, then the frequency band is determined to be a noise concentration band.

[0228] In some embodiments, preset conditions may be set by a technician based on experience.

[0229] In some embodiments, the control device can perform spectral analysis on the acquired environmental noise signal to obtain a noise spectrum. Then, based on the noise spectrum, the frequency bands where noise energy is concentrated are identified.

[0230] For example, the identification process can be achieved by setting a fixed energy threshold. The control device compares the energy of each frequency point in the spectrum with the preset threshold, and identifies the range of consecutive frequency points with energy exceeding the threshold as a noise concentration band. In some embodiments, the energy threshold can be set by a technician based on experience.

[0231] For example, control equipment can use peak detection algorithms to identify noise concentration bands. The control equipment searches for local energy peaks on the noise spectrum and identifies the region within a certain bandwidth surrounding each peak as a noise concentration band. This method can more flexibly adapt to different types of noise spectrum characteristics.

[0232] Spectrum analysis is the process of decomposing a complex signal into its constituent frequency components. For example, spectrum analysis can be achieved using the Fast Fourier Transform (FFT) algorithm.

[0233] In some embodiments, the identification of noise-concentrated frequency bands can also be achieved through other signal processing techniques, such as wavelet transform analysis, or by classifying and identifying spectral patterns through machine learning models to more accurately locate noise sources.

[0234] In some embodiments, after identifying the noise concentration band, the control device determines the corresponding frequency response parameters for each band, for example, by configuring one or more digital filters (such as peak filters).

[0235] For example, if indoor noise is identified as primarily concentrated in the low-frequency range of 80Hz-150Hz, the control device generates a peak filter. The filter's center frequency is set to the peak energy frequency of that range (e.g., 100Hz), its bandwidth (Q value) is set to cover the range, and its gain is set to a positive value (e.g., +6dB). In this way, the energy of the target audio signal in this frequency range can be increased to counteract the masking effect of noise.

[0236] The target audio signal refers to the audio signal played based on the audio data built into or received by the speaker.

[0237] For example, when multiple noise-concentrated frequency bands are detected, the control device generates multiple peak filters, each corresponding to a different noise-concentrated frequency band. Furthermore, the filter gain can be proportional to the noise intensity of the corresponding frequency band. For instance, low-frequency noise has higher energy, so its corresponding filter gain can be set to +6dB; while high-frequency noise has lower energy, so its gain can be set to +3dB, achieving more refined compensation.

[0238] In some embodiments, in addition to using a peak filter, the frequency response parameters can also be determined by adjusting the gain of the corresponding frequency band in a graphic equalizer (Graphic EQ) or parametric equalizer (Parametric EQ). The adjustment method is not limited to increasing the gain; in certain specific scenarios, it can also involve moderately attenuating certain interfering frequency bands.

[0239] In some embodiments of this specification, the volume is adaptively adjusted through a mapping function, and the frequency response is precisely adjusted based on the noise-concentrated frequency bands identified by spectrum analysis. This solution can intelligently cope with environmental noise, not only ensuring the overall audibility of audio under different noise levels, but also effectively combating noise masking by specifically enhancing the energy of certain frequency bands. This significantly enhances the clarity and penetration of speech or music in noisy environments, thereby improving the consistency of the listening experience.

[0240] In some embodiments of this specification, ambient noise is monitored in real time by a reference microphone and quantified as background noise sound pressure level, enabling dynamic and adaptive adjustment of the audio adjustment scheme. When ambient noise increases, the control device automatically increases the volume or adjusts the response of specific frequency bands to ensure the clarity and intelligibility of the audio content; when the environment becomes quiet, it returns to normal levels, avoiding excessive volume that may disturb the user. This scheme achieves intelligent matching of audio playback effects with the environment, requiring no manual intervention from the user, and significantly improves the listening experience in changing environments.

[0241] The basic concepts have been described above. Obviously, for those skilled in the art, the detailed disclosure above is merely illustrative and does not constitute a limitation of this specification. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and corrections to this specification. Such modifications, improvements, and corrections are suggested in this specification and therefore remain within the spirit and scope of the exemplary embodiments described herein.

[0242] Furthermore, this specification uses specific terms to describe embodiments thereof. For example, "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Moreover, certain features, structures, or characteristics in one or more embodiments of this specification can be appropriately combined.

[0243] Furthermore, unless expressly stated in the claims, the order of processing elements and sequences, the use of numbers and letters, or other names described in this specification are not intended to limit the order of the processes and methods described herein. Although various examples have been discussed in the foregoing disclosure of some embodiments of the invention that are currently considered useful, it should be understood that such details are for illustrative purposes only, and the appended claims are not limited to the disclosed embodiments; rather, the claims are intended to cover all modifications and equivalent combinations that conform to the spirit and scope of the embodiments described herein. For example, while the system components described above can be implemented using hardware devices, they can also be implemented solely using software solutions, such as installing the described system on existing servers or mobile devices.

[0244] Similarly, it should be noted that, in order to simplify the description disclosed herein and thus aid in the understanding of one or more embodiments of the invention, the foregoing description of embodiments in this specification may sometimes combine multiple features into a single embodiment, drawing, or description thereof. However, this method of disclosure does not imply that the subject matter of this specification requires more features than those mentioned in the claims. In fact, the embodiments contain fewer features than all the features of a single embodiment disclosed above.

[0245] In some embodiments, numbers describing the quantity of components and attributes are used. It should be understood that such numbers used in the description of embodiments are modified in some examples with the terms "approximately," "approximately," or "generally." Unless otherwise stated, "approximately," "approximately," or "generally" indicates that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, which may be changed depending on the characteristics required by individual embodiments. In some embodiments, numerical parameters should take into account specified significant digits and employ a general method of digit reservation. Although the numerical ranges and parameters used to confirm their breadth of range in some embodiments of this specification are approximate values, in specific embodiments, such values ​​are set as precisely as feasible.

[0246] For each patent, patent application, patent application publication, and other material such as articles, books, specifications, publications, and documents referenced in this specification, the entire contents of which are incorporated herein by reference. This excludes historical application documents that are inconsistent with or conflict with the content of this specification, as well as documents that limit the broadest scope of the claims in this specification (currently or subsequently appended to this specification). It should be noted that in the event of any inconsistency or conflict between the descriptions, definitions, and / or terminology used in the supplementary materials to this specification and the content of this specification, the descriptions, definitions, and / or terminology used in this specification shall prevail.

[0247] Finally, it should be understood that the embodiments described in this specification are merely illustrative of the principles of the embodiments described herein. Other variations may also fall within the scope of this specification. Therefore, alternative configurations of the embodiments described herein are intended to be illustrative rather than limiting, and should be considered consistent with the teachings of this specification. Accordingly, the embodiments described herein are not limited to those explicitly introduced and described herein.

Claims

1. A speaker sound field adaptive method, characterized in that, include: Obtain the speaker's evaluation parameters; Based on the evaluation parameters, the current placement status of the speaker is determined; Based on the historical placement status and the current placement status, determine whether the placement status of the speaker has changed; In response to a change in the placement of the speaker: Based on the evaluation parameters, the acoustic feature identifiers and corresponding audio adjustment schemes for the current placement state are determined through an acoustic environment model, wherein the acoustic environment model is a machine learning model. Based on the aforementioned audio adjustment scheme, the parameters of the speaker's audio processing unit are adjusted.

2. The method according to claim 1, characterized in that, The state parameters are obtained through the attitude sensing unit. The state parameters include the three-dimensional attitude angle and absolute orientation of the speaker. The attitude sensing unit includes an inertial measurement unit and a magnetic field sensor. The inertial measurement unit includes a three-axis gyroscope and a three-axis accelerometer. The environmental distance data is acquired through an environmental sensing unit, which includes an infrared sensor array distributed around the periphery of the speaker enclosure. The infrared sensor array is configured to construct a polar coordinate distance map centered on the speaker, and the environmental distance data is determined based on the polar coordinate distance map.

3. The method according to claim 1, characterized in that, The state parameters include the speaker's three-dimensional attitude angle and absolute orientation; Determining whether the speaker's placement has changed based on its historical placement and current placement includes: Based on the historical placement status, a historical baseline sequence is determined; Based on the current placement state, determine the current state sequence; By comparing the current state sequence with the historical baseline sequence, and in response to the change in at least one dimension of the three-dimensional attitude angle, the absolute orientation, or the environmental distance data exceeding its corresponding preset tolerance threshold, it is determined that the placement state of the speaker has changed.

4. The method according to claim 3, characterized in that, The acoustic environment model is a neural network model, which includes a detection layer and a compensation layer connected in sequence. The detection layer is configured to determine the acoustic feature identifier in the current placement state based on the current state sequence and the historical reference sequence; The compensation layer is configured to determine the audio adjustment scheme based on the current state sequence, the historical reference sequence, and the acoustic feature identifiers in the current placement state.

5. The method according to claim 1, characterized in that, The method further includes: Acquire recorded data of the speaker in a specific placement state, the recorded data including the state parameters corresponding to the specific placement state, the environmental distance data, and the scheme adjustment data; Based on the recorded data, a preference prediction model is trained; the preference prediction model is a machine learning model. In response to the fact that the matching degree between the current placement state and the specific placement state meets the preset conditions, the corrected audio adjustment scheme is determined based on the recorded data and the preference prediction model.

6. A speaker sound field adaptive system, characterized in that, The system includes a control device, which includes a data processing unit, a logic judgment unit, and a status control unit. The data processing unit is configured to acquire evaluation parameters for the speaker; The logic judgment unit is configured to determine the current placement state of the speaker based on the evaluation parameters; Based on the historical placement status and the current placement status, determine whether the placement status of the speaker has changed; The status control unit is configured as follows: In response to a change in the placement of the speaker: Based on the evaluation parameters, the acoustic feature identifiers and corresponding audio adjustment schemes for the current placement state are determined through an acoustic environment model, wherein the acoustic environment model is a machine learning model. Based on the aforementioned audio adjustment scheme, the parameters of the speaker's audio processing unit are adjusted.

7. The system according to claim 6, characterized in that, The data processing unit includes an attitude perception unit and an environment perception unit; The attitude sensing unit is configured to acquire the state parameters, which include the three-dimensional attitude angle and absolute orientation of the speaker. The attitude sensing unit includes an inertial measurement unit and a magnetic field sensor. The inertial measurement unit includes a three-axis gyroscope and a three-axis accelerometer. The environmental sensing unit is configured to acquire the environmental distance data. The environmental sensing unit includes an infrared sensor array distributed around the periphery of the speaker enclosure. The infrared sensor array is configured to construct a polar coordinate distance map centered on the speaker. The environmental distance data is determined based on the polar coordinate distance map.

8. The system according to claim 6, characterized in that, The state parameters include the speaker's three-dimensional attitude angle and absolute orientation, and the control device is further configured to: Based on the historical placement status, a historical baseline sequence is determined; Based on the current placement state, determine the current state sequence; By comparing the current state sequence with the historical baseline sequence, in response to the change in at least one dimension of the three-dimensional attitude angle, the absolute orientation, or the environmental distance data exceeding its corresponding preset tolerance threshold, it is determined that the placement state of the speaker has changed.

9. The system according to claim 8, characterized in that, The acoustic environment model is a neural network model, which includes a detection layer and a compensation layer connected in sequence. The detection layer is configured to determine the acoustic feature identifier in the current placement state based on the current state sequence and the historical reference sequence; The compensation layer is configured to determine the audio adjustment scheme based on the current state sequence, the historical reference sequence, and the acoustic feature identifiers in the current placement state.

10. The system according to claim 6, characterized in that, The control device is further configured to: Acquire recorded data of the speaker in a specific placement state, the recorded data including the state parameters corresponding to the specific placement state, the environmental distance data, and the scheme adjustment data; Based on the recorded data, a preference prediction model is trained; the preference prediction model is a machine learning model. In response to the fact that the matching degree between the current placement state and the specific placement state meets the preset conditions, the corrected audio adjustment scheme is determined based on the recorded data and the preference prediction model.