Calibration method and system for playing parameters
By acquiring audio data through an audio monitoring device, analyzing spectral characteristics, and adjusting playback parameters, the problem of sound quality imbalance in indoor acoustic environments is solved, enabling real-time calibration and improvement of audio quality.
Patent Information
- Application Number
- CN202511039125.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-10-31
AI Technical Summary
In indoor acoustic environments, existing technologies struggle to dynamically calibrate and optimize indoor acoustics as room layouts and usage scenarios change, leading to unbalanced and inaccurate sound quality.
Audio data is acquired through an audio monitoring device, its spectral characteristics are analyzed, it is determined whether the calibration conditions are met, and the calibration parameters are determined based on the spectral characteristics. The playback parameters of the audio system are then automatically adjusted to achieve real-time calibration.
It enables real-time calibration of the audio system, improves audio propagation quality and listening experience, reduces unnecessary computation and energy consumption, and adapts to changes in room acoustics.
Smart Images

Figure CN120877779A_ABST
Abstract
Description
Technical Field
[0001] This manual relates to the field of acoustic calibration technology, and in particular to a method and system for calibrating playback parameters. Background Technology
[0002] The size, shape, and surface materials of a space typically affect sound quality. For example, different surface materials reflect and absorb sound at different frequencies. In indoor environments with high sound quality requirements, such as recording studios and playback rooms, it is necessary to optimize and adjust the room's acoustics to obtain more balanced and accurate high-quality audio. When the indoor environment changes, such as changes in room layout, usage, or noise levels, further calibration of the room's acoustics is required. Therefore, how to dynamically calibrate and optimize indoor acoustics has become an urgent problem to be solved.
[0003] Therefore, it is desirable to propose a method and system for calibrating playback parameters that can perform real-time calibration of indoor acoustics to improve the propagation quality of audio and the listening experience. Summary of the Invention
[0004] One embodiment of this specification provides a method for calibrating playback parameters, comprising: acquiring first audio data corresponding to a first playback audio through an audio monitoring device; determining a first spectral feature based on the first audio data; determining whether the first spectral feature meets a first calibration condition; in response to the first spectral feature meeting the first calibration condition, determining a first calibration parameter based on the first spectral feature; and calibrating the playback parameters of an audio system based on the first calibration parameter.
[0005] One embodiment of this specification provides a playback parameter calibration system, comprising: an acquisition module configured to acquire first audio data corresponding to a first playback audio through an audio monitoring device; a first determination module configured to determine a first spectral feature based on the first audio data; a judgment module configured to determine whether the first spectral feature meets a first calibration condition; and a second determination module configured to determine a first calibration parameter based on the first spectral feature in response to the first spectral feature meeting the first calibration condition; and to calibrate the playback parameters of the audio system based on the first calibration parameter.
[0006] One embodiment of this specification provides a playback parameter calibration device, characterized in that the device includes at least one processor and at least one memory; the at least one memory is used to store computer instructions; the at least one processor is used to execute at least a portion of the computer instructions to implement the playback parameter calibration method.
[0007] One embodiment of this specification provides a computer-readable storage medium that stores computer instructions. When a computer reads the computer instructions from the storage medium, the computer executes the playback parameter calibration method. Attached Figure Description
[0008] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein:
[0009] Figure 1 This is a schematic diagram illustrating an application scenario of a playback parameter calibration system according to some embodiments of this specification;
[0010] Figure 2 This is an exemplary block diagram of a calibration system for playback parameters according to some embodiments of this specification;
[0011] Figure 3 This is an exemplary flowchart of a method for calibrating playback parameters according to some embodiments of this specification;
[0012] Figure 4 This is an exemplary flowchart illustrating the determination of a second calibration parameter according to some embodiments of this specification;
[0013] Figure 5 This is an exemplary schematic diagram of a calibration judgment model shown in some embodiments of this specification. Detailed Implementation
[0014] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of this specification. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.
[0015] It should be understood that the terms “system,” “device,” “unit,” and / or “module” used herein are one way to distinguish different components, elements, parts, sections, or assemblies at different levels. However, if other terms can achieve the same purpose, they may be replaced by other expressions.
[0016] Unless the context clearly indicates an exception, words such as "a," "an," "a kind," and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of explicitly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.
[0017] Flowcharts are used in this specification to illustrate the operations performed by the system according to embodiments of this specification. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the steps can be processed in reverse order or simultaneously. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.
[0018] Figure 1 This is a schematic diagram illustrating an application scenario of a playback parameter calibration system according to some embodiments of this specification. For example... Figure 1 As shown, the application scenario 100 of the playback parameter calibration system may include an audio monitoring device 110, a processor 120, a storage device 130, a network 140, a user terminal 150, and an audio system 160, etc. The application scenario 100 of the playback parameter calibration system can be applied to various environments requiring high-quality audio, such as recording scenarios and playback scenarios. An area refers to a specific spatial area, such as an indoor scene. An indoor scene can be a room, such as at least one of a recording studio, a cinema, or a concert hall. By using the playback parameter calibration system, the influence of the area on sound reflection, reverberation, and frequency response can be eliminated or reduced, thereby obtaining a more balanced and accurate sound performance.
[0019] The audio monitoring device 110 refers to the device configured to monitor relevant audio data in the monitoring area.
[0020] In some embodiments, the audio monitoring device 110 may include at least one microphone, an audio sensor, a monitoring device, etc. A microphone is a sound input device, and the audio monitoring device 110 can acquire sound data sequences through the microphone. The type of microphone may include at least one of condenser microphones, microelectromechanical systems (MEMS) microphones, electret microphones, etc. Different microphones may be located at different locations within the area. Each element in the audio monitoring device 110 can represent a microphone and its corresponding location.
[0021] In some embodiments, at least one microphone can be represented by a microphone array. A microphone array refers to a combination of multiple microphones arranged in a specific configuration. In some embodiments, the arrangement of the microphones in the microphone array can be determined based on the room structure. For example, multiple microphones can be arranged at equal intervals on the floor and walls of a room, according to the room's height, length, and width. For instance, if a wall in the room is 2m high and 4m wide, the microphones in the microphone array could be arranged as follows: one microphone is placed at each of the four corners of the wall, and one microphone is placed at the center of the top and bottom edges of the wall, thus forming a square microphone array with a side length of 2m.
[0022] The monitoring device is used to monitor the operational status of a microphone array and audio sensors in a current scene. In some embodiments, the monitoring device can transmit the operational status of the microphone array and audio sensors to the processor 120 via network 140. For example, microphone data in the microphone array {(microphone number 1, position data 1, signal transmission interrupted), ... (microphone number n, position data n, signal transmission normal)}.
[0023] The processor 120 can be used to manage data resources and process data and / or information from at least one component or external data source involved in the application scenario 100 of the playback parameter calibration system. The processor 120 can execute program instructions based on this data, information, and / or processing results to perform one or more functions described in this specification. For example, the processor 120 can acquire first audio data corresponding to a first playback audio through the audio monitoring device 110; and determine a first spectral characteristic based on the first audio data. As another example, the processor 120 can determine whether the first spectral characteristic meets a first calibration condition: in response to the first spectral characteristic meeting the first calibration condition, determine a first calibration parameter based on the first spectral characteristic; and calibrate the playback parameters of the audio system 160 based on the first calibration parameter.
[0024] In some embodiments, processor 120 may include one or more sub-processing devices (e.g., a single-core processing device or a multi-core multi-chip processing device). By way of example only, processor 120 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC) microprocessor, or any combination thereof.
[0025] In some embodiments, the processor 120 may include one or more of the acquisition module 210, the first determination module 220, the judgment module 230, and the second determination module 240 in the playback parameter calibration system 200. Further details can be found in [link to documentation]. Figure 2 Related descriptions.
[0026] Storage device 130 can be used to store data and / or instructions. For example, storage device 130 can be used to store user preference data input by the user through user terminal 150. As another example, storage device 130 can also be used to store first audio data acquired by processor 120.
[0027] In some embodiments, storage device 130 and processor 120 can communicate via network 140, and storage device 130 may also be part of the processor. In some embodiments, storage device 130 may include random access memory (RAM), read-only memory (ROM), mass storage, or any combination thereof. For a description of the first audio data, see [link to documentation]. Figure 3 And its contents.
[0028] Network 140 can connect various components of the playback parameter calibration system and / or connect the system to external resources. Network 140 enables communication between the components, as well as with other parts outside the playback parameter calibration system. For example, processor 120 can acquire first audio data in a scene area from audio monitoring device 110 via network 140, etc.
[0029] In some embodiments, network 140 can be any one or more of wired or wireless networks. For example, network 140 may include cable networks, fiber optic networks, or any combination thereof. Network connections between components can be made using one or more of the methods described above. In some embodiments, the network can be a point-to-point, shared, centralized, or other topologies, or a combination of multiple topologies.
[0030] User terminal 150 refers to one or more terminal devices or software used by a user. In some embodiments, user terminal 150 may be one or any combination of other devices with input and / or output functions, such as mobile device 150-1, tablet computer 150-2, and laptop computer 150-3. In some embodiments, the user of user terminal 150 may be one or more users. "User" may refer to the administrator or operator of the calibration system for playback parameters.
[0031] Audio system 160 is a system for playing audio. For example, it is a system for controlling the playback of speakers or audio equipment. In some embodiments, the audio system can be configured in application scenario 100 of a playback parameter calibration system to play first playback audio, second playback audio, composite playback audio, corrected composite audio, etc. In some embodiments, a user can control the audio system through a user terminal, such as switching the playing audio or adjusting the volume of the playing audio. In some embodiments, processor 120 can use the audio currently being played by the audio system as the first playback audio.
[0032] It should be noted that the above description of the application scenario 100 of the playback parameter calibration system is for illustrative purposes only and is not intended to limit the scope of this specification. Various modifications and variations can be made based on this specification by those skilled in the art. However, these changes and modifications do not depart from the scope of this specification.
[0033] Figure 2 This is an exemplary block diagram of a calibration system for playback parameters according to some embodiments of this specification.
[0034] In some embodiments, such as Figure 2 As shown, the playback parameter calibration system 200 may include an acquisition module 210, a first determination module 220, a judgment module 230, and a second determination module 240.
[0035] In some embodiments, the acquisition module 210 may be configured to acquire first audio data corresponding to the first played audio through an audio monitoring device.
[0036] In some embodiments, the first determining module 220 may be configured to determine a first spectral feature based on the first audio data.
[0037] In some embodiments, the determination module 230 may be configured to determine whether the first spectral feature satisfies the first calibration condition.
[0038] In some embodiments, the second determining module 240 may be configured to determine a first calibration parameter based on the first spectral feature in response to the first spectral feature satisfying the first calibration condition; and to calibrate the playback parameters of the audio system based on the first calibration parameter.
[0039] In some embodiments, the second determining module 240 may be further configured to: generate a test audio signal; generate a composite playback audio based on the test audio signal and a second playback audio in response to the first spectral feature satisfying the first calibration condition; acquire composite audio data corresponding to the playback audio through the audio monitoring device; determine a second spectral feature and a test spectral feature based on the composite audio data; determine whether the second spectral feature and the test spectral feature satisfy the second calibration condition; determine a second calibration parameter based on the second spectral feature and the test spectral feature in response to the second calibration condition; and calibrate the playback parameters of the audio system based on the second calibration parameter. For further explanation of determining the second calibration parameter, please refer to [link to relevant documentation]. Figure 4 , Figure 5 And its related descriptions.
[0040] In some embodiments, the second determining module 240 may be further configured to determine an initial test signal based on the first spectral characteristics and background noise data; and generate the test audio signal based on the initial test signal and room structure data. For further explanation regarding the generation of the audio test signal, please refer to... Figure 4 And its related descriptions.
[0041] In some embodiments, the second determining module 240 may be further configured to determine the calibration effect score of the candidate calibration parameters based on the candidate calibration parameters with a preset computational load, the first spectral feature, and historical vibration data; and to determine the first calibration parameter based on the calibration effect score. For further explanation of determining the first calibration parameter, please refer to [link to relevant documentation]. Figure 3 And its related descriptions.
[0042] It should be noted that the above description of the playback parameter calibration system 200 and its modules is for ease of description only and should not limit this specification to the scope of the illustrated embodiments. It is understood that those skilled in the art, after understanding the principle of the system, may arbitrarily combine the various modules or construct subsystems connected to other modules without departing from this principle. In some embodiments, Figure 2 The acquisition module 210, the first determination module 220, the judgment module 230, and the second determination module 240 disclosed herein can be different modules within a single system, or a single module can implement the functions of the two aforementioned modules. For example, the modules can share a single storage module, or each module can have its own separate storage module. Such variations are all within the scope of protection of this specification.
[0043] Figure 3 This is an exemplary flowchart of a method for calibrating playback parameters according to some embodiments of this specification. Figure 3 As shown, process 300 includes the following steps. In some embodiments, process 300 may be performed by a playback parameter calibration system 200.
[0044] Step 310: Obtain the first audio data corresponding to the first played audio using the audio monitoring device. For instructions on the audio monitoring device, please refer to [link to instructions]. Figure 1 And its contents.
[0045] The first audio being played refers to the audio that the user is currently playing. For example, the song currently being played by the audio system. In some embodiments, the acquisition module can acquire the first audio being played based on the audio system. For more information about the audio system, please refer to [link to relevant documentation]. Figure 1 Related descriptions.
[0046] The first audio data refers to the acoustic data related to the response in the room when the first audio is played. In some embodiments, the acquisition module can collect the acoustic data related to the room in real time when the first audio is played alone through an audio monitoring device, and use it as the first audio data.
[0047] Step 320: Determine the first spectral feature based on the first audio data.
[0048] The first spectral feature refers to the spectral features corresponding to the first audio data. For example, the first spectral feature may include at least one of the following: spectral envelope, spectral center, spectral bandwidth, spectral roll statistics, Mel-Frequency Cepstral Coefficients (MFCCs) of the first audio data.
[0049] In some embodiments, the first determining module 220 may determine the first spectral feature based on the first audio data through various methods. For example, the first determining module 220 may determine the first spectral feature through a preset spectral algorithm, etc. The preset spectral algorithm refers to a preset algorithm for extracting the first spectral feature, such as at least one of the following: Fast Fourier Transform (FFT), spectral analysis algorithm, etc.
[0050] In some embodiments, the first determining module 220 may also determine the first spectral feature in any feasible manner.
[0051] Step 330: Determine whether the first spectral feature meets the first calibration condition.
[0052] The first calibration condition refers to a preset condition used to determine whether the playback parameters of the audio system need to be calibrated. For example, the first calibration condition could be a first similarity lower than a first similarity threshold. Here, the first similarity refers to the similarity between the acoustic fingerprint data corresponding to the first spectral feature and historical fingerprint data.
[0053] In some embodiments, the judgment module may determine a first similarity threshold based on empirical presets. Playback parameters refer to the settings parameters of the audio system when playing audio. In some embodiments, playback parameters may include at least one of equalization (EQ) settings, speaker output parameters, tone control (TC) parameters, and reverberation control parameters.
[0054] The equalizer settings parameters refer to the parameters of the equalizer in the audio system. These parameters can be used to adjust the gain at different frequencies to compensate for the room's acoustic characteristics that may cause excessive amplification or attenuation of specific frequencies. Speaker output parameters refer to the output level of the speakers in the audio system; tone control parameters adjust the tone of the audio, and can be used to adjust the gain of bass and treble frequencies to suit the room's acoustic characteristics; reverberation control parameters adjust the parameters of the reverberation processor, and can be used to simulate or compensate for the room's natural reverberation characteristics, making the sound clearer or fuller.
[0055] Acoustic fingerprint data refers to data extracted from the first spectral feature using a preset fingerprint algorithm. Historical fingerprint data refers to acoustic fingerprint data calculated and stored from the last time the playback parameters of the audio system were calibrated, up to the current time point. The calculation may include: when the playback parameters of the audio system were last calibrated, the acquisition module can acquire the first audio data at that time; the first determination module can determine the first spectral feature at that time based on the first audio data; and the judgment module can extract data from the first spectral feature at that time using a preset fingerprint algorithm and determine the extracted data as historical fingerprint data. The judgment module can store this historical fingerprint data in a storage device.
[0056] In some embodiments, the determination module may generate acoustic fingerprint data corresponding to the current first spectral feature based on the first spectral feature using a preset fingerprint algorithm. The preset fingerprint algorithm is an algorithm used to extract acoustic fingerprint data. For example, the preset fingerprint algorithm may include at least one of the following: Chromaprint algorithm, Shazam algorithm, ACRCloud, Constant-Q Transform (CQT).
[0057] The first similarity threshold refers to the maximum value of the first similarity when the audio system needs calibration. In some embodiments, the first similarity threshold can be determined based on an empirical preset.
[0058] In some embodiments, the first similarity threshold may also be related to room type. A room refers to the area monitored by the audio monitoring device. Room type can include various types, such as open space, enclosed space, etc.
[0059] In some embodiments, when the room type is an open space or public place, because the space is large (or relatively spacious), the user's auditory sensitivity to changes in room acoustics is low. Therefore, the first similarity threshold can be reduced according to a preset percentage. Conversely, when the room type is a closed space, compared to an open space, the closed space is smaller (or the room layout is more compact), and the user's auditory sensitivity to changes in room acoustics is high. Therefore, the first similarity threshold can be increased according to a preset percentage.
[0060] In some embodiments, the room type can be obtained in a variety of ways. For example, the room type can be obtained based on user input on a user terminal. Another example is that the room type can be obtained by analyzing images captured by devices such as cameras installed in the room.
[0061] In some embodiments, the determining module can determine whether the first spectral feature meets the first calibration condition. If so, step 340 and / or step 420 are executed; if not, the playback parameters of the audio system are not calibrated. When the first spectral feature meets the first calibration condition, step 340 can be executed first, followed by step 420.
[0062] For steps 310-330, the system executes the detection at a preset frequency, either periodically or in real-time. This allows for the periodic or real-time detection of changes in room acoustics and the determination of whether the audio system's playback parameters need calibration. The preset detection frequency refers to the frequency at which the system detects changes in room acoustics. This preset frequency can be determined based on empirical settings, such as 12 times per hour.
[0063] Step 340: In response to the first spectral feature satisfying the first calibration condition, determine the first calibration parameter based on the first spectral feature.
[0064] When the first spectral feature meets the first calibration condition, it indicates that the acoustic fingerprint data corresponding to the first spectral feature is significantly different from the historical fingerprint data, and that the room layout and other factors have changed. Therefore, it is necessary to calibrate the playback parameters of the audio system.
[0065] The first calibration parameter refers to the playback parameters used to calibrate the audio system during the first playback of audio. For example, the first calibration parameter may include at least one of the following: equalization (EQ) parameters, speaker output parameters, tone control (TC) parameters, reverberation control parameters, etc.
[0066] In some embodiments, the second determining module can determine the first calibration parameter based on the first spectral feature using various methods. For example, the second determining module can construct a first calibration vector based on the acoustic fingerprint data, historical fingerprint data, and current playback parameters corresponding to the first spectral feature, and based on the first calibration vector, query a first database to select the first reference parameter corresponding to the first reference vector with the highest similarity to the first calibration vector as the first calibration parameter corresponding to the first calibration vector. The similarity calculation can include Euclidean distance, cosine similarity, etc.
[0067] The current playback parameters refer to the current playback parameters of the audio system, which the second determination module can obtain by reading the settings of the audio system. The first calibration vector can be in the form of a three-dimensional vector, such as (x, y, z), where x represents the acoustic fingerprint data corresponding to the first spectral feature, y represents historical fingerprint data, and z represents the current playback parameters; the first database can be constructed based on historical data, or it can be manually modified and supplemented.
[0068] The first database includes acoustic fingerprint data corresponding to historical first spectral features, historical fingerprint data, a first reference vector constructed from historical playback parameters, and a first reference parameter corresponding to the first reference vector. The second determining module can determine the actual calibration data corresponding to the first reference vector in the historical data as the first reference parameter.
[0069] In some embodiments, the second determining module may further determine the calibration effect score of the candidate calibration parameters based on the preset calculation of candidate calibration parameters, the first spectral characteristics and historical vibration data; and determine the first calibration parameter based on the calibration effect score.
[0070] Candidate calibration parameters refer to candidate playback parameters used to calibrate an audio system. In some embodiments, a set of candidate calibration parameters may include multiple playback parameters, such as at least one of equalization (EQ) settings, speaker output parameters, tone control (TC) parameters, and reverberation control parameters.
[0071] In some embodiments, the second determining module can randomly adjust the current playback parameters to generate multiple sets of candidate calibration parameters with a preset computational load. For example, for different types of calibration parameters in each set of candidate calibration parameters, the second determining module can adjust the values of the corresponding type of calibration parameters in the current playback parameters by multiple randomly generated percentages to obtain multiple sets of candidate calibration parameters.
[0072] For example, the current playback parameters include equalizer setting parameters and speaker output parameters. The second determining module can generate random adjustment percentages for these two parameters respectively. For example, the equalizer setting parameters are adjusted to 98% of their original values, and the speaker output parameters are adjusted to 102% of their original values. The adjusted equalizer setting parameters and speaker output parameters are used as a set of candidate calibration parameters.
[0073] The preset computational load is the preset number of candidate calibration parameter sets required. In some embodiments, the preset computational load can be determined based on empirical presets. For example, the preset computational load can be 20 sets, 40 sets, etc.
[0074] Historical vibration data refers to vibration sensing data from historical data. In some embodiments, the second determining module can acquire vibration sensing data from a historical period of time relative to the current time point and use it as historical vibration data. In some embodiments, vibration sensing data can be acquired based on vibration sensors configured in the room. For a description of vibration sensing data, see [link to documentation]. Figure 4 And its contents.
[0075] The calibration effect score refers to the score of the calibration effect of the candidate calibration parameters on the audio system. In some embodiments, the second determining module can characterize the calibration effect score in multiple ways. For example, the calibration effect score can be a value between 0 and 100, with a higher value indicating a better calibration effect corresponding to the candidate calibration parameter. Using different candidate calibration parameters to calibrate the playback parameters of the audio system will result in different sound quality of the audio system after calibration, and thus different calibration effects.
[0076] In some embodiments, the second determining module may determine the calibration effect score of the candidate calibration parameters based on the candidate calibration parameters, the first spectral characteristics, and historical vibration data through a variety of methods.
[0077] For example, the second determining module can determine the calibration effect score corresponding to the candidate calibration parameters, the first spectral characteristics, and the historical vibration data through the first preset table.
[0078] The first preset table reflects the relationship between candidate calibration parameters, first spectral characteristics, historical vibration data, and calibration effect scores. The second determination module can determine the first preset table based on experience.
[0079] In some embodiments, the second determining module may determine the calibration effect score of the candidate calibration parameter based on the candidate calibration parameter, the first spectral feature and historical vibration data in any feasible manner, and determine the candidate calibration parameter with the highest calibration effect score among the candidate calibration parameters with preset calculation as the first calibration parameter.
[0080] In some embodiments, the second determining module can determine the calibration effect score of the group of candidate calibration parameters based on each group of candidate calibration parameters, acoustic fingerprint data, historical fingerprint data, historical vibration data, environmental sensing data, current playback parameters, and room structure data, through an effect determination model.
[0081] An effectiveness determination model is a model used to determine the calibration effectiveness score. The effectiveness determination model can be a machine learning model, such as any one or a combination of neural networks (NN), convolutional neural networks (CNN), etc.
[0082] In some embodiments, the input to the effect determination model may include at least one of candidate calibration parameters, acoustic fingerprint data, historical fingerprint data, historical vibration data, environmental sensor data, current playback parameters, and room structure data, and the output may include a calibration effect score. In some embodiments, the second determination module may input each set of candidate calibration parameters into the effect determination model to determine the calibration effect score corresponding to each set of candidate calibration parameters.
[0083] Environmental sensing data refers to data related to the room environment. In some embodiments, environmental sensing data may include at least one of ambient temperature and ambient humidity. Different ambient temperatures and humidity levels can also affect audio quality, therefore environmental sensing data needs to be taken into account to more accurately determine the calibration effect score. In some embodiments, the second determining module can acquire environmental sensing data based on devices such as hygrometers and thermometers in the room.
[0084] Room structure data refers to data related to the structural characteristics of a room. In some embodiments, room structure data may include at least one of the following: room size, room shape, ceiling height, whether there are windows, furniture layout, etc. The shape of the room can be represented by complexity. The greater the difference between the room shape and the standard room shape, the more complex the room shape, and the higher the corresponding complexity. The shape of the standard room can be preset. Among them, the room size, room shape, ceiling height, and other data related to the initial construction layout of the room are usually fixed and can be determined by user input. The second determination module can also obtain data by taking images once through a camera or other device installed in the room and analyzing the images. The second determination module can compare the shape of the room with the shape of the standard room, determine the differences between the two, and label the complexity of the room based on the differences. For data that is prone to change, such as whether there are windows or the furniture layout, the second determination module can take images multiple times through a camera or other device installed in the room and analyze the images to obtain this type of room structure data. The second determination module can also determine this type of room structure data by obtaining user terminal input.
[0085] In some embodiments, the performance determination model can be trained using a large number of first training samples and first labels corresponding to the first training samples. In some embodiments, multiple first training samples with first labels can be input into the initial performance determination model. A loss function is constructed using the first labels and the results of the initial performance determination model. The parameters of the initial performance determination model are iteratively updated based on the loss function using gradient descent or other methods. When a preset condition is met, the model training is complete, and the trained performance determination model is obtained. The preset condition may be that the loss function converges, the number of iterations reaches a threshold, etc.
[0086] Each training sample in the first training set can include sample calibration parameters, sample acoustic fingerprint data, sample historical fingerprint data, sample historical vibration data, sample environmental sensor data, sample current playback parameters, and sample room structure data. The first training sample can be obtained from historical data. The first label is the sample calibration effect score corresponding to the first training sample.
[0087] In some embodiments, the first label can be obtained manually or by other methods. Manual labeling can be achieved by: after the second determining module calibrates the audio system using sample calibration parameters, the audio system plays audio for the experimenters to listen to and score; or by arranging different experimenters at different locations in the room to listen and score simultaneously, and taking the average score as the first label.
[0088] The trained recommendation model can quickly calculate the calibration effect score of the candidate calibration parameters, so as to judge the calibration effect of the candidate calibration parameters and further improve the efficiency of determining the first calibration parameter.
[0089] In some embodiments of this specification, the calibration effect score is determined based on the effect determination model, which has higher accuracy and efficiency, and is beneficial for subsequently determining the first calibration parameters that are more in line with the actual room conditions based on the calibration effect score.
[0090] In some embodiments, the input to the effectiveness determination model may further include test fingerprint data, historical test fingerprint data, and user preference data, and the output of the effectiveness determination model is a calibration effectiveness score for the candidate calibration parameters. The second determination module may obtain a second calibration parameter based on the calibration effectiveness score. Further explanation of the second calibration parameter can be found in [link to documentation]. Figure 4 The relevant description of step 460.
[0091] In some embodiments, when the input to the effect determination model includes test fingerprint data, historical test fingerprint data, and user preference data, the first training sample may further include sample test fingerprint data, sample historical test fingerprint data, and sample user preference data. The first label is the sample calibration effect score corresponding to the first training sample. Further explanation of the first label can be found in the relevant description above in step 340.
[0092] Test fingerprint data refers to the acoustic fingerprint data corresponding to the test audio signal. For an explanation of the test audio signal, see [link to documentation]. Figure 4Step 410 and its contents. In some embodiments, the second determining module may generate test fingerprint data corresponding to the test audio signal based on the test audio signal using a preset fingerprint algorithm. For details regarding the preset fingerprint algorithm, please refer to the relevant content of step 330 above, which will not be repeated here. Historical test fingerprint data refers to the test fingerprint data calculated and stored from the last time the playback parameters of the audio system were calibrated, based on historical data.
[0093] User preference data refers to data related to users' auditory preferences. Different users have different preferences for sound quality characteristics. For example, some users care more about the mixing effect of vocals and background music than the clarity of vocals. Therefore, the second determination module can filter out candidate calibration parameters that better meet the user's needs based on user preference data.
[0094] In some embodiments, user preference data may include the degree of preference for sound quality characteristics such as vocal clarity, mixing effects, and bass clarity. In some embodiments, the second determining module may acquire user preference data based on user input. In some embodiments, the second determining module may obtain user preference data based on analysis of the user's historical access data. Historical access data refers to relevant data from the user's historical access times, such as access data for songs, etc. Historical access data of different users may be pre-stored in the storage device.
[0095] In some embodiments, the number of candidate calibration parameters can be a preset number. The preset number refers to the number of candidate calibration parameters used to determine the second calibration parameter. In some embodiments, the second determination module can output a calibration effect score corresponding to each set of candidate calibration parameters based on the preset number and through an effect determination model. In some embodiments, the preset number can be determined based on empirical presets, and the preset number is less than a preset computational load.
[0096] In some embodiments, the preset quantity is also negatively correlated with the first similarity or the second similarity. For an explanation of the first similarity, see the relevant content in step 330 above; for an explanation of the second similarity, see [link to relevant content]. Figure 4 Step 450 and its contents.
[0097] The second determination module randomly generates candidate calibration parameters based on the current playback parameters. If the values of the first similarity or the second similarity are relatively large (but both are less than the first similarity threshold and the second similarity threshold), it indicates that the acoustic fingerprint data in the room does not change much. In this case, the second determination module only needs to generate a small number of candidate calibration parameters and input them into the effect determination model to determine the first or second calibration parameter with better calibration effect. Conversely, if the values of the first similarity or the second similarity are too small, it indicates that the acoustic fingerprint data in the room changes significantly. In this case, it is necessary to generate a large number of candidate calibration parameters and filter them based on the effect determination model to improve the calibration effect of the finally determined first or second calibration parameters.
[0098] In some embodiments, the second determining module may sort the calibration effect scores corresponding to a preset number of candidate calibration parameters and select the candidate calibration parameter with the highest calibration effect score as the second calibration parameter. For a description of the second calibration parameter, see [link to documentation]. Figure 4 Step 460 and its contents.
[0099] In some embodiments of this specification, by using test fingerprint data, historical test fingerprint data, user preference data, and a preset number of candidate calibration parameters as input to the effect determination model, more candidate calibration parameters that meet user needs can be obtained. Furthermore, the candidate calibration parameters are screened, and the candidate calibration parameter with the highest calibration effect score is used as the second calibration parameter, thereby improving the efficiency of obtaining the second calibration parameter.
[0100] In some embodiments, the second determining module can input the candidate calibration parameters with preset calculations into the effect determining model, and simultaneously input acoustic fingerprint data, historical fingerprint data, historical vibration data, environmental sensing data, current playback parameters and room structure data. It can also sort the multiple calibration effect scores output by the effect determining model, select the candidate calibration parameter with the highest calibration effect score, and determine it as the first calibration parameter.
[0101] In some embodiments, the second determining module may further input test fingerprint data, historical test fingerprint data, and user preference data simultaneously into the effect determination model to obtain the calibration effect score corresponding to the candidate calibration parameter output by the effect determination model. The second determining module may sort the multiple calibration effect scores corresponding to a preset number of candidate calibration parameters, select the candidate calibration parameter with the highest calibration effect score, and determine it as the second calibration parameter.
[0102] In some embodiments of this specification, candidate calibration parameters are scored using multiple methods based on candidate calibration parameters, first spectral characteristics, and historical vibration data. The candidate calibration parameter with the highest calibration effect is selected as the first calibration parameter, so that the determined first calibration parameter conforms to the actual room conditions and improves the calibration effect of the first calibration parameter on the audio system.
[0103] Step 350: Based on the first calibration parameters, calibrate the playback parameters of the audio system.
[0104] In some embodiments, the second determining module may adjust and replace the playback parameters of the audio system based on the first calibration parameters to complete the calibration of the playback parameters of the audio system.
[0105] In some embodiments of this specification, by determining whether the first spectral characteristics meet the first calibration conditions, it is determined whether the room layout will significantly affect the sound quality. For cases where the sound quality change is minor, unnecessary calibration can be avoided, reducing computation and energy consumption. For cases where the sound quality change is significant, the first calibration parameters can be determined to calibrate the playback parameters of the audio system, achieving real-time calibration of the playback parameters without requiring manual adjustment by the user, which is beneficial for improving audio propagation quality and listening experience.
[0106] Figure 4 This is an exemplary flowchart illustrating the determination of a second calibration parameter according to some embodiments of this specification. Figure 4 As shown, process 400 includes the following steps. In some embodiments, process 400 may be performed by a playback parameter calibration system 200.
[0107] Step 410: Generate a test audio signal.
[0108] For an explanation of the first spectral characteristics and the test audio signal, please refer to [link / reference]. Figure 3 And its contents.
[0109] The test audio signal is an audio signal used to test changes in room characteristics. In some embodiments, the test audio signal may include one or a combination of pure tone (single-frequency sound), broadband noise, narrowband noise, sweep signal, pulse signal, etc. For example, the test audio signal may be in the form of {(pure tone, 5s), (broadband noise, 10s)}, where the above expression means that the test audio signal is a combination of 5 seconds of pure tone followed by 10 seconds of broadband noise.
[0110] For example, pure tones can be used to test the response at specific frequencies, such as testing a room's response to 1500Hz sound; broadband noise and narrowband noise can be used to test audio that requires a longer coverage time, such as audio with an overall length greater than 60 seconds; swept signals can be used to quickly measure the response of audio within a specific frequency range, such as testing a room's response to sound in the 100Hz-2000Hz range; and pulse signals can be used to test a room's response to instantaneous changes in sound pressure.
[0111] In some embodiments, the second determining module can generate multiple test audio signals using various methods. For example, the multiple test audio signals can be pre-set to be generated. The second determining module can store the pre-generated multiple test audio signals in a storage device. The second determining module can select any one of them as the test audio signal to be used in this instance. For example, if the first spectral characteristics show that the audio frequencies are concentrated in the range of 100Hz-2000Hz, the second determining module can use a swept frequency signal as the test audio signal.
[0112] In some embodiments, the second determining module may further determine an initial test signal based on the first spectral characteristics and background noise data; and generate a test audio signal based on the initial test signal and room structure data. For a description of the first spectral characteristics, see [link to documentation]. Figure 3 Step 320 and its contents.
[0113] Background noise data refers to other data unrelated to the first played audio. For example, at least one of the following: noise coming from outside the room, or noise caused by people moving inside the room. In some embodiments, the second determining module may acquire background noise data based on an audio monitoring device. See [link to audio monitoring device description] for details. Figure 1 And its contents.
[0114] The initial test signal refers to the candidate test audio signal. In some embodiments, different initial test signals have different frequency modulation parameters, lengths, etc. The frequency modulation parameters refer to the characteristic parameters of the initial test signal, such as amplitude, waveform, and frequency. The frequency modulation parameters of the initial test signal are particularly important for testing room characteristic changes; therefore, the frequency modulation parameters of the initial test signal determined by the second determining module must be within a preset frequency modulation range.
[0115] The preset frequency range can be determined based on experience, including the corresponding numerical ranges of amplitude and frequency, as well as the waveform type. For example, amplitude reflects the intensity and loudness of the initial test signal. If the amplitude is too small, the initial test signal may not be accurately captured by the audio monitoring device in a large space or a room with high background noise. If the amplitude is too large, it may cause distortion of the initial test signal, damage to the equipment, and affect the customer's listening experience. Therefore, the amplitude needs to be controlled within the preset frequency range.
[0116] For example, a waveform describes how the frequency, loudness, etc. of the initial test signal change over time. Different types of waveforms can be used for different measurement purposes, which are related to the measurement purpose of the test audio signal. Therefore, the waveform needs to conform to the measurement purpose.
[0117] For example, a sine wave has only one frequency component, making it suitable for frequency response measurements and easy to analyze. Pulse signals can be used to measure a room's impulse response, i.e., the room's reaction to instantaneous changes in sound pressure. Noise signals (such as white noise or pink noise), containing a wide frequency band, can be used to comprehensively assess a room's frequency response. Sweep signals can reflect continuously varying frequencies from low to high or from high to low, making them suitable for quickly measuring room responses across the entire frequency range.
[0118] For example, frequency reflects the periodicity and pitch of the initial test signal and is crucial for detecting various reflection, absorption and resonance characteristics in a room. For instance, low-frequency sound waves travel further than high-frequency sound waves and are more easily affected by room size and object layout, while high-frequency sound waves are more easily absorbed and scattered by surface materials. Therefore, the frequency needs to be controlled within the frequency range corresponding to the preset tuning range.
[0119] In some embodiments, the second determining module can determine the initial test signal through various methods based on the first spectral features and background noise data. In some embodiments, the second determining module can construct an initial test vector (x, y) based on the first spectral features and background noise data. Based on the initial test vector, by querying a test database, the reference test signal corresponding to the second reference vector with the highest similarity to the initial test vector is used as the initial test signal corresponding to the initial test vector. The test database can be constructed based on historical data, or it can be manually modified and supplemented.
[0120] The test database includes a second reference vector constructed based on historical first spectral characteristics and historical background noise data, as well as the reference test signal corresponding to the second reference vector. The second determination module can determine the actual test signal that meets the preset test conditions corresponding to the second reference vector in the historical data as the reference test signal. The preset test conditions refer to the conditions for determining the reference test signal that are preset in advance, for example, the preset test condition is that the actual test signal is within a preset frequency modulation range.
[0121] In some embodiments, the second determining module can generate a test audio signal in various ways based on the initial test signal and room structure data. For example, if the room structure data indicates that the room's area or volume is greater than a room area threshold or room volume threshold, the second determining module can increase the amplitude of the initial test signal and decrease its frequency, using the adjusted initial test signal as the test audio signal. The increase or decrease in amplitude can be positively correlated with the difference between the room's area or volume and the room area threshold or room volume threshold. The room's area or volume can be calculated based on the room's dimensions. The room area threshold or room volume threshold can be preset in advance.
[0122] For example, if the room structure data represents a high degree of complexity in terms of room shape, the second determining module can extend the duration of the initial test signal and increase the waveform types of the initial test signal, using the adjusted initial test signal as the test audio signal. The extended duration of the initial test signal can be positively correlated with the complexity; for example, the greater the complexity, the longer the extended duration of the initial test signal. For a description of the complexity of room structure data, room dimensions, and room shape, see [link to documentation]. Figure 3 And its contents.
[0123] In some embodiments of this specification, an initial test signal is determined based on the first spectral characteristics and background noise data, and a test audio signal is generated based on the initial test signal and room structure data. By taking into account the room structure and background noise data, the generated test audio signal can more accurately test the room acoustics according to the actual situation of the room, which is beneficial for generating a second calibration parameter with better calibration effect based on the test audio signal.
[0124] In some embodiments, during the playback of composite audio data, the second determining module can dynamically adjust subsequent test audio signals based on vibration sensing data. For an explanation of the composite audio data, see the relevant description of step 420 below. Subsequent test audio signals refer to the portion of the test audio signal that has not yet been played. Subsequent test audio signals can also refer to the corrected test signal generated in step 420 below.
[0125] Vibration sensing data refers to data related to the vibrations caused by the sound waves of played audio. Examples include the amplitude and frequency of the vibrations caused by the sound waves. In some embodiments, vibration sensing data can be acquired based on vibration sensors configured within the room.
[0126] In some embodiments, in response to the vibration sensing data exceeding a first amplitude threshold during the playback of composite audio data, the second determining module can adjust subsequent test audio signals based on a first preset amplitude. For example, the amplitude of subsequent test audio signals can be adjusted to 80% of the original amplitude. The first amplitude threshold can be determined empirically. The first preset amplitude is related to the magnitude of the vibration sensing data; the larger the vibration sensing data value, the larger the first preset amplitude, meaning the larger the vibration sensing data value, the smaller the amplitude of the adjusted test audio signal.
[0127] In some embodiments, in response to vibration sensing data falling below a second amplitude threshold during playback of composite audio data, the second determining module may adjust the test audio signal based on a second preset amplitude. For example, the amplitude of subsequent test audio signals may be adjusted to 120% of the original amplitude.
[0128] The second amplitude threshold can be determined empirically. The second preset amplitude is related to the magnitude of the vibration sensing data; the smaller the vibration sensing data value, the larger the second preset amplitude, meaning the smaller the vibration sensing data value, the larger the amplitude of the adjusted test audio signal. The value of the second amplitude threshold is less than the first amplitude threshold.
[0129] In some embodiments of this specification, during the playback of composite audio data, the subsequent test audio signal is dynamically adjusted based on vibration sensing data. This makes the determined subsequent test audio signal more adaptable to changes in room conditions, which is beneficial for determining more accurate second calibration parameters.
[0130] Step 420: In response to the first spectral feature satisfying the first calibration condition, a composite playback audio is generated based on the test audio signal and the second playback audio.
[0131] For instructions on how to determine if the first spectral feature meets the first calibration condition, please refer to [link to documentation]. Figure 3 Step 330 and its contents. In some embodiments, in response to the first spectral feature satisfying the first calibration condition, the second determining module may also generate a test audio signal to further detect the impact of room feature changes on audio quality, which is beneficial for further calibration of the playback parameters of the audio system.
[0132] The second playback audio refers to the audio that the audio system is playing when the test audio signal is played. The playback order of the second playback audio is after the first playback audio. In some embodiments, the second determining module may obtain the second playback audio based on the audio system.
[0133] Composite playback audio refers to audio that combines a test audio signal and a second playback audio signal. In some embodiments, the second determining module can enable the audio system to simultaneously play the test audio signal and the second playback audio signal, and acquire the composite playback audio.
[0134] In some embodiments, the second determining module may also determine the embedding time point based on vibration sensing data; and generate composite playback audio based on the embedding time point, the test audio signal, and the second playback audio.
[0135] Some test audio signals may be outside the range of human hearing when played alone, but the test audio signals will interfere with and superimpose with sounds that are within the range of human hearing. Therefore, in order to reduce the impact on the user's hearing, the second determining module can transmit test audio signals when the amplitude in the room is small (i.e., the amplitude of the second played audio is small).
[0136] The embedding point refers to a time when the vibration amplitude in the room is relatively small. In some embodiments, the second determining module can acquire vibration sensing data and use the time point when the value is less than a third amplitude threshold as the embedding point. The third amplitude threshold can be determined based on an empirical preset. In some embodiments, if the second determining module detects that the user has reduced the output power of the audio system (e.g., lowered the volume) at a certain time, the second determining module can use that time point as the embedding point.
[0137] In some embodiments, the second determining module may simultaneously play the test audio signal and the second playback audio at the embedding time point to generate composite playback audio.
[0138] In some embodiments of this specification, by setting an embedding time point and adding a test audio signal at the embedding time point to generate composite playback audio, the impact of the test audio signal on the user's hearing can be reduced, thereby improving the user experience.
[0139] In some embodiments, the second determining module can enable the audio system to play the composite playback audio generated in step 420 above, and obtain the composite spectrum features corresponding to the composite playback audio based on the audio monitoring device, and separate the test spectrum features from them to determine whether the test spectrum features meet the correction conditions.
[0140] The test spectral characteristics are the spectral characteristics corresponding to the test audio data during composite audio playback. The test audio data is the data corresponding to the test audio signal. The method for obtaining the composite spectral characteristics corresponding to the composite audio playback is similar to the method for determining the first spectral characteristics based on the first audio data; see [link to documentation] for details. Figure 3 Related descriptions.
[0141] In some embodiments, if the test spectrum characteristics meet the correction conditions, it indicates that the quality of the currently obtained composite playback audio is poor, and the calibration effect of the second calibration parameters obtained based on the composite playback audio is also poor. Therefore, the second determination module needs to further correct the composite playback audio. If the test spectrum characteristics do not meet the correction conditions, it indicates that the quality of the currently obtained composite playback audio is good, and step 430 can continue.
[0142] In some embodiments, in response to the test spectrum characteristics satisfying the correction conditions, the second determining module may determine a corrected test signal based on the test spectrum characteristics; and generate a corrected composite audio based on the corrected test signal and the second played audio.
[0143] For more information on the characteristics of the test spectrum, please refer to [link / reference]. Figure 4 The above description.
[0144] Correction conditions refer to the conditions under which the test audio signal needs to be corrected. For example, correction conditions may include at least one of the following: the test spectral features cannot be separated from the composite spectral features, or some frequencies of the separated test spectral features are missing spectral features. Here, composite spectral features refer to the spectral features corresponding to the composite audio data.
[0145] In some embodiments, the correction conditions may be determined based on manual presets. For further explanation of composite audio data, see step 430 and its contents below; for further explanation of separating composite spectral features, see step 440 and its contents below.
[0146] In some embodiments, when the test spectrum feature cannot be separated from the composite spectrum feature, the second determining module can determine that the test spectrum feature meets the correction condition. In some embodiments, the second determining module can test the test spectrum feature, and when the test reveals that some frequency-corresponding spectrum features are missing in the test spectrum feature separated from the composite spectrum feature, the second determining module can determine that the test spectrum feature meets the correction condition.
[0147] The corrected test signal refers to the corrected test audio signal.
[0148] In some embodiments, the second determining module can determine the corrected test signal based on the test spectrum characteristics through various methods. For example, when the test spectrum characteristics cannot be separated from the composite spectrum characteristics, the second determining module can increase the amplitude of the test audio signal or change the waveform of the test audio signal to obtain the corrected test signal. As another example, when some frequencies in the separated test spectrum characteristics are missing, the second determining module can increase the frequency or amplitude of the missing frequencies to obtain the corrected test signal.
[0149] Corrected composite audio refers to the corrected composite playback audio. In some embodiments, the second determining module can cause the audio system to simultaneously play the corrected test signal and the second playback audio, and use the composite audio as the corrected composite audio.
[0150] In some embodiments, after one round of correction, the second determining module can cause the audio system to play the generated corrected composite audio, and obtain the composite spectral characteristics corresponding to the corrected composite audio based on the audio monitoring device, and separate the test spectral characteristics from them, determine whether the test spectral characteristics meet the correction conditions, and then correct the corrected test signal again until its corresponding test spectral characteristics no longer meet the correction conditions, and generate the corrected composite audio based on the corrected test signal. The frequency modulation parameters of the corrected test signal must still meet the preset frequency modulation range.
[0151] In some embodiments of this specification, by setting correction conditions for the test spectrum characteristics, the actual correction effect of the finally determined composite playback audio on the audio system can be further improved, ensuring reliability.
[0152] Step 430: Obtain composite audio data corresponding to the composite audio playback through the audio monitoring device.
[0153] Composite audio data refers to the acoustic data related to the room's response when playing composite audio. In some embodiments, the second determining module can acquire composite audio data based on an audio monitoring device while playing composite audio.
[0154] Step 440: Based on the composite audio data, determine the second spectral feature and the test spectral feature.
[0155] The second spectral feature refers to the spectral feature corresponding to the second audio data of the second played audio. The second audio data refers to the acoustic data related to the response in the room when the second played audio is played. In some embodiments, the second determining module can acquire the second audio data based on the audio monitoring device and acquire the second spectral feature based on the second audio data. The acquisition method is similar to the method of acquiring the first spectral feature based on the first audio data, and will not be described again here. Please refer to [link to relevant documentation]. Figure 3 And its contents.
[0156] In some embodiments, the second determining module may also obtain composite spectral features based on composite audio data, and the method of obtaining the composite spectral features is similar to that of obtaining the first spectral features based on the first audio data.
[0157] The second determining module can further separate the composite spectral features, extracting the second spectral features and the test spectral features respectively. The method by which the second determining module separates the composite spectral features can include at least one or a combination of blind source separation (BSS) algorithms, correlation analysis algorithms, etc. For example, the second determining module can perform correlation analysis on the test audio signal and the composite audio signal to find the position and intensity of the test audio signal in the composite audio signal, and separate the corresponding portion of the composite spectral features to obtain the test spectral features.
[0158] Step 450: Determine whether the second spectral feature and the test spectral feature meet the second calibration conditions.
[0159] The second calibration condition refers to the preset conditions used for a secondary determination of whether the playback parameters need to be calibrated. That is, after determining that the playback parameters need to be calibrated using the first calibration condition, the second determination module can again determine whether the second spectral characteristics and the test spectral characteristics meet the second calibration condition, and whether the playback parameters of the audio system need to be calibrated.
[0160] For example, the second calibration condition could be a second similarity lower than a second similarity threshold. Here, the second similarity is the weighted sum of similarity 1 and similarity 2. Similarity 1 refers to the similarity between the acoustic fingerprint data corresponding to the second spectral feature and the historical fingerprint data 'a'. Historical fingerprint data 'a' refers to the historical fingerprint data extracted from the second spectral feature in the historical data, calculated and stored from the last time the playback parameters of the audio system were calibrated. Similarity 2 refers to the similarity between the acoustic fingerprint data corresponding to the test spectral feature and the historical test fingerprint data 'a'. Historical test fingerprint data 'a' refers to the historical test fingerprint data extracted from the test spectral feature in the historical data, calculated and stored from the last time the playback parameters of the audio system were calibrated.
[0161] The weighting coefficients can be determined based on empirical presets. The calculation methods for the acoustic fingerprint data corresponding to the second spectral feature and the test spectral feature are similar to those for the acoustic fingerprint data corresponding to the first spectral feature, and will not be repeated here. Please refer to [link to relevant documentation]. Figure 3 And its contents.
[0162] The second similarity threshold can be determined based on an empirical preset. In some embodiments, the second similarity threshold can also be positively correlated with the room area. The room area can be obtained based on user input on the user terminal. In some embodiments, the second similarity threshold can also be positively correlated with the time interval between the current playback parameter calibration and the previous playback parameter calibration. In some embodiments, the second similarity threshold is less than the first similarity threshold. For an explanation of the first similarity threshold, see [link to documentation]. Figure 3 And its contents.
[0163] Step 460: In response to the second spectral feature and the test spectral feature satisfying the second calibration condition, determine the second calibration parameter based on the second spectral feature and the test spectral feature.
[0164] The second calibration parameters refer to the playback parameters used to calibrate the audio system during the second playback audio. Examples include equalization (EQ) settings, speaker output parameters, tone control (TC) parameters, and reverb control parameters. For explanations of these playback parameters, please refer to [link to documentation / reference]. Figure 3 And its contents.
[0165] In some embodiments, the second determining module can determine the second calibration parameters based on the second spectral features and the test spectral features using various methods. For example, the second determining module can construct a second calibration vector based on the acoustic fingerprint data corresponding to the second spectral features, the acoustic fingerprint data corresponding to the test spectral features, historical fingerprint data 'a', historical test fingerprint data 'a', and the current playback parameters. Based on the second calibration vector, by querying the second database, the second reference parameter corresponding to the third reference vector with the highest similarity to the second calibration vector is used as the second calibration parameter corresponding to the second calibration vector. The second calibration vector can be in the form of a four-dimensional vector, such as (a, b, c, d); the second database contains preset playback parameters corresponding to different second calibration vectors, and the second database can be determined based on preset parameters. The construction method of the second database is similar to that of the first database; see [link to documentation] for more details. Figure 3 Related descriptions.
[0166] Step 470: Based on the second calibration parameter, calibrate the playback parameters of the audio system.
[0167] In some embodiments, the second determining module may adjust and replace the playback parameters of the audio system based on the second calibration parameters to complete the calibration of the audio system.
[0168] In some embodiments, the second determining module may also determine whether it is necessary to determine the second calibration parameter based on the test audio signal, the second played audio, the second spectral characteristics, the test spectral characteristics, and room structure data, using a calibration judgment model. For further details, please refer to [link to relevant documentation]. Figure 5 And its contents.
[0169] In some embodiments of this specification, a test audio signal is generated and fused with a second playback audio signal to obtain a composite audio signal. The composite audio signal is then further evaluated to determine whether the audio system needs to be calibrated. This avoids unnecessary calibration, reduces computation and energy consumption, and allows for more accurate dynamic detection of changes in room acoustics, improving calibration efficiency and user experience.
[0170] Figure 5 This is an exemplary schematic diagram of a calibration judgment model shown in some embodiments of this specification.
[0171] In some embodiments, the second determining module may determine whether it is necessary to determine the second calibration parameter by means of a calibration judgment model. The calibration judgment model 500 refers to a model used for calibrating and determining whether the audio system needs to determine the second calibration parameter. The calibration judgment model can be a machine learning model, such as any one or a combination of Neural Networks (NN), Recurrent Neural Networks (RNN), and Convolutional Neural Networks (CNN).
[0172] In some embodiments, the calibration decision model may include a prediction layer and a decision layer.
[0173] In some embodiments, the second determining module may determine global spectral features 530-1 and global test features 530-2 based on the test audio signal 510-1, the second playback audio 510-2, the second spectral feature 510-3, the test spectral feature 510-4, and the room structure data 510-5, through the prediction layer 520 of the calibration judgment model; and determine whether it is necessary to determine the second calibration parameter 550 based on the global spectral feature 530-1, the global test feature 530-2, the historical global spectral feature 530-3, the historical global test feature 530-4, the historical background noise data 530-5, and the current background noise data 530-6, through the judgment layer 540 of the calibration judgment model.
[0174] The prediction layer 520 of the calibration judgment model is the part of the calibration judgment model used to predict global spectral features and global test features.
[0175] In some embodiments, the input to the prediction layer 520 of the calibration judgment model may include a test audio signal 510-1, a second playback audio signal 510-2, a second spectral feature 510-3, a test spectral feature 510-4, and room structure data 510-5, and the output may include a global spectral feature 530-1 and a global test feature 530-2.
[0176] For more information on the test audio signal, the second playback audio, the second spectral characteristics, the test spectral characteristics, and the room structure data, please refer to [link to relevant documentation]. Figure 3 And related content.
[0177] Global spectral features refer to the predicted spectral features corresponding to multiple locations within the room when the second audio is played. For example, global spectral features can be represented in the form of {(room location 1, spectral feature 1), (room location 2, spectral feature 2)...}.
[0178] Global test characteristics refer to the spectral characteristics of multiple locations within a room when a test audio signal is played.
[0179] In some embodiments, the prediction layer of the calibration judgment model can be trained using a large number of first audio samples and labels corresponding to the first audio samples.
[0180] In some embodiments, each group of first audio samples may include: a sample test audio signal, a sample second playback audio, sample test spectral features, sample second playback spectral features, sample room structure data, etc. The tag corresponding to the first audio sample may include the actually acquired global spectral features of the sample, the global test spectral features of the sample, etc.
[0181] The first audio sample can be obtained based on historical data. The tag corresponding to the first audio sample can be obtained by using audio monitoring devices pre-placed at multiple locations in the experimental room to collect composite audio data from each location, and to acquire the sample test spectral characteristics and sample second spectral characteristics at each location. These are then combined to form the sample global spectral characteristics and sample global test characteristics. For an explanation of determining the test spectral characteristics and second spectral characteristics based on composite audio data, please refer to... Figure 4 And its contents.
[0182] The judgment layer 540 of the calibration judgment model refers to the part of the calibration judgment model used to determine whether a second calibration parameter needs to be determined.
[0183] In some embodiments, the inputs to the judgment layer 540 of the calibration judgment model may include global spectral features 530-1, global test features 530-2, historical global spectral features 530-3, historical global test features 530-4, historical background noise data 530-5, and current background noise data 530-6, and the output may include whether a second calibration parameter needs to be determined 550.
[0184] If the calibration judgment model determines that it is required, then the second spectral feature and the test spectral feature meet the second calibration conditions, and the second determination module needs to determine the second calibration parameters; if the calibration judgment model determines that it is not required, then the second spectral feature and the test spectral feature do not meet the second calibration conditions, and the second determination module does not need to determine the second calibration parameters.
[0185] Historical global spectral characteristics refer to the spectral characteristics of multiple locations within the room when playing the second audio recording, calculated and stored from the time of the last calibration of the audio system's playback parameters at the current point in time. For example, historical global playback spectral characteristics can be represented in the form of {(room location 1, spectral characteristic 1), (room location 2, spectral characteristic 2)...}. In some embodiments, the second determining module can obtain historical global spectral characteristics from a storage device via a network.
[0186] Historical global test features refer to the spectral characteristics of multiple positions corresponding to the playback of test audio by the audio system, calculated and stored from the last time the playback parameters of the audio system were calibrated. In some embodiments, the processor can obtain historical global test features from a storage device via a network.
[0187] Current background noise data refers to data related to background noise at the current time or in the near future. For example, at least one of the following: noise coming from outside the room at the current time, or noise caused by people's activities inside the room. In some embodiments, the second determining module may acquire current background noise data based on an audio monitoring device.
[0188] Historical background noise data refers to data related to background noise over a period of time relative to the current point in time. In some embodiments, the second determining module may acquire historical background noise data at a historical time based on an audio monitoring device and store it in a storage device. The second determining module may also acquire historical background noise data from the storage device via a network.
[0189] In some embodiments, the judgment layer of the calibration judgment model can be trained using a large number of second audio samples and labels corresponding to the second audio samples. In some embodiments, the second audio samples may include: global spectral features, global test features, and background noise data of the samples at a second time point, as well as historical global spectral features, historical global test features, and historical background noise data of the samples at a first time point. The first time point is earlier than the second time point. The labels corresponding to the second audio samples may include "needed" (requiring determination of the second calibration parameters) and "not needed" (not requiring determination of the second calibration parameters). The second audio samples can be obtained based on historical data. The second audio samples and the labels corresponding to the second audio samples can be obtained through the following steps:
[0190] S1, Play the sample audio under the original room characteristics, and obtain the historical global spectrum features and historical global test features of the sample at the first time by obtaining the label corresponding to the first audio sample as described above, and simultaneously collect the historical background noise data of the sample at the first time.
[0191] Original room features refer to the room structure data before the audio sample is played. Original room features may include at least one of the following: room size, room shape, ceiling height, whether there are windows, room layout, etc.
[0192] The sample playback audio can be any pre-stored playback audio or test audio.
[0193] S2, change room characteristics (e.g., move furniture, change window open / close status, etc.) or not change room characteristics, play sample audio, and obtain the global spectral characteristics and global test characteristics of the sample at the second time point using the method described above for obtaining the label corresponding to the first audio sample. Simultaneously collect the current background noise data of the sample at the second time point. The first time point is earlier than the second time point.
[0194] S3: Based on whether the room characteristics were actually changed in S2, label the label corresponding to the second audio sample as "needs" or "does not need". For example, if the room characteristics are changed, the label corresponding to the second audio sample is "needs to determine the second calibration parameter". If the room characteristics are not changed, the label corresponding to the second audio sample is "does not need to determine the second calibration parameter".
[0195] In some embodiments, the original room features can be changed multiple times to obtain more diverse samples. For changes to the same room features, different samples can be used to play audio before and after the change (e.g., different test audio signals and second playback audio). Correspondingly, the global playback spectrum features and global test spectrum features obtained by the second determining module through the audio system will also be different, thereby obtaining more diverse samples, which helps to improve the generalization ability of the calibration judgment model and improve the accuracy of the judgment. There can be multiple original room features, such as changing the room scene or changing the furniture layout. Increasing the original room features helps to improve the diversity of samples and expand the coverage of samples.
[0196] In some embodiments, the input to the judgment layer 540 of the calibration judgment model may also include current environmental data and historical environmental data, current audio system data and historical audio system data.
[0197] Current environmental data refers to data on various parameters or conditions related to the environment in which the audio system is located at the current time or in the near future. In some embodiments, current environmental data may be current ambient temperature and current ambient humidity, etc. Historical environmental data refers to data on various parameters or conditions related to the environment in which the audio system is located over a historical period of time relative to the current point in time. Current environmental data and historical environmental data can be acquired by the second determining module through the audio monitoring device.
[0198] Current audio system data refers to various parameters or performance data related to the current state of the audio system at the current time. Historical audio system data refers to various parameters or performance data related to the historical state of the audio system over a period of time relative to the current point in time.
[0199] In some embodiments, the current audio system data may include at least one of playback device specifications, current playback parameters, etc. The playback device specifications may include device version type, impedance, rated power, etc. The current audio system data may be acquired by the second determining module through an audio detection device. Playback device specifications can be obtained by pre-checking device parameters. For more information on current playback parameters, please refer to [link to relevant documentation]. Figure 3 And related content.
[0200] In some embodiments, when the input to the judgment layer 540 of the calibration judgment model includes current environmental data, historical environmental data, current audio system data, and historical audio system data, the corresponding second audio sample may further include sample environmental data and sample audio system data at a second time, as well as sample historical environmental data and sample historical audio system data at a first time. The first time is earlier than the second time.
[0201] In some embodiments of this specification, by considering the influence of environmental factors, such as ambient temperature and humidity, on audio, the reliability of the calibration judgment model's output can be improved. By inputting current audio system data and historical audio system data into the judgment layer of the calibration judgment model, the influence of playback devices and playback parameters on audio can be considered, and further determination can be made as to whether a second calibration parameter is needed. This increases the factors for judging audio changes, making the judgment of changes more detailed and helping to improve the accuracy of subsequent audio calibration.
[0202] In some embodiments of this specification, the global spectral characteristics and global test characteristics are predicted by the prediction layer of the calibration judgment model to obtain room acoustic characteristics that can cover the entire room range. Then, based on the global spectral characteristics and global test characteristics, it is determined whether a second calibration parameter needs to be determined, so as to accurately judge the acoustic characteristics of the entire room range and accurately obtain the overall room acoustic calibration requirements.
[0203] One or more embodiments of this specification also provide a playback parameter calibration extension device, the device including at least one processor and at least one memory; the at least one memory is used to store computer instructions; the at least one processor is used to execute at least a portion of the computer instructions to implement the playback parameter calibration method as described in any of the above embodiments.
[0204] One or more embodiments of this specification also provide a computer-readable storage medium that stores computer instructions. When a computer reads the computer instructions in the storage medium, the computer runs a calibration method for playback parameters as described in any of the above embodiments.
[0205] The basic concepts have been described above. Obviously, for those skilled in the art, the detailed disclosure above is merely illustrative and does not constitute a limitation of this specification. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and corrections to this specification. Such modifications, improvements, and corrections are suggested in this specification and therefore remain within the spirit and scope of the exemplary embodiments described herein.
[0206] Furthermore, this specification uses specific terms to describe embodiments thereof. For example, "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Moreover, certain features, structures, or characteristics in one or more embodiments of this specification can be appropriately combined.
[0207] Furthermore, unless expressly stated in the claims, the order of processing elements and sequences, the use of numbers and letters, or other names described in this specification are not intended to limit the order of the processes and methods described herein. Although various examples have been discussed in the foregoing disclosure of some embodiments of the invention that are currently considered useful, it should be understood that such details are for illustrative purposes only, and the appended claims are not limited to the disclosed embodiments; rather, the claims are intended to cover all modifications and equivalent combinations that conform to the spirit and scope of the embodiments described herein. For example, while the system components described above can be implemented using hardware devices, they can also be implemented solely using software solutions, such as installing the described system on existing servers or mobile devices.
[0208] Similarly, it should be noted that, in order to simplify the description disclosed herein and thus aid in the understanding of one or more embodiments of the invention, the foregoing description of embodiments in this specification may sometimes combine multiple features into a single embodiment, drawing, or description thereof. However, this method of disclosure does not imply that the subject matter of this specification requires more features than those mentioned in the claims. In fact, the embodiments contain fewer features than all the features of a single embodiment disclosed above.
[0209] In some embodiments, numbers describing the quantity of components and attributes are used. It should be understood that such numbers used in the description of embodiments are modified in some examples with the terms "approximately," "approximately," or "generally." Unless otherwise stated, "approximately," "approximately," or "generally" indicates that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, which may be changed depending on the characteristics required by individual embodiments. In some embodiments, numerical parameters should take into account specified significant digits and employ a general method of digit reservation. Although the numerical ranges and parameters used to confirm their breadth of range in some embodiments of this specification are approximate values, in specific embodiments, such values are set as precisely as feasible.
[0210] For each patent, patent application, patent application publication, and other material such as articles, books, specifications, publications, and documents referenced in this specification, the entire contents of which are incorporated herein by reference. This excludes historical application documents that are inconsistent with or conflict with the content of this specification, as well as documents that limit the broadest scope of the claims in this specification (currently or subsequently appended to this specification). It should be noted that in the event of any inconsistency or conflict between the descriptions, definitions, and / or terminology used in the supplementary materials to this specification and the content of this specification, the descriptions, definitions, and / or terminology used in this specification shall prevail.
[0211] Finally, it should be understood that the embodiments described in this specification are merely illustrative of the principles of the embodiments described herein. Other variations may also fall within the scope of this specification. Therefore, alternative configurations of the embodiments described herein are intended to be illustrative rather than limiting, and should be considered consistent with the teachings of this specification. Accordingly, the embodiments described herein are not limited to those explicitly introduced and described herein.
Claims
1. A method for calibrating playback parameters, characterized in that, include: The first audio data corresponding to the first played audio is obtained through an audio monitoring device; Based on the first audio data, a first spectral feature is determined; Determine whether the first spectral feature meets the first calibration condition: In response to the first spectral feature satisfying the first calibration condition, a first calibration parameter is determined based on the first spectral feature; Based on the first calibration parameter, the playback parameters of the audio system are calibrated.
2. The method according to claim 1, characterized in that, The method further includes: Generate test audio signals; In response to the first spectral feature satisfying the first calibration condition, a composite playback audio is generated based on the test audio signal and the second playback audio. The audio monitoring device acquires composite audio data corresponding to the composite playback audio. Based on the composite audio data, the second spectral feature and the test spectral feature are determined; Determine whether the second spectral feature and the test spectral feature satisfy the second calibration condition; In response to the second spectral feature and the test spectral feature satisfying the second calibration condition, a second calibration parameter is determined based on the second spectral feature and the test spectral feature; The playback parameters of the audio system are calibrated based on the second calibration parameter.
3. The method according to claim 2, characterized in that, The generated test audio signal includes: Based on the first spectral characteristics and background noise data, the initial test signal is determined; The test audio signal is generated based on the initial test signal and room structure data.
4. The method according to claim 1, characterized in that, The determination of the first calibration parameter based on the first spectral feature includes: Based on the candidate calibration parameters with preset calculations, the first spectral characteristics, and historical vibration data, the calibration effect score of the candidate calibration parameters is determined; Based on the calibration effect score, the first calibration parameter is determined.
5. A calibration system for playback parameters, characterized in that, include: The acquisition module is configured to acquire the first audio data corresponding to the first played audio through the audio monitoring device; The first determining module is configured to determine a first spectral feature based on the first audio data; The judgment module is configured to determine whether the first spectral feature meets the first calibration condition: The second determining module is configured as follows: In response to the first spectral feature satisfying the first calibration condition, a first calibration parameter is determined based on the first spectral feature; Based on the first calibration parameter, the playback parameters of the audio system are calibrated.
6. The system according to claim 5, characterized in that, The second determining module is further configured as follows: Generate test audio signals; In response to the first spectral feature satisfying the first calibration condition, a composite playback audio is generated based on the test audio signal and the second playback audio. The audio monitoring device acquires composite audio data corresponding to the composite playback audio. Based on the composite audio data, the second spectral feature and the test spectral feature are determined; Determine whether the second spectral feature and the test spectral feature satisfy the second calibration condition; In response to the second spectral feature and the test spectral feature satisfying the second calibration condition, a second calibration parameter is determined based on the second spectral feature and the test spectral feature; The playback parameters of the audio system are calibrated based on the second calibration parameter.
7. The system according to claim 6, characterized in that, The second determining module is further configured as follows: Based on the first spectral characteristics and background noise data, the initial test signal is determined; The test audio signal is generated based on the initial test signal and room structure data.
8. The system according to claim 5, characterized in that, The second determining module is further configured as follows: Based on the candidate calibration parameters with preset calculations, the first spectral characteristics, and historical vibration data, the calibration effect score of the candidate calibration parameters is determined; Based on the calibration effect score, the first calibration parameter is determined.
9. A device for calibrating playback parameters, characterized in that, The device includes at least one processor and at least one memory; The at least one memory is used to store computer instructions; The at least one processor is configured to execute at least a portion of the computer instructions to implement the method as described in any one of claims 1 to 4.
10. A computer-readable storage medium, characterized in that, The storage medium stores computer instructions. When the computer reads the computer instructions from the storage medium, the computer executes the method as described in any one of claims 1 to 4.