Acoustic model adjustment method and device, electronic equipment and storage medium

By obtaining and analyzing environmental acoustic data in the karaoke system, extracting acoustic feature parameters and adjusting the acoustic model, the problem of manual adaptation and training in the existing karaoke system is solved, and high-quality audio output and dynamic adaptation to environmental changes are achieved.

CN120048277APending Publication Date: 2025-05-27GUANGZHOU AUTOMOBILE GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510021656.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing karaoke system requires manual adaptation and acoustic training, which is time-consuming and difficult to maintain consistency, and cannot continuously optimize the sound effects to cope with cockpit hardware aging and environmental changes.

Method used

By obtaining the ambient acoustic data collected by the microphone in the target environment when the speaker plays the sweep reference signal, multiple sets of acoustic feature parameters are extracted, and the preset acoustic model is adjusted based on these parameters and preset scoring thresholds, the target acoustic model is obtained to adapt to different environments and microphone characteristics.

Benefits of technology

It realizes the output of high-quality audio that fits the environment in any environment, enhances the user's karaoke experience, and can be dynamically adjusted to adapt to environmental changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048277A_ABST
    Figure CN120048277A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an acoustic model adjustment method and device, equipment and a storage medium. The method comprises the following steps: acquiring environmental acoustic data collected by a microphone in a target environment when a loudspeaker plays a frequency sweeping reference signal; extracting a plurality of groups of acoustic characteristic parameters from the environmental acoustic data; model parameters of a preset acoustic model are adjusted based on a preset score threshold value and the multiple sets of acoustic feature parameters, a target acoustic model is obtained, and the preset acoustic model is a corresponding relation between the parameters of the multiple acoustic features and the sound effect score established based on the multiple sets of acoustic sample feature data. Through adoption of the method, personalized optimization can be performed on the preset acoustic model according to different environments and microphone characteristics to obtain the target acoustic model, so that when a user karaoke or plays audio and video in the target environment, high-quality audio fitting the environment can be output based on the target acoustic model, and the user experience is improved. And thus, the user experience is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more specifically, to an acoustic model adjustment method, device, electronic device, and storage medium. Background Art

[0002] Currently, with the development of automotive intelligence, mainstream vehicle manufacturers pay more and more attention to in-vehicle entertainment experience. Among them, the in-vehicle karaoke function, as an important in-vehicle entertainment method, has become one of the key research objects of vehicle manufacturers.

[0003] However, due to differences in different vehicle models and hardware configurations, existing karaoke systems usually need to rely on manual adaptation and acoustic tuning one by one, which is time-consuming and difficult to maintain consistency. In addition, over time, the aging of cockpit hardware and environmental changes will further affect the sound effect, resulting in a gradual decline in the karaoke experience. The current system is difficult to cope with these dynamic changes and cannot continuously optimize the sound effect. Summary of the Invention

[0004] In view of this, the embodiments of this application propose an acoustic model adjustment method, device, electronic device, and storage medium, which can realize personalized optimization of a preset acoustic model for different environments and microphone characteristics to obtain a target acoustic model, so that when a user sings karaoke or plays audio and video in any environment, high-quality audio that fits the environment can be output based on the target acoustic model, thereby enhancing the user experience.

[0005] In a first aspect, the embodiments of this application provide an acoustic model adjustment method, the method includes: obtaining environmental acoustic data collected by a microphone in a target environment when a speaker plays a swept-frequency reference signal; obtaining multiple groups of acoustic feature parameters according to the swept-frequency signal, the environmental acoustic data, and the performance parameters of the microphone when collecting the acoustic feature data, each group of acoustic feature parameters includes parameters of multiple acoustic features, and the multiple acoustic feature parameters include at least one microphone feature and at least one environmental acoustic feature; adjusting the model parameters of a preset acoustic model based on the multiple groups of acoustic feature parameters and a preset scoring threshold to obtain a target acoustic model, where the preset acoustic model is a correspondence relationship between the parameters of multiple acoustic features and the sound effect score established based on multiple groups of acoustic sample feature data, each group of the acoustic sample feature data includes a group of sample acoustic feature parameters and the actual sound effect score corresponding to the group of sample acoustic feature parameters, and the preset scoring threshold is used to indicate the expected sound effect score of the karaoke audio output when singing karaoke through the microphone and the speaker.

[0006] Second aspect, embodiments of the present application provide an acoustic model adjustment device, the acoustic model adjustment device includes: a data acquisition module, configured to acquire environmental acoustic data collected by a microphone in a target environment when a speaker plays a swept-frequency reference signal; a feature extraction module, configured to obtain multiple sets of acoustic feature parameters according to the swept-frequency signal, the environmental acoustic data, and performance parameters of the microphone when collecting the acoustic feature data, each set of acoustic feature parameters including parameters of multiple acoustic features, the multiple acoustic feature parameters including at least one microphone feature and at least one environmental acoustic feature; a model adjustment module, configured to adjust model parameters of a preset acoustic model based on the multiple sets of acoustic feature parameters and a preset scoring threshold to obtain a target acoustic model, where the preset acoustic model is a correspondence relationship between parameters of multiple acoustic features and a sound effect score established based on multiple sets of acoustic sample feature data, each set of the acoustic sample feature data including a set of sample acoustic feature parameters and an actual sound effect score corresponding to the set of sample acoustic feature parameters, and the preset scoring threshold is used to indicate an expected sound effect score of a KTV audio output when using the microphone and the speaker for KTV singing.

[0007] In an implementable manner, the acoustic model adjustment device further includes: a model acquisition module, a scoring acquisition module, a loss acquisition module, and a parameter adjustment module. The model acquisition module is configured to acquire an initial machine learning model; the scoring acquisition module is configured to input multiple sets of acoustic sample feature data into the initial machine learning model respectively to obtain a first predicted sound effect score for each set of the acoustic sample feature data; the loss acquisition module is configured to obtain a first model loss based on the first predicted sound effect score of each set of the acoustic sample feature data and the actual sound effect score label of each set of the acoustic sample feature data; the parameter adjustment module is configured to adjust model parameters of the initial machine learning model based on the first model loss to minimize the first model loss, and when a first iteration end condition is reached, obtain the preset acoustic model.

[0008] In an implementable manner, the model adjustment module is further configured to input each set of the acoustic feature parameters into the preset acoustic model to obtain a second predicted sound effect score corresponding to each set of the acoustic feature parameters; obtain a second model loss based on the preset scoring threshold and the second predicted sound effect score, and adjust model parameters of the preset acoustic model based on the second model loss to minimize the second model loss, and when a second iteration end condition is reached, obtain the target acoustic model.

[0009] In an implementable manner, the preset acoustic model is a multiple linear regression model, the multiple linear regression model includes an independent variable representing the sound effect score and a dependent variable representing the acoustic feature parameters, and the model parameters of the preset acoustic model include regression coefficients corresponding to each of the dependent variables.

[0010] In one possible implementation, the microphone feature includes a microphone gain; the environmental acoustic feature includes a distortion level, and the environmental acoustic feature further includes at least one of an environmental frequency response, a signal-to-noise ratio, and a loudness.

[0011] In one possible implementation, the acoustic model adjustment device further includes a device parameter adjustment module, configured to adjust the performance parameters of the microphone and the speaker respectively based on the model parameters corresponding to various acoustic features in the target acoustic model.

[0012] In one possible implementation, the acoustic model adjustment device further includes a data merging module and an audio playback control module. The data acquisition module is further configured to acquire first background music data and first user karaoke data collected by the microphone when the performance parameters have been adjusted. The data merging module is configured to merge the first user karaoke data with the first background music data to obtain a first karaoke audio. The audio playback control module is configured to control the speaker to play the first karaoke audio when the performance parameters have been adjusted.

[0013] In one possible implementation, the data acquisition module is further configured to, in response to a switching instruction, acquire second user karaoke data collected by the microphone when the performance parameters are restored to the unadjusted state. The data merging module is further configured to merge the second user karaoke data with second background music data to obtain a second karaoke audio. The audio playback control module is further configured to control the speaker to play the second karaoke audio when the performance parameters are restored to the unadjusted state.

[0014] In a third aspect, an embodiment of the present application provides an electronic device, including a processor and a memory; one or more programs are stored in the memory and configured to be executed by the processor to implement the above method.

[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which program code is stored. When the program code is run by a processor, the above method is executed.

[0016] In a fifth aspect, an embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device obtains the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above method.

[0017] An acoustic model adjustment method, device, electronic device, and storage medium provided by an embodiment of the present application obtain environmental acoustic data collected by a microphone in a target environment when a speaker plays a swept reference signal, and extract multiple groups of acoustic feature parameters from the environmental acoustic data. Since each group of acoustic feature parameters includes parameters of multiple acoustic features, that is, it means that the characteristics of sound in the target environment are described from multiple angles. Therefore, when subsequently adjusting the model parameters of a preset acoustic model based on the multiple groups of acoustic feature parameters and a preset scoring threshold to obtain a target acoustic model, the initial acoustic model parameters are adjusted according to the acoustic characteristics of the target environment, so that the finally obtained target acoustic model can better adapt to the target environment to maintain high-quality audio output in the target environment. That is, by adopting the above method, when a user sings KTV or plays audio and video in any environment, the above method can be used to obtain a target acoustic model to output high-quality audio that fits the environment by using the target acoustic model, thereby enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0019] Figure 1 Shows a schematic flowchart of an acoustic model adjustment method provided by an embodiment of the present application;

[0020] Figure 2 Shows another schematic flowchart of an acoustic model adjustment method provided by an embodiment of the present application;

[0021] Figure 3 Shows Figure 1 Another schematic flowchart of step S130 in

[0022] Figure 4 Shows another schematic flowchart of an acoustic model adjustment method provided by an embodiment of the present application;

[0023] Figure 5 Shows another schematic flowchart of an acoustic model adjustment method provided by an embodiment of the present application;

[0024] Figure 6 Shows a flowchart block diagram of an acoustic model adjustment method provided by an embodiment of the present application;

[0025] Figure 7 Shows a structural block diagram of a vehicle provided by an embodiment of the present application;

[0026] Figure 8 The connection block diagram of an acoustic model adjustment device proposed by an embodiment of the present application is shown;

[0027] Figure 9 The structural block diagram of an electronic device for executing the method of an embodiment of the present application is shown. Detailed implementation manners

[0028] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.

[0029] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will recognize that the technical solutions of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be used. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.

[0030] The block diagrams shown in the accompanying drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0031] The flowcharts shown in the accompanying drawings are only illustrative and do not necessarily include all the content and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0032] It should be noted that: "a plurality of" mentioned in this article refers to two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0033] In addition, it should be noted that in the embodiments of the present application, the collection, use, processing, and storage of application information are all subject to the user's permission and need to comply with the regulations of the region where it is located.

[0034] An acoustic model adjustment method provided by this application can be applied to an electronic device, which can be a server, a terminal device, a vehicle, or a combination of one or more of the above.

[0035] In some embodiments, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0036] The terminal device can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart voice interaction device, a smart home appliance, a vehicle terminal, etc., but is not limited thereto.

[0037] Figure 1 Specifically, the acoustic model adjustment method of this application is shown. This method can be applied to an electronic device and includes:

[0038] Step S110: Obtain environmental acoustic data collected by a microphone in a target environment when a speaker plays a swept-frequency reference signal.

[0039] Among them, the target environment can be an interior environment of a vehicle, an indoor environment, or an open square, etc., which can be set according to actual needs. The specific types of the microphone and the speaker can vary according to the different target environments, and no specific limitation is made here.

[0040] Exemplarily, if the target environment is the interior environment of a vehicle, the microphone can be a built-in microphone configured in the vehicle for telephone calls or voice assistants, or a microphone connected through a data transmission interface, such as a capacitive or dynamic microphone. Correspondingly, the speaker can be a speaker configured in the vehicle, such as a speaker set in the front door, rear door, roof, etc., or a speaker added by the user. If the target environment is a square, the microphone can be a wireless microphone, and the speaker can be a loudspeaker or a broadcast system, etc. If the target environment is an indoor environment, the microphone can be a wired microphone, and the speaker can be a surround sound system set indoors.

[0041] It is worth mentioning that regardless of the environment, during the process of the microphone collecting environmental acoustic data, the external noise is relatively low (for example, the external noise is lower than a preset noise threshold, such as 20 dB or 30 dB, etc.), and the speaker needs to be able to generate a high enough sound pressure level without distortion, while the microphone has to accurately capture the sound emitted by the speaker. In addition, the positions of the microphone and the speaker can be fixed.

[0042] In an implementable manner, the target environment is the interior environment of a vehicle, and the microphone may include a far-field voice microphone built into the vehicle.

[0043] A sweep signal is a test signal with continuously changing frequency, usually scanned gradually from low frequency to high frequency. Its frequency can increase linearly (increase uniformly) or logarithmically (increase according to musical octave intervals), and maintain a relatively constant amplitude throughout the frequency range to ensure consistent energy of each frequency component. Its covered frequency band usually includes the entire audible frequency band (e.g., 20 Hz to 20 kHz), and in some cases, also includes the ultrasonic frequency band.

[0044] Step S120: Obtain multiple groups of acoustic feature parameters based on the sweep signal, the environmental acoustic data, and the performance parameters of the microphone when collecting the acoustic feature data.

[0045] Among them, each group of acoustic feature parameters includes parameters of multiple acoustic features, and the multiple acoustic feature parameters include at least one microphone feature and at least one environmental acoustic feature.

[0046] Among them, a group of acoustic features can be determined based on environmental acoustic data within a period of time, or based on acoustic feature data at a fixed frequency. There is no specific limitation here, and it can be set according to the requirements of the mobile phone.

[0047] Among them, the extraction methods and extraction sources of parameters of different types of acoustic features may be different. Therefore, the above step S120 may include: obtaining extraction methods corresponding to multiple acoustic feature types, and using the extraction methods corresponding to multiple acoustic feature types to obtain multiple groups of acoustic feature data based on the sweep signal, the environmental acoustic data, and the performance parameters when collecting the acoustic feature data. Each group of acoustic feature data includes acoustic feature data corresponding to multiple acoustic feature types respectively.

[0048] In an implementable manner, the microphone feature includes microphone gain. That is, the performance parameters of the microphone when collecting the acoustic feature data include the microphone gain of the microphone when collecting the acoustic feature data. Then, obtaining multiple groups of acoustic feature parameters based on the sweep signal, the environmental acoustic data, and the performance parameters of the microphone when collecting the acoustic feature data includes:

[0049] Obtain the microphone gains at multiple specified moments according to the performance parameters of the microphone when collecting the acoustic feature data. Among them, the microphone gains at multiple specified moments can be the same or different.

[0050] In an implementable manner, the environmental acoustic feature parameters include distortion, where the distortion is calculated based on the frequencies of the swept-frequency reference signal at multiple specified moments and the frequencies of the environmental acoustic data.

[0051] In some embodiments, the acoustic feature parameters further include environmental frequency response, signal-to-noise ratio, and loudness. Among them, the above acoustic feature parameters can be obtained based on one or more of the swept-frequency signal, the environmental acoustic data, and the microphone, etc. For specific implementation methods, reference can be made to the relevant technologies, and no specific limitations are made in this embodiment.

[0052] Step S130: Adjust the model parameters of the preset acoustic model based on the multiple sets of acoustic feature parameters and a preset scoring threshold to obtain a target acoustic model.

[0053] Among them, the preset acoustic model is the correspondence between the parameters of multiple acoustic features and the sound effect score established based on multiple sets of acoustic sample feature data. Each set of acoustic sample feature data includes a set of sample acoustic feature parameters and the actual sound effect score corresponding to this set of sample acoustic feature parameters. The preset scoring threshold is used to indicate the expected sound effect score of the karaoke audio output when singing karaoke through the microphone and the speaker.

[0054] It is worth mentioning that in the karaoke scene, the karaoke audio that users hope to hear is a sound with clear sound, no distortion, and good sense of space, and the corresponding expected sound effect score is usually high. Exemplarily, when the value range of the sound effect frequency division is 1 - 10 points, the value of the preset scoring threshold can be between 8 - 10 points. When the value range of the sound effect frequency division is 1 - 100 points, the value of the preset scoring threshold can be between 80 - 100 points, and it can be set according to actual needs.

[0055] Among them, the preset acoustic model can be obtained by training a machine learning model using each set of acoustic sample feature data, and it can accurately predict the audio effect corresponding to the collected acoustic data according to the acoustic data. Among them, the machine learning model can be a multiple linear regression model, a support vector machine, a random forest, or a neural network model (such as a multi-layer perceptron MLP, a convolutional neural network CNN), etc., which can be selected according to actual needs.

[0056] Among them, please refer to Figure 2 , the preset acoustic model is trained in the following way:

[0057] Step S210: Obtain an initial machine learning model.

[0058] Among them, the type of the initial machine learning model can be referred to the foregoing description. It is worth mentioning that the obtained initial machine learning model has completed the initialization of model parameters to ensure that model training starts from a reasonable starting point.

[0059] Step S220: Input multiple groups of acoustic sample feature data into the initial machine learning model respectively to obtain the first predicted sound effect score of each group of the acoustic sample feature data.

[0060] Among them, multiple groups of acoustic sample feature data can be extracted from multiple audio segments, and each audio segment corresponds to an actual sound effect score set by a user or a professional scoring system. The above-mentioned multiple audio segments can include audio segments obtained when different users sing KTV with different microphones and speakers.

[0061] Specifically, the above step S220 can be to use the initial machine learning model to perform vector transformation on each group of sample acoustic feature data. For a neural network model, each group of sample acoustic feature data can be converted into a two-dimensional tensor; for a linear regression model, each group of sample acoustic feature data is converted into a feature vector. The initial machine learning model makes a prediction based on the converted vector to obtain the first predicted sound effect score.

[0062] In an implementable manner, if the above-mentioned machine learning model is a neural network model with a feature extraction network and a classification network, then the above step S220 can be: using the feature extraction network to extract features from each group of sample acoustic feature data to obtain a sample acoustic feature vector, and using the classification network to perform classification calculation based on the sample acoustic feature vector to obtain the first predicted sound effect score.

[0063] In another implementable manner, if the above-mentioned machine learning model is a multiple linear regression model, the multiple linear regression model includes independent variables representing sound effect scores and dependent variables representing acoustic feature parameters, and there is a linear relationship between the sound effect score and multiple acoustic feature parameters. Each of the dependent variables in the linear relationship corresponds to a regression coefficient; the correspondence between the independent variable and the dependent variable is (the multiple linear regression model can be represented by the following formula): y = β 0 + β 1 .x 1 + β 2 .x 2 +…+ β n .x n+∈, where y is the independent variable representing the score, xi is the dependent variable representing the acoustic feature parameter, βi is the regression coefficient indicating the influence degree of each independent variable on the dependent variable. ∈ is the error term representing the part that the model cannot explain, and it is usually assumed to follow a normal distribution with a mean of zero. Then the above step S220 can be: converting each group of acoustic sample feature data into a vector of a specified length; calculating the first predicted sound effect score based on the vector corresponding to each group of acoustic features using the above formula.

[0064] Step S230: Obtain the first model loss based on the first predicted sound effect score of each group of the acoustic sample feature data and the actual sound effect score label of each group of the acoustic sample feature data.

[0065] Among them, the first model loss can be calculated based on the first sound effect score corresponding to each group of acoustic sample feature data and the actual sound effect score label. The loss function used for loss calculation can include but is not limited to any one of the cross-entropy loss function, mean square error loss function, etc., and can be selected according to the specific type of the initial machine learning model, which is not specifically limited here.

[0066] Exemplarily, if the initial machine learning model is a neural network model, then for the first predicted sound effect score of each group of the acoustic sample feature data and the actual sound effect score label corresponding to each group of the acoustic sample feature data, the cross-entropy loss function or mean square error loss function can be used for loss calculation to obtain the first model loss; if the initial machine learning model is a multiple linear regression model, then for the first predicted sound effect score of each group of the acoustic sample feature data and the actual sound effect score label corresponding to each group of the acoustic sample feature data, the least squares method can be used for loss calculation to obtain the first model loss.

[0067] It is worth mentioning that for the neural network model, the gradient descent method or Adam optimizer can be used for backpropagation to update the model parameters. For the multiple linear regression model, the normal equation can be directly solved or a regularization term can be introduced to ensure the generalization ability of the model.

[0068] Step S240: Adjust the model parameters of the initial machine learning model based on the first model loss to minimize the first model loss, and obtain a preset acoustic model until the first iteration end condition is reached.

[0069] Among them, the first iteration end condition can be that the number of iterations reaches a first preset number or the first model loss is less than a preset loss threshold.

[0070] In a specific implementation scenario, if the preset acoustic model is a multiple linear regression model, the preset acoustic model is established in the following manner: Obtain an initial multiple linear regression model, where the initial regression model includes an independent variable representing the sound effect score and a dependent variable representing the acoustic feature parameters, and there is a linear relationship between the sound effect score and multiple acoustic feature parameters. Each dependent variable in the linear relationship corresponds to a regression coefficient; Input each group of the acoustic sample feature data into the initial multiple linear regression model to obtain a predicted sound effect score; Based on the predicted sound effect score and the actual sound effect score label corresponding to each group of the acoustic sample feature data, obtain a first model loss. Adjust the regression coefficients in the initial multiple linear regression model based on the first model loss to minimize the first model loss. When the first iteration end condition is reached, obtain the preset acoustic model.

[0071] Please refer to Figure 3 , the above step S130 includes:

[0072] Step S132: Input each group of the acoustic feature parameters into the preset acoustic model to obtain a second predicted sound effect score corresponding to each group of the acoustic feature parameters.

[0073] Step S134: Based on a preset score threshold and the second predicted sound effect score, obtain a second model loss. Adjust the model parameters of the preset acoustic model based on the second model loss to minimize the second model loss. When the second iteration end condition is reached, obtain the target acoustic model.

[0074] Among them, the specific implementation principles of steps S132 and S134 can refer to the specific descriptions of steps S230 and S240 above, and will not be elaborated here one by one.

[0075] By obtaining the environmental acoustic data collected by a microphone in a target environment when a speaker plays a swept-frequency reference signal, and obtaining multiple groups of acoustic feature parameters based on the swept-frequency signal, the environmental acoustic data, and the performance parameters of the microphone when collecting the acoustic feature data, these parameters can reflect the characteristics of the microphone in the target environment and the spatial acoustic characteristics of the target environment. Subsequently, since the preset acoustic model can correctly evaluate the audio data effect, by training the preset acoustic model based on multiple groups of acoustic sample feature data and setting a preset score threshold, it is possible to realize that the adjustment direction of the model parameters is to ensure that the audio output in the target environment meets the user's expectations. Therefore, by adopting the above method, it is possible to perform personalized optimization on the preset acoustic model according to different environments and microphone characteristics to obtain the target acoustic model, so that when the user sings KTV or plays audio and video in the target environment, high-quality audio that fits the environment can be output based on the target acoustic model, thereby enhancing the user experience.

[0076] To achieve high-quality KTV audio output through a microphone and a speaker in a target environment, in one implementable manner, the method further includes: adjusting the performance parameters of the microphone and the speaker respectively based on the model parameters corresponding to various acoustic features in the target acoustic model.

[0077] Exemplarily, if the microphone characteristic parameters include microphone gain, the method further includes: adjusting the performance parameters of the microphone and the speaker respectively based on the model parameters corresponding to various acoustic features in the target acoustic model.

[0078] Specifically, a target gain can be determined according to the model parameters and the current microphone gain, and the target gain is used as the new current microphone gain.

[0079] In one implementable manner, the environmental acoustic characteristic parameters include distortion. The method can further determine a target frequency point or a target time period with distortion greater than a preset threshold. If a target time period with distortion greater than the preset distortion threshold is obtained, the target frequency point with distortion greater than the preset threshold is obtained according to the frequency of the audio played within the target time period. Adjust the power output of the speaker at the target frequency point, or compensate the output at the target frequency point through the equalizer of the microphone.

[0080] If the environmental acoustic characteristic parameters include signal-to-noise ratio, the microphone can be replaced or the noise reduction method or the noise reduction parameters of the audio collected by the microphone can be adjusted according to the model parameters corresponding to the signal-to-noise ratio. If the environmental acoustic characteristic includes loudness, the power of the speaker can be adjusted according to the model parameters corresponding to the loudness.

[0081] In one implementable manner, please refer to Figure 4 , the method further includes:

[0082] Step S140: Obtain the first background music data and the first user KTV data collected by the microphone with the adjusted performance parameters.

[0083] Among them, the first background music data can be obtained from a pre-stored audio file or in real time from an online platform through streaming.

[0084] It is worth mentioning that the first background music data and the first user KTV data are time-aligned.

[0085] Step S150: Merge the first user KTV data with the first background music data to obtain the first KTV audio.

[0086] Specifically, a professional audio processing software or hardware device can be used to mix the first user KTV data and the first background music data.

[0087] Step S160: Control the speaker to play the first karaoke audio with the adjusted performance parameters.

[0088] Since the microphone ensures that the user's singing voice collected is as pure as possible and meets the expected sound effect standard when the performance parameters are adjusted; the speaker guarantees the best sound quality output when the performance parameters are adjusted, and the sound heard by the user is clearer and fuller. Based on this, by adopting the above steps S140 - S160, high-quality karaoke audio can be output, thereby effectively improving the user's karaoke experience.

[0089] To facilitate the user to compare the karaoke effects before and after adjusting the performance parameters of the microphone and the speaker using the target acoustic model, in an implementable manner of this application, please refer to Figure 5 , the method further includes:

[0090] Step S170: In response to a switching instruction, obtain the second user karaoke data collected by the microphone when it returns to the unadjusted performance parameters.

[0091] Step S180: Merge the second user karaoke data with the second background music data to obtain a second karaoke audio.

[0092] Step S190: Control the speaker to play the second karaoke audio when it returns to the unadjusted performance parameters.

[0093] For the specific implementation principle of the above steps S170 - S190, please refer to the specific description of steps S140 - S160 in the foregoing embodiments, and details will not be repeated here.

[0094] To facilitate the user to intuitively understand the sound effects when karaoking before and after adjusting the performance parameters of the microphone and the speaker, the method further includes generating a first signal-to-noise ratio curve and a first frequency response curve based on the first karaoke audio, generating a second signal-to-noise ratio curve and a second frequency response curve based on the second karaoke audio, and controlling the display screen to display the first signal-to-noise ratio curve, the second signal-to-noise ratio curve, the first frequency response curve, and the second frequency response curve.

[0095] It is worth mentioning that there can also be other comparison methods, using a sound effect scoring model to score the first karaoke audio and the second karaoke audio respectively to obtain the respective scoring results of the first karaoke audio and the second karaoke audio, and controlling the display to show the respective scoring results of the first karaoke audio and the second karaoke audio.

[0096] Please refer to Figure 6 and Figure 7 , this application embodiment also provides an acoustic model adjustment method, which is applied to the processor of a vehicle, and the method includes the following stages:

[0097] I. Adaptive verification mode startup:

[0098] When the vehicle's processor starts the adaptive verification mode, run the karaoke software and control the speaker to play the sweep reference signal.

[0099] II. Ambient acoustic data collection:

[0100] Collect the in-vehicle ambient acoustic data through the integrated or external microphone in the vehicle.

[0101] III. Ambient acoustic data analysis:

[0102] The processor executes the method steps such as the foregoing steps S110 - S130 to achieve personalized optimization of the preset acoustic model for different environments and microphone characteristics to obtain the target acoustic model.

[0103] IV. Performance optimization:

[0104] Based on the model parameters corresponding to various acoustic features in the target acoustic model, adjust the performance parameters of the microphone and the speaker respectively.

[0105] V. Effect verification:

[0106] The processor executes the foregoing steps S140 - S190 to achieve subjective experience verification of the karaoke effect when dynamically switching between adjusting the parameters of the speaker and the microphone using the target acoustic model and not using the target acoustic model.

[0107] It should be understood that although the steps in the flowcharts involved in the above embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0108] Please refer to Figure 8, Another embodiment of the present application provides an acoustic model adjustment device 300. The acoustic model adjustment device 300 includes: a data acquisition module 310, configured to acquire environmental acoustic data collected by a microphone in a target environment when a speaker plays a swept reference signal; a feature extraction module 320, configured to obtain multiple groups of acoustic feature parameters according to the swept signal, the environmental acoustic data, and the performance parameters of the microphone when collecting the acoustic feature data. Each group of acoustic feature parameters includes parameters of multiple acoustic features, and the multiple acoustic feature parameters include at least one microphone feature and at least one environmental acoustic feature; a model adjustment module 330, configured to adjust the model parameters of a preset acoustic model based on the multiple groups of acoustic feature parameters and a preset scoring threshold to obtain a target acoustic model, where the preset acoustic model is a correspondence between the parameters of multiple acoustic features and the sound effect score established based on multiple groups of acoustic sample feature data. Each group of the acoustic sample feature data includes a group of sample acoustic feature parameters and the actual sound effect score corresponding to the group of sample acoustic feature parameters, and the preset scoring threshold is used to indicate the expected sound effect score of the karaoke audio output when karaoke is performed through the microphone and the speaker.

[0109] In an implementable manner, the acoustic model adjustment device 300 further includes: a model acquisition module, a scoring acquisition module, a loss acquisition module, and a parameter adjustment module. The model acquisition module is configured to acquire an initial machine learning model; the scoring acquisition module is configured to input multiple groups of acoustic sample feature data into the initial machine learning model respectively to obtain predicted sound effect scores; the loss acquisition module is configured to obtain a first model loss based on the first predicted sound effect score of each group of the acoustic sample feature data and the actual sound effect score label of each group of the acoustic sample feature data; the parameter adjustment module is configured to adjust the model parameters of the initial machine learning model based on the first model loss to minimize the first model loss until a first iteration end condition is reached, and then obtain the preset acoustic model.

[0110] In an implementable manner, the model adjustment module 330 is further configured to input each group of the acoustic feature parameters into the preset acoustic model to obtain a second predicted sound effect score corresponding to each group of the acoustic feature parameters; obtain a second model loss based on the preset scoring threshold and the second predicted sound effect score, and adjust the model parameters of the preset acoustic model based on the second model loss to minimize the second model loss until a second iteration end condition is reached, and then obtain the target acoustic model.

[0111] In an implementable manner, the preset acoustic model is a multiple linear regression model. The multiple linear regression model includes an independent variable representing the sound effect score and a dependent variable representing the acoustic feature parameters, and the model parameters of the preset acoustic model include the regression coefficients corresponding to each of the dependent variables.

[0112] In one implementable manner, the microphone feature includes a microphone gain; the environmental acoustic feature includes a distortion degree, and the environmental acoustic feature further includes at least one of an environmental frequency response, a signal-to-noise ratio, and a loudness.

[0113] In one implementable manner, the acoustic model adjustment device 300 further includes a device parameter adjustment module, configured to adjust the performance parameters of the microphone and the speaker respectively based on the model parameters corresponding to various acoustic features in the target acoustic model.

[0114] In one implementable manner, the acoustic model adjustment device 300 further includes a data merging module and an audio playback control module. The data acquisition module 310 is further configured to acquire first background music data and first user karaoke data collected by the microphone under the adjusted performance parameters; the data merging module is configured to merge the first user karaoke data with the first background music data to obtain a first karaoke audio; the audio playback control module is configured to control the speaker to play the first karaoke audio under the adjusted performance parameters.

[0115] In one implementable manner, the data acquisition module 310 is further configured to, in response to a switching instruction, acquire second user karaoke data collected by the microphone when the performance parameters are restored to the unadjusted state; the data merging module is further configured to merge the second user karaoke data with the second background music data to obtain a second karaoke audio; the audio playback control module is further configured to control the speaker to play the second karaoke audio when the performance parameters are restored to the unadjusted state.

[0116] Each module in the above device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules. It should be noted that the device embodiments in this application correspond to the foregoing method embodiments. The specific principles in the device embodiments can refer to the content in the foregoing method embodiments and will not be elaborated here.

[0117] Next, Figure 9 an electronic device provided by this application will be described.

[0118] Please refer to Figure 9 , based on the acoustic model adjustment method provided in the foregoing embodiments, another electronic device 100 provided in the embodiments of this application includes a processor 102 that can execute the foregoing method. The electronic device 100 can be a server, a terminal device, or a vehicle, and the terminal device can be a device such as a smart phone, a tablet computer, a computer, or a portable computer.

[0119] In one aspect of the present application, the electronic device is a vehicle.

[0120] The electronic device 100 further includes a memory 104. Among them, a program that can execute the content in the foregoing embodiments is stored in the memory 104, and the processor 102 can execute the program stored in the memory 104.

[0121] Among them, the processor 102 may include one or more cores for processing data and a message matrix unit. The processor 102 connects various parts within the entire electronic device 100 using various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 104, and by calling data stored in the memory 104, the processor 102 performs various functions of the electronic device 100 and processes data. Optionally, the processor 102 may be implemented in at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 102 may integrate a combination of one or several of a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing display content; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor 102 and may be implemented separately through a communication chip.

[0122] The memory 104 may include Random Access Memory (RAM) and may also include Read-Only Memory. The memory 104 can be used to store instructions, programs, codes, code sets, or instruction sets. The memory 104 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for implementing at least one function, instructions for implementing the following method embodiments, etc. The data storage area may also store data obtained during the use of the electronic device 100.

[0123] In an implementable manner, the electronic device 100 includes a microphone and a speaker (not shown in the figure), and the microphone and the speaker are respectively connected to the processor 102.

[0124] The electronic device 100 may further include a network module and a screen. The network module is used to receive and send electromagnetic waves, realize the mutual conversion between electromagnetic waves and electrical signals, so as to communicate with a communication network or other devices, such as communicating with an audio playback device. The network module may include various existing circuit elements for performing these functions, such as antennas, radio frequency transceivers, digital signal processors, encryption / decryption chips, subscriber identity module (SIM) cards, memories, and so on. The network module can communicate with various networks such as the Internet, enterprise intranets, wireless networks or communicate with other devices through wireless networks. The aforementioned wireless networks may include cellular phone networks, wireless local area networks or metropolitan area networks. The screen can display interface content and perform data interaction, such as displaying the aforementioned interface and triggering operations through the screen, etc.

[0125] In some embodiments, the electronic device 100 may further include: a peripheral interface 106 and at least one peripheral device. The processor 102, the memory 104 and the peripheral interface 106 may be connected by a bus or signal lines. Each peripheral device can be connected to the peripheral interface through a bus, signal lines or circuit board. Specifically, the peripheral devices include at least one of a radio frequency component 108, a positioning component 112, a camera 114, an audio component 116, a display screen 118, and a power supply 122, etc.

[0126] The peripheral interface 106 can be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 102 and the memory 104. In some embodiments, the processor 102, the memory 104 and the peripheral interface 106 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 102, the memory 104 and the peripheral interface 106 can be implemented on a separate chip or circuit board, and the embodiments of the present application do not limit this.

[0127] The radio frequency component 108 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency component 108 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency component 108 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency component 108 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency component 108 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, various generations of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency component 108 may further include a circuit related to NFC (Near Field Communication), which is not limited in this application.

[0128] The positioning component 112 is used to locate the current geographical location of the electronic device to implement navigation or LBS (Location-Based Service). The positioning component 112 can be a positioning component based on the US GPS (Global Positioning System), the Beidou system, or the Galileo system.

[0129] The camera 114 is used to capture images or videos. Optionally, the camera 114 includes a front camera and a rear camera. Generally, the front camera is disposed on the front panel of the electronic device 100, and the rear camera is disposed on the back of the electronic device 100. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, to implement the function of background blurring by fusing the main camera and the depth-of-field camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera 114 may further include a flash. The flash can be a single-color temperature flash or a two-color temperature flash. The two-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.

[0130] The audio component 116 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 102 for processing, or input to the radio frequency component 108 to implement voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the electronic device 100. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signals from the processor 102 or the radio frequency component 108 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio component 114 may further include a headphone jack.

[0131] The display screen 118 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 118 is a touch display screen, the display screen 118 also has the ability to collect touch signals on or above the surface of the display screen 118. The touch signals can be input to the processor 102 as control signals for processing. At this time, the display screen 118 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 118, which is arranged on the front panel of the electronic device 100; in other embodiments, there may be at least two display screens 118, which are respectively arranged on different surfaces of the electronic device 100 or in a foldable design; in still other embodiments, the display screen 118 may be a flexible display screen, which is arranged on a curved surface or a folding surface of the electronic device 100. Even, the display screen 118 can be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 118 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0132] The power supply 122 is used to supply power to each component in the electronic device 100. The power supply 122 may be alternating current, direct current, a disposable battery, or a rechargeable battery. When the power supply 122 includes a rechargeable battery, the rechargeable battery may be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0133] The embodiment of the present application also provides a structural block diagram of a computer-readable storage medium. Program code is stored in the computer-readable medium, and the program code can be called by a processor to execute the method described in the above method embodiment.

[0134] The computer-readable storage medium may be an electronic memory such as a flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, a hard disk, or a ROM. Optionally, the computer-readable storage medium includes a non-transitory computer-readable storage medium. The computer-readable storage medium has a storage space for program code for executing any method step in the above method. These program codes can be read out from or written into one or more computer program products. The program code can be compressed in an appropriate form, for example.

[0135] The embodiment of the present application also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method described in the above various optional implementation manners.

[0136] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An acoustic model adjustment method, characterized in that: The method comprises: Acquire environmental acoustic data collected by a microphone in a target environment when a frequency sweep reference signal is played by a speaker; Obtaining multiple groups of acoustic feature parameters according to the swept frequency signal, the environmental acoustic data, and the performance parameters of the microphone when collecting the acoustic feature data, each group of acoustic feature parameters includes parameters of multiple acoustic features, and the multiple acoustic features include at least one microphone feature and at least one environmental acoustic feature; The model parameters of the preset acoustic model are adjusted based on the preset score threshold and the multiple groups of acoustic feature parameters to obtain a target acoustic model, wherein the preset acoustic model is a correspondence between parameters of multiple acoustic features and sound effect scores established based on multiple groups of acoustic sample feature data, each group of the acoustic sample feature data includes a group of sample acoustic feature parameters and an actual sound effect score corresponding to the group of sample acoustic feature parameters, and the preset score threshold is used to indicate the expected sound effect score of the karaoke audio output when karaoke is performed through the microphone and the speaker.

2. The method according to claim 1, characterized in that Before adjusting the model parameters of the preset acoustic model based on the multiple groups of acoustic feature parameters and the preset scoring threshold to obtain the target acoustic model, the method further includes: Get an initial machine learning model; Inputting multiple groups of acoustic sample feature data into the initial machine learning model respectively to obtain a first predicted sound effect score for each group of acoustic sample feature data; Obtaining a first model loss based on a first predicted sound effect score of each group of acoustic sample feature data and an actual sound effect score label of each group of acoustic sample feature data; Based on the first model loss, the model parameters of the initial machine learning model are adjusted to minimize the first model loss until the first iteration end condition is reached to obtain a preset acoustic model.

3. The method according to claim 1, characterized in that The adjusting the model parameters of the preset acoustic model based on the multiple groups of acoustic feature parameters to obtain the target acoustic model includes: Inputting each group of the acoustic feature parameters into the preset acoustic model to obtain a second predicted sound effect score corresponding to each group of the acoustic feature parameters; A second model loss is obtained based on a preset score threshold and the second predicted sound effect score, and model parameters of the preset acoustic model are adjusted based on the second model loss to minimize the second model loss until the second iteration end condition is reached, thereby obtaining a target acoustic model.

4. The method according to any one of claims 1 to 3, characterized in that: The preset acoustic model is a multiple linear regression model, which includes independent variables representing sound effect scores and dependent variables representing acoustic feature parameters. The model parameters of the preset acoustic model include regression coefficients corresponding to each of the dependent variables.

5. The method according to any one of claims 1 to 3, characterized in that: The microphone characteristic includes microphone gain; the environmental acoustic characteristic includes distortion, and the environmental acoustic characteristic further includes at least one of environmental frequency response, signal-to-noise ratio and loudness.

6. The method according to claims 1-3, characterized in that: After adjusting the model parameters of the preset acoustic model based on the preset scoring threshold and the multiple groups of acoustic feature parameters to obtain the target acoustic model, the method further includes: The performance parameters of the microphone and the speaker are adjusted based on the model parameters corresponding to the various acoustic features in the target acoustic model.

7. The method according to any one of claim 6, characterized in that After adjusting the model parameters of the preset acoustic model based on the preset scoring threshold and the multiple groups of acoustic feature parameters to obtain the target acoustic model, the method further includes: Acquire first background music data and first user karaoke data collected by the microphone with performance parameters adjusted; Merging the first user karaoke data with the first background music data to obtain a first karaoke audio; The speaker is controlled to play the first karaoke audio with the performance parameters adjusted.

8. The method according to claim 7, characterized in that After controlling the speaker to play the first karaoke audio with the performance parameters adjusted, the method further includes: In response to the switching instruction, obtaining karaoke data of a second user collected by the microphone when the microphone is restored to the state where the performance parameters are not adjusted; Merging the second user karaoke data with the second background music data to obtain a second karaoke audio; The speaker is controlled to play the second karaoke audio when the performance parameters are restored to the unadjusted state.

9. An acoustic model adjustment device, characterized in that: The device comprises: A data acquisition module, used to acquire environmental acoustic data collected by a microphone in a target environment when a speaker plays a frequency sweep reference signal; a feature extraction module, configured to obtain a plurality of groups of acoustic feature parameters according to the frequency sweep signal, the environmental acoustic data, and the performance parameters of the microphone when collecting the acoustic feature data, each group of acoustic feature parameters including parameters of a plurality of acoustic features, wherein the plurality of acoustic feature parameters include at least one microphone feature and at least one environmental acoustic feature; A model adjustment module is used to adjust the model parameters of a preset acoustic model based on a preset score threshold and the multiple groups of acoustic feature parameters to obtain a target acoustic model, wherein the preset acoustic model is a correspondence between parameters of multiple acoustic features and sound effect scores established based on multiple groups of acoustic sample feature data, each group of the acoustic sample feature data includes a group of sample acoustic feature parameters and an actual sound effect score corresponding to the group of sample acoustic feature parameters, and the preset score threshold is used to indicate the expected sound effect score of the karaoke audio output when karaoke is performed through the microphone and the speaker.

10. An electronic device, characterized in that: include: one or more processors; Memory; One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the method according to any one of claims 1 to 8.

11. The electronic device according to claim 9, characterized in that: The electronic device is a vehicle, and the electronic device further includes a microphone and a speaker, wherein the microphone and the speaker are respectively connected to the processor.

12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, and the program codes can be called by a processor to execute the method according to any one of claims 1 to 8.