An audio sound effect generalization tuning method, medium, device, and program product.

By recognizing user comments in the tuning instructions and generating sound effect adjustment parameters using a sound effect mapping model, the problem of audio tuning relying on professionals in existing technologies is solved, enabling personalized audio sound effect adjustment and improving sound quality and flexibility.

CN119649779BActive Publication Date: 2026-04-07TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In existing technologies, audio tuning relies on professionals, making it difficult to meet users' needs for personalized adjustments based on music playback scenarios and personal preferences.

Method used

By recognizing user comments in the tuning instructions, sound effect adjustment parameters are generated using a sound effect mapping model, enabling personalized sound effect adjustment of audio.

Benefits of technology

It enables intelligent sound tuning based on user needs and preferences, improving sound quality and personalization to meet the sound needs of different scenarios and listeners.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119649779B_ABST
    Figure CN119649779B_ABST
Patent Text Reader

Abstract

This application provides a method, medium, device, and program product for generalized audio sound effect tuning, relating to the field of audio processing. The method includes: upon receiving a tuning instruction for a target audio, identifying the tuning instruction to obtain user feedback information on the target audio, and determining the sound effect parameter vector of the target audio; inputting the user feedback information into a sound effect mapping model to obtain sound effect adjustment parameters corresponding to the user feedback information; and applying the sound effect adjustment parameters to perform generalized tuning on the sound effect parameter vector. By applying this application, users can intelligently tune audio using text information according to their needs and preferences, improving sound quality and personalization, helping to meet the sound effect needs of different scenarios and listeners, and bringing greater flexibility and innovation to the fields of music production and audio processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing, and in particular to an audio sound effect generalization tuning method, medium, device and program product. Background Technology

[0002] With the widespread adoption of streaming music platforms, music has become an integral part of more and more people's lives. In audio processing, generalized tuning involves adjusting the overtone components of an audio signal. Overtones are a crucial component of timbre; by enhancing or weakening the frequency and amplitude of certain overtones, the characteristics and expressiveness of the timbre can be altered. However, current audio tuning relies on professional adjustments, and when music streaming platforms do not offer music that has undergone professional generalized tuning, it often fails to meet users' musical needs.

[0003] For example, depending on the music playback scenario and personal preferences, there may be many other subjective tuning options. For instance, in a car setting, some users might prefer to enhance the low frequencies of the music, while others might prefer music that emphasizes vocals. Summary of the Invention

[0004] The purpose of this application is to provide an audio sound effect generalization tuning method, a computer-readable storage medium, an electronic device, and a computer program product that can adjust sound effects according to user needs and achieve personalized audio sound effects.

[0005] To address the aforementioned technical problems, this application provides an audio sound effect generalization tuning method, the specific technical solution of which is as follows:

[0006] When a tuning instruction for a target audio is received, the tuning instruction is identified to obtain user comments on the target audio, and the sound effect parameter vector of the target audio is determined.

[0007] The user review information is input into the sound effect mapping model to obtain the sound effect adjustment parameters corresponding to the user review information;

[0008] The sound effect adjustment parameters are applied to the sound effect parameter vector for generalized tuning.

[0009] Optionally, obtaining user feedback information on the target audio by recognizing the tuning command includes:

[0010] The tuning command is analyzed to obtain the text features;

[0011] Identify the semantic information corresponding to the text features, and use the semantic information as user comment information for the target audio.

[0012] Optionally, obtaining user feedback information on the target audio by recognizing the tuning command includes:

[0013] If the tuning instruction includes quantization information, the target audio is identified based on the quantization information group to obtain an objective sound effect evaluation result; the objective sound effect evaluation result is used to guide objective sound effect tuning;

[0014] If the tuning instruction includes non-quantized information, the target audio is identified based on the non-quantized information group to obtain a subjective sound effect evaluation result; the subjective sound effect evaluation result is used to guide subjective sound effect tuning.

[0015] Optionally, before inputting the user review information into the sound effect mapping model, the method further includes:

[0016] Obtain an audio combination; the audio combination includes standard audio and processed audio; the processed audio is obtained by processing the standard audio through a sound effects processing link;

[0017] Obtain annotation information after annotating the audio combination; the annotation information includes non-quantized annotation information obtained based on the non-quantized information group annotation, and quantized annotation information obtained based on the quantized information group annotation;

[0018] The annotation information and the corresponding audio combination are correlated and analyzed to generate the sound effect mapping model, which contains the audio processing mapping relationship between the annotation information and the audio combination.

[0019] Optionally, performing correlation analysis on the annotation information and the corresponding audio combinations to generate the sound effect mapping model containing the audio processing mapping relationship between the annotation information and the audio combinations includes:

[0020] Extract the audio effect processing features corresponding to the audio effect processing link in the audio combination;

[0021] Extract the text features from the annotation information;

[0022] The text features are matched with the annotation information to determine the association rules between the sound effect processing features and the annotation information; the association rules are used to indicate the text features matched by the sound effect processing features.

[0023] A sound effect mapping model is obtained by using a custom data structure to maintain the audio combination, the annotation information, and the association rules.

[0024] Optionally, obtaining the quantization annotation information from the annotation information includes:

[0025] The audio combination is evaluated using a multi-dimensional sound effect scoring model with reference to a quantitative information group to obtain quantitative annotation information.

[0026] The multi-dimensional sound effect scoring model is generated as follows:

[0027] Obtain the evaluation results of the effects of several sound effect dimensions included in the audio combination;

[0028] The multi-dimensional sound effect scoring model is generated based on the evaluation results and the evaluation methods corresponding to each dimension of sound effect.

[0029] Optionally, inputting the user review information into the sound effect mapping model to obtain the sound effect adjustment parameters corresponding to the user review information includes:

[0030] The user review information is determined to include multiple dimensions of sound effect review information;

[0031] Determine the corresponding sound effect adjustment methods for each dimension of sound effect review information;

[0032] Convert each of the aforementioned sound effect adjustment methods into a corresponding sound effect adjustment vector;

[0033] By combining the sound effect adjustment vectors corresponding to the sound effects in each dimension, a sound effect adjustment parameter vector is obtained.

[0034] Optionally, if the tuning instruction includes audio playback environment information, after obtaining the tuning instruction, the method further includes:

[0035] Determine the scene sound effect information that matches the target audio playback environment information;

[0036] Accordingly, after inputting the user review information into the sound effect mapping model to obtain the sound effect adjustment parameters corresponding to the user review information, the method further includes:

[0037] The sound effect adjustment parameters are corrected based on the scene sound effect information to obtain sound effect correction parameters;

[0038] Correspondingly, applying the sound effect adjustment parameters to generalize and tune the sound effect parameter vector includes:

[0039] The sound effect correction parameters are applied to generalize and tune the sound effect parameter vector.

[0040] Optionally, after generalizing the sound effect parameter vector using the sound effect adjustment parameters, the method further includes:

[0041] Obtain the evaluation data of the target audio;

[0042] The evaluation data will be used as the results of public user reviews.

[0043] The results of the public user reviews are input into the sound effect mapping model to obtain the second sound effect adjustment parameters;

[0044] The sound effect parameter vector is generalized and tuned according to the second sound effect adjustment parameter.

[0045] This application also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the audio sound effect generalization tuning method described above.

[0046] This application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor, when calling the computer program in the memory, implements the steps of the audio sound effect generalization tuning method as described above.

[0047] This application also provides a computer program product, including a computer program that, when executed, implements the steps of the audio sound effect generalization tuning method described above.

[0048] This application provides a method for generalizing audio sound effects tuning, comprising: when receiving a tuning instruction for a target audio, identifying the tuning instruction to obtain user comment information on the target audio, and determining the sound effect parameter vector of the target audio; inputting the user comment information into a sound effect mapping model to obtain sound effect adjustment parameters corresponding to the user comment information; and applying the sound effect adjustment parameters to perform generalized tuning on the sound effect parameter vector.

[0049] After obtaining the audio tuning command, this application determines the user's comment information and the audio effect parameter vector corresponding to the tuning command, thereby converting the user's comment information into corresponding audio effect adjustment parameters. Users can intelligently tune the audio by issuing command information according to their own needs and preferences, improving the sound effect quality and personalization, which helps to meet the sound effect needs of different scenarios and listeners, and brings greater flexibility and innovation to the fields of music production and audio processing.

[0050] This application also provides a computer-readable storage medium, an electronic device, and a computer program product, which have the above-mentioned beneficial effects, and will not be repeated here. Attached Figure Description

[0051] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0052] Figure 1 A flowchart illustrating an audio sound effect generalization tuning method provided in an embodiment of this application;

[0053] Figure 2 This is a schematic diagram illustrating the audio pair generation process provided in an embodiment of this application;

[0054] Figure 3 This is a schematic diagram of the audio-to-text annotation process provided in the embodiments of this application;

[0055] Figure 4 A five-dimensional schematic diagram of the sound effect multi-dimensional scoring model provided in the embodiments of this application;

[0056] Figure 5 This is a schematic diagram of the model application structure provided in the embodiments of this application;

[0057] Figure 6 This is a structural diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0058] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0059] Generalized tuning is a sound processing technique designed to enable audio devices or systems to adapt to different acoustic environments and listening preferences. By broadening the target audience and usage scenarios of an audio system, making it not limited to specific sound settings or user configurations, generalized tuning enhances the overall sound quality experience and achieves the best listening effect.

[0060] See Figure 1 , Figure 1 A flowchart of an audio sound effect generalization tuning method provided in this application embodiment is included, the method comprising:

[0061] S101: When a tuning instruction for a target audio is received, the tuning instruction is identified to obtain user comments on the target audio, and the sound effect parameter vector of the target audio is determined.

[0062] S102: Input the user review information into the sound effect mapping model to obtain the sound effect adjustment parameters corresponding to the user review information;

[0063] S103: Apply the sound effect adjustment parameters to generalize and tune the sound effect parameter vector.

[0064] First, the tuning instruction is obtained. The method of obtaining this instruction is not limited here; it can include direct and indirect tuning instructions. Direct tuning instructions refer to instructions that contain sound effect tuning-related parameters, including but not limited to text-based tuning information input by the user, or tuning instructions that can be processed and converted into other formats containing text information, such as voice tuning instructions containing user comments. In a feasible practical application scenario, the user's voice instructions can be obtained through a sound pickup device such as a smart speaker, and then parsed to obtain text-based tuning instructions.

[0065] Indirect tuning commands are commands that do not contain audio effect tuning parameters, but can be converted into direct tuning commands through command recognition. For example, indirect tuning commands may contain audio playback environment information, mood information, or music preferences. After semantic recognition, the indirect tuning commands are converted into corresponding direct tuning commands. This conversion process can be predefined to determine the conversion relationship between different types of indirect and direct tuning commands, and different types of tuning commands can have different conversion relationships. For example, for audio playback environment information, there may be a conversion relationship between weather information and direct tuning commands, or a conversion relationship between playback environment and direct tuning commands; for mood information, a conversion relationship between mood information and direct tuning commands can be used.

[0066] It's easy to understand that if a text-based tuning instruction is used, corresponding text content recognition is needed to achieve generalized tuning effects. The tuning instruction can be parsed to obtain text features, and then the semantic information corresponding to these features can be used as user commentary on the target audio.

[0067] There are no restrictions on how to perform text content recognition. Pre-trained text recognition models such as BERT (Bidirectional Encoder Representations from Transformers), GPT (Generative Pre-trained Transformer), and ELMo (Embeddings from Language Models) can be used to understand the text content of tone instructions in order to determine the user comment information contained therein.

[0068] In step S102, in addition to determining the user comment information contained in the tuning instruction, it is also necessary to determine the current sound effect parameter vector of the corresponding target audio. It is easy to understand that the tuning instruction has a corresponding target audio. This target audio can be included within the tuning instruction or determined based on the instruction acquisition time. In one feasible implementation, the target audio playing at the same time can be determined based on the acquisition time of the tuning instruction. For example, when music is playing on a smart speaker or other audio device, if a tuning instruction is acquired, the target audio can be determined to be the currently playing music. The determined target audio should at least contain unique identification features, such as the name of the target audio, artist information, album, etc. There are no limitations on how the sound effect parameter vector is obtained; after determining the target audio, it can be obtained from a local audio database or from an online source.

[0069] The target audio's sound effect parameter vector refers to the current sound effect parameter vector of the target audio. Using vector-based sound effect parameters makes it easier to directly generalize and tune the target audio, improving tuning efficiency. User reviews include subjective evaluations of the target audio, such as comments on parameters like current audio quality, loudness, and timbre. The sound effect parameter vector represents the objective parameters of the audio as described by the user, primarily including timbre, number of channels, sound quality, loudness, and dynamics. Dynamics refers to the changes in loudness during the playback of the target audio, such as crescendo or diminuendo. It should be noted that the target audio in this application is not limited to a specific song; it can be any audio segment or a long audio file containing several audio segments—in short, any length and type of audio intended for the user.

[0070] When inputting user feedback information into the sound effect mapping model to obtain the corresponding sound effect adjustment parameters, the following process can be used:

[0071] The first step is to determine the multiple dimensions of sound effect review information included in the user review information;

[0072] The second step is to determine the corresponding sound effect adjustment methods for each dimension of sound effect evaluation information.

[0073] The third step is to convert each of the aforementioned sound effect adjustment methods into corresponding sound effect adjustment vectors.

[0074] The fourth step is to combine the sound effect adjustment vectors corresponding to the sound effects of each dimension to obtain the sound effect adjustment parameter vector.

[0075] First, determine the sound effect evaluation information for each dimension, and then determine the corresponding sound effect adjustment method. For example, if the sound effect evaluation information for the reverb dimension is too low, the corresponding sound effect adjustment method is to increase the reverb.

[0076] Audio effect adjustment parameters can be represented as a multi-dimensional vector, where each component of the vector represents a specific audio effect parameter. For example, suppose we have the following audio effect parameters: volume, low-frequency gain, mid-frequency gain, high-frequency gain, reverb depth, and delay time. Then the audio effect adjustment parameter vector can be represented as:

[0077] [Volume, Low-frequency gain, Mid-frequency gain, High-frequency gain, Reverb depth, Delay time];

[0078] In practical applications, the specific values ​​of each component will vary depending on the sound effect adjustment method. For example, if you need to increase the volume and adjust the equalizer settings to enhance the low-frequency effect, while adding some reverb and delay, the sound effect adjustment parameter vector can be:

[0079] [0.75, 0.1, 0, -0.2, 0.5, 0.3];

[0080] At this point, 0.75 represents a volume increase to 75%, 0.1 represents a 10% increase in low-frequency gain, no change in mid-frequency gain (0), a 20% decrease in high-frequency gain (-0.2), a reverb depth of 50% (0.5), and a delay time of 0.3 seconds. It should be noted that the units of the sound effect adjustment vectors corresponding to different sound effects may vary.

[0081] It should be noted that the dimensions and components of the sound effect adjustment parameter vector can vary depending on the specific sound effect processing system or software. Furthermore, the numerical range within the vector may also differ depending on the system, and can be set by those skilled in the art based on the vector corresponding to the sound effect adjustment vector.

[0082] After obtaining the audio tuning instruction in this embodiment, the user's comment information and the audio effect parameter vector corresponding to the tuning instruction are determined. The user's comment information is then converted into the corresponding audio effect parameter vector. Users can use text information to intelligently tune the audio according to their needs and preferences, thereby improving the sound quality and personalization. This helps to meet the sound effect needs of different scenarios and listeners, bringing greater flexibility and innovation to the fields of music production and audio processing.

[0083] Based on the above embodiments, in one feasible implementation, the user review information may include at least one of quantized information and non-quantized information. In this case, when identifying a tuning instruction and obtaining user review information for the target audio, if the tuning instruction includes quantized information, the target audio can be identified based on the quantized information group to obtain an objective sound effect review result. This objective sound effect review result is used to guide objective sound effect tuning.

[0084] Quantitative information refers to the subjective description of sound effect quantification information, which is strongly correlated with sound effects and includes, but is not limited to, loudness, equalization, compression, delay, and reverberation. Quantitative information can be labeled based on a multi-dimensional sound effect scoring model through questionnaires or Q&A. In this case, the objective sound effect evaluation result is represented as the instruction in the tuning command that is directly related to the sound effect adjustment. Adjusting the quantitative information can directly affect the sound effect of the audio. For quantitative information, a questionnaire can be used to obtain users' subjective evaluations. For example, users can be asked: "What is the reverberation like in this song?" If the answer received is: "The reverberation of this song is too low and can be adjusted much higher," then the objective sound effect evaluation result can be that the reverberation of the song is low. When adjusting the sound effect based on the objective sound effect evaluation result, the reverberation of the song can be enhanced.

[0085] If the tuning command includes non-quantized information, the target audio can be identified based on the non-quantized information group to obtain a subjective sound effect evaluation result.

[0086] Non-quantitative information is weakly correlated with sound effects and mainly consists of music style information, such as the emotional type of the music, the genre of the music, and whether the music contains vocals. Accordingly, when identifying the target audio based on the non-quantitative information group to obtain a subjective sound effect evaluation result, this subjective sound effect evaluation result can include comments on music style information. When using the subjective sound effect evaluation to guide subjective sound effect tuning, adjustments can be made according to the music style.

[0087] The quantitative and non-quantitative information groups mentioned above serve as a reference when determining user review information, that is, to determine whether the user review information contains information of the corresponding sound effect type in the group.

[0088] Before determining quantized and non-quantized information, the types of quantized and non-quantized information to be determined can be identified through annotation. Specifically, audio combinations can be obtained first, and the standard audio and processed audio in the audio combinations can be annotated in pairs to obtain annotation information. This annotation information includes non-quantized information groups and quantized information groups for the audio combinations. See also... Figure 2 , Figure 2This diagram illustrates the audio pair generation process provided in this application embodiment. The audio pair consists of standard audio (SQ, StereoQuality) and processed audio (PA, Professional Audio). Standard audio typically refers to the original audio file without processing or compression, providing the purest and closest sound quality experience to the original. Processed audio refers to the audio file obtained after undergoing a series of audio effect processing steps. Specifically, processed audio can be obtained from standard audio through an audio effect processing chain, which consists of several audio effect processing steps, such as loudness, equalization, compression, delay, and reverb. Audio effect parameter vectors can be applied in the audio effect processing chain; by adjusting the audio effect parameter vectors corresponding to each part of the audio effect processing steps, the processed audio can be obtained.

[0089] See afterward. Figure 3 , Figure 3 This is a schematic diagram of the audio-to-text annotation process provided in this application embodiment. Text annotation is performed on audio combinations formed by pairing standard audio and processed audio, resulting in non-quantized audio information and quantized sound effect information. The annotation information after annotating the audio combinations is obtained. Specifically, the annotator can obtain the annotation information of the processed audio by referring to the standard audio. This annotation information consists of two parts: non-quantized information and quantized information of the audio combinations.

[0090] Finally, a correlation analysis is performed on the annotation information and the corresponding audio combinations to generate a sound effect mapping model that includes the audio processing mapping relationship between the annotation information and the audio combinations. Specifically, this may include the following steps:

[0091] Step 1: Extract the audio effect processing features corresponding to the audio effect processing link in the audio combination;

[0092] The second step is to extract the text features from the annotation information.

[0093] The third step is to match the text features with the annotation information to determine the association rules between the sound effect processing features and the annotation information; the association rules are used to indicate the text features matched by the sound effect processing features.

[0094] Fourth step: Use a custom data structure to maintain the audio combination, the annotation information and the association rules to obtain the sound effect mapping model.

[0095] During association analysis, the standard audio and its corresponding processed audio in the audio combination are compared. The differences in textual features of the annotation information are used to reflect the differences in sound effects between the processed audio obtained after the standard audio is processed through the sound effect processing link, that is, the differences between the sound effect processing features corresponding to the sound effect processing link. By analyzing and summarizing several audio combinations, the audio processing mapping relationship between the annotation information and the audio combination is determined, that is, the association rule between the sound effect processing link contained in the audio combination and the corresponding annotation information, thus obtaining the sound effect mapping model.

[0096] It is easy to understand that the text features and corresponding semantic information obtained by the pre-trained text recognition model can include both objective and subjective sound effect evaluation results.

[0097] Based on the above embodiments, in one feasible implementation, when obtaining sound effect quantification information, a multi-dimensional sound effect scoring model can be applied to evaluate the sound effect of the audio combination by referring to the quantification information group, thereby obtaining quantification annotation information. By using the multi-dimensional sound effect scoring model, the sound effect processing link used by the audio combination can be directly mapped to the corresponding sound effect parameters, thereby simplifying the sound effect generalization tuning process and improving the efficiency of sound effect generalization tuning.

[0098] The specific method for generating the multi-dimensional sound effect scoring model is not limited here. In one feasible implementation, the evaluation results of the effects of several sound effect dimensions contained in the audio combination can be obtained. Based on the evaluation results and the evaluation methods corresponding to each dimension of sound effect, a multi-dimensional sound effect scoring model can be generated. Subsequently, the quantitative information group of the multi-dimensional sound effect scoring model can be directly collected and applied to the audio combination to label the quantitative information of the sound effect effect. The evaluation method used to obtain the evaluation results is not limited here; evaluation methods such as user subjective ratings and user questionnaires can be used.

[0099] In the multi-dimensional sound effect scoring model, each dimension corresponds to some important parameters in the sound effect. For example... Figure 4 As shown, Figure 4 This diagram illustrates the five dimensions of the multi-dimensional sound effect scoring model provided in this application, listing some of the sound effect dimensions, including sound quality, equalization, reverberation, dynamics, and loudness. For each dimension of any audio combination, the multi-dimensional sound effect scoring model is applied to ask questions, collecting user feedback information through a questionnaire.

[0100] For quantitative information in user reviews, quantitative information groups can be used to compare the quantitative information and directly determine the corresponding sound effect adjustment parameters.

[0101] For non-quantifiable information in user reviews, such as adjusting the emotional tone of audio, it typically involves adjusting at least one sound effect. Taking adjusting the emotional tone of audio as an example, changing an audio track from upbeat to melancholic can be achieved by comprehensively adjusting multiple sound effect parameters, including rhythm, harmony, melody, timbre, dynamics, and reverb. In terms of rhythm, slowing down the tempo can lengthen the note values, making the music sound slower and heavier, thus creating a sense of sadness. In terms of harmony, using minor keys instead of major keys, introducing more minor chords, or lowering the tonality of the music usually makes the music sound more melancholic and sad. In terms of melody, simplifying the melodic line, reducing leaping notes, and using more descending and continuous bass melodies can make the music sound heavier and sadder. In terms of timbre, choosing warm and soft timbres, such as strings and piano, can reduce brightness and cheerfulness, increasing the melancholic atmosphere. In terms of loudness, reducing the dynamic range of the music and lowering the overall volume can make the music sound more subtle and restrained, thus creating a sad mood. Reverb adjustment, by increasing the reverb effect, can simulate the effect of playing in a larger space (such as a church or cave), making the music sound more spacious, melodious, and melancholic.

[0102] Therefore, non-quantitative information such as user reviews can also be transformed into sound effect reviews that include several dimensions.

[0103] It can be seen that the embodiments of this application can not only process quantized text information that is directly related to sound effect parameters, but also process non-quantized information that is weakly related to sound effect parameters, thereby expanding the application scope of sound effect tuning and meeting more types of sound effect tuning needs.

[0104] Based on the above embodiments, as a preferred embodiment, after obtaining the tuning instruction, if the tuning instruction is an indirect tuning instruction, such as audio playback environment information, the scene sound effect information adapted to the audio playback environment information can be determined, and the sound effect adjustment parameters can be corrected according to the scene sound effect information to obtain sound effect correction parameters, so as to apply the sound effect correction parameters to generalize the sound effect parameter vector for tuning.

[0105] The audio playback environment information can include spatial information, weather information, etc. For example, if a user prefers to listen to upbeat music on a cloudy day, based on the sound effect adjustment parameters obtained in the above embodiment, these parameters can be further modified according to the scene sound effect information to obtain sound effect correction parameters, thereby performing generalized tuning. Furthermore, if the user's comment information for the tuning command is invalid—that is, it has no relevance to the sound effect and only contains audio playback environment information—then the sound effect adjustment parameters corresponding to the scene sound effect information can be directly determined, thereby performing generalized tuning. Of course, users can preset sound effect adjustment parameters for different environments, so that after detecting keywords in the audio playback environment information, the appropriate sound effect adjustment parameters can be directly applied for generalized tuning. For example, different sound effect adjustment parameters can be adapted for open environments and bathrooms.

[0106] Based on the above embodiments, as a preferred embodiment, after applying the sound effect adjustment parameters to generalize and tune the sound effect parameter vector, evaluation data of the target audio can also be obtained; the evaluation data is used as the result of public user comments, and the result of public user comments is input into the sound effect mapping model to obtain the second sound effect adjustment parameters, and the sound effect parameter vector is generalized and tuned according to the second sound effect adjustment parameters.

[0107] If the target audio is a song, in addition to obtaining the song's sound effect parameter vector, related comments and other information can also be extracted. This embodiment limits the evaluation data to be obtained after generalizing the sound effect parameter vector using sound effect adjustment parameters. In other embodiments of this application, the evaluation data can be obtained synchronously with the sound effect parameter vector. Furthermore, after generalizing the sound effect parameter vector using sound effect adjustment parameters, the sound effect parameter vector is further generalized and adjusted based on a second sound effect adjustment parameter corresponding to the results of general user reviews.

[0108] If the song's comments contain other users' suggestions for adjusting the sound effects, further tuning can be done based on those suggestions to meet more diverse tuning needs and achieve a better listening experience.

[0109] In the practical application of this application, see Figure 5 , Figure 5 This is a schematic diagram of the model application structure provided in this application embodiment. For the acquired text, it is used as a tuning instruction and input into a pre-generated sound effect parameter model. The sound effect parameter generation model consists of two parts: a text understanding module that processes the text to obtain text features, and a sound effect mapping module that maps the text features obtained by the text understanding module into sound effect parameters, ultimately obtaining a sound effect adjustment parameter vector. The text understanding module can use a pre-trained text understanding model; the final model can obtain a corresponding sound effect adjustment parameter vector from any input text (Chinese).

[0110] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed, can implement the steps of the sound effect generalization tuning method provided in the above embodiments. The storage medium may include various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0111] This application also provides a computer program product, including a computer program that, when executed, implements the steps of the audio sound effect generalization tuning method as described in the above embodiments.

[0112] This application also provides an electronic device, see [link to document]. Figure 6 The present application provides a structural diagram of an electronic device, such as... Figure 6 As shown, it may include a processor 610 and a memory 620.

[0113] The processor 610 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor 610 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 610 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 610 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor 610 may also include an AI (Artificial Intelligence) processor, which handles computational operations related to machine learning.

[0114] The memory 620 may include one or more computer-readable storage media, which may be non-transitory. The memory 620 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In this embodiment, the memory 620 is used to store at least the following computer program 621, which, after being loaded and executed by the processor 610, is capable of implementing the relevant steps in the methods executed by the electronic device side as disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 620 may also include an operating system 622 and data 623, etc., and the storage method may be temporary storage or permanent storage. The operating system 622 may include Windows, Linux, Android, etc.

[0115] In some embodiments, the electronic device may further include a display screen 630, an input / output interface 640, a communication interface 650, a sensor 660, a power supply 670, and a communication bus 680.

[0116] certainly, Figure 6 The structure of the electronic device shown does not constitute a limitation on the electronic device in the embodiments of this application. In practical applications, the electronic device may include more than [other components]. Figure 6 More or fewer components as shown, or combinations of certain components.

[0117] The various embodiments in the specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0118] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. It should be noted that those skilled in the art can make several improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of this application.

[0119] It should also be noted that, in this specification, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

Claims

1. An audio sound effect generalization tuning method, characterized in that, include: When a tuning instruction for a target audio is received, the tuning instruction is identified to obtain user comments on the target audio, and the sound effect parameter vector of the target audio is determined. The user review information is input into the sound effect mapping model to obtain the sound effect adjustment parameters corresponding to the user review information; The sound effect adjustment parameters are applied to the sound effect parameter vector for generalized tuning; The user review information includes at least one of objective sound effect review results and subjective sound effect review results; the process of identifying the tuning command to obtain user review information for the target audio includes: If the tuning instruction includes quantization information, the target audio is identified based on the quantization information group to obtain an objective sound effect evaluation result; the objective sound effect evaluation result is used to guide objective sound effect tuning; the quantization information is a subjective description of sound effect quantization information; If the tuning instruction includes non-quantized information, the target audio is identified based on the non-quantized information group to obtain a subjective sound effect evaluation result; the subjective sound effect evaluation result is used to guide subjective sound effect tuning; the non-quantized information is music style information; Before inputting the user review information into the sound effect mapping model, the process also includes: Obtain an audio combination; the audio combination includes standard audio and processed audio; the processed audio is obtained by processing the standard audio through a sound effects processing link; Obtain annotation information after annotating the audio combination; the annotation information includes non-quantized annotation information obtained based on the non-quantized information group annotation, and quantized annotation information obtained based on the quantized information group annotation; The annotation information and the corresponding audio combination are correlated and analyzed to generate the sound effect mapping model, which contains the audio processing mapping relationship between the annotation information and the audio combination.

2. The sound effect generalization tuning method according to claim 1, characterized in that, Recognizing the tuning command to obtain user feedback information on the target audio includes: The tuning command is analyzed to obtain the text features; Identify the semantic information corresponding to the text features, and use the semantic information as user comment information for the target audio.

3. The sound effect generalization tuning method according to claim 1, characterized in that, The association analysis of the annotation information and the corresponding audio combinations is performed to generate the sound effect mapping model, which includes the audio processing mapping relationship between the annotation information and the audio combinations. Extract the audio effect processing features corresponding to the audio effect processing link in the audio combination; Extract the text features from the annotation information; The text features are matched with the annotation information to determine the association rules between the sound effect processing features and the annotation information; the association rules are used to indicate the text features matched by the sound effect processing features. A sound effect mapping model is obtained by using a custom data structure to maintain the audio combination, the annotation information, and the association rules.

4. The sound effect generalization tuning method according to claim 1, characterized in that, Obtaining the quantization annotation information from the annotation information includes: The audio combination is evaluated using a multi-dimensional sound effect scoring model with reference to a quantitative information group to obtain quantitative annotation information. The multi-dimensional sound effect scoring model is generated as follows: Obtain the evaluation results of the effects of several sound effect dimensions included in the audio combination; The multi-dimensional sound effect scoring model is generated based on the evaluation results and the evaluation methods corresponding to each dimension of sound effect.

5. The sound effect generalization tuning method according to claim 1, characterized in that, Inputting the user review information into the sound effect mapping model yields the sound effect adjustment parameters corresponding to the user review information, including: The user review information is determined to include multiple dimensions of sound effect review information; Determine the corresponding sound effect adjustment methods for each dimension of sound effect review information; Convert each of the aforementioned sound effect adjustment methods into a corresponding sound effect adjustment vector; By combining the sound effect adjustment vectors corresponding to the sound effects in each dimension, a sound effect adjustment parameter vector is obtained.

6. The sound effect generalization tuning method according to any one of claims 1-5, characterized in that, If the tuning instruction includes audio playback environment information, after obtaining the tuning instruction, it also includes: Determine the scene sound effect information that matches the target audio playback environment information; After inputting the user review information into the sound effect mapping model to obtain the sound effect adjustment parameters corresponding to the user review information, the method further includes: The sound effect adjustment parameters are corrected based on the scene sound effect information to obtain sound effect correction parameters; The generalization tuning of the sound effect parameter vector by applying the sound effect adjustment parameters includes: The sound effect correction parameters are applied to generalize and tune the sound effect parameter vector.

7. The sound effect generalization tuning method according to any one of claims 1-5, characterized in that, After applying the sound effect adjustment parameters to generalize and tune the sound effect parameter vector, the method further includes: Obtain the evaluation data of the target audio; The evaluation data will be used as the results of public user reviews. The results of the public user reviews are input into the sound effect mapping model to obtain the second sound effect adjustment parameters; The sound effect parameter vector is generalized and tuned according to the second sound effect adjustment parameter.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the audio sound effect generalization tuning method as described in any one of claims 1 to 7.

9. An electronic device, characterized in that, The device includes a memory and a processor, wherein the memory stores a computer program, and the processor, when calling the computer program in the memory, implements the steps of the audio sound effect generalization tuning method as described in any one of claims 1 to 7.

10. A computer program product, characterized in that, Includes a computer program, which, when executed, implements the steps of the audio sound effect generalization tuning method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Sound effect adjusting method based on voice control, medium, device and computing equipment

    CN109147739A

  • System and method of acoustically controlling equalizer in natural language and computer readable storage medium

    US20190385603A1