Audio adjustment method, computer device, readable storage medium and program product
By obtaining sound effect adjustment instructions and identifying the sound effect adjustment intention, and using the sound effect effector to adjust the audio signal, the low accuracy problem caused by user differences in traditional audio adjustment methods is solved, and personalized audio adjustment is achieved.
Patent Information
- Application Number
- CN202510404315.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-29
AI Technical Summary
Traditional audio adjustment methods are limited by the differences in user's audio knowledge and aesthetic standards, which makes audio adjustment difficult and low accuracy.
By obtaining sound effect adjustment instructions, identifying the sound effect adjustment intention, including sound effect categories and parameters, and using the sound effector to adjust the audio signal to achieve personalized and accurate adjustments.
No need for users to have professional knowledge, accurately analyze the sound effect adjustment categories and parameters based on user instructions, realize personalized and accurate audio adjustments, and improve the accuracy of audio adjustments.
Smart Images

Figure CN120388549A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of audio processing, and particularly to an audio adjustment method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Art
[0002] In the field of music, audio adjustment refers to a series of operations for modifying, enhancing, or creating audio signals using various techniques and tools. Through audio adjustment, the quality, expressiveness, and artistic effect of music works can be improved. Making corresponding adjustments to the audio based on user needs can significantly enhance the sound quality and listening experience, strengthen the musical expressiveness, optimize the experience in different scenarios, meet personalized needs, and improve user satisfaction.
[0003] When traditional techniques are used for audio adjustment, they are limited by requirements such as music literacy in audio knowledge. Moreover, differences in user aesthetic standards / user preferences lead to differences in audio adjustment requirements, increasing the difficulty of audio adjustment and being unfavorable for improving the accuracy of audio adjustment. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide an audio adjustment method, apparatus, computer device, computer-readable storage medium, and computer program product that can improve the accuracy of audio adjustment.
[0005] In a first aspect, an embodiment of the present application provides an audio adjustment method, including:
[0006] Obtaining a sound effect adjustment instruction input for the audio to be adjusted;
[0007] Performing intent recognition on the sound effect adjustment instruction to determine a sound effect adjustment intent for the audio to be adjusted; the sound effect adjustment intent includes a sound effect category to be adjusted and sound effect adjustment parameters for the sound effect category to be adjusted;
[0008] Adjusting the audio signal of the audio to be adjusted according to the sound effect category to be adjusted and the sound effect adjustment parameters associated with the sound effect category to be adjusted.
[0009] In a second aspect, an embodiment of the present application further provides an audio adjustment apparatus, including:
[0010] An instruction obtaining module, configured to obtain a sound effect adjustment instruction input for the audio to be adjusted;
[0011] An instruction recognition module, configured to perform intent recognition on the sound effect adjustment instruction to determine a sound effect adjustment intent for the audio to be adjusted; the sound effect adjustment intent includes a sound effect category to be adjusted and sound effect adjustment parameters for the sound effect category to be adjusted;
[0012] The sound effect adjustment module is used to adjust the audio signal of the audio to be adjusted according to the sound effect category to be adjusted and the sound effect adjustment parameters associated with the sound effect category to be adjusted.
[0013] In a third aspect, an embodiment of the present application further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the steps of the above audio adjustment method are implemented.
[0014] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the steps of the above audio adjustment method are implemented.
[0015] In a fifth aspect, the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by the processor, the steps of the above audio adjustment method are implemented.
[0016] The above audio adjustment method, device, computer device, computer-readable storage medium, and computer program product identify the intent of the sound effect adjustment instruction for the audio to be adjusted, analyze the sound effect adjustment category and sound effect adjustment parameters, and adjust the audio signal of the audio in combination with the sound effect adjustment category and sound effect adjustment parameters, so as to meet the user's audio adjustment requirements for the audio to be adjusted. It does not require the user to have relevant professional knowledge, can accurately analyze the sound effect adjustment category and adjustment parameters based on the user's instruction, realize personalized and accurate adjustment of the sound effect of the audio, and thus improve the accuracy of audio adjustment. Description of the Drawings
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required to be used in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 It is an application environment diagram of an audio adjustment method in an embodiment;
[0019] Figure 2 It is a flowchart of an audio adjustment method in an embodiment;
[0020] Figure 3 It is a schematic diagram of a vector clustering space in an embodiment;
[0021] Figure 4Schematic diagram of an effect classifier in an embodiment;
[0022] Figure 5 Schematic diagram of a parameter vectorization in an embodiment;
[0023] Figure 6 Structural block diagram of an audio adjustment device in an embodiment;
[0024] Figure 7 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0025] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0026] It can be understood that the terms "first", "second", etc. used in the present application may be used herein to describe various technical features, but these technical features are not limited by these terms. These terms are only used to distinguish the first technical feature from another technical feature.
[0027] The audio adjustment method provided by the embodiments of the present application can be applied to an application environment as shown in Figure 1 wherein, the terminal 102 communicates with the server 104 through a network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or placed in the cloud or other network servers. The terminal 102 obtains a sound effect adjustment instruction for the audio to be adjusted; the terminal 102 performs intent recognition on the sound effect adjustment instruction to determine the sound effect adjustment intent for the audio to be adjusted; the sound effect adjustment intent includes the sound effect category to be adjusted and the sound effect adjustment parameters for the sound effect category to be adjusted; the terminal 102 adjusts the audio signal of the audio to be adjusted according to the sound effect category to be adjusted and the sound effect adjustment parameters associated with the sound effect category to be adjusted. Among them, the terminal 102 can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart vehicle-mounted devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted device can be a virtual reality (VR) device, an augmented reality (AR) device, smart glasses, etc. The server 104 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0028] In an exemplary embodiment, as Figure 2 shown, an audio adjustment method is provided. Taking the application of this method to a terminal as an example, it includes the following steps S202 to S206. Among them:
[0029] Step S202: Obtain a sound effect adjustment instruction for the audio to be adjusted.
[0030] Among them, the audio to be adjusted can refer to the audio that needs to have its sound effects adjusted. In practical applications, the audio to be adjusted can include, but is not limited to, the audio sung by the user, the audio in the preset audio database, etc.
[0031] Among them, the sound effect adjustment instruction can refer to an instruction / command input by the user for adjusting the audio signal of the audio. In practical applications, the manifestation form of the sound effect adjustment instruction can include, but is not limited to, the voice form, the text form, etc.
[0032] As an example, a sound effect processor can be used to adjust the sound effects of the audio. However, a sound effect processor often requires the user to have relevant professional knowledge (such as audio knowledge and the usage method of the sound effect processor, etc.). There is a problem that the sound effect adjustment is inaccurate when the user directly uses the sound effect processor to adjust the sound effects of the audio. In order to accurately adjust the sound effects of the audio based on the user's audio adjustment requirements, the terminal can first obtain the sound effect adjustment instruction input by the user for the audio to be adjusted, and through parsing and analyzing the sound effect adjustment instruction, adjust the audio signal of the audio to be adjusted through the sound effect processor until the user's audio adjustment requirements are met.
[0033] Step S204: Perform intent recognition on the sound effect adjustment instruction to determine the sound effect adjustment intent for the audio to be adjusted.
[0034] Among them, the sound effect adjustment intent can include the sound effect category to be adjusted and the sound effect adjustment parameters for the sound effect category to be adjusted.
[0035] Among them, the sound effect category to be adjusted is the sound effect category obtained after performing intent recognition on the sound effect adjustment instruction. Essentially, it can be understood as the sound effect attribute of the audio signal, representing the sound effect attribute that the user expects to adjust for the audio signal to be adjusted. The sound effect category to be adjusted can correspond to one or more sound effect processors. Specifically, it can be understood that the sound effect category to be adjusted is realized by the processing of the audio signal attributes by one or more different sound effect processors. In practical applications, a sound effect processor can refer to a device or software used to process and improve the audio signal. The sound effect processor can enhance, modify, or create specific attributes of the audio signal. Sound effect processors are widely used in fields such as music production, recording, and live performances.
[0036] In a specific implementation, the sound effect categories to be adjusted may include but are not limited to timbre, surround, and energy, etc. Among them, when the sound effect category is timbre, the equalizer sound effect processor can perform equalization processing on the audio to be adjusted, and the pitch shifter sound effect processor can perform pitch shifting processing on the audio to be adjusted; when the sound effect category is surround, the reverb sound effect processor can perform reverb processing on the audio to be adjusted, and the delay sound effect processor can perform delay processing on the audio to be adjusted; when the sound effect category is energy, the compressor sound effect processor can perform compression processing on the audio to be adjusted, and the automatic gain sound effect processor can perform automatic gain processing on the audio to be adjusted.
[0037] Among them, the sound effect adjustment parameter can refer to the way (direction) and amplitude (degree) of sound effect adjustment when adjusting the sound effect category to be adjusted. For example, the sound effect adjustment parameter being A can indicate that the sound effect category to be adjusted (such as timbre) of the audio to be adjusted is enhanced by X levels, and the sound effect adjustment parameter being B can indicate that the sound effect category to be adjusted (such as timbre) of the audio to be adjusted is weakened by Y levels.
[0038] As an example, after the terminal obtains the sound effect adjustment instruction for the audio to be adjusted, the terminal can perform intent recognition on the sound effect adjustment instruction and analyze the sound effect adjustment intent for the audio to be adjusted from the sound effect adjustment instruction. In practical applications, the terminal can convert the sound effect adjustment instruction into text, which can be used as a (rough or vague) description of the user's sound effect adjustment intent by the terminal, and the terminal can analyze this text to determine the sound effect category to be adjusted and the sound effect adjustment parameter for the sound effect category to be adjusted.
[0039] Step S206: Adjust the audio signal of the audio to be adjusted according to the sound effect category to be adjusted and the sound effect adjustment parameter associated with the sound effect category to be adjusted.
[0040] As an example, the terminal can generate a control instruction for the sound effect processor corresponding to the sound effect category to be adjusted according to the sound effect category to be adjusted and the sound effect adjustment parameter associated with the sound effect category to be adjusted, and send this control instruction to the sound effect processor corresponding to the sound effect category to be adjusted. The sound effect processor corresponding to the sound effect category to be adjusted adjusts the audio signal of the audio to be adjusted based on the control instruction to obtain the adjusted audio, and the adjusted audio can meet the user's sound effect adjustment requirements / audio adjustment requirements.
[0041] In the above audio adjustment method, by performing intent recognition on the sound effect adjustment instruction for the audio to be adjusted, analyzing the sound effect adjustment category and the sound effect adjustment parameters, and combining the sound effect adjustment category and the sound effect adjustment parameters to adjust the audio signal of the audio, the audio adjustment requirements of the user for the audio to be adjusted can be met. There is no need to require the user to have relevant professional knowledge, and it can accurately analyze the adjustment category and adjustment parameters of the sound effect based on the user instruction, realize accurate and personalized adjustment of the sound effect of the audio, and further improve the accuracy of audio adjustment.
[0042] In an exemplary embodiment, the performing intent recognition on the sound effect adjustment instruction to determine the sound effect adjustment intent for the audio to be adjusted includes: recognizing the text related to the sound effect category in the sound effect adjustment instruction as the object text, and recognizing the text related to the adjustment parameters of the sound effect category in the sound effect adjustment instruction as the processing parameter text; in the case of recognizing the object text, obtaining, from a plurality of pre-set sound effect text tags, the sound effect text tag that satisfies a preset similarity degree with the object text as the target sound effect text tag, and determining the sound effect category represented by the target sound effect text tag as the sound effect category to be adjusted; in the case of recognizing the processing parameter text, obtaining, from a plurality of pre-set parameter tags, the parameter tag that satisfies a preset similarity degree with the processing parameter text as the target parameter tag, and determining the sound effect adjustment parameter represented by the target parameter tag as the sound effect adjustment parameter of the sound effect category to be adjusted.
[0043] Among them, the object text may refer to the description information in the sound effect adjustment instruction for characterizing the user's sound effect category (such as timbre, surround, and energy, etc.), and this description information is the information input by the user, and the description terms are relatively colloquial and usually do not belong to standard description terms.
[0044] Among them, the processing parameter text may refer to the description information in the sound effect adjustment instruction for characterizing the user's sound effect adjustment parameters (such as adjustment direction, adjustment amplitude, etc.), and this description information is the information input by the user, and the description terms are relatively colloquial and usually do not belong to standard description terms.
[0045] Among them, the sound effect tag corresponding to the object text may refer to the pre-set information for characterizing the category of the audio sound effect, and the sound effect tag is a standard sound effect category description term.
[0046] Among them, the parameter tag corresponding to the processing parameter text may refer to the pre-set information for characterizing the adjustment direction and adjustment amplitude of the sound effect category described by the sound effect tag, and the parameter tag is a standardized term.
[0047] As an example, the terminal can convert the sound effect adjustment instruction into text form to obtain the sound effect adjustment instruction text. For example, when the sound effect adjustment instruction is in voice form, the terminal can analyze the sound effect adjustment instruction (such as the user voice signal) through natural language processing (NLP) algorithms, identify language patterns and features, and obtain the sound effect adjustment instruction text. When the sound effect adjustment instruction is in text form, the sound effect adjustment instruction in text form can be used as the sound effect adjustment instruction text. Then the terminal can perform semantic understanding on the sound effect adjustment instruction text (such as syntactic analysis or context correlation analysis, etc.), analyze the sound effect category (i.e., the object of action) and the sound effect adjustment parameters (i.e., the processing parameters) that the sound effect adjustment instruction text needs to adjust. When the object of action and the processing parameters are both specified / included in the sound effect adjustment instruction text, the terminal can determine the object of action text and the processing parameter text identified in the sound effect adjustment instruction by performing semantic understanding on the sound effect adjustment instruction text. When the object of action text is recognized, the terminal can obtain the sound effect text label that meets the preset similarity degree with the object of action text from a plurality of pre-set sound effect text label sets as the target sound effect text label. The sound effect text label can represent / describe the sound effect category. Therefore, the terminal can use the sound effect category represented / described by the target sound effect text label as the sound effect category to be adjusted. When the processing parameter text is recognized, the terminal can obtain the parameter label that meets the preset similarity degree with the processing parameter text from the pre-set parameter label set as the target parameter label. The parameter label can represent / describe the sound effect adjustment parameters (such as the method / direction of sound effect adjustment and the degree / amplitude of sound effect adjustment). Therefore, the terminal can use the sound effect adjustment parameters represented / described by the target parameter label as the sound effect adjustment parameters of the sound effect category to be adjusted.
[0048] In this embodiment, by determining the object of action and the processing parameters based on the sound effect adjustment instruction, and determining the sound effect text label that meets the preset similarity degree with the object of action text as the sound effect category to be adjusted, and determining the parameter label that meets the preset similarity degree with the processing parameter text as the sound effect adjustment parameters of the sound effect category to be adjusted, it is possible to accurately obtain the sound effect category to be adjusted and the sound effect adjustment parameters of the sound effect category to be adjusted by performing semantic understanding on the sound effect adjustment instruction, improve the accuracy of the sound effect category to be adjusted and the sound effect adjustment parameters, so as to accurately adjust the sound effect of the audio to be adjusted and improve the accuracy of audio adjustment.
[0049] In some embodiments, obtaining a target sound effect text label that meets a preset similarity degree with the text of the object of action includes: mapping the feature vector of the text of the object of action to a first vector clustering space, where the first vector clustering space includes at least one first vector clustering cluster, and the first vector clustering cluster is a vector clustering cluster in the first vector clustering space whose clustering center represents a sound effect category or a sound effect processor; each first vector clustering cluster corresponds to a candidate sound effect text label; determining the first similarity information between the feature vector of the object of action and each first vector clustering cluster; and screening out the target sound effect text label from each candidate sound effect text label according to whether the first similarity information meets the preset similarity degree.
[0050] Among them, the feature vector of the text of the object of action may refer to the information representing the text of the object of action in numerical form. In practical applications, the feature vector of the text of the object of action may include, but is not limited to, the embedding of the text of the object of action.
[0051] Among them, the first vector clustering space may refer to a pre-set space / model that contains several first vector clustering clusters. The first vector clustering cluster may refer to each group / cluster of vectors obtained by dividing vectors according to a preset similarity metric. Each first vector clustering cluster has its corresponding clustering center. In practical applications, the first vector clustering space may include first vector clustering clusters.
[0052] Among them, the first vector clustering cluster may refer to a vector clustering cluster in the vector clustering space whose clustering center represents a sound effect category / sound effect processor. In practical applications, each first vector clustering cluster corresponds to a candidate sound effect label, that is, the candidate sound effect text labels corresponding to each first vector clustering cluster can represent the sound effect category represented by the clustering center of the first vector clustering cluster or the attributes of the audio signal processed by the sound effect processor. The candidate sound effect text label may refer to a pre-set text / text label for representing a sound effect category.
[0053] Among them, the first similarity information may refer to the information representing the similarity between the feature vector of the text of the object of action and each first vector clustering cluster in the first vector clustering space. In practical applications, the first similarity information may include the distance (such as Euclidean distance, etc.) between the feature vector of the text of the object of action and each first vector clustering cluster in the first vector clustering space.
[0054] As an example, the terminal can first obtain the feature vector of the text of the object of action, and then map the feature vector of the text of the object of action to a preset first vector clustering space. Since the preset first vector clustering space contains a number of preset first vector clustering clusters, the terminal can calculate the distance between the feature vector of the text of the object of action and each first vector clustering cluster in the first vector clustering space as the first similarity information one by one. After that, the terminal can analyze whether the first similarity information meets the preset similarity degree. For example, the terminal can compare each first similarity information pairwise, and determine the information with the largest represented similarity (such as the smallest distance) as the first target similarity information. The terminal can use the first vector clustering cluster corresponding to the first target similarity information as the first target vector clustering cluster. Since each first vector clustering cluster corresponds to a candidate sound effect label, the terminal can use the candidate sound effect text label corresponding to the first target vector clustering cluster as the sound effect label that matches the text of the object of action, so as to screen out the target sound effect text label from each candidate sound effect text label.
[0055] In this embodiment, by analyzing the similarity between the feature vector of the text of the object of action and each first vector clustering cluster in the first vector clustering space, the target sound effect text label is accurately screened out from the candidate sound effect text labels, which can improve the accuracy of the sound effect text label, thereby improving the accuracy of the sound effect category to be adjusted, so as to accurately adjust the sound effect of the audio to be adjusted and improve the accuracy of the audio adjustment.
[0056] In some embodiments, according to whether the first similarity information meets the preset similarity degree, screening out the target sound effect text label from each candidate sound effect text label includes: when the first similarity information is greater than or equal to the first preset threshold, screening out the target sound effect text label from each candidate sound effect text label according to the maximum value of the first similarity information.
[0057] Wherein, the first preset threshold may be information used to determine whether the feature vector of the text of the object of action matches the first target vector clustering cluster.
[0058] As an example, the terminal can first compare the first similarity information with the first preset threshold, determine the first similarity information that is greater than or equal to the first preset threshold, and determine the first similarity information with the largest value among the first similarity information that is greater than or equal to the first preset threshold as the maximum value of the first similarity information. At this time, the terminal can use the candidate sound effect text label corresponding to the maximum value of the first similarity information as the target sound effect text label corresponding to the text of the object of action.
[0059] In this embodiment, based on the size relationship between the first similarity information and the first preset threshold, the target sound effect text label of the text of the object of action is accurately screened out from the candidate sound effect text labels, which can improve the accuracy of sound effect label recognition.
[0060] In some embodiments, the above method further includes: when each first similarity information is less than a first preset threshold, generating feedback information indicating that the recognition of the target sound effect text label fails.
[0061] As an example, if each first similarity information is less than the first preset threshold, the terminal can determine that the sound effect category to be adjusted cannot be recognized at this time. The terminal can generate feedback information indicating that the recognition of the target sound effect text label fails and display the feedback information to the user, and the feedback information can be used to prompt the user to adjust the sound effect adjustment instruction.
[0062] In this embodiment, when the target sound effect text label of the object text to be processed cannot be recognized, generating feedback information indicating that the recognition of the sound effect text label fails can timely feedback the situation of incorrect recognition of the sound effect label to the user and ensure the effective implementation of the recognition of the sound effect adjustment intention.
[0063] In some embodiments, obtaining a parameter label that meets a preset similarity degree with the processing parameter text as the target parameter label includes: mapping the feature vector of the processing parameter text to a second vector clustering space, where the second vector clustering space includes at least one second vector clustering cluster, and the second vector clustering cluster is a vector clustering cluster in the second vector clustering space whose clustering center represents an adjustment parameter; each second vector clustering cluster corresponds to a candidate parameter label; the candidate parameter label is used to represent the direction and amplitude of the sound effect adjustment; determining the second similarity information between the feature vector of the processing parameter and each second vector clustering cluster; and screening out the target parameter label from each candidate parameter label according to whether the second similarity information meets the preset similarity degree.
[0064] Among them, the processing parameter text may refer to information (rough or fuzzy) describing the adjustment parameters for the audio sound effect by the user in the sound effect adjustment instruction. For example, the sound effect adjustment instruction can be expressed as "I want to increase the space where the music is located", "space" can be the object to be processed, and "increase" can be the processing parameter.
[0065] Among them, the feature vector of the processing parameter text may refer to information representing the processing parameter text in numerical form. In practical applications, the feature vector of the processing parameter text may include, but is not limited to, the embedding of the processing parameter text.
[0066] Among them, the second vector clustering space may refer to a pre-set space / model containing several second vector clustering clusters. The second vector clustering cluster may refer to each group / cluster of vectors obtained by dividing the vectors into several groups / clusters according to a preset similarity metric. Each second vector clustering cluster has its corresponding clustering center. In practical applications, the second vector clustering space may include the first vector clustering cluster.
[0067] Among them, the second vector clustering cluster may refer to the vector clustering cluster in the second vector clustering space where the clustering center represents the sound effect adjustment parameter. In practical applications, each second vector clustering cluster corresponds to a candidate parameter label, that is, the candidate parameter label corresponding to each second vector clustering cluster can represent the adjustment parameter represented by the clustering center of the second vector clustering cluster. The candidate parameter label may refer to the pre-set text used to represent the parameter label, and the candidate parameter label can be used to represent the direction and amplitude of the sound effect adjustment.
[0068] Among them, the second similarity information may refer to the information representing the similarity between the feature vector of the processing parameter and each second vector clustering cluster in the vector clustering space. In practical applications, the second similarity information may include the distance (such as Euclidean distance, etc.) between the feature vector of the processing parameter text and each second vector clustering cluster in the vector clustering space. As an example, after determining the sound effect category to be adjusted according to the sound effect label corresponding to the object text, the terminal can obtain the feature vector of the processing parameter text and map the feature vector of the processing parameter text to the pre-set second vector clustering space. Since the pre-set second vector clustering space contains a pre-set number of second vector clustering clusters, the terminal can calculate the distance between the feature vector of the processing parameter text and each second vector clustering cluster in the second vector clustering space one by one as the second similarity information. Then, the terminal can analyze whether it meets the pre-set similarity degree. For example: the terminal can compare each second similarity information pairwise, and determine the information with the largest represented similarity (such as the smallest distance) as the second target similarity information. The terminal can use the second vector clustering cluster corresponding to the second target similarity information as the second target vector clustering cluster. Since each second vector clustering cluster corresponds to a candidate parameter label, therefore, the terminal can use the candidate parameter label corresponding to the second target vector clustering cluster as the parameter label corresponding to the processing parameter text, so as to screen out the target parameter label from each candidate parameter label.
[0069] In this embodiment, by analyzing the similarity between the feature vector of the processing parameter text and each second vector clustering cluster in the second vector clustering space, the target parameter label is accurately screened out from the candidate parameter labels, which can improve the accuracy of the target parameter label, thereby improving the accuracy of the sound effect adjustment parameter, so as to accurately adjust the sound effect of the audio to be adjusted and improve the accuracy of the audio adjustment.
[0070] In some embodiments, screening out the target parameter label from each candidate parameter label according to whether the second similarity information meets the pre-set similarity degree includes: when the second similarity information is greater than or equal to the second preset threshold, screening out the target parameter label from each candidate parameter label according to the maximum value of the second similarity information; the parameter described by the target parameter label is used to adjust the sound effect parameter of the audio to be adjusted on the basis of the current sound effect parameter of the audio to be adjusted for the sound effect category to be adjusted.
[0071] Among them, the second preset threshold may refer to information used to determine whether the feature vector of the processing parameter text matches the second target vector clustering cluster.
[0072] As an example, the terminal may first compare the second similarity information with the second preset threshold, determine the second similarity information greater than or equal to the second preset threshold, and determine the second similarity information with the largest value from the second similarity information greater than or equal to the first preset threshold as the maximum value of the second similarity information. At this time, the terminal may use the candidate parameter label corresponding to the maximum value of the second similarity information as the target parameter label corresponding to the processing parameter text. It can be understood that if the sound effect adjustment instruction describes both the direction and magnitude of the sound effect adjustment at the same time, the terminal may accurately cluster based on the processing parameter text in the sound effect adjustment instruction in the second vector clustering space, so as to screen out the accurate parameter label from the candidate parameter labels.
[0073] In this embodiment, based on the magnitude relationship between the second similarity information and the second preset threshold, accurately screening out the target parameter label corresponding to the processing parameter from the candidate parameter labels can improve the accuracy of parameter label recognition.
[0074] In some embodiments, the above method further includes: generating feedback information indicating that the recognition of the target parameter label fails in the case where each second similarity information is less than the second preset threshold.
[0075] As an example, if each second similarity information is less than the second preset threshold, the terminal may determine that the sound effect adjustment parameter of the sound effect category to be adjusted cannot be recognized at this time. The terminal may generate feedback information indicating that the recognition of the target parameter label fails and display the feedback information to the user. It can be understood that if the sound effect adjustment instruction only describes the direction or magnitude of the sound effect adjustment, the terminal cannot accurately cluster based on the processing parameter text in the sound effect adjustment instruction in the second vector clustering space. The terminal may generate feedback information indicating that the recognition of the target parameter label fails, and this feedback information may prompt the user that the parameter label recognition fails (such as reporting an error) or prompt the user to adjust the sound effect adjustment instruction.
[0076] In this embodiment, when the parameter label corresponding to the processing parameter text cannot be recognized, generating feedback information indicating that the recognition of the target parameter label fails can timely feedback the situation of parameter label recognition error to the user and ensure the effective implementation of the sound effect adjustment intention recognition.
[0077] In practical applications, by pre-analyzing user input and combining the analysis results, an identification model can be trained and a first vector clustering space and a second vector clustering space can be constructed. The identification model can be used to determine the similarity information or distance between the feature vector of the user input text and each vector clustering cluster in the first vector clustering space and the second vector clustering space. In specific implementation, the user input can be used as a training sample, and the corresponding sound effect category and sound effect adjustment parameters of the training sample can be known or pre-labeled (i.e., the true sound effect category and the true sound effect adjustment parameters). The true sound effect category can be used as the sound effect text label, and the true sound effect adjustment parameters can characterize the specific direction and specific amplitude of the sound effect adjustment. The specific direction and specific amplitude of the sound effect adjustment characterized by the true sound effect adjustment parameters can be used as the parameter label. The training sample is input into the identification model to be trained. The identification model to be trained can analyze the training sample, determine and output the corresponding sound effect category and sound effect adjustment parameters of the training sample. The corresponding sound effect category and sound effect adjustment parameters of the training sample output by the identification model can be used as prediction information. The terminal can train the identification model to be trained based on the difference between the true sound effect category and the sound effect category in the prediction information, and the difference between the true sound effect adjustment parameters and the sound effect adjustment parameters in the prediction information, to obtain the trained identification model (the trained identification model can be used as the pre-trained identification model).
[0078] The identification model can be used to analyze the vector clustering cluster corresponding to the object text in the sound effect adjustment instruction in the first vector clustering space and determine the sound effect text label. The identification model can be used to analyze the vector clustering cluster corresponding to the processing parameter text in the sound effect adjustment instruction in the second vector clustering space and determine the parameter label. For example: when the sound effect adjustment instruction is "I want more space", the object is "space" and the processing parameter is "more". "Space" will be clustered under the sound effect text label of space, so the sound effect processor related to space will be selected for audio processing. The processing parameter is "more". The terminal can also obtain the sound effect parameters of the current sound effect processor related to space of the audio to be adjusted. If the feature vector of the processing parameter has the maximum similarity with the vector clustering cluster representing the two-unit translation distance in the second vector clustering space, that is, the feature vector of the processing parameter is close to the two-unit translation distance in the scale, and the feature vector of the processing parameter has the maximum similarity with the vector clustering cluster representing an increase in the second vector clustering space, that is, the feature vector of the processing parameter is close to the parameter increase, at this time the terminal can determine that the processing parameter is to adjust the gear up by 0.2.
[0079] When it is necessary to identify the sound effect text label corresponding to the object text and the parameter label corresponding to the processing parameter text, the terminal can input the feature vector of the object text into the pre-trained recognition model. The sound effect category output by the pre-trained recognition model can be used as the sound effect category to be adjusted, and the sound effect adjustment parameter output by the pre-trained recognition model can be used as the sound effect adjustment parameter of the sound effect category to be adjusted.
[0080] In some embodiments, intent recognition is performed on the sound effect adjustment instruction to determine the sound effect adjustment intent for the audio to be adjusted, including: identifying the text related to the sound effect category in the sound effect adjustment instruction as the object text, performing semantic understanding on the sound effect adjustment instruction, and identifying the text related to the adjustment parameter of the sound effect category in the sound effect adjustment instruction as the processing parameter text identified in the sound effect adjustment instruction; in the case where the processing parameter text is recognized but the object text is not recognized, among the multiple pre-set parameter labels, obtaining the parameter label corresponding to the processing parameter text that meets the preset similarity degree as the target parameter label, and determining the sound effect adjustment parameter represented by the target parameter label as the sound effect adjustment parameter of the sound effect category to be adjusted; according to the preset correspondence relationship between the parameter label and the sound effect text label, determining the sound effect category represented by the sound effect text label corresponding to the selected target parameter label as the sound effect category to be adjusted.
[0081] Among them, the preset correspondence relationship can refer to the information used to represent the mapping relationship between the parameter label and the sound effect text label. In practical applications, each parameter label can have a corresponding sound effect text label. For example, when the parameter label is warm, bright, sharp, etc., the sound effect text label corresponding to the parameter label can be timbre / equalization / pitch shift; when the parameter label is large space, surround, depth, etc., the sound effect text label corresponding to the parameter label can be surround / reverb / delay; when the parameter label is increase volume, etc., the sound effect text label corresponding to the parameter label can be energy / compression / automatic gain.
[0082] As an example, the terminal can perform semantic understanding on the sound effect adjustment instruction text (such as syntactic analysis or context correlation analysis, etc.), analyze the sound effect category (i.e., the object of action) and the sound effect adjustment parameters (i.e., the processing parameters) required to be adjusted by the sound effect adjustment instruction text. When only the processing parameters are specified / included in the sound effect adjustment instruction text, the terminal can determine the processing parameter text recognized in the sound effect adjustment instruction by performing semantic understanding on the sound effect adjustment instruction text. The terminal can obtain the target parameter label corresponding to the processing parameter text from the preset parameter label set. The parameter label can characterize / describe the sound effect adjustment parameter (such as the method / direction of sound effect adjustment and the degree / amplitude of sound effect adjustment). Therefore, the terminal can use the sound effect adjustment parameter characterized / described by the target parameter label corresponding to the processing parameter text as the sound effect adjustment parameter of the sound effect category to be adjusted. At this time, in order to determine the sound effect category to be adjusted, the terminal can obtain the preset correspondence between the parameter label and the sound effect text label, and based on the preset correspondence between the parameter label and the sound effect text label, filter out the sound effect text label corresponding to the target parameter label from the candidate sound effect text labels. The terminal can determine the sound effect category described by the filtered sound effect text label corresponding to the target parameter label as the sound effect category to be adjusted.
[0083] In this embodiment, by using the preset correspondence to determine the sound effect text label corresponding to the target parameter label as the sound effect category to be adjusted, the sound effect text label of the sound effect adjustment instruction can be accurately obtained by using the parameter label, improving the accuracy of the sound effect text label, thereby improving the accuracy of the sound effect category to be adjusted, so as to perform accurate sound effect adjustment on the audio to be adjusted and improve the accuracy of audio adjustment.
[0084] In some embodiments, identifying the text related to the sound effect category in the sound effect adjustment instruction as the object of action text, and identifying the text related to the adjustment parameter of the sound effect category in the sound effect adjustment instruction as the processing parameter text includes: obtaining a preset regular expression; wherein, the regular expression is used to find the preset object of action and the preset processing parameter in the text; matching the sound effect adjustment instruction according to the regular expression to identify the text related to the sound effect category in the sound effect adjustment instruction as the object of action text and identify the text related to the adjustment parameter of the sound effect category in the sound effect adjustment instruction as the processing parameter text.
[0085] Wherein, the preset regular expression may refer to the information that the regular expression is used to find the preset object of action and the preset processing parameter in the text. In practical applications, the preset regular expression may include all or part of the preset object of action text and the preset processing parameter text.
[0086] As an example, in order to determine the sound effect category to be adjusted and the sound effect adjustment parameters for the sound effect category to be adjusted, the terminal can generate a preset regular expression according to the preset object text and the preset processing parameter text. Then, the terminal can use the preset regular expression to perform text matching on the sound effect adjustment instruction word by word / phrase by phrase, search whether the sound effect adjustment instruction contains the preset object text and the preset processing parameter text, and determine the preset object text and the preset processing parameter text contained in the sound effect adjustment instruction. Then, the terminal can use the preset object text contained in the sound effect adjustment instruction as the object text related to the sound effect category identified from the sound effect adjustment instruction, and use the preset processing parameter text contained in the sound effect adjustment instruction as the processing parameter text related to the adjustment parameters of the sound effect category identified. In practical applications, the terminal can also use methods such as pattern matching to determine the object text and the processing parameter text in the sound effect adjustment instruction.
[0087] In this embodiment, by performing regular matching on the sound effect adjustment instruction using a regular expression and querying the object and the processing parameters that match the sound effect adjustment instruction, it is possible to improve the accurate identification of the object text and the processing parameter text, improve the accuracy of the sound effect text label and the parameter label, thereby improving the accuracy of the sound effect category to be adjusted and the sound effect adjustment parameters, so as to perform accurate sound effect adjustment on the audio to be adjusted and improve the accuracy of audio adjustment.
[0088] In some embodiments, the above method further includes: determining user sound effect adjustment preference information according to the sound effect adjustment intention for the audio to be adjusted; in the case where no sound effect adjustment instruction for the new audio to be adjusted is obtained, adjusting the audio signal of the new audio to be adjusted according to the user sound effect adjustment preference information.
[0089] Among them, the user sound effect adjustment preference information may refer to the information representing the personal preferences, interests, habits, and tendencies shown by the user during the process of performing sound effect adjustment operations on the audio.
[0090] As an example, due to the different preferences and aesthetics of different users, there are differences in the audio adjustment requirements of users for audio. For each user, the terminal can obtain the audio effect adjustment intention for each audio to be adjusted during the process of the user adjusting the audio effects of at least one audio to be adjusted, and determine the user audio effect adjustment preference information by analyzing the audio effect adjustment intention for each audio to be adjusted. When the user needs to adjust the audio effects of a new audio to be adjusted, before the user issues an audio effect adjustment instruction for the new audio to be adjusted, the terminal can adjust the audio signal of the new audio to be adjusted based on the user audio effect adjustment preference information, so as to achieve adaptive audio effect adjustment without the need for user instructions. It can be understood that if the user issues an audio effect adjustment instruction for the new audio to be adjusted, the terminal needs to determine the audio effect adjustment intention for the new audio to be adjusted according to the audio effect adjustment instruction for the new audio to be adjusted, so as to adjust the audio signal of the new audio to be adjusted.
[0091] In this embodiment, determining the user audio effect adjustment preference based on the audio effect adjustment intention for the audio to be adjusted and using the user audio effect adjustment preference to adjust the audio effects can achieve personalized adjustment of the audio effects and improve the accuracy and flexibility of audio adjustment.
[0092] In some embodiments, in the traditional technology, adjusting the audio effects through an audio effect processor often requires the user to understand relevant audio knowledge and usage methods, which requires a relatively high professional quality of the user. For users, different preferences bring different aesthetic standards, resulting in different listening requirements and audio effect adjustment requirements, so there will be a large gap between the subjective description of the user and the actual change of the audio effects. Therefore, it is difficult to modify the audio based on the audio effect processor. At the same time, in the traditional technology, the music generation and music editing technologies based on natural language processing technology can change the music content through deep learning methods by user input, but this is not to perform effect processing on the existing audio, and there are problems such as large models that cannot respond in real time on the upper end. At the same time, there is a huge unexplainability in the process of mapping the input semantic information of the user to the music change, and the effect is often uncontrollable. To solve the problems existing in the traditional technology when implementing audio effect adjustment, an audio adjustment method is provided. First, the terminal can model the audio effect processor and the parameters of the audio effect processor to obtain a preset vector clustering space, such as Figure 3As shown, a schematic diagram of a vector clustering space is provided. Then, the terminal projects the user's input instruction (such as a sound effect adjustment instruction) into a pre-designed model (such as a preset vector clustering space), selects the parameter pair of the parameter coordinate point (such as a sound effect label and a parameter label) closest to the user's input instruction, and performs real-time processing on the audio, thereby realizing the adjustment of the audio sound effect. If the user's input instruction only changes slightly in a certain dimension, the terminal can change the processing effect / sound effect adjustment effect by parameter translation in that dimension. In specific implementation, for the descriptive words representing degree in the user input, such as "turn it up a bit", the terminal can first locate the specific effector according to the user instruction, obtain the current effect parameter, and then adjust the effect by a unit length.
[0093] In practical applications, to model the sound effectors and their parameters, the sound effectors can be classified and preset by labels and scales first, as Figure 4 As shown, a schematic diagram of effector classification is provided. The terminal can adapt the existing effector names to labels. Here, the adaptation can be manually labeled by relevant experts. The parameters of each effector will be uniformly defined as n gears, and the definition standard is based on the minimum perceptible change of the human ear. The final processing parameter range will be uniformly standardized to [-1, 1], and the unit interval is 0.1. After the terminal obtains the feature vector embedding of the sample text (including the sound effect category to be adjusted and the adjustment parameters), the terminal can cluster the feature vector embedding of the sample text to obtain the vector clustering cluster corresponding to the sound effect category in the vector clustering space. For example, descriptive words such as warm, bright, and sharp for timbre will be clustered into the timbre label (timbre - frequnecy), and the corresponding effectors are equalization, pitch change, etc.; descriptive words related to space such as large space, surround, and depth will be clustered into the surround label (space / surround), and the corresponding ones are reverberation, delay, etc.; expressions related to energy such as increasing volume will be clustered into the energy label (energy / dynamic), that is, corresponding to compression, automatic gain, etc. For inputs that cannot be clustered into the timbre label, surround label, and energy label, the terminal can perform extreme value elimination. In specific implementation, the clustering labels and sample texts can be expanded through a thesaurus of synonyms.
[0094] In practical applications, the user's input can be located at a certain point in the vector clustering space. Repeated interactions will translate this point, and finally, the user's input will be located at the accurate parameter of a certain sound effector as the output, as Figure 5As shown, a schematic diagram of parameter vectorization is provided. The terminal can adjust the sound effects of a standard audio based on different parameters corresponding to different user inputs to obtain different versions of sample audio. Then, the terminal can use the standard audio and different versions of sample audio with different parameters as a sample dataset to be fed into a parameter vectorization model for learning, analyze the correlation between user inputs and parameters, achieve parameter vectorization, and cluster the parameters to obtain the vector cluster corresponding to the parameters in the vector clustering space.
[0095] After constructing the vector clustering space, when the user needs to adjust the sound effects of an audio, the terminal can obtain the sound effect adjustment instruction for the audio from the user and convert the sound effect adjustment instruction into a sound effect adjustment instruction text. Then, the terminal can use word segmentation technology to split the sound effect adjustment instruction text into words or phrases. After that, the terminal can utilize a predefined series of keywords or phrases, such as music, space, etc., and match the split words or phrases through regular expressions or pattern matching, etc., to identify the object of action (such as an effector or a sound effect label). The terminal can, according to the object of action, perform syntactic analysis or context correlation analysis on the sound effect adjustment instruction text, etc., to determine the descriptive words or phrases associated with the object of action, such as warm, louder, etc., as adjustment parameters. The terminal can also associate the identified object of action and adjustment parameters and output the recognition result. Taking the sound effect adjustment instruction text "I want the music to be warmer" as an example for illustration, the object of action is music, the adjustment parameter is warm, and the output recognition result can be "Object of action - music, adjustment parameter - warm", thus splitting the corresponding text input by the user into the object of action and adjustment parameters.
[0096] The terminal can extract the embedding of the target object and, through a pre-constructed vector clustering space, obtain the clustering cluster in which the embedding of the target object is located in the vector clustering space. The clustering cluster in which the embedding of the target object is located in the vector clustering space can be used as a sound effect label, and the sound effect label can be used to characterize the effector on which the sound effect adjustment operation acts. The terminal can also extract the embedding of the adjustment parameter and, through the pre-constructed vector clustering space, obtain the clustering cluster in which the embedding of the adjustment parameter is located in the vector clustering space. The clustering cluster in which the embedding of the adjustment parameter is located in the vector clustering space can be used as a parameter label, thereby converting the parameter in text form into a vector parameter. Taking the sound effect adjustment instruction text "I want more space" as an example for illustration, the target object is "space" and the adjustment parameter is "more". "Space" will be clustered under the label of space, that is, the sound effect label is "space". The processing parameter is "more". If the adjustment parameter at this time is 0, the feature vector of the adjustment parameter at this time is close to the two-unit translation distance representing the scale in the vector clustering space. Therefore, the adjustment gear for the sound effect adjustment operation for "space" can be adjusted upward by 0.2. It can be understood that for inputs in the user input that have nothing to do with sound effect adjustment (such as sound effect labels and parameter labels), a preset distance threshold is set. When the distance between the feature vector of a certain text in the user input and each clustering cluster in the vector clustering space is greater than the preset distance threshold, the terminal can generate an identification result of "unrecognized".
[0097] In specific implementation, if the user input cannot be divided into two parts: the target object and the adjustment parameter, the terminal can first identify the adjustment parameter, and then, according to the vector distance calculation, match the clustering cluster with the closest feature vector of the text in the user input in the clustering cluster representing the sound effect label in the vector clustering space, and use the sound effect category corresponding to the closest-matched clustering cluster as the target object. Since there is an association / correspondence relationship between the target object and the adjustment parameter, a mapping relationship between the target object and the adjustment parameter can also be pre-constructed. When the user input cannot be divided into two parts: the target object and the adjustment parameter, the terminal can first identify the adjustment parameter, and then, according to the pre-constructed mapping relationship between the target object and the adjustment parameter, determine the target object associated with the adjustment parameter. Taking the sound effect adjustment instruction text "I want it to be more enthusiastic" as an example for illustration, at this time, the terminal cannot identify a clear target object. The terminal can identify the adjustment parameter as "enthusiastic + more". By matching the clustering cluster in the vector clustering space and calculating the word vector distance (such as the Euclidean distance) between the feature vector of the text and the clustering cluster, the terminal can use the sound effect label "timbre" corresponding to the matched clustering cluster as the target object. The Euclidean distance only needs to consider the angle between the vectors. Assuming that S1 is the first word segment and S2 is the second word segment, the calculation expression of the Euclidean distance can be expressed as:
[0098]
[0099] Among them, (x1, x2, x3) can represent the coordinates of the first word segment in the vector clustering space, (y1, y2, y3) can represent the coordinates of the second word segment in the vector clustering space, and d can represent the Euclidean distance between the first word segment and the second word segment.
[0100] The terminal can also query in the preset mapping relationship that the sound effect label corresponding to the adjustment parameter is "timbre", and determine that the object of action corresponding to the adjustment parameter is "timbre".
[0101] After obtaining the object of action and the adjustment parameter, the terminal can convert the object of action and the adjustment parameter into actual DSP effector signals, so as to perform sound effect adjustment processing on the audio signal of the audio. The DSP processing code is written in the front end. After receiving the instruction sent by the background, the front end performs sound effect processing in real time, and the finally output signal will be the signal processed after the user issues the instruction.
[0102] In this embodiment, by designing the mapping method from user description to audio parameters, the gap between subjective description and actual change of sound effect is solved to a certain extent. At the same time, the sound effect parameters of the effector improve the controllability and interpretability of the sound effect adjustment process. At the same time, the sound effect processing based on the effector is based on signal processing, its kernel is lightweight enough, and the demand for computing resources is wider, with the possibility of real-time upper end.
[0103] It should be understood that although the steps in the flowcharts involved in the above-mentioned embodiments are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.
[0104] Based on the same inventive concept, the embodiment of the present application also provides an audio adjustment device for implementing the above-mentioned audio adjustment method. The implementation solution provided by this device to solve the problem is similar to the implementation solution described in the above method. Therefore, the specific limitations in one or more embodiments of the audio adjustment device provided below can refer to the limitations on the audio adjustment method in the above text, and will not be repeated here.
[0105] In an exemplary embodiment, such asFigure 6 As shown, an audio adjustment device is provided, including: an instruction acquisition module 602, an instruction recognition module 604, and a sound effect adjustment module 606, where:
[0106] The instruction acquisition module 602 is configured to acquire a sound effect adjustment instruction for the audio to be adjusted.
[0107] The instruction recognition module 604 is configured to perform intention recognition on the sound effect adjustment instruction to determine the sound effect adjustment intention for the audio to be adjusted; the sound effect adjustment intention includes the sound effect category to be adjusted and the sound effect adjustment parameters for the sound effect category to be adjusted.
[0108] The sound effect adjustment module 606 is configured to adjust the audio signal of the audio to be adjusted according to the sound effect category to be adjusted and the sound effect adjustment parameters associated with the sound effect category to be adjusted.
[0109] In one exemplary embodiment, the instruction recognition module 604 is specifically further configured to perform semantic understanding on the sound effect adjustment instruction, identify the text related to the sound effect category in the determined sound effect adjustment instruction as the target text of the object, and identify the text related to the adjustment parameters of the sound effect category in the sound effect adjustment instruction as the processing parameter text; in the case of identifying the target text of the object, obtain the sound effect text label corresponding to the sound effect text label that satisfies the preset similarity degree with the target text of the object from a plurality of preset sound effect text labels as the target sound effect label, and determine the sound effect category represented by the target sound effect label as the sound effect category to be adjusted; in the case of identifying the processing parameter text, obtain the parameter label corresponding to the processing parameter text that satisfies the preset similarity degree in a plurality of preset parameter labels as the target parameter label, and determine the sound effect adjustment parameter represented by the target parameter label as the sound effect adjustment parameter of the sound effect category to be adjusted.
[0110] In one exemplary embodiment, the instruction recognition module 604 is specifically further configured to map the feature vector of the target text of the object to a first vector clustering space, the first vector clustering space includes at least one first vector clustering cluster, and the first vector clustering cluster is a vector clustering cluster in the first vector clustering space whose clustering center represents a sound effect category or a sound effect effector; each of the first vector clustering clusters corresponds to a candidate sound effect label; determine the first similarity information between the feature vector of the target text of the object and each of the first vector clustering clusters; and screen out the target sound effect label from each of the candidate sound effect labels according to whether the first similarity information satisfies the preset similarity degree.
[0111] In one exemplary embodiment, the instruction recognition module 604 is further specifically configured to, when the first similarity information is greater than or equal to a first preset threshold, screen out the target sound effect label sound effect text label from each of the candidate sound effect label sound effect text labels according to the maximum value of the first similarity information.
[0112] In one exemplary embodiment, the device further includes a first feedback module, which is specifically configured to generate first feedback information indicating that the recognition of the target sound effect label sound effect text label fails when each of the first similarity information is less than the first preset threshold.
[0113] In one exemplary embodiment, the instruction recognition module 604 is further specifically configured to map the feature vector of the processing parameter text to a second vector clustering space, where the second vector clustering space includes at least one second vector clustering cluster, and the second vector clustering cluster is a vector clustering cluster in the second vector clustering space whose clustering center represents an adjustment parameter; each of the second vector clustering clusters corresponds to a candidate parameter label; determine the second similarity information between the feature vector of the processing parameter text and each of the second vector clustering clusters; and screen out the target parameter label from each of the candidate parameter labels according to whether the second similarity information meets the preset similarity degree.
[0114] In one exemplary embodiment, the instruction recognition module 604 is further specifically configured to, when the second similarity information is greater than or equal to a second preset threshold, screen out the target parameter label from each of the candidate parameter labels according to the maximum value of the second similarity information; the sound effect adjustment parameter described and represented by the target parameter label is used to adjust the sound effect parameters on the basis of the current sound effect parameters of the to-be-adjusted audio for the to-be-adjusted sound effect category.
[0115] In one exemplary embodiment, the device further includes a second feedback module, which is specifically configured to generate second feedback information indicating that the recognition of the target parameter label fails when each of the second similarity information is less than the second preset threshold.
[0116] In one exemplary embodiment, the instruction recognition module 604 is further specifically configured to recognize the text related to the sound effect category from the sound effect adjustment instruction as the target object text, perform semantic understanding on the sound effect adjustment instruction, and recognize the text related to the adjustment parameter of the sound effect category from the sound effect adjustment instruction as the processing parameter text identified in the sound effect adjustment instruction; in the case where the processing parameter text is recognized but the target object text is not recognized, among the multiple preset parameter tags, obtain the parameter tag corresponding to the preset similarity degree with the processing parameter text as the target parameter tag, and determine the sound effect adjustment parameter characterized by the target parameter tag as the sound effect adjustment parameter of the sound effect category to be adjusted; according to the preset correspondence between the parameter tag and the sound effect tag (sound effect text tag), determine the sound effect category characterized by the sound effect tag (sound effect text tag) corresponding to the selected target parameter tag as the sound effect category to be adjusted.
[0117] In one exemplary embodiment, the instruction recognition module 604 is further specifically configured to obtain a preset regular expression; wherein, the regular expression is used to find a preset target object and a preset processing parameter in the text; match the sound effect adjustment instruction according to the regular expression to determine the target object text and the processing parameter text in the sound effect adjustment instruction, so as to recognize the text related to the sound effect category from the sound effect adjustment instruction as the target object text and recognize the text related to the adjustment parameter of the sound effect category as the processing parameter text.
[0118] Each module in the above audio adjustment device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective modules.
[0119] In an exemplary embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 7As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it implements an audio adjustment method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the outer shell of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0120] Those skilled in the art can understand that Figure 7 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0121] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0122] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0123] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0124] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0125] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0126] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.
[0127] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application shall be subject to the appended claims.
Claims
1. An audio adjustment method, characterized in that, The method includes: Obtaining a sound effect adjustment instruction for an audio input to be adjusted; Performing intent recognition on the sound effect adjustment instruction to determine a sound effect adjustment intent for the audio to be adjusted; the sound effect adjustment intent includes a sound effect category to be adjusted and sound effect adjustment parameters for the sound effect category to be adjusted; Adjusting the audio signal of the audio to be adjusted according to the sound effect category to be adjusted and the sound effect adjustment parameters associated with the sound effect category to be adjusted.
2. The method according to claim 1, wherein The performing intent recognition on the sound effect adjustment instruction to determine a sound effect adjustment intent for the audio to be adjusted includes: Identifying text related to a sound effect category in the sound effect adjustment instruction as an object text, and identifying text related to adjustment parameters of the sound effect category in the sound effect adjustment instruction as a processing parameter text; When the object text is identified, obtaining, from a plurality of pre-set sound effect text tags, a sound effect text tag that meets a preset similarity degree with the object text as a target sound effect text tag, and determining the sound effect category represented by the target sound effect text tag as the sound effect category to be adjusted; When the processing parameter text is identified, obtaining, from a plurality of pre-set parameter tags, a parameter tag that meets a preset similarity degree with the processing parameter text as a target parameter tag, and determining the sound effect adjustment parameter represented by the target parameter tag as the sound effect adjustment parameter of the sound effect category to be adjusted.
3. The method according to claim 2, wherein The obtaining a sound effect text tag that meets a preset similarity degree with the object text as a target sound effect text tag includes: Mapping a feature vector of the object text to a first vector clustering space, the first vector clustering space including at least one first vector clustering cluster, and the first vector clustering cluster being a vector clustering cluster in the first vector clustering space where a clustering center represents a sound effect category or a sound effect effector; each of the first vector clustering clusters corresponds to a candidate sound effect text tag; Determining first similarity information between the feature vector of the object text and each of the first vector clustering clusters; and screening out the target sound effect text tag from each of the candidate sound effect text tags according to whether the first similarity information meets the preset similarity degree.
4. The method according to claim 3, characterized in that, The screening out the target sound effect text tag from each of the candidate sound effect text tags according to whether the first similarity information meets the preset similarity degree includes: When the first similarity information is greater than or equal to a first preset threshold, screening out the target sound effect text tag from each of the candidate sound effect text tags according to the maximum value of the first similarity information.
5. The method according to claim 4, wherein The method further includes: Generating feedback information indicating that the recognition of the target sound effect text tag fails when each of the first similarity information is less than the first preset threshold.
6. The method according to claim 2, wherein The obtaining a parameter tag that meets a preset similarity degree with the processing parameter text as a target parameter tag includes: Map the feature vector of the processing parameter text to a second vector clustering space, where the second vector clustering space includes at least one second vector clustering cluster, and the second vector clustering cluster is a vector clustering cluster in the second vector clustering space where the clustering center represents an adjustment parameter; each of the second vector clustering clusters corresponds to a candidate parameter label; Determine the second similarity information between the feature vector of the processing parameter text and each of the second vector clustering clusters; based on whether the second similarity information meets the preset similarity degree, screen out the target parameter label from each of the candidate parameter labels.
7. The method according to claim 6, wherein The screening out the target parameter label from each of the candidate parameter labels according to whether the second similarity information meets the preset similarity degree includes: When the second similarity information is greater than or equal to a second preset threshold, screen out the target parameter label from each of the candidate parameter labels according to the maximum value of the second similarity information.
8. The method according to claim 7, characterized in that The method further includes: When each of the second similarity information is less than the second preset threshold, generate feedback information indicating that the recognition of the target parameter label fails.
9. The method according to claim 1, characterized in that The intention recognition of the sound effect adjustment instruction to determine the sound effect adjustment intention for the audio to be adjusted includes: Identify the text related to the sound effect category in the sound effect adjustment instruction as the object text, and identify the text related to the adjustment parameter of the sound effect category in the sound effect adjustment instruction as the processing parameter text; When the processing parameter text is recognized but the object text is not recognized, among the multiple parameter labels set in advance, obtain the parameter label that meets the preset similarity degree with the processing parameter text as the target parameter label, and determine the sound effect adjustment parameter represented by the target parameter label as the sound effect adjustment parameter of the sound effect category to be adjusted; According to the preset correspondence relationship between the parameter label and the sound effect text label, determine the sound effect category represented by the sound effect text label corresponding to the target parameter label as the sound effect category to be adjusted.
10. The method according to claim 2 or 9, characterized in that, The identifying the text related to the sound effect category in the sound effect adjustment instruction as the object text, and identifying the text related to the adjustment parameter of the sound effect category in the sound effect adjustment instruction as the processing parameter text includes: Obtain a preset regular expression; wherein, the regular expression is used to find a preset object and a preset processing parameter in the text; Match the sound effect adjustment instruction according to the regular expression to identify the text related to the sound effect category in the sound effect adjustment instruction as the object text and identify the text related to the adjustment parameter of the sound effect category in the sound effect adjustment instruction as the processing parameter text.
11. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 10.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 10.
13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 10.