An effecter debugging method and device based on semantic understanding
Patent Information
- Application Number
- CN202611172574.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-04
- Publication Date
- 2026-10-09
AI Technical Summary
[0005]本申请提供了一种基于语义理解的效果器调试方法及装置,解决了调音器调音效率较低的问题
1、过将用户的语音指令上传至云端进行语义解析并获取对应的DSP滤波系数,使得用户无需掌握音频参数知识即可通过自然语言描述完成效果器调试,降低了效果器的使用门槛;同时,云端处理完成后直接在本地对原始音频流进行卷积处理并输出试听片段,用户可在试听后确认是否完成调试,形成完整的语音交互调试闭环,提高了调试效率。
Smart Images

Figure CN122888933A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of effects tuning, specifically to an effects tuning method and apparatus based on semantic understanding. Background Technology
[0002] Effects pedals are important devices used in musical instrument performance to adjust and modify the tone. By adjusting the various parameters of the effects pedal, the performer can obtain the desired sound effect to meet the tone requirements of different performance scenarios.
[0003] Adjusting effects processors primarily relies on manual operation. Users need to adjust parameters such as equalization, dynamics, and reverb one by one using physical knobs or buttons on the device. This adjustment method requires users to have a certain level of knowledge of audio parameters in order to translate subjective listening impressions into specific parameter values. However, for ordinary users who lack professional knowledge, it is often difficult to accurately determine which parameters need to be adjusted and to what extent. Moreover, the adjustment process requires repeated trial and error, resulting in a high barrier to entry and low efficiency.
[0004] Therefore, there is an urgent need for an effect debugging method and device based on semantic understanding. Summary of the Invention
[0005] This application provides a semantic understanding-based effects tuning method and apparatus, which solves the problem of low tuning efficiency of tuners.
[0006] This application provides a semantic understanding-based effects debugging method in its first aspect. The method includes: in response to a debugging start command, acquiring the user's debugging voice and the original audio stream, and uploading the debugging voice to a cloud server; acquiring DSP filtering coefficients issued by the cloud server, wherein the DSP filtering coefficients are obtained by: a semantic parsing module performing speech recognition and intent understanding on the debugging voice to obtain structured control tags, and performing mapping transformation on the structured control tags through a mapper and a coefficient solving module to obtain DSP filtering coefficients; the cloud server including a semantic parsing module, a mapper, and a coefficient solving module; performing convolution processing on the original audio stream according to the DSP filtering coefficients to obtain a debugging audio stream; playing a first preset duration listening segment of the debugging audio stream through a speaker device, while continuously judging whether the user has issued a debugging completion command within a second preset duration; if a debugging completion command exists, then confirming that the original audio stream has been debugged.
[0007] By adopting the above technical solution, the user's voice commands are uploaded to the cloud for semantic parsing and the corresponding DSP filtering coefficients are obtained. This allows users to complete the effects debugging through natural language description without having to master audio parameter knowledge, thus lowering the barrier to entry for using effects. At the same time, after cloud processing, the original audio stream is directly convolved locally and a preview clip is output. Users can confirm whether the debugging is complete after listening, forming a complete voice interaction debugging loop and improving debugging efficiency.
[0008] Optionally, after responding to the debug start command, the method further includes: determining whether an audio source is connected; if an audio source is connected, obtaining a first test audio from the audio source and playing the first test audio through a speaker; if an audio source is not connected, determining whether automatic playback conditions are met; if automatic playback conditions are met, obtaining a second test audio from the local cache and playing the second test audio through a speaker, wherein the automatic playback conditions are having permission to automatically play historical audio and having historically played audio stored in the local cache, and the second test audio being the audio most recently played in the historically played audio; if the automatic playback conditions are not met, playing a preset third test audio through a speaker.
[0009] By adopting the above technical solution, after debugging starts, it automatically determines whether an audio source is currently connected. If an audio source is connected, it directly obtains external audio as a test reference. If not connected, it prioritizes reusing the most recently played historical audio in the local cache. Only when none of the conditions are met will the preset audio be used. This allows users to obtain a suitable debugging reference benchmark without manually selecting test audio, reducing the number of steps in the debugging preparation process.
[0010] Optionally, the debug start command is obtained in the following ways: in response to the user's operation of opening the effect, the debug start command is obtained to continuously detect the input of debug speech and raw audio stream for a first preset time period; the ambient noise level is continuously acquired and monitored, and when the ambient noise level is lower than a preset noise threshold, the system automatically switches to wake-up-free mode for a second preset time period, and wake-up-free mode is regarded as receiving the debug start command.
[0011] By adopting the above technical solution, it is possible to directly enter the debugging state by physically opening the effects unit, while continuously monitoring the ambient noise level. When the environment is quiet, it automatically switches to wake-up-free mode to receive debugging commands, so that users can easily start the debugging process in different usage scenarios, taking into account both operational certainty and usage flexibility.
[0012] Optionally, after the semantic parsing module performs speech recognition and intent understanding on the debugging speech, the method further includes: determining whether there is a non-parametric descriptive intent in the debugging speech; if there is a non-parametric descriptive intent, obtaining the audio template corresponding to the non-parametric descriptive intent from the preset audio template library, so as to obtain structured control tags based on the audio template; the non-parametric descriptive intent includes scene description and music style description.
[0013] By adopting the above technical solution, when the user's debugging voice contains non-parametric intentions such as scene description or music style description, the ambiguous style requirements are transformed into structured control tags that can be processed later through a preset audio template library. This allows the user to obtain the desired style effect configuration without giving specific parameter instructions, further reducing the requirements for the user's professional knowledge.
[0014] Optionally, after confirming that the original audio stream has been debugged, the method further includes: obtaining the audio style features of the original audio stream, wherein the audio style features include at least one of spectral features, rhythmic features and harmonic features; constructing the correspondence between the audio style features, structured control labels and DSP filtering coefficients, and storing the correspondence in an audio preference library.
[0015] By adopting the above technical solution, after debugging, the stylistic features such as the spectrum features, rhythm features, and harmony features of the current audio are extracted, and a corresponding relationship is established between them and the structured control labels and DSP filtering coefficients and stored in the audio preference library. This allows the generation of parameters by using historical preference data during subsequent debugging, and the personalization and accuracy of parameter recommendations are gradually improved as the number of times they are used increases.
[0016] Optionally, the semantic parsing module performs speech recognition and intent understanding on the debugging speech to obtain structured control labels. Specifically, this includes: performing speech recognition on the debugging speech to obtain recognized text; classifying the recognized text by intent to obtain target parameter category, adjustment polarity, and adjustment amplitude. The target parameter category is the type of effect parameter that the user intends to adjust. The effect parameter type includes at least one of equalization parameters, dynamic processing parameters, and reverberation parameters. The adjustment polarity is the adjustment direction of the target parameter category. The adjustment direction includes enhancement and attenuation. The adjustment amplitude is the degree level corresponding to the adjustment direction. A structured control label containing the target parameter category, adjustment polarity, and adjustment amplitude is generated.
[0017] By adopting the above technical solution, the user's natural language instructions are parsed into structured control labels with three dimensions: target parameter category, adjustment polarity, and adjustment range by classifying the recognized text. This clarifies the three core elements of the user's intention: what to adjust, in which direction to adjust, and by how much. This provides a precise input basis for subsequent mapping transformations and avoids parameter deviations caused by direct mapping of fuzzy semantics.
[0018] Optionally, the structured control labels are mapped and transformed using a mapper and a coefficient solving module to obtain DSP filter coefficients. Specifically, the mapper matches the target frequency response curve corresponding to the structured control label in a preset acoustic curve library; the coefficient solving module obtains the current filter coefficient state of the effect unit, inputs the filter coefficient state into the acoustic simulation model, and obtains the current frequency response curve; the coefficient solving module iteratively solves the filter coefficient state with the optimization objective of minimizing the deviation between the current frequency response curve and the target frequency response curve to obtain the DSP filter coefficients.
[0019] By adopting the above technical solution, the structured control labels are matched to the target frequency response curve through the mapper, and then the coefficient solving module calculates the current frequency response curve based on the acoustic simulation model. The DSP filter coefficients are obtained by iteratively solving with the goal of minimizing the deviation between the two. The user's subjective listening description is converted into accurate filter parameters based on the acoustic optimization target, which avoids the problem of inconsistent listening effect and user expectations caused by simple table lookup mapping and improves the accuracy of parameter generation.
[0020] In a second aspect, this application provides an effector debugging device based on semantic understanding, the device including an acquisition unit and a processing unit;
[0021] The acquisition unit is used to respond to the debugging start command, acquire the user's debugging voice and the raw audio stream, and upload the debugging voice to the cloud server; it is also used to acquire the DSP filtering coefficients issued by the cloud server. The DSP filtering coefficients are obtained in the following way: the semantic parsing module performs speech recognition and intent understanding on the debugging voice to obtain structured control tags, and the structured control tags are mapped and transformed through the mapper and coefficient solving module to obtain the DSP filtering coefficients. The cloud server includes the semantic parsing module, the mapper, and the coefficient solving module.
[0022] The processing unit is used to perform convolution processing on the original audio stream according to the DSP filtering coefficients to obtain the debugging audio stream; it is also used to play a test segment of the debugging audio stream for a first preset duration through a speaker device, and at the same time continuously determine whether the user issues a debugging completion command within a second preset duration; it is also used to confirm that the original audio stream has been debugged if a debugging completion command exists.
[0023] Optionally, the acquisition unit is used to determine whether an audio source is connected. If an audio source is connected, the first test audio is acquired from the audio source and played through the speaker. If no audio source is connected, the unit determines whether the automatic playback conditions are met. If the automatic playback conditions are met, the second test audio is acquired from the local cache and played through the speaker. The automatic playback conditions are having permission to automatically play historical audio and having historical audio stored in the local cache. The second test audio is the audio among the historical audio that is closest in playback time to the present. The processing unit is used to play a preset third test audio through the speaker if the automatic playback conditions are not met.
[0024] Optionally, the acquisition unit is used to respond to the user's operation of opening the effects unit, acquire the debugging start command, continuously detect the input of debugging voice and raw audio stream, and the detection duration is a first preset time period; continuously acquire and monitor the ambient noise level, and when the ambient noise level is lower than a preset noise threshold, automatically switch to wake-up-free mode and continue for a second preset time period, and wake-up-free mode is regarded as receiving the debugging start command.
[0025] Optionally, the processing unit is used to determine whether there is a non-parametric descriptive intent in the debugging speech. If there is a non-parametric descriptive intent, the audio template corresponding to the non-parametric descriptive intent is obtained from the preset audio template library, so as to obtain the structured control label based on the audio template. The non-parametric descriptive intent includes scene description and music style description.
[0026] Optionally, the acquisition unit is used to acquire the audio style features of the original audio stream, which include at least one of spectral features, rhythmic features, and harmonic features; the processing unit is used to construct the correspondence between the audio style features, structured control labels, and DSP filtering coefficients, and store the correspondence in the audio preference library.
[0027] Optionally, the processing unit performs speech recognition on the debugging speech to obtain recognized text, performs intent classification on the recognized text, and obtains the target parameter category, adjustment polarity, and adjustment amplitude. The target parameter category is the type of effect parameter that the user intends to adjust. The effect parameter type includes at least one of equalization parameter, dynamic processing parameter, and reverberation parameter. The adjustment polarity is the adjustment direction of the target parameter category. The adjustment direction includes enhancement and attenuation. The adjustment amplitude is the degree level corresponding to the adjustment direction. The processing unit generates a structured control label containing the target parameter category, adjustment polarity, and adjustment amplitude.
[0028] Optionally, the acquisition unit is used by the coefficient solving module to acquire the current filter coefficient state of the effect, input the filter coefficient state into the acoustic simulation model, and obtain the current frequency response curve; the processing unit is used by the mapper to match the target frequency response curve corresponding to the structured control label in the preset acoustic curve library; the coefficient solving module takes minimizing the deviation between the current frequency response curve and the target frequency response curve as the optimization objective, and iteratively solves the filter coefficient state to obtain the DSP filter coefficient.
[0029] In a third aspect, this application provides an electronic device including a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the first aspect or any possible implementation of the first aspect.
[0030] In a fourth aspect, this application provides a computer-readable storage medium storing a computer program, which is executed by a processor as described in the first aspect or any possible implementation thereof.
[0031] In summary, one or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: 1. By uploading users' voice commands to the cloud for semantic parsing and obtaining the corresponding DSP filtering coefficients, users can complete effects debugging through natural language description without needing to master audio parameter knowledge, thus lowering the barrier to entry for using effects. At the same time, after cloud processing, the original audio stream is directly convolved locally and a preview clip is output. Users can confirm whether the debugging is complete after listening, forming a complete voice interaction debugging loop and improving debugging efficiency.
[0032] 2. When the system recognizes that the user's debugging voice contains non-parametric intentions such as scene descriptions or music style descriptions, it transforms the vague style requirements into structured control tags that can be processed later through a preset audio template library. This allows the user to obtain the desired style without having to give specific parameter instructions, further reducing the requirements for the user's professional knowledge.
[0033] 3. After debugging, extract the stylistic features such as spectral features, rhythmic features, and harmonic features of the current audio, and establish a correspondence between them and the structured control labels and DSP filtering coefficients. Store them in the audio preference library so that historical preference data can be used to assist in parameter generation during subsequent debugging. As the number of times it is used increases, the personalization and accuracy of parameter recommendations will be gradually improved. Attached Figure Description
[0034] Figure 1 This is a flowchart illustrating an effect debugging method based on semantic understanding provided in an embodiment of this application.
[0035] Figure 2 This is a schematic diagram of the structure of an effects debugging device based on semantic understanding provided in an embodiment of this application.
[0036] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0037] Explanation of reference numerals in the attached drawings: 201, acquisition unit; 202, processing unit; 300, electronic device; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. Detailed Implementation
[0038] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0039] In the description of the embodiments of this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.
[0040] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0041] Adjusting effects processors primarily relies on manual operation, requiring users to manually adjust parameters such as equalization, dynamics, and reverb using physical knobs or buttons on the device. This method demands a certain level of audio parameter knowledge from the user to translate subjective listening impressions into specific parameter values. However, for ordinary users lacking professional expertise, accurately determining which parameters need adjustment and to what extent is often difficult. Furthermore, the adjustment process involves repeated trial and error, resulting in a high barrier to entry and low efficiency. Therefore, this embodiment provides an effects processor adjustment method and apparatus based on semantic understanding.
[0042] This application provides a semantic understanding-based effects debugging method for reference. Figure 1 , Figure 1 This is a flowchart illustrating a semantic understanding-based effects debugging method provided in an embodiment of this application, applied to effects processors. The method includes steps S101 to S105.
[0043] S101: In response to the debug start command, obtain the user's debug voice and the original audio stream, and upload the debug voice to the cloud server.
[0044] In the above steps, when using the effects unit, the user can initiate the debugging process by speaking a preset wake-up word or pressing a specific button on the effects unit. Upon receiving the debugging start command, the effects unit activates its microphone to capture the user's debugging voice, and simultaneously acquires the currently input raw audio stream through the instrument input interface. The raw audio stream can be analog or digital audio signals input in real-time via an audio cable from instruments such as guitars and basses. The effects unit compresses and encodes the captured debugging voice and uploads it to the cloud server via its built-in wireless communication module. Debugging voice commands, for example, might be phrases like "Please soften the high notes" or "Give me a rock tone suitable for stage performance." Along with uploading the voice data, the effects unit can also send the current device identifier and user identity information so that the cloud server can recognize and access the corresponding user preference data.
[0045] In one possible implementation, after responding to the debug start command, the method further includes: determining whether an audio source is connected; if an audio source is connected, obtaining a first test audio from the audio source and playing the first test audio through a speaker; if an audio source is not connected, determining whether automatic playback conditions are met; if automatic playback conditions are met, obtaining a second test audio from a local cache and playing the second test audio through a speaker, wherein the automatic playback conditions are having permission to automatically play historical audio and having historically played audio stored in the local cache, and the second test audio being the audio most recently played among the historically played audio; if the automatic playback conditions are not met, playing a preset third test audio through a speaker.
[0046] Specifically, after debugging begins, the effect processor first checks if an external audio source is connected, such as a mobile phone, player, or computer connected via audio cable or Bluetooth. If an external audio source is detected, it directly obtains the first test audio from that source and plays it through the speaker. The first test audio can be accompaniment music or a reference track currently playing on the external device, which the user can use to adjust their instrument's tone during debugging. If no external audio source is connected, the effect processor further checks if the automatic playback conditions are met. The automatic playback conditions are: the user has authorized the automatic playback of historical audio, and historical playback audio is stored in the local cache. Historical playback audio refers to audio files that the user has played previously while using the effect processor. When the automatic playback conditions are met, the effect processor selects the most recent historical playback audio from the local cache as the second test audio. When the automatic playback conditions are not met, the effect processor plays a factory-preset third test audio, such as a sweep signal covering multiple frequency bands, with a sweep range covering 20Hz to 20kHz, so that the user can fully perceive the changes in each frequency band when adjusting equalizer parameters. Through the above three-level judgment, users can automatically obtain suitable reference audio without manually selecting test audio after debugging starts.
[0047] In one possible implementation, the debug start command is obtained in the following way: in response to the user's operation of opening the effects unit, the debug start command is obtained to continuously detect the input of debug speech and raw audio stream for a first preset time period; the ambient noise level is continuously acquired and monitored, and when the ambient noise level is lower than a preset noise threshold, the system automatically switches to wake-up-free mode for a second preset time period, and wake-up-free mode is regarded as receiving the debug start command.
[0048] Specifically, the debugging start command is obtained in two ways. The first method is physical switch triggering: After the user turns on the effect's power switch or presses the dedicated debugging button, the effect immediately receives the debugging start command and continuously detects microphone and instrument input for a first preset time period, such as 30 seconds. During this time, the user can issue a debugging voice at any time. If no valid voice is detected after the first preset time period, the effect automatically exits the debugging standby state. The second method is wake-up-free mode: The effect continuously monitors the current ambient noise level through its built-in microphone, expressed in sound pressure level (dB). When the ambient noise level is consistently below a preset noise threshold, such as below 40dB, it indicates a quiet environment, and the effect automatically switches to wake-up-free mode. In this mode, the debugging start command is considered received, and the process continues for a second preset time period, such as 60 seconds. The user can directly state their debugging needs without physical operation or uttering a wake-up word. Wake-up-free mode is particularly suitable for scenarios where users are playing instruments with both hands and cannot easily operate the device. If the ambient noise level rises above the preset noise threshold in wake-free mode, the effect will exit wake-free mode and return to normal standby mode to avoid accidental triggering in noisy environments.
[0049] S102. Obtain the DSP filtering coefficients issued by the cloud server. The DSP filtering coefficients are obtained in the following way: the semantic parsing module performs speech recognition and intent understanding on the debugging speech to obtain structured control tags. The structured control tags are then mapped and transformed by the mapper and coefficient solving module to obtain the DSP filtering coefficients. The cloud server includes a semantic parsing module, a mapper, and a coefficient solving module.
[0050] In the above steps, after receiving the debugging voice, the cloud server's internal semantic parsing module automatically performs speech recognition and intent understanding, parsing the user's natural language commands into structured debugging intent information, i.e., structured control tags. The structured control tags extract three dimensions of information from the user's voice: "what parameters need to be adjusted," "in which direction to adjust," and "to what extent to adjust." Then, the cloud server further processes the structured control tags through a mapper and coefficient solving module. Combining this with the current parameter state of the effects unit, it transforms the debugging intent into a set of DSP filter coefficients that can be directly used for audio processing, and sends these DSP filter coefficients to the effects unit. DSP filter coefficients are a set of numerical parameters describing the audio filtering effect, specifically including filter type, center frequency, gain value, and quality factor (Q value). For example, a typical set of DSP filter coefficients might include a peak filter with a center frequency of 2kHz, a gain of -3dB, and a Q value of 0.8, indicating attenuation processing around 2kHz. Another example is a set of DSP filter coefficients used to achieve treble enhancement, which might include a high-shelf filter with a cutoff frequency of 6kHz and a gain of +4dB. When the cloud server sends DSP filter coefficients, it can encapsulate them in a structured data format. The effect processor can then parse the data and load it into the DSP processing unit without requiring any intermediate conversion by the user.
[0051] In one possible implementation, the semantic parsing module performs speech recognition and intent understanding on the debugging speech to obtain structured control labels. Specifically, this includes: performing speech recognition on the debugging speech to obtain recognized text; classifying the recognized text by intent to obtain target parameter category, adjustment polarity, and adjustment amplitude. The target parameter category is the type of effect parameter that the user intends to adjust, and the effect parameter type includes at least one of equalization parameter, dynamic processing parameter, and reverberation parameter. The adjustment polarity is the adjustment direction of the target parameter category, and the adjustment direction includes enhancement and attenuation. The adjustment amplitude is the degree level corresponding to the adjustment direction. A structured control label containing the target parameter category, adjustment polarity, and adjustment amplitude is then generated.
[0052] Specifically, the semantic parsing module performs speech recognition and intent understanding on debugging speech as follows: First, the debugging speech is decoded using a pre-trained acoustic model and a language model. The acoustic model converts the speech signal into a phoneme sequence, and the language model converts the phoneme sequence into grammatically and semantically correct text, ultimately yielding the recognized text. Then, intent classification is performed on the recognized text, which can be achieved using a pre-trained intent classification model. The intent classification model takes the recognized text as input and outputs the corresponding target parameter category, modulation polarity, and modulation amplitude. The training process of the intent classification model includes: collecting a large number of debugging speech samples, labeling each sample with the corresponding target parameter category, modulation polarity, and modulation amplitude as training labels; inputting the labeled samples into the initial classification model for training, optimizing the model parameters using the cross-entropy loss function until the model converges. The trained intent classification model can automatically classify unlabeled user debugging speech. The target parameter category indicates the type of effect parameter adjusted by the user's intent, which can be classified as equalization parameters, dynamic processing parameters, or reverberation parameters. For example, if a user says "the bass is too heavy," the intent classification model outputs the target parameter category as the equalization parameter; if a user says "the sound isn't tight enough," the intent classification model outputs the target parameter category as the dynamic processing parameter. Adjustment polarity is divided into two categories: enhancement and attenuation. For example, "too heavy" corresponds to attenuation, and "not enough" corresponds to enhancement. Adjustment amplitude indicates the degree of adjustment, which can be preset to three levels: mild, moderate, and significant. For example, "too" can be identified by the model as the significant level, and "slightly" as the mild level. The semantic parsing module integrates the above three fields into a structured control label, which serves as the input for subsequent mapping transformations.
[0053] In one possible implementation, the structured control labels are mapped and transformed using a mapper and a coefficient solving module to obtain DSP filter coefficients. Specifically, the mapper matches the target frequency response curve corresponding to the structured control label in a preset acoustic curve library; the coefficient solving module obtains the current filter coefficient state of the effect unit, inputs the filter coefficient state into the acoustic simulation model, and obtains the current frequency response curve; the coefficient solving module iteratively solves the filter coefficient state with the optimization objective of minimizing the deviation between the current frequency response curve and the target frequency response curve to obtain the DSP filter coefficients.
[0054] Specifically, after receiving the structured control tag, the mapper matches it against a pre-defined acoustic curve library. This library stores multiple target frequency response curves, each corresponding to a combination of target parameter type, adjustment polarity, and adjustment amplitude. The target frequency response curve describes the desired shape of the spectral response, with frequency on the x-axis and relative gain on the y-axis. The pre-defined acoustic curve library is constructed by professional sound engineers who mark ideal frequency response curves corresponding to different levels and directions for various parameter adjustment needs, forming a standardized curve template library. For example, when the target parameter type is equalization, the adjustment polarity is attenuation, and the adjustment amplitude is significant, the corresponding target frequency response curve can exhibit a sloping shape in the low-frequency range, decreasing by 6dB per octave starting from 100Hz. The coefficient solving module obtains the current filter coefficient status of the effects unit, including parameters such as the current filter type, center frequency, gain, and Q value. The current filter coefficient status is input into a pre-built acoustic simulation model, which calculates the overall frequency response by cascading the transfer functions of each filter. The cascading transfer function involves multiplying the gain values of the amplitude-frequency response of each filter at the corresponding frequency points to obtain the overall frequency response under the current filter coefficient state, i.e., the current frequency response curve. The current frequency response curve and the target frequency response curve are represented using the same frequency and gain coordinates. The coefficient solving module quantifies the deviation between the current frequency response curve and the target frequency response curve into a loss function value. The loss function is defined as the sum of the squares of the gain differences between the two curves at discrete frequency points. The coefficient solving module uses an iterative optimization algorithm to adjust the filter coefficient state with the goal of minimizing the loss function value. The iterative optimization algorithm can be the gradient descent method. In each iteration, the partial derivative of the loss function with respect to each filter coefficient is calculated, and each filter coefficient is updated along the descent direction of the partial derivative until the loss function value converges to below the preset convergence threshold. The converged filter coefficients are the final generated DSP filter coefficients. These DSP filter coefficients enable the frequency response output of the effect unit to approximate the target frequency response curve as closely as possible, thereby meeting the user's auditory needs expressed through speech.
[0055] In one possible implementation, after the semantic parsing module performs speech recognition and intent understanding on the debugging speech, the method further includes: determining whether there is a non-parametric descriptive intent in the debugging speech; if there is a non-parametric descriptive intent, obtaining the audio template corresponding to the non-parametric descriptive intent from a preset audio template library, so as to obtain a structured control label based on the audio template; the non-parametric descriptive intent includes scene description and music style description.
[0056] Specifically, after performing speech recognition and intent understanding on the debugging speech, the semantic parsing module first determines whether it contains non-parametric descriptive intent. Non-parametric descriptive intent refers to situations where the user doesn't provide specific parameter adjustment instructions but instead expresses their needs using scene descriptions or music style descriptions, such as saying "a warm tone suitable for listening before bed" or "blues style suitable for street performances." When a non-parametric descriptive intent is detected, the semantic parsing module searches a preset audio template library. This library stores multiple audio templates, each containing a scene description or music style description and a set of preset audio processing parameters. The audio template library can be constructed by collecting parameter schemes configured by numerous professional sound engineers for different scenes and styles. For example, the audio template corresponding to "warm tone" might include a +3dB gain boost to the low-frequency range and a -2dB attenuation to frequencies above 2kHz; the audio template corresponding to "blues style" might include a slight overload effect and an EQ configuration that emphasizes the mid-frequency range. After the semantic parsing module matches the corresponding audio template, it extracts the parameter configuration information from the template and generates structured control tags based on this information. This enables subsequent mapping transformations to generate DSP filter coefficients that match the non-parametric description intent.
[0057] S103. Perform convolution processing on the original audio stream according to the DSP filtering coefficients to obtain the debug audio stream.
[0058] In the above steps, after receiving the DSP filter coefficients from the cloud server, the effects unit loads them into its internal DSP processing unit. The DSP processing unit uses the DSP filter coefficients as convolution kernels to perform real-time convolution operations on the currently input raw audio stream, obtaining a debug audio stream processed by the effects unit. Specifically, the convolution processing involves multiplying and adding the impulse response corresponding to the DSP filter coefficients with each sample point of the raw audio stream, thereby altering the spectral characteristics of the raw audio. When the DSP filter coefficients switch from one set of parameters to another, the DSP processing unit can perform cross-gradient processing on the two sets of filter coefficients at the moment of switching to avoid audio clicking or popping sounds caused by sudden parameter changes. It should be noted that, limited by the size of the effects unit, in this solution, the specific adjustment of the DSP filter coefficients is performed by the cloud server, and the effects unit directly receives and uses the adjusted DSP filter coefficients.
[0059] S104. Play a test segment of the debugging audio stream for a first preset duration through the speaker device, and at the same time continuously determine whether the user issues a debugging completion command within a second preset duration.
[0060] In the above steps, the effects unit outputs the debug audio stream to the connected speaker for playback for a first preset duration, such as a 5-second sample clip. The speaker can be the effects unit's built-in speaker, an external speaker, or headphones. The user listens to the sample clip to determine if the current effect meets their needs. While the sample clip is playing, the effects unit initiates a waiting window with a second preset duration, for example, continuously detecting whether the user issues a debug completion command via microphone within 10 seconds. The debug completion command can be a spoken confirmation phrase, such as "That's it," "It's okay," or "Save," or it can be a specific confirmation button pressed by the user. When detecting a debug completion command, the effects unit can perform local endpoint detection on the captured speech, extract valid speech segments, and upload them to the cloud server for semantic analysis to confirm whether the speech constitutes a debug completion command. If no speech activity is detected, the effects unit automatically exits the debug state or replays the sample clip after the waiting window ends.
[0061] S105. If a debugging completion command exists, then the debugging of the original audio stream is confirmed to be complete.
[0062] In the above steps, if a user issues a debugging completion command within the second preset time period, the effects unit confirms the end of the current debugging process and saves the current DSP filter coefficients to the local storage module as the user's effect parameters for later use. During saving, the DSP filter coefficients can be associated with information such as the current timestamp and structured control tags for easy viewing and management by the user later. If no debugging completion command is detected within the second preset time period, the effects unit can continue to collect the user's next debugging voice segment and repeat the above process, or it can automatically restore the original parameter state before the start of this debugging session to prevent the user from losing the original effect settings without confirmation.
[0063] In one possible implementation, after confirming that the original audio stream has been debugged, the method further includes: obtaining the audio style features of the original audio stream, the audio style features including at least one of spectral features, rhythmic features and harmonic features; constructing the correspondence between the audio style features, structured control labels and DSP filtering coefficients, and storing the correspondence in an audio preference library.
[0064] Specifically, once the debugging completion command is confirmed, the effects processor extracts the audio style features of the current original audio stream. These features include spectral characteristics, rhythmic characteristics, and harmonic characteristics. Spectral characteristics are obtained by calculating the spectral centroid and bandwidth of the audio signal; the spectral centroid reflects the overall brightness of the timbre, with higher centroid values indicating a brighter timbre. Rhythmic characteristics are obtained by detecting the starting point intensity and beat interval of the audio signal; the starting point intensity reflects the dynamics of the performance, and the beat interval reflects the tempo. Harmonious characteristics are obtained by calculating the distribution of a chromatogram; the concentration of the chromatogram distribution reflects the tonality of the music. The effects processor establishes a ternary correspondence between the extracted audio style features, the structured control tags generated during the current debugging process, and the final determined DSP filter coefficients, and uploads this correspondence to the cloud server, storing it in the audio preference library associated with the user. For example, if a user modifies a powerful, brightly spectral performance into a rock timbre with slight compression and reverb during a debugging session, the system associates and stores the spectral and rhythmic characteristics of that audio with the structured control tag for "rock style" and the corresponding DSP filter coefficients. The data accumulated in the audio preference library can be used as a reference for the cloud server in subsequent debugging. When the user inputs an original audio stream with similar historical audio style characteristics again, the cloud server can first search for the matching correspondence in the audio preference library to help generate DSP filtering coefficients that are more in line with the user's personal preferences.
[0065] This application also provides a semantic understanding-based effects debugging device, see reference Figure 2 The device is an effects unit, which includes an acquisition unit 201 and a processing unit 202.
[0066] The acquisition unit 201 is used to respond to the debugging start command, acquire the user's debugging voice and the original audio stream, and upload the debugging voice to the cloud server; it is also used to acquire the DSP filtering coefficients issued by the cloud server. The DSP filtering coefficients are obtained in the following way: the semantic parsing module performs speech recognition and intent understanding on the debugging voice to obtain structured control tags, and performs mapping transformation on the structured control tags through the mapper and coefficient solving module to obtain the DSP filtering coefficients. The cloud server includes the semantic parsing module, the mapper and the coefficient solving module.
[0067] The processing unit 202 is used to perform convolution processing on the original audio stream according to the DSP filtering coefficients to obtain the debugging audio stream; it is also used to play a listening segment of the debugging audio stream for a first preset duration through a speaker device, and at the same time continuously judge whether the user issues a debugging completion command within a second preset duration; it is also used to confirm that the original audio stream has been debugged if a debugging completion command exists.
[0068] In one possible implementation, the acquisition unit 201 is used to determine whether an audio source is connected. If an audio source is connected, the first test audio is acquired from the audio source and played through the speaker. If no audio source is connected, the automatic playback condition is determined. If the automatic playback condition is met, the second test audio is acquired from the local cache and played through the speaker. The automatic playback condition is having permission to automatically play historical audio and having historical audio stored in the local cache. The second test audio is the audio among the historical audio that is closest in playback time to the present. The processing unit 202 is used to play a preset third test audio through the speaker if the automatic playback condition is not met.
[0069] In one possible implementation, the acquisition unit 201 is used to acquire a debug start command in response to the user's operation of opening the effects unit, to continuously detect the input of debug voice and raw audio stream, and the detection duration is a first preset time period; to continuously acquire and monitor the ambient noise level, and when the ambient noise level is lower than a preset noise threshold, to automatically switch to wake-up-free mode and continue for a second preset time period, and wake-up-free mode is regarded as receiving the debug start command.
[0070] In one possible implementation, the processing unit 202 is used to determine whether there is a non-parametric description intent in the debugging voice. If there is a non-parametric description intent, the processing unit 202 obtains the audio template corresponding to the non-parametric description intent from the preset audio template library, so as to obtain the structured control label based on the audio template. The non-parametric description intent includes scene description and music style description.
[0071] In one possible implementation, the acquisition unit 201 is used to acquire the audio style features of the original audio stream, the audio style features including at least one of spectral features, rhythm features and harmony features; the processing unit 202 is used to construct the correspondence between the audio style features, structured control labels and DSP filter coefficients, and store the correspondence in the audio preference library.
[0072] In one possible implementation, the processing unit 202 is used to perform speech recognition on the debugging speech to obtain recognized text, classify the recognized text according to intent, and obtain target parameter category, adjustment polarity, and adjustment amplitude. The target parameter category is the type of effect parameter that the user intends to adjust. The effect parameter type includes at least one of equalization parameter, dynamic processing parameter, and reverberation parameter. The adjustment polarity is the adjustment direction of the target parameter category. The adjustment direction includes enhancement and attenuation. The adjustment amplitude is the degree level corresponding to the adjustment direction. The processing unit 202 generates a structured control label containing the target parameter category, adjustment polarity, and adjustment amplitude.
[0073] In one possible implementation, the acquisition unit 201 is used by the coefficient solving module to acquire the current filter coefficient state of the effect unit, input the filter coefficient state into the acoustic simulation model, and obtain the current frequency response curve; the processing unit 202 is used by the mapper to match the target frequency response curve corresponding to the structured control label in the preset acoustic curve library; the coefficient solving module takes minimizing the deviation between the current frequency response curve and the target frequency response curve as the optimization objective, and iteratively solves the filter coefficient state to obtain the DSP filter coefficient.
[0074] It should be noted that the above embodiments of the apparatus are only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0075] This application also provides an electronic device. (See reference...) Figure 3 , Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 300 may include: at least one processor 301, at least one communication bus 302, at least one user interface 303, a network interface 304, and a memory 305.
[0076] The communication bus 302 is used to enable communication between these components.
[0077] The user interface 303 may include a display screen and a camera. Optionally, the user interface 303 may also include a standard wired interface and a wireless interface.
[0078] The network interface 304 may include standard wired interfaces and wireless interfaces (such as Wi-Fi interfaces).
[0079] The processor 301 may include one or more processing cores. The processor 301 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 305, and by calling data stored in memory 305. Optionally, the processor 301 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 301 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 301 and may be implemented as a separate chip.
[0080] The memory 305 may include random access memory (RAM) or read-only memory. Optionally, the memory 305 may include a non-transitory computer-readable storage medium. The memory 305 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. The memory 305 may also be at least one storage device located remotely from the aforementioned processor 301. (Refer to...) Figure 3 The memory 305, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for an effects debugging method based on semantic understanding.
[0081] exist Figure 3In the illustrated electronic device 300, the user interface 303 is mainly used to provide an input interface for the user and acquire user input data; while the processor 301 can be used to call an application program of a semantic understanding-based effects debugging method stored in the memory 305. When executed by one or more processors 301, the electronic device 300 performs one or more methods as described in the above embodiments. It should be noted that, for the foregoing method embodiments, for the sake of simplicity, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0082] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0083] This application also provides a computer-readable storage medium storing instructions. When executed by one or more processors, these instructions cause an electronic device to perform one or more of the methods described in the above embodiments.
[0084] In the various embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some service interface; the indirect coupling or communication connection between apparatuses or units may be electrical or other forms.
[0085] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0086] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0087] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, portable hard drives, magnetic disks, or optical disks.
[0088] The above description is merely an exemplary embodiment of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure.
[0089] This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are to be considered exemplary only, and the scope of this disclosure is defined by the claims.
Claims
1. A method for debugging effects based on semantic understanding, characterized in that, The method includes: In response to the debug start command, the system acquires the user's debug voice and the original audio stream, and uploads the debug voice to the cloud server. The DSP filtering coefficients issued by the cloud server are obtained in the following way: the semantic parsing module performs speech recognition and intent understanding on the debugging speech to obtain structured control tags, and the structured control tags are mapped and transformed through a mapper and a coefficient solving module to obtain the DSP filtering coefficients. The cloud server includes the semantic parsing module, the mapper and the coefficient solving module. The original audio stream is convolved according to the DSP filtering coefficients to obtain the debug audio stream; Play a preview segment of the debugging audio stream for a first preset duration using a speaker device, and continuously determine whether the user issues a debugging completion command within a second preset duration. If the debugging completion instruction exists, it confirms that the original audio stream has been debugged.
2. The method according to claim 1, characterized in that, Following the response to the debug start command, the method further includes: Determine whether an audio source is connected. If the audio source is connected, obtain the first test audio from the audio source and play the first test audio through the speaker device. If the audio source is not connected, determine whether the automatic playback conditions are met. If the automatic playback conditions are met, obtain the second test audio from the local cache and play the second test audio through the speaker. The automatic playback conditions are having permission to automatically play historical audio and having historical audio stored in the local cache. The second test audio is the audio among the historical audio that is closest to the current playback time. If the automatic playback conditions are not met, a preset third test audio will be played through the speaker device.
3. The method according to claim 1, characterized in that, The debug startup command is obtained in the following way: In response to the user's operation of opening the effects unit, the debug start command is obtained to continuously detect the input of the debug voice and the original audio stream for a first preset time period. The ambient noise level is continuously acquired and monitored. When the ambient noise level is lower than a preset noise threshold, the system automatically switches to wake-up-free mode and continues for a second preset time period. The wake-up-free mode is regarded as receiving the debugging start command.
4. The method according to claim 1, characterized in that, After the semantic parsing module performs speech recognition and intent understanding on the debugging speech, the method further includes: Determine whether there is a non-parametric descriptive intent in the debugging voice. If the non-parametric descriptive intent exists, obtain the audio template corresponding to the non-parametric descriptive intent from the preset audio template library, and obtain the structured control tag based on the audio template. The non-parametric descriptive intent includes scene description and music style description.
5. The method according to claim 1, characterized in that, After confirming that the original audio stream has been debugged, the method further includes: The audio style features of the original audio stream are obtained, and the audio style features include at least one of spectral features, rhythmic features, and harmonic features; Construct the correspondence between the audio style features, the structured control tags, and the DSP filtering coefficients, and store the correspondence in the audio preference library.
6. The method according to claim 1, characterized in that, The semantic parsing module performs speech recognition and intent understanding on the debugging speech to obtain structured control tags, specifically including: The debugging speech is subjected to speech recognition to obtain recognized text. The recognized text is then classified according to intent to obtain target parameter category, adjustment polarity, and adjustment amplitude. The target parameter category is the type of effect parameter that the user intends to adjust. The effect parameter type includes at least one of equalization parameter, dynamic processing parameter, and reverberation parameter. The adjustment polarity is the adjustment direction of the target parameter category. The adjustment direction includes enhancement and attenuation. The adjustment amplitude is the degree level corresponding to the adjustment direction. Generate the structured control label containing the target parameter category, the adjustment polarity, and the adjustment magnitude.
7. The method according to claim 1, characterized in that, The process of mapping and transforming the structured control labels using a mapper and a coefficient solving module to obtain the DSP filter coefficients specifically includes: The mapper matches the target frequency response curve corresponding to the structured control label in a preset acoustic curve library; The coefficient solving module obtains the current filter coefficient state of the effect unit, inputs the filter coefficient state into the acoustic simulation model, and obtains the current frequency response curve; The coefficient solving module takes minimizing the deviation between the current frequency response curve and the target frequency response curve as the optimization objective, and iteratively solves the filter coefficient state to obtain the DSP filter coefficient.
8. A semantic understanding-based effects debugging device, characterized in that, The device includes an acquisition unit (201) and a processing unit (202): The acquisition unit (201) is used to respond to the debugging start command, acquire the user's debugging voice and the original audio stream, and upload the debugging voice to the cloud server; The acquisition unit (201) is also used to acquire the DSP filtering coefficients issued by the cloud server. The DSP filtering coefficients are obtained in the following way: the semantic parsing module performs speech recognition and intent understanding on the debugging speech to obtain structured control tags, and the structured control tags are mapped and transformed through the mapper and the coefficient solving module to obtain the DSP filtering coefficients. The cloud server includes the semantic parsing module, the mapper and the coefficient solving module. The processing unit (202) is used to perform convolution processing on the original audio stream according to the DSP filtering coefficients to obtain a debug audio stream; The processing unit (202) is also used to play a listening segment of the debugging audio stream for a first preset duration through a speaker device, and at the same time continuously determine whether the user issues a debugging completion command within a second preset duration. The processing unit (202) is also configured to confirm that the original audio stream has been debugged if the debugging completion instruction exists.
9. An electronic device, characterized in that, The device includes a processor (301), a memory (305), a user interface (303), and a network interface (304). The memory (305) is used to store instructions. The user interface (303) and the network interface (304) are used to communicate with other devices. The processor (301) is used to execute the instructions stored in the memory (305) to cause the electronic device (300) to perform the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7 above.