Call voice restoration method and electronic equipment

By repairing the voice signal of the call in the time frequency domain multiple times, the voice repair model is used to solve the call quality problems caused by signal drop, and improve the call quality and user experience.

CN120356477APending Publication Date: 2025-07-22HONOR DEVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410058863.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-15
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In forests, mountains, basements and other environments, the call signal of electronic devices will drop sharply, resulting in poor sound quality, resulting in sound distortion, discontinuity, live noise, and even loss of words, affecting the normal call of users and leading to the loss of key information.

Method used

By performing multiple gradual repairs to the voice signal to be repaired in the time-frequency domain, a high-quality voice signal is generated using the preset voice repair model, including drift coefficient calculation, diffusion coefficient calculation, gradient calculation, inverse drift coefficient calculation, inverse diffusion coefficient calculation and prediction module.

Benefits of technology

It improves the call quality of electronic devices in environments with poor signal, improves the user's call experience, and ensures the complete transmission of key information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356477A_ABST
    Figure CN120356477A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of voice restoration, and provides a conversation voice restoration method and electronic equipment, and the method comprises the steps: obtaining a to-be-restored conversation voice signal and a sampling total step number N corresponding to the to-be-restored conversation voice signal in a conversation process of the electronic equipment; initializing the sampling time and the sampling step number n to determine an initial value of the sampling time and an initial value of the sampling step number n; performing feature transformation on the to-be-repaired call voice signal to determine a first time-frequency domain signal corresponding to the to-be-repaired call voice signal; according to the initial value of the sampling time, the initial value of the sampling step number n and the total sampling step number, repairing the first time-frequency domain signal for N times by using a preset voice repairing model so as to generate a repaired second time-frequency domain signal; and performing feature inverse transformation on the second time-frequency domain signal to generate a restored call voice signal. Therefore, the call quality of the electronic equipment in a poor-signal environment is improved, and the call experience of a user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of voice repair, and particularly relates to a method for repairing call voices, an electronic device, and a computer-readable storage medium. Background Art

[0002] For electronic devices with call functions such as mobile communication devices, when the electronic device is in an environment such as a forest, a mountain, or a basement, the signal will drop sharply, resulting in poor call sound quality. The communication system will generate voice distortion, discontinuity, charged noise, and even missing words, which can easily lead to the loss of key call information, causing the user to be unable to have a normal conversation and affecting the user experience. Summary of the Invention

[0003] The embodiments of this application provide a method for repairing call voices, an electronic device, and a computer-readable storage medium, which can solve the problem that when the electronic device is in an environment such as a forest, a mountain, or a basement, the signal drops sharply, resulting in poor call sound quality. The communication system will generate voice distortion, discontinuity, charged noise, and even missing words, which can easily lead to the loss of key call information, causing the user to be unable to have a normal conversation and affecting the user experience.

[0004] In a first aspect, the embodiments of this application provide a method for repairing call voices, including: during the call process of the electronic device, obtaining the call voice signal to be repaired and the total number of sampling steps N corresponding to the call voice signal to be repaired, where N is a positive integer; initializing the sampling time and the sampling step n to determine the initial value of the sampling time and the initial value of the sampling step n, where the initial value of the sampling time is a first preset value and the initial value of the sampling step n is 1; performing a feature transformation on the call voice signal to be repaired to determine the first time-frequency domain signal corresponding to the call voice signal to be repaired; according to the initial value of the sampling time, the initial value of the sampling step n, and the total number of sampling steps, using a preset voice repair model to perform N repairs on the first time-frequency domain signal to generate a repaired second time-frequency domain signal; performing an inverse feature transformation on the second time-frequency domain signal to generate a repaired call voice signal.

[0005] In this way, through the sampling time and the preset voice repair model, the call voice signal to be repaired is gradually repaired multiple times in the time-frequency domain to repair the low-quality call voice signal, thereby improving the quality of the call voice signal, improving the call quality of the electronic device in a poor signal environment, and improving the user's call experience.

[0006] In a possible implementation manner of the first aspect, the above-mentioned using a preset voice repair model to perform N repairs on the first time-frequency domain signal according to the initial value of the sampling time, the initial value of the sampling step, and the total number of sampling steps to generate a repaired second time-frequency domain signal includes:

[0007] Input the nth sampling time, the first time-frequency domain signal, and the (n - 1)th repaired time-frequency domain signal into a preset speech repair model to perform the nth repair on the first time-frequency domain signal to generate the nth repaired time-frequency domain signal, where n is an integer greater than or equal to 1 and less than or equal to N. The above (n - 1)th repaired time-frequency domain signal is the time-frequency domain signal generated after performing the (n - 1)th repair on the first time-frequency domain signal. When n = 1, the above nth sampling time is the initial value of the sampling time, and the above (n - 1)th repaired time-frequency domain signal is the first time-frequency domain signal;

[0008] Update the sampling time according to the sampling step number n and the total number of sampling steps N to generate the (n + 1)th sampling time;

[0009] Update the sampling step number to n + 1;

[0010] When the sampling step number is less than or equal to N, input the (n + 1)th sampling time, the first time-frequency domain signal, and the nth repaired time-frequency domain signal into a preset speech repair model to generate the (n + 1)th repaired time-frequency domain signal;

[0011] When the sampling step number is greater than N, determine the nth repaired time-frequency domain signal as the second time-frequency domain signal.

[0012] Optionally, in another possible implementation manner of the first aspect, the above updating the sampling time according to the sampling step number n and the total number of sampling steps N to generate the (n + 1)th sampling time includes:

[0013] Determine a candidate sampling time according to the sampling step number n and the total number of sampling steps N;

[0014] When the candidate sampling time is greater than a second preset value, determine the candidate sampling time as the (n + 1)th sampling time, where the second preset value is less than the first preset value;

[0015] When the candidate sampling time is less than or equal to the second preset value, determine the second preset value as the (n + 1)th sampling time.

[0016] Optionally, in yet another possible implementation manner of the first aspect, the above preset speech repair model includes a drift coefficient calculation module, a diffusion coefficient calculation module, a gradient calculation module, an inverse drift coefficient calculation module, an inverse diffusion coefficient calculation module, and a prediction module; correspondingly, the above inputting the nth sampling time, the first time-frequency domain signal, and the (n - 1)th repaired time-frequency domain signal into a preset speech repair model to perform the nth repair on the first time-frequency domain signal to generate the nth repaired time-frequency domain signal includes:

[0017] Input the nth sampling time, the first time-frequency domain signal, and the (n - 1)th repaired time-frequency domain signal into the drift coefficient calculation module to determine the nth drift coefficient corresponding to the first time-frequency domain signal;

[0018] Input the nth sampling time into the diffusion coefficient calculation module to determine the nth diffusion coefficient corresponding to the first time-frequency domain signal;

[0019] Input the nth sampling time, the first time-frequency domain signal, and the (n - 1)th repaired time-frequency domain signal into the gradient calculation module to determine the nth gradient corresponding to the first time-frequency domain signal;

[0020] Input the nth drift coefficient, the nth diffusion coefficient, and the nth gradient into the inverse drift coefficient calculation module to determine the nth inverse drift coefficient corresponding to the first time-frequency domain signal;

[0021] Determine the nth inverse diffusion coefficient according to the nth diffusion coefficient;

[0022] Obtain Gaussian noise;

[0023] Input the (n - 1)th repaired time-frequency domain signal, the nth inverse drift coefficient, the nth inverse diffusion coefficient, and the Gaussian noise into the prediction module to generate the nth repaired time-frequency domain signal.

[0024] Optionally, in another possible implementation manner of the first aspect, the above-mentioned preset speech repair model further includes a variance calculation module and a correction module; correspondingly, after the above-mentioned input of the (n - 1)th repaired time-frequency domain signal, the nth inverse drift coefficient, the nth inverse diffusion coefficient, and the Gaussian noise into the prediction module to generate the nth repaired time-frequency domain signal, it further includes:

[0025] Input the nth repaired time-frequency domain signal into the variance calculation module to determine the variance of the nth repaired time-frequency domain signal;

[0026] Input the nth repaired time-frequency domain signal and the variance of the nth repaired time-frequency domain signal into the correction module to correct the nth repaired time-frequency domain signal.

[0027] Optionally, in another possible implementation manner of the first aspect, the above-mentioned obtaining Gaussian noise includes:

[0028] Generate Gaussian noise according to a preset random seed.

[0029] Optionally, in another possible implementation manner of the first aspect, the above-mentioned method further includes:

[0030] Update the preset random seed when the electronic device restarts.

[0031] Second aspect, an embodiment of the present application provides a repair device for call voice, including: a first acquisition module, configured to acquire a call voice signal to be repaired and the total number of sampling steps N corresponding to the call voice signal to be repaired during a call of an electronic device, where N is a positive integer; an initialization module, configured to initialize the sampling time and the sampling step number n to determine the initial value of the sampling time and the initial value of the sampling step number n, where the initial value of the sampling time is a first preset value, and the initial value of the sampling step number n is 1; a feature transformation module, configured to perform feature transformation on the call voice signal to be repaired to determine a first time-frequency domain signal corresponding to the call voice signal to be repaired; a repair module, configured to perform N repairs on the first time-frequency domain signal by using a preset voice repair model according to the initial value of the sampling time, the initial value of the sampling step number n, and the total number of sampling steps to generate a repaired second time-frequency domain signal; a feature inverse transformation module, configured to perform feature inverse transformation on the second time-frequency domain signal to generate a repaired call voice signal.

[0032] In a possible implementation manner of the second aspect, the above-mentioned repair module includes:

[0033] A first repair unit, configured to input the nth sampling time, the first time-frequency domain signal, and the (n - 1)th repaired time-frequency domain signal into a preset voice repair model to perform the nth repair on the first time-frequency domain signal to generate the nth repaired time-frequency domain signal, where n is an integer greater than or equal to 1 and less than or equal to N, the above-mentioned (n - 1)th repaired time-frequency domain signal is the time-frequency domain signal generated after performing the (n - 1)th repair on the first time-frequency domain signal, when n = 1, the above-mentioned nth sampling time is the initial value of the sampling time, and the above-mentioned (n - 1)th repaired time-frequency domain signal is the first time-frequency domain signal;

[0034] A first update unit, configured to update the sampling time according to the sampling step number n and the total number of sampling steps N to generate the (n + 1)th sampling time;

[0035] A second update unit, configured to update the sampling step number to n + 1;

[0036] A second repair unit, configured to, when the sampling step number is less than or equal to N, input the (n + 1)th sampling time, the first time-frequency domain signal, and the nth repaired time-frequency domain signal into a preset voice repair model to generate the (n + 1)th repaired time-frequency domain signal;

[0037] A first determination unit, configured to, when the sampling step number is greater than N, determine the nth repaired time-frequency domain signal as the second time-frequency domain signal.

[0038] Optionally, in another possible implementation manner of the second aspect, the above-mentioned first update unit is specifically configured to:

[0039] Determine the candidate sampling time according to the sampling step number n and the total sampling step number N;

[0040] In the case where the candidate sampling time is greater than the second preset value, determine the candidate sampling time as the (n + 1)-th sampling time, where the second preset value is less than the first preset value;

[0041] In the case where the candidate sampling time is less than or equal to the second preset value, determine the second preset value as the (n + 1)-th sampling time.

[0042] Optionally, in another possible implementation manner of the second aspect, the above-mentioned preset voice repair model includes a drift coefficient calculation module, a diffusion coefficient calculation module, a gradient calculation module, an inverse drift coefficient calculation module, an inverse diffusion coefficient calculation module, and a prediction module; correspondingly, the above-mentioned first repair unit is specifically used for:

[0043] Input the n-th sampling time, the first time-frequency domain signal, and the (n - 1)-th repaired time-frequency domain signal into the drift coefficient calculation module to determine the n-th drift coefficient corresponding to the first time-frequency domain signal;

[0044] Input the n-th sampling time into the diffusion coefficient calculation module to determine the n-th diffusion coefficient corresponding to the first time-frequency domain signal;

[0045] Input the n-th sampling time, the first time-frequency domain signal, and the (n - 1)-th repaired time-frequency domain signal into the gradient calculation module to determine the n-th gradient corresponding to the first time-frequency domain signal;

[0046] Input the n-th drift coefficient, the n-th diffusion coefficient, and the n-th gradient into the inverse drift coefficient calculation module to determine the n-th inverse drift coefficient corresponding to the first time-frequency domain signal;

[0047] Determine the n-th inverse diffusion coefficient according to the n-th diffusion coefficient;

[0048] Obtain Gaussian noise;

[0049] Input the (n - 1)-th repaired time-frequency domain signal, the n-th inverse drift coefficient, the n-th inverse diffusion coefficient, and the Gaussian noise into the prediction module to generate the n-th repaired time-frequency domain signal.

[0050] Optionally, in another possible implementation manner of the second aspect, the above-mentioned preset voice repair model further includes a variance calculation module and a correction module; correspondingly, the above-mentioned first repair unit is further used for:

[0051] Input the n-th repaired time-frequency domain signal into the variance calculation module to determine the variance of the n-th repaired time-frequency domain signal;

[0052] Input the nth repaired time-frequency domain signal and the variance of the nth repaired time-frequency domain signal into the calibration module to calibrate the nth repaired time-frequency domain signal.

[0053] Optionally, in another possible implementation manner of the second aspect, the above first repair unit is further configured to:

[0054] Generate Gaussian noise according to a preset random seed.

[0055] Optionally, in another possible implementation manner of the second aspect, the above device further includes:

[0056] An update module, configured to update the preset random seed when the electronic device restarts.

[0057] In a third aspect, an embodiment of the present application provides an electronic device, including: one or more processors, and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, and the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the electronic device to execute the call voice repair method as described above.

[0058] In a fourth aspect, an embodiment of the present application provides a chip system, applied to an electronic device, the chip system includes one or more processors, and the one or more processors are used to call computer instructions to enable the electronic device to execute the call voice repair method as described above.

[0059] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, the computer-readable storage medium includes instructions, when the instructions run on an electronic device, enabling the electronic device to execute the call voice repair method as described above.

[0060] In a sixth aspect, an embodiment of the present application provides a computer program product, when the computer program product runs on an electronic device, enabling the electronic device to execute the call voice repair method as described above.

[0061] The technical effects obtained in the above second aspect, third aspect, fourth aspect, fifth aspect and sixth aspect are similar to the technical effects obtained by the corresponding technical means in the above first aspect and second aspect, and will not be elaborated here. Description of the Drawings

[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0063] Figure 1 is a schematic flowchart of a method for repairing call voice provided by an embodiment of the present application;

[0064] Figure 2 is a schematic flowchart of a method for repairing call voice provided by another embodiment of the present application;

[0065] Figure 3 is a structural block diagram of a preset voice repair module provided by an embodiment of the present application;

[0066] Figure 4 is a structural block diagram of another preset voice repair module provided by an embodiment of the present application;

[0067] Figure 5 is a training flowchart of a gradient calculation module provided by an embodiment of the present application;

[0068] Figure 6 is a schematic structural diagram of a call voice repair device provided by an embodiment of the present application;

[0069] Figure 7 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0070] The following describes in detail the call voice repair method, device, electronic device, storage medium, and computer program provided by the present application with reference to the accompanying drawings.

[0071] Please refer to Figure 1 , Figure 1 is a schematic flowchart of a call voice repair method provided by an embodiment of the present application. The method may include the following parts or all of the content:

[0072] Step 101, during a call of an electronic device, obtain a call voice signal to be repaired and the total number of sampling steps N corresponding to the call voice signal to be repaired, where N is a positive integer.

[0073] It should be noted that the call voice repair method of the embodiments of the present application can be applied to any voice repair scenario. For example, it can be applied to any call scenario to repair damaged call voices in a low-signal environment, thereby improving call quality.

[0074] Among them, the call voice signal to be repaired may refer to any voice signal received from the communication counterpart during a call of an electronic device.

[0075] Among them, the total number of sampling steps N may refer to the number of times of repairing the call voice signal to be repaired.

[0076] As a possible implementation, during a call on an electronic device, each received call voice signal can be determined as a call voice signal to be repaired, so as to repair all the received call voice signals during the call, thereby improving the call quality during the entire call process.

[0077] As a possible implementation, during a call on an electronic device, the communication signal of the electronic device can be detected to determine the communication signal strength of the electronic device. When the current communication signal strength of the electronic device is less than the signal strength threshold, it can be determined that the current communication quality of the electronic device is poor, that is, the currently received call voice signal may be of poor quality, which will affect the user's call quality. Therefore, the currently received call voice signal can be determined as a call voice signal to be repaired, so as to perform targeted repair on the low-quality call voice signals received when the communication signal strength of the electronic device is weak, so as to ensure the communication quality while minimizing the resource occupancy and call delay of voice repair.

[0078] As a possible implementation, during a call on an electronic device, the signal quality of the currently received call voice signal of the electronic device can be evaluated to determine the signal quality parameter of the currently received call voice signal. When the signal quality parameter of the currently received call voice signal meets the repair condition, it can be determined that the quality of the currently received call voice signal is poor. Therefore, the currently received call voice signal can be determined as a call voice signal to be repaired, so as to perform targeted repair on the call voice signals with low quality, so as to ensure the communication quality while minimizing the resource occupancy and call delay of voice repair.

[0079] For example, the above signal quality parameter can be the signal-to-noise ratio. Therefore, during a call on an electronic device, the signal-to-noise ratio of the currently received call voice signal of the electronic device can be determined, and when the signal-to-noise ratio of the currently received call voice signal is less than the signal-to-noise ratio threshold, the currently received call voice signal can be determined as a call voice signal to be repaired.

[0080] It should be noted that the above example is only exemplary and should not be regarded as a limitation to this application. In actual use, the type of call voice to be repaired can be determined according to actual needs and specific application scenarios, and the embodiments of this application do not make any limitations thereto.

[0081] As a possible implementation, the call voice repair method according to the embodiments of the present application can be applied to any call scenario of an electronic device. In any call scenario of the electronic device, the call voice signal to be repaired can be obtained in any of the foregoing manners to repair the call quality of the electronic device in any scenario. For example, satellite call scenarios, ordinary phone call scenarios, third-party software call scenarios (such as making calls through third-party social software), etc.

[0082] As a possible implementation, the call voice repair method according to the embodiments of the present application can also be applied to specific call scenarios where the signal quality is usually poor, such as satellite call scenarios. Thus, when it is determined that the electronic device is in a specific call scenario, the call voice signal to be repaired can be obtained in any of the foregoing manners; when the electronic device is in other call scenarios, the call voice signal to be repaired may not be obtained, so as to perform targeted repair on call scenarios where the signal quality is usually poor, while ensuring the call quality and minimizing the resource consumption of voice repair as much as possible.

[0083] It should be noted that the above-listed application scenarios are only exemplary and should not be regarded as a limitation to the present application. In actual use, the applicable call scenarios can be set according to actual needs, and the embodiments of the present application do not make any limitations thereto.

[0084] As a possible implementation, the total number of sampling steps N can be preset. Therefore, after the call voice signal to be repaired is obtained, the preset total number of sampling steps N can be directly obtained. For example, N can be set to 2, 3, 4, etc.

[0085] As a possible implementation, since the damaged degrees of the call voice signals to be repaired may be different, the appropriate total number of sampling steps N can also be determined in real time according to the damaged degree of the call voice signal to be repaired, so as to further improve the voice repair effect. For example, the signal-to-noise ratio of the voice signal to be repaired can be determined. When the signal-to-noise ratio of the call voice signal to be repaired is relatively high, it can be determined that the damaged degree of the call voice signal to be repaired is relatively low, and thus the total number of sampling steps N can be determined as a smaller value, so as to reduce the number of voice repair times while ensuring the voice repair quality and further reducing the computational complexity of voice repair; when the signal-to-noise ratio of the call voice signal to be repaired is relatively low, it can be determined that the damaged degree of the call voice signal to be repaired is relatively high, and thus the total number of sampling steps N can be determined as a larger value, so as to further improve the voice repair quality by increasing the number of repair times for the call voice signal to be repaired.

[0086] It should be noted that the above-listed methods for determining the total number of sampling steps N are only exemplary and should not be regarded as a limitation to this application. In actual use, the appropriate method for determining the total number of sampling steps N and the specific value of N can be selected according to actual needs and specific application scenarios, and the embodiments of this application do not make any limitations in this regard.

[0087] Step 102: Initialize the sampling time and the number of sampling steps n to determine the initial value of the sampling time and the initial value of the number of sampling steps n.

[0088] Among them, the initial value of the sampling time can be a first preset value, and the initial value of the number of sampling steps n can be 1.

[0089] Among them, the sampling time can be used to represent the current degree of repair of the call voice signal to be repaired.

[0090] Among them, the number of sampling steps can be used to represent the current number of repairs of the call voice signal to be repaired. For example, when n is 1, it means that the call voice signal to be repaired is currently being repaired for the first time; when n is 2, it means that the call voice signal to be repaired is currently being repaired for the second time.

[0091] In the embodiments of this application, before repairing the call voice signal to be repaired, the sampling time can be initialized to the first preset value, and the number of sampling steps can be initialized to 1.

[0092] For example, assuming that it is preset that the sampling time changes within the first preset range, the sampling time can be initialized to the maximum value of the first preset range, or the sampling time can be initialized to a value close to the maximum value of the first preset range. For example, if the first preset range is [0.003, 1], the sampling time can be initialized to 1, or the sampling time can also be initialized to a value close to 1, such as 0.99, 0.999, 0.9999, etc.

[0093] It should be noted that the above example is only exemplary and should not be regarded as a limitation to this application. In actual use, the initial value of the sampling time can be determined according to actual needs and specific application scenarios, and the embodiments of this application do not make any limitations in this regard.

[0094] Step 103: Perform feature transformation on the call voice signal to be repaired to determine the first time-frequency domain signal corresponding to the call voice signal to be repaired.

[0095] In the embodiments of this application, feature transformation can be performed on the call voice signal to be repaired to determine the frequency domain characteristics of the call voice signal to be repaired, so as to repair the call voice signal to be repaired in the time-frequency domain.

[0096] As a possible implementation, the voice signal of the call to be repaired can be subjected to short-time Fourier transform and amplitude transformation to generate the time-frequency domain signal of the voice signal of the call to be repaired.

[0097] Step 104: According to the initial value of the sampling time, the initial value of the sampling step number n, and the total number of sampling steps, use a preset voice repair model to repair the first time-frequency domain signal N times to generate a repaired second time-frequency domain signal.

[0098] Among them, the preset voice repair model can be a deep learning model that has been pre-trained for repairing damaged voice signals. As an example, the preset voice repair model can be a deep learning model based on a diffusion model and has learned the probability distributions of a large number of voice signals (i.e., the essential characteristics of voice signals), so that it can effectively repair low-quality voice signals and improve the reliability of voice repair.

[0099] In the embodiment of the present application, when the first time-frequency domain signal corresponding to the voice signal of the call to be repaired is repaired for the first time, the initial value of the sampling time and the first time-frequency domain signal can be input into the preset voice repair model to repair the first time-frequency domain signal for the first time. Then, the sampling time and the sampling step number can be updated, and the updated sampling time, the first time-frequency domain signal, and the repair result after the first repair can be input into the preset voice repair model for the second repair until the sampling step number n reaches the total number of sampling steps N, that is, it can be determined that the N repairs of the first time-frequency domain signal are completed, thus completing the repair of the voice signal of the call to be repaired.

[0100] As a possible implementation, the first time-frequency domain signal can be repaired N times in the following way:

[0101] Input the nth sampling time, the first time-frequency domain signal, and the (n - 1)th repaired time-frequency domain signal into the preset voice repair model to repair the first time-frequency domain signal for the nth time to generate the nth repaired time-frequency domain signal, where n is an integer greater than or equal to 1 and less than or equal to N, and the (n - 1)th repaired time-frequency domain signal is the time-frequency domain signal generated after the (n - 1)th repair of the first time-frequency domain signal. When n = 1, the nth sampling time is the initial value of the sampling time, and the (n - 1)th repaired time-frequency domain signal is the first time-frequency domain signal;

[0102] Update the sampling time according to the sampling step number n and the total number of sampling steps N to generate the (n + 1)th sampling time;

[0103] Update the sampling step number to n + 1;

[0104] When the number of sampling steps is less than or equal to N, the (n + 1)-th sampling time, the first time-frequency domain signal, and the n-th restored time-frequency domain signal are input into a preset voice restoration model to generate the (n + 1)-th restored time-frequency domain signal;

[0105] When the number of sampling steps is greater than N, the n-th restored time-frequency domain signal is determined as the second time-frequency domain signal.

[0106] It can be understood that, as Figure 2 described above, after obtaining the voice signal x(t) of the call to be restored, the sampling time s and the number of sampling steps n can be initialized (for example, initializing the sampling time s to 1 and the number of sampling steps n to 1), and the voice signal x(t) of the call to be restored is subjected to feature transformation to generate the first time-frequency domain signal X(f, t); when performing the first restoration on the first time-frequency domain signal, the sampling time s = 1 and the first time-frequency domain signal X(f, t) can be input into a preset voice restoration model to generate the first restored time-frequency domain signal; then, the value of the sampling time s can be updated according to the number of sampling steps n and the total number of sampling steps N, and the sampling time n is updated to 2; then, it is judged whether the sampling time n is greater than the total number of sampling steps N. If the sampling time n is less than or equal to the total number of sampling steps N, it can be determined that the N-time restoration of the first time-frequency domain signal has not been completed yet, and then the updated sampling time s, the first time-frequency domain signal, and the first restored time-frequency domain signal are input into a preset voice restoration model to perform the second restoration on the first time-frequency domain signal and generate the second restored time-frequency domain signal, and the above restoration process is continued to be repeated until it is determined that the sampling time n is greater than the total number of sampling steps N, then it is determined that the N-time restoration of the first time-frequency domain signal has been completed, and thus the restoration process of the first time-frequency domain signal can be ended.

[0107] As a possible implementation manner, when updating the sampling time s, the ratio of the number of sampling steps n to the total number of sampling steps N can be determined as the updated sampling time s.

[0108] As a possible implementation manner, the sampling time s can also be changed within a preset range, that is, in a possible implementation manner of the embodiment of the present application, the above update of the sampling time according to the number of sampling steps n and the total number of sampling steps N to generate the (n + 1)-th sampling time includes:

[0109] Determine a candidate sampling time according to the number of sampling steps n and the total number of sampling steps N;

[0110] When the candidate sampling time is greater than a second preset value, the candidate sampling time is determined as the (n + 1)-th sampling time, where the second preset value is less than the first preset value;

[0111] When the candidate sampling time is less than or equal to the second preset value, the second preset value is determined as the (n + 1)-th sampling time.

[0112] As an example, assume that the sampling time varies uniformly within the first preset range [a, b], where a is the second preset value, b is the third preset value, the third preset value is greater than the second preset value, and the first preset value is the third preset value or the first preset value is very close to the third preset value; then the sampling time s can be updated according to the following formula:

[0113] s = max(b - n / N, a)

[0114] For example, as Figure 2 shown, assume a = 0.003 and b = 1, then the sampling time s can be updated by the following formula:

[0115] s = max(1 - n / N, 0.003)

[0116] As a possible implementation, as Figure 3 shown, the preset voice repair model may include a drift coefficient calculation module, a diffusion coefficient calculation module, a gradient calculation module, an inverse drift coefficient calculation module, an inverse diffusion coefficient calculation module, and a prediction module. Then, inputting the n-th sampling time, the first time-frequency domain signal, and the (n - 1)-th repaired time-frequency domain signal into the preset voice repair model to perform the n-th repair on the first time-frequency domain signal to generate the n-th repaired time-frequency domain signal includes:

[0117] Inputting the n-th sampling time, the first time-frequency domain signal, and the (n - 1)-th repaired time-frequency domain signal into the drift coefficient calculation module to determine the n-th drift coefficient corresponding to the first time-frequency domain signal;

[0118] Inputting the n-th sampling time into the diffusion coefficient calculation module to determine the n-th diffusion coefficient corresponding to the first time-frequency domain signal;

[0119] Inputting the n-th sampling time, the first time-frequency domain signal, and the (n - 1)-th repaired time-frequency domain signal into the gradient calculation module to determine the n-th gradient corresponding to the first time-frequency domain signal;

[0120] Inputting the n-th drift coefficient, the n-th diffusion coefficient, and the n-th gradient into the inverse drift coefficient calculation module to determine the n-th inverse drift coefficient corresponding to the first time-frequency domain signal;

[0121] Determining the n-th inverse diffusion coefficient according to the n-th diffusion coefficient;

[0122] Obtaining Gaussian noise;

[0123] Input the (n - 1)-th repaired time-frequency domain signal, the n-th inverse drift coefficient, the n-th inverse diffusion coefficient, and Gaussian noise into the prediction module to generate the n-th repaired time-frequency domain signal.

[0124] Among them, the drift coefficient can be used to represent the difference between the first time-frequency domain signal and the repaired time-frequency signal.

[0125] Among them, the diffusion coefficient can be used to represent the repair strength when repairing the first time-frequency domain signal.

[0126] Among them, the gradient can be used to measure the probability distribution of the high-quality speech signal based on which the first time-frequency domain signal is repaired, can characterize the essential features of the high-quality speech signal, and provide a repair basis for the n-th repair of the first time-frequency domain signal.

[0127] Among them, the inverse drift coefficient can be used to characterize the regular information when converting the low-quality speech signal into a high-quality speech signal.

[0128] Among them, the Gaussian noise can be Gaussian noise with a mean of zero and a variance of 1.

[0129] As a possible implementation, the n-th drift coefficient can be determined according to the stochastic differential equation through the following formula:

[0130]

[0131] where f n is the n-th drift coefficient, X(f, t) is the first time-frequency domain signal, is the (n - 1)-th repaired time-frequency domain signal, s is the sampling time; when n = 1, that is, the first drift coefficient is 0.

[0132] As a possible implementation, a diffusion equation can be predefined, and the n-th sampling time can be substituted into the predefined diffusion equation to determine the n-th diffusion coefficient. For example, the predefined diffusion equation can be g(s)=2 s , where s is the sampling time.

[0133] As a possible implementation, in order to further improve the reliability of speech repair, the diffusion equation can also be determined in real time according to the damaged degree of the call speech signal to be repaired, and the n-th sampling time can be substituted into the diffusion equation to determine the n-th diffusion coefficient. It can be understood that the greater the damaged degree of the call speech signal to be repaired, the greater the repair strength can be used to repair the call speech signal to be repaired, so as to further improve the efficiency of speech repair and reduce the communication delay. Therefore, the signal-to-noise ratio of the call speech signal to be repaired can be determined, and the diffusion equation can be determined according to the signal-to-noise ratio of the call speech signal to be repaired.

[0134] For example, when the signal-to-noise ratio of the voice signal of the call to be repaired is relatively large, it can be determined that the damage degree of the voice signal of the call to be repaired is relatively small, so that the diffusion equation can be determined as g(s)=2 s ; when the signal-to-noise ratio of the voice signal of the call to be repaired is relatively small, it can be determined that the damage degree of the voice signal of the call to be repaired is relatively large, so that the diffusion equation can be determined as g(s)=3 s .

[0135] It should be noted that the above examples are only exemplary and should not be regarded as a limitation of this application. In actual use, the diffusion equation can be determined according to actual needs and specific application scenarios, and the embodiments of this application do not make any limitations thereto.

[0136] As a possible implementation manner, the gradient calculation module can be a pre-trained deep neural network, so that the nth sampling time, the first time-frequency domain signal, and the (n - 1)th repaired time-frequency domain signal can be input into the gradient calculation module to determine the nth gradient corresponding to the first time-frequency domain signal.

[0137] For example, the gradient calculation module can be a noise condition scoring network, a U-Net network, a convolutional neural network (CRN), a convolutional recurrent network (CRN), etc., but is not limited thereto.

[0138] As a possible implementation manner, the nth inverse drift coefficient can be determined by the following formula:

[0139] f′ n =-f n +g 2 (s)grad

[0140] where f′ n is the nth inverse drift coefficient, f n is the nth drift coefficient, g(s) is the nth diffusion coefficient, and grad is the nth gradient.

[0141] As a possible implementation manner, the nth diffusion coefficient can be determined as the nth inverse diffusion coefficient. That is, when the nth diffusion coefficient is g(s)=2 s , the nth inverse diffusion coefficient can be determined as 2 s .

[0142] As a possible implementation manner, Gaussian noise with a mean of zero and a variance of 1 can be generated randomly, and the nth repaired time-frequency domain signal can be generated by the following formula:

[0143]

[0144] Among them, is the nth repaired time-frequency domain signal, is the (n - 1)th repaired time-frequency domain signal, f′ n is the nth inverse drift coefficient, g(s) is the nth inverse diffusion coefficient, and z is Gaussian noise.

[0145] As a possible implementation, a random seed can be preset, and Gaussian noise can be generated according to the preset random seed.

[0146] As a possible implementation, the preset random seed can also be updated when the electronic device restarts to ensure the randomness of Gaussian noise.

[0147] Furthermore, since there may be certain errors in the repaired time-frequency domain signal generated by the prediction module, the repaired time-frequency domain signal generated by the prediction module can also be corrected to ensure that the repair direction of speech repair is correct and to further improve the reliability of speech repair. That is, in a possible implementation manner of the embodiment of the present application, as Figure 4 shown, the above-mentioned preset speech repair model can also include a variance calculation module and a correction module; correspondingly, after inputting the (n - 1)th repaired time-frequency domain signal, the nth inverse drift coefficient, the nth inverse diffusion coefficient, and Gaussian noise into the prediction module to generate the nth repaired time-frequency domain signal, it can also include:

[0148] Input the nth repaired time-frequency domain signal into the variance calculation module to determine the variance of the nth repaired time-frequency domain signal;

[0149] Input the nth repaired time-frequency domain signal and the variance of the nth repaired time-frequency domain signal into the correction module to correct the nth repaired time-frequency domain signal.

[0150] As a possible implementation, the variance of the nth repaired time-frequency domain signal can be calculated through a variance calculation formula, and the nth repaired time-frequency domain signal and the variance of the nth repaired time-frequency domain signal can be processed using annealed Langevin dynamics to correct the nth repaired time-frequency domain signal.

[0151] Step 105, perform an inverse feature transform on the second time-frequency domain signal to generate a repaired call voice signal.

[0152] Among them, the inverse feature transform can refer to the inverse operation of the feature transform mentioned in the foregoing steps, which can convert the time-frequency domain signal into a time domain signal.

[0153] In an embodiment of the present application, after determining the repaired second time-frequency domain signal, the inverse feature transformation can be performed on the second time-frequency domain signal to convert the second time-frequency domain signal to the time domain, thereby generating the repaired call voice signal, and the repaired call voice signal can be played through the electronic device to enable the user to make a normal call.

[0154] The training process of the gradient calculation module in the preset voice repair model in the embodiment of the present application is described below:

[0155] As Figure 5 shown, the gradient calculation module includes a damage simulation module, a sampling time generation module, a Gaussian noise generation module, a sample generation module, a deep neural network module, and a loss function calculation module.

[0156] Among them, the damage simulation module may include some operators, such as low-quality encoding operators, amplitude attenuation operators, phase distortion operators, etc. This module can randomly select several operators to perform damage simulation on the clean lossless voice signal and output the damaged low-quality voice signal.

[0157] The sampling time generation module does not require input information. The function of this module is to randomly generate a number between t_min and t_max, where t_min is the minimum value, such as 0.003, and t_max is the maximum value, such as 1. The sampling time generated by this module is output to the sample generation module.

[0158] The Gaussian noise generation module can generate Gaussian noise with a mean of zero and a variance of 1.

[0159] The sample generation module can generate samples according to x = μ + σz, where μ is the mean of the noise signal generated after feature transformation, σ is the standard deviation of the noise signal generated after feature transformation, and z is Gaussian noise. Repeating the foregoing process can generate multiple samples to generate a training sample set containing multiple samples. In actual use, the specific method of sample generation can be determined according to actual needs and specific application scenarios, and the embodiment of the present application does not limit this.

[0160] The deep neural network module can be a noise prediction model. After inputting the generated samples, sampling time, and the signal after feature transformation into the noise prediction model, this module can output the gradient of the samples.

[0161] The loss function calculation module can determine the loss value of this training according to the training result of each time, so as to update the parameters of the deep neural network module according to the loss value until the loss value meets the preset conditions, and then it can be determined that the deep neural network module has been trained, and the trained deep neural network is used as the gradient calculation module in the embodiment of the present application. The calculation formula of the loss function can be Among them, σ is the standard deviation of the sample, grad is the gradient of the sample, and z is Gaussian noise.

[0162] The call voice repair method provided by the embodiments of the present application performs multiple progressive repairs on the call voice signal to be repaired in the time-frequency domain through the sampling time and a preset voice repair model, so as to repair the low-quality call voice signal, improve the quality of the call voice signal, thereby improving the call quality of the electronic device in an environment with poor signals and enhancing the user's call experience.

[0163] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0164] Corresponding to the call voice repair method described in the foregoing embodiments, Figure 6 The structural block diagram of the call voice repair device provided by the embodiments of the present application is shown. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown.

[0165] Refer to Figure 6 , the device 60 includes:

[0166] The first acquisition module 61 is configured to acquire the call voice signal to be repaired and the total number of sampling steps N corresponding to the call voice signal to be repaired during the call process of the electronic device, where N is a positive integer;

[0167] The initialization module 62 is configured to initialize the sampling time and the sampling step number n to determine the initial value of the sampling time and the initial value of the sampling step number n. The initial value of the sampling time is a first preset value, and the initial value of the sampling step number n is 1;

[0168] The feature transformation module 63 is configured to perform feature transformation on the call voice signal to be repaired to determine the first time-frequency domain signal corresponding to the call voice signal to be repaired;

[0169] The repair module 64 is configured to perform N repairs on the first time-frequency domain signal by using a preset voice repair model according to the initial value of the sampling time, the initial value of the sampling step number n, and the total number of sampling steps to generate a repaired second time-frequency domain signal;

[0170] The feature inverse transformation module 65 is configured to perform feature inverse transformation on the second time-frequency domain signal to generate a repaired call voice signal.

[0171] In actual use, the call voice repair device provided by the embodiments of the present application can be configured in any electronic device to execute the foregoing call voice repair method.

[0172] The call voice repair device provided by the embodiment of the present application performs multiple step-by-step repairs on the call voice signal to be repaired in the time-frequency domain through the sampling time and a preset voice repair model, so as to repair the low-quality call voice signal, improve the quality of the call voice signal, thereby improving the call quality of the electronic device in an environment with poor signal, and enhancing the user's call experience.

[0173] In a possible implementation manner of the present application, the above-mentioned repair module 64 includes:

[0174] The first repair unit is configured to input the nth sampling time, the first time-frequency domain signal, and the (n - 1)th repaired time-frequency domain signal into a preset voice repair model to perform the nth repair on the first time-frequency domain signal, so as to generate the nth repaired time-frequency domain signal, where n is an integer greater than or equal to 1 and less than or equal to N. The above-mentioned (n - 1)th repaired time-frequency domain signal is the time-frequency domain signal generated after the (n - 1)th repair on the first time-frequency domain signal. When n = 1, the above-mentioned nth sampling time is the initial value of the sampling time, and the above-mentioned (n - 1)th repaired time-frequency domain signal is the first time-frequency domain signal;

[0175] The first update unit is configured to update the sampling time according to the sampling step number n and the total sampling step number N to generate the (n + 1)th sampling time;

[0176] The second update unit is configured to update the sampling step number to n + 1;

[0177] The second repair unit is configured to, when the sampling step number is less than or equal to N, input the (n + 1)th sampling time, the first time-frequency domain signal, and the nth repaired time-frequency domain signal into a preset voice repair model to generate the (n + 1)th repaired time-frequency domain signal;

[0178] The first determination unit is configured to, when the sampling step number is greater than N, determine the nth repaired time-frequency domain signal as the second time-frequency domain signal.

[0179] Further, in another possible implementation manner of the present application, the above-mentioned first update unit is specifically configured to:

[0180] Determine a candidate sampling time according to the sampling step number n and the total sampling step number N;

[0181] When the candidate sampling time is greater than the second preset value, determine the candidate sampling time as the (n + 1)th sampling time, where the second preset value is less than the first preset value;

[0182] When the candidate sampling time is less than or equal to the second preset value, determine the second preset value as the (n + 1)th sampling time.

[0183] Further, in another possible implementation manner of the present application, the above-mentioned preset voice repair model includes a drift coefficient calculation module, a diffusion coefficient calculation module, a gradient calculation module, an inverse drift coefficient calculation module, an inverse diffusion coefficient calculation module, and a prediction module; correspondingly, the above-mentioned first repair unit is specifically used for:

[0184] Input the nth sampling time, the first time-frequency domain signal, and the (n - 1)th repaired time-frequency domain signal into the drift coefficient calculation module to determine the nth drift coefficient corresponding to the first time-frequency domain signal;

[0185] Input the nth sampling time into the diffusion coefficient calculation module to determine the nth diffusion coefficient corresponding to the first time-frequency domain signal;

[0186] Input the nth sampling time, the first time-frequency domain signal, and the (n - 1)th repaired time-frequency domain signal into the gradient calculation module to determine the nth gradient corresponding to the first time-frequency domain signal;

[0187] Input the nth drift coefficient, the nth diffusion coefficient, and the nth gradient into the inverse drift coefficient calculation module to determine the nth inverse drift coefficient corresponding to the first time-frequency domain signal;

[0188] Determine the nth inverse diffusion coefficient according to the nth diffusion coefficient;

[0189] Obtain Gaussian noise;

[0190] Input the (n - 1)th repaired time-frequency domain signal, the nth inverse drift coefficient, the nth inverse diffusion coefficient, and the Gaussian noise into the prediction module to generate the nth repaired time-frequency domain signal.

[0191] Further, in another possible implementation manner of the present application, the above-mentioned preset voice repair model further includes a variance calculation module and a correction module; correspondingly, the above-mentioned first repair unit is further used for:

[0192] Input the nth repaired time-frequency domain signal into the variance calculation module to determine the variance of the nth repaired time-frequency domain signal;

[0193] Input the nth repaired time-frequency domain signal and the variance of the nth repaired time-frequency domain signal into the correction module to correct the nth repaired time-frequency domain signal.

[0194] Further, in another possible implementation manner of the present application, the above-mentioned first repair unit is further used for:

[0195] Generate Gaussian noise according to a preset random seed.

[0196] Further, in another possible implementation manner of the present application, the above-mentioned device 60 further includes:

[0197] An update module, configured to update a preset random seed when the electronic device restarts.

[0198] It should be noted that, for the information interaction, execution process, etc. between the above-mentioned devices / units, since they are based on the same concept as the method embodiments of this application, for their specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details will not be elaborated here.

[0199] Those skilled in the art can clearly understand that, for the sake of convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiments, and details will not be elaborated here.

[0200] To implement the above embodiments, this application also proposes an electronic device.

[0201] Figure 7 It is a schematic structural diagram of an electronic device according to an embodiment of this application.

[0202] See Figure 7, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. Among them, the sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0203] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than those illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0204] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0205] Among them, the controller may be the nerve center and command center of the electronic device 100. The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching instructions and executing instructions.

[0206] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory may hold the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from this memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0207] The wireless communication function of the electronic device 100 may be implemented by the antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modulation and demodulation processor, and baseband processor, etc.

[0208] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example, the antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.

[0209] The mobile communication module 150 can provide solutions for wireless communications such as 2G / 3G / 4G / 5G applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves through the antenna 1, filter, amplify, etc. the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor and convert it into electromagnetic waves through the antenna 1 for radiation. In some embodiments, at least some functional modules of the mobile communication module 150 may be provided in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be provided in the same device.

[0210] The modulation and demodulation processor may include a modulator and a demodulator. Among them, the modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. Subsequently, the demodulator transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, receiver 170B, etc.), or displays an image or video through the display screen 194. In some embodiments, the modulation and demodulation processor may be an independent device. In some other embodiments, the modulation and demodulation processor may be independent of the processor 110 and be provided in the same device as the mobile communication module 150 or other functional modules.

[0211] The wireless communication module 160 may provide solutions for wireless communications applied to the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite systems (GNSSs), frequency modulation (FM), near field communication (NFC), infrared (IR), and the like. The wireless communication module 160 may be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency-modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 may also receive signals to be sent from the processor 110, frequency-modulate them, amplify them, and convert them into electromagnetic waves through the antenna 2 for radiation.

[0212] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling electronic device 100 to communicate with a network and other devices via wireless communication technologies. The wireless communication technologies may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. GNSS may include Global Positioning System (GPS), Global Navigation Satellite System (GLONASS), BeiDou Navigation Satellite System (BDS), Quasi-Zenith Satellite System (QZSS), and / or Satellite Based Augmentation Systems (SBAS).

[0213] Internal memory 121 may be used to store computer-executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of electronic device 100 by running the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.). The data storage area may store data created during the use of electronic device 100 (such as audio data, a phone book, etc.). In addition, internal memory 121 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, a flash memory device, Universal Flash Storage (UFS), etc.

[0214] The electronic device 100 can implement audio functions such as music playback and recording through the audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor, etc.

[0215] The microphone 170C, also known as a "microphone" or "transmitter", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak close to the microphone 170C with their mouth to input the sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In some other embodiments, the electronic device 100 can be provided with two microphones 170C, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C to collect sound signals, reduce noise, identify the sound source, and implement functions such as directional recording.

[0216] It should be noted that for the implementation process and technical principle of the electronic device in this embodiment, refer to the foregoing explanation of the method for repairing call voices in the embodiments of the present application, which will not be elaborated here.

[0217] The embodiments of the present application also provide a chip system applied to an electronic device. The chip system includes one or more processors, and the one or more processors are used to call computer instructions to enable the electronic device to implement the steps in the above-mentioned method embodiments.

[0218] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium includes instructions that, when running on the electronic device, enable the electronic device to implement the steps in the above-mentioned method embodiments.

[0219] The embodiments of the present application also provide a computer program product that, when running on the electronic device, enables the electronic device to implement the steps in the above-mentioned method embodiments.

[0220] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the device / electronic device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0221] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not described in detail or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0222] In the above embodiments, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are proposed to thoroughly understand the embodiments of this application. However, those skilled in the art should understand that this application can also be implemented in other embodiments without these specific details. In other cases, the detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of this application.

[0223] It should be understood that when used in the specification and claims of this application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0224] It should also be understood that the term "and / or" used in the specification and claims of this application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0225] As used in the specification of this application and the appended claims, the term "if" can be construed, depending on the context, as "when", "once", "in response to determining", or "in response to detecting". Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be construed, depending on the context, to mean "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]".

[0226] In addition, in the description of the specification of this application and the appended claims, the terms "first", "second", "third", etc. are used only for descriptive distinction and should not be construed as indicating or implying relative importance.

[0227] Reference to "one embodiment" or "some embodiments" or the like described in the specification of this application means that a particular feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized.

[0228] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled artisans can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0229] In the embodiments provided in this application, it should be understood that the disclosed apparatus / electronic device and method can be implemented in other ways. For example, the apparatus / electronic device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.

[0230] The unit described as the separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0231] The above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A method for repairing call voice, characterized in that, Including: During a call of an electronic device, obtain a call voice signal to be repaired and the total number of sampling steps N corresponding to the call voice signal to be repaired, where N is a positive integer; Initialize the sampling time and the sampling step number n to determine the initial value of the sampling time and the initial value of the sampling step number n, where the initial value of the sampling time is a first preset value, and the initial value of the sampling step number n is 1; Perform a feature transformation on the call voice signal to be repaired to determine a first time-frequency domain signal corresponding to the call voice signal to be repaired; According to the initial value of the sampling time, the initial value of the sampling step number n, and the total number of sampling steps N, use a preset voice repair model to repair the first time-frequency domain signal N times to generate a repaired second time-frequency domain signal; Perform an inverse feature transformation on the second time-frequency domain signal to generate a repaired call voice signal.

2. The method according to claim 1, characterized in that The step of using the preset voice repair model to repair the first time-frequency domain signal N times according to the initial value of the sampling time, the initial value of the sampling step number n, and the total number of sampling steps N to generate a repaired second time-frequency domain signal includes: Input the nth sampling time, the first time-frequency domain signal, and the (n - 1)th repaired time-frequency domain signal into the preset voice repair model to perform the nth repair on the first time-frequency domain signal to generate the nth repaired time-frequency domain signal, where n is an integer greater than or equal to 1 and less than or equal to N, and the (n - 1)th repaired time-frequency domain signal is the time-frequency domain signal generated after performing the (n - 1)th repair on the first time-frequency domain signal. When n = 1, the nth sampling time is the initial value of the sampling time, and the (n - 1)th repaired time-frequency domain signal is the first time-frequency domain signal; Update the sampling time according to the sampling step number n and the total number of sampling steps N to generate the (n + 1)th sampling time; Update the sampling step number to n + 1; When the sampling step number is less than or equal to N, input the (n + 1)th sampling time, the first time-frequency domain signal, and the nth repaired time-frequency domain signal into the preset voice repair model to generate the (n + 1)th repaired time-frequency domain signal; When the sampling step number is greater than N, determine the nth repaired time-frequency domain signal as the second time-frequency domain signal.

3. The method according to claim 2, characterized in that, The step of updating the sampling time according to the sampling step number n and the total number of sampling steps N to generate the (n + 1)th sampling time includes: Determine a candidate sampling time according to the sampling step number n and the total number of sampling steps N; When the candidate sampling time is greater than a second preset value, determine the candidate sampling time as the (n + 1)th sampling time, where the second preset value is less than the first preset value; When the candidate sampling time is less than or equal to the second preset value, determine the second preset value as the (n + 1)th sampling time.

4. The method according to claim 2, wherein The preset voice repair model includes a drift coefficient calculation module, a diffusion coefficient calculation module, a gradient calculation module, an inverse drift coefficient calculation module, an inverse diffusion coefficient calculation module, and a prediction module. The step of inputting the nth sampling time, the first time-frequency domain signal into the preset voice repair model to perform the nth repair on the first time-frequency domain signal to generate the nth repaired time-frequency domain signal includes: Inputting the nth sampling time, the first time-frequency domain signal, and the (n - 1)th repaired time-frequency domain signal into the drift coefficient calculation module to determine the nth drift coefficient corresponding to the first time-frequency domain signal; Inputting the nth sampling time into the diffusion coefficient calculation module to determine the nth diffusion coefficient corresponding to the first time-frequency domain signal; Inputting the nth sampling time, the first time-frequency domain signal, and the (n - 1)th repaired time-frequency domain signal into the gradient calculation module to determine the nth gradient corresponding to the first time-frequency domain signal; Inputting the nth drift coefficient, the nth diffusion coefficient, and the nth gradient into the inverse drift coefficient calculation module to determine the nth inverse drift coefficient corresponding to the first time-frequency domain signal; Determining the nth inverse diffusion coefficient according to the nth diffusion coefficient; Obtaining Gaussian noise; Inputting the (n - 1)th repaired time-frequency domain signal, the nth inverse drift coefficient, the nth inverse diffusion coefficient, and the Gaussian noise into the prediction module to generate the nth repaired time-frequency domain signal.

5. The method according to claim 4, characterized in that, The preset voice repair model further includes a variance calculation module and a correction module. After inputting the (n - 1)th repaired time-frequency domain signal, the nth inverse drift coefficient, the nth inverse diffusion coefficient, and the Gaussian noise into the prediction module to generate the nth repaired time-frequency domain signal, it further includes: Inputting the nth repaired time-frequency domain signal into the variance calculation module to determine the variance of the nth repaired time-frequency domain signal; Inputting the nth repaired time-frequency domain signal and the variance of the nth repaired time-frequency domain signal into the correction module to correct the nth repaired time-frequency domain signal.

6. The method according to claim 4 or 5, characterized in that, The step of obtaining Gaussian noise includes: Generating the Gaussian noise according to a preset random seed.

7. The method according to claim 6, wherein The method further includes: Updating the preset random seed when the electronic device restarts.

8. An electronic device, characterized in that, The electronic device includes: one or more processors, and a memory; The memory is coupled to the one or more processors. The memory is used to store computer program code, and the computer program code includes computer instructions. The one or more processors call the computer instructions to cause the electronic device to execute the method according to any one of claims 1 - 7.

9. A chip system, characterized in that, The chip system is applied to an electronic device. The chip system includes one or more processors, and the one or more processors are used to call computer instructions to cause the electronic device to execute the method according to any one of claims 1 - 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes instructions that, when executed on an electronic device, cause the electronic device to perform the method according to any one of claims 1-7.