Federal learning-based user speech enhancement method and device, equipment and medium

Through federated learning technology, the personalized audio-enhanced neural network model processes noise and echoes in VoLTE technology, solving the limitations of VoLTE in complex noise environments and echo cancellation, and significantly improving the quality of voice communication and user experience.

CN120201129APending Publication Date: 2025-06-24NANJING AIPULU SATELLITE COMMUNICATION TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510318573.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

VoLTE technology has limitations in dealing with complex noise environments and echo cancellation, and cannot effectively personalize the user experience, resulting in a decrease in intelligibility of voice signals.

Method used

The user voice enhancement method based on federated learning is adopted. By receiving call request information, retrieving user data, determining that the user activates the audio enhancement service, the preheating process is performed, the parameters of the personalized audio enhancement neural network model are obtained, and the audio data packets are received through the user audio enhancement channel for enhanced processing.

Benefits of technology

It realizes personalized audio enhancement, accurately recognizes voice data, accurately eliminates noise and echoes, and improves voice communication quality and user experience effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120201129A_ABST
    Figure CN120201129A_ABST
Patent Text Reader

Abstract

The invention discloses a user speech enhancement method and device based on federal learning, equipment and a medium, and can be applied to the technical field of communication. After the call request information forwarded by the IP multimedia system is received, the user data corresponding to the user mobile number is retrieved according to the call request information, the user mobile number is determined to open the audio enhancement service according to the user data, the first state code is sent to the IP multimedia system, and the preheating process is executed. After a personalized audio enhancement neural network model is established from a local network parameter database and a user audio enhancement channel is established in a preheating process, a current to-be-processed audio data packet is received through the user audio enhancement channel, and then enhancement processing is performed on the current to-be-processed audio data packet through the personalized audio enhancement neural network model. Therefore, the personalized audio enhancement neural network model can accurately eliminate noise and echo, and the voice communication quality and the user experience effect are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technologies, and in particular, to a user voice enhancement method, apparatus, device, and medium based on federated learning. Background Art

[0002] In related technologies, VoLTE (Voice over LTE), as a high-definition voice call solution in a 4G network environment, provides users with a higher-quality voice service experience in mobile communications. However, the noise reduction algorithm of VoLTE is usually based on a preset noise model, and when encountering unforeseen noise types, its noise reduction performance will drop significantly. In the actual application process, the echo cancellation technology of VoLTE may not be able to completely eliminate echoes, especially when the network latency is high or the layout of the speaker and microphone of the terminal device is unreasonable, which will lead to voice latency and distortion, thereby reducing the intelligibility of the voice signal. In addition, the noise reduction and echo cancellation algorithms of VoLTE are usually designed based on a general model and cannot be personalized according to the specific needs of each user, thus limiting the user experience effect.

[0003] In summary, the technical problems existing in the related technologies need to be improved. Summary of the Invention

[0004] The main purpose of the embodiments of this application is to propose a user voice enhancement method, apparatus, device, and medium based on federated learning, which can effectively improve the quality of voice communication and the user experience effect.

[0005] To achieve the above object, on the one hand, an embodiment of this application proposes a user voice enhancement method based on federated learning, and the method includes the following steps:

[0006] Receiving call request information forwarded by an IP multimedia system, where the IP multimedia system interacts with a user terminal corresponding to a user mobile number;

[0007] Retrieving user data corresponding to the user mobile number according to the call request information;

[0008] Determining that the audio enhancement service has been activated for the user mobile number according to the user data, sending a first status code to the IP multimedia system and executing a warm-up process, where the first status code is used to indicate that the call request information has been received and processed;

[0009] The warm-up process includes:

[0010] Obtain the parameters of the personalized audio enhancement neural network model from the local network parameter database to form the personalized audio enhancement neural network model, and establish a user audio enhancement channel. The parameters of the personalized audio enhancement neural network model are dynamically trained through the federated learning algorithm;

[0011] Receive the current audio data packet to be processed corresponding to the user mobile number forwarded by the IP multimedia system through the user audio enhancement channel, so as to perform enhancement processing on the current audio data packet to be processed through the personalized audio enhancement neural network model.

[0012] In some embodiments, the performing enhancement processing on the current audio data packet to be processed through the personalized audio enhancement neural network model includes:

[0013] Perform preprocessing on the current audio data packet to be processed, the audio input signal;

[0014] Perform Fourier transform on the audio input signal to obtain a complex spectrum;

[0015] Perform enhancement processing of neural network forward propagation on the complex spectrum.

[0016] In some embodiments, the method further includes the following steps:

[0017] Determine that the user mobile number has not subscribed to the audio enhancement service according to the user data, and send a second status code to the IP multimedia system. The second status code is used to indicate that the call request information has been received but refused to be executed.

[0018] In some embodiments, the process of dynamically training the parameters of the personalized audio enhancement neural network model through the federated learning algorithm includes:

[0019] Perform homomorphic encryption on the current audio data packet to be processed to obtain data to be trained;

[0020] Perform local training on the personalized audio enhancement neural network model through the data to be trained to obtain parameters of the model to be aggregated;

[0021] Send all the parameters of the model to be aggregated to the federated learning center server for aggregated training to obtain optimized model parameters;

[0022] Update the parameters of the personalized audio enhancement neural network model in the local network parameter database according to the optimized model parameters in the federated learning center server.

[0023] In some embodiments, the process of performing aggregated training according to all the parameters of the model to be aggregated is as follows:

[0024]

[0025] In the formula, θ t+1 represents the optimized model parameters after the (t + 1)-th training aggregation; D represents the data volume of all models under the current federated learning; D i represents the data volume of the i-th model to be aggregated participating in the aggregation; represents the optimized model parameters corresponding to the i-th model to be aggregated after the t-th aggregation training; K represents the total number of all models to be aggregated participating in the aggregation; i represents the serial number of the model to be aggregated; t represents the number of training times.

[0026] In some embodiments, the method further includes the following steps:

[0027] When the current audio data packet to be processed is not received within the first preset duration, a query request is sent to the IP multimedia system, and the query request is used to query whether the current call has ended;

[0028] When the reply information sent by the IP multimedia system based on the query request is not received within the second preset duration, the user audio enhancement channel is released;

[0029] When the call end information sent by the IP multimedia system based on the query request is received within the second preset duration, the user audio enhancement channel is released.

[0030] In some embodiments, receiving, through the user audio enhancement channel, the current audio data packet to be processed corresponding to the user mobile number forwarded by the IP multimedia system, so as to perform enhancement processing on the current audio data packet to be processed through the personalized audio enhancement neural network model, includes:

[0031] When the current audio data packet to be processed is received within the first preset duration, it is determined whether a call cancellation instruction is carried in the current audio data packet to be processed;

[0032] If it is determined that the call cancellation instruction is carried, the user audio enhancement channel is released;

[0033] If it is determined that the call cancellation instruction is not carried, enhancement processing is performed on the current audio data packet to be processed through the personalized audio enhancement neural network model.

[0034] To achieve the above object, another aspect of the embodiments of the present application proposes a user voice enhancement device based on federated learning, and the device includes:

[0035] A first module, configured to receive call request information forwarded by an IP multimedia system, where the IP multimedia system interacts with a user terminal corresponding to a user mobile number;

[0036] A second module, configured to retrieve user data corresponding to the user mobile number according to the call request information;

[0037] A third module, configured to determine that the audio enhancement service has been enabled for the user mobile number according to the user data, send a first status code to the IP multimedia system and execute a warm-up process, where the first status code is used to indicate that the call request information has been received and processed;

[0038] The warm-up process includes:

[0039] Obtain parameters of a personalized audio enhancement neural network model from a local network parameter database to form a personalized audio enhancement neural network model, and establish a user audio enhancement channel, where the parameters of the personalized audio enhancement neural network model are dynamically trained through a federated learning algorithm;

[0040] Receive, through the user audio enhancement channel, the current audio data packet to be processed corresponding to the user mobile number forwarded by the IP multimedia system, so as to perform enhancement processing on the current audio data packet to be processed through the personalized audio enhancement neural network model.

[0041] To achieve the above object, another aspect of the embodiments of the present application provides an electronic device, including:

[0042] At least one processor;

[0043] At least one memory, configured to store at least one program;

[0044] When the at least one program is executed by the at least one processor, the at least one processor implements the above method.

[0045] To achieve the above object, another aspect of the embodiments of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above method is implemented.

[0046] The embodiments of the present application at least include the following beneficial effects: The present application provides a user voice enhancement method, device, equipment and medium based on federated learning. After receiving the call request information forwarded by the IP multimedia system, the solution retrieves the user data corresponding to the user mobile number according to the call request information, determines that the audio enhancement service has been enabled for the user mobile number according to the user data, sends a first status code to the IP multimedia system and executes a warm-up process, so as to obtain the parameters of the personalized audio enhancement neural network model from the local network parameter database during the warm-up process to form a personalized audio enhancement neural network model and establish a user audio enhancement channel, and then receives the current audio data packet corresponding to the user mobile number forwarded by the IP multimedia system through the user audio enhancement channel, so as to perform enhancement processing on the current audio data packet to be processed through the personalized audio enhancement neural network model, so that the personalized audio enhancement neural network model can accurately identify the voice data of the current call, and then can accurately eliminate noise and echo, effectively improving the voice communication quality and user experience effect. Description of the Drawings

[0047] Figure 1 is a flowchart of the user voice enhancement method based on federated learning provided by the embodiments of the present application;

[0048] Figure 2 is a schematic diagram of the interaction scenario of the user voice enhancement method based on federated learning provided by the embodiments of the present application;

[0049] Figure 3 is a schematic diagram of the modules of VEAS and the federated learning center server provided by the embodiments of the present application;

[0050] Figure 4 is a schematic diagram of the working mechanism of VEAS provided by the embodiments of the present application;

[0051] Figure 5 is a complete flowchart of the user voice enhancement method based on federated learning provided by the embodiments of the present application;

[0052] Figure 6 is a schematic diagram of the structure of the user voice enhancement device based on federated learning provided by the embodiments of the present application;

[0053] Figure 7 is a schematic diagram of the hardware structure of the electronic device provided by the embodiments of the present application. Detailed Embodiments

[0054] To make the objectives, technical solutions, and advantages of this application clearer and more understandable, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely used to explain this application and are not intended to limit this application. When the following description involves the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of this application. They are merely examples of devices and methods that are consistent with some aspects of the embodiments of this application.

[0055] It can be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if", "when" as used herein may be interpreted as "when...", "while...", or "in response to determining".

[0056] The terms "at least one", "multiple", "each", "any one", etc. used in this application, at least one includes one, two, or more than two, multiple includes two or more than two, each refers to each of the corresponding multiple, and any one refers to any one of the multiple.

[0057] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0058] In related technologies, the VoLTE (Voice over LTE) technology has many drawbacks in noise reduction, especially in dealing with complex scenarios, echo cancellation, data personalization optimization, privacy protection, and complex noise environments. These drawbacks may have a significant impact on the user experience in practical applications.

[0059] First of all, in dealing with complex noise environments, the VoLTE technology has obvious limitations. The noise reduction algorithm of VoLTE is usually based on a preset noise model. When encountering unforeseen noise types, its performance will drop significantly. For example, in a high-reverberation meeting room or a noisy public place, the noise reduction effect of VoLTE may not meet the user's needs, resulting in "mechanical sounds" or other unnatural sound effects during the call. In addition, the VoLTE technology also has difficulties in dealing with complex human voice scenarios. For example, in a scenario where multiple people are speaking simultaneously, VoLTE may not be able to effectively distinguish the voices of different speakers, thus affecting the call quality.

[0060] Secondly, in terms of echo cancellation, VoLTE technology also faces challenges. Although VoLTE adopts technologies such as acoustic echo cancellation (AEC), its effect is affected by the characteristics of the echo path and environmental noise. In practical applications, the echo cancellation technology of VoLTE may not be able to completely eliminate echoes, especially when the network latency is high or the layout of the speaker and microphone of the terminal device is unreasonable. This will not only cause delays and distortions in the sound, but also may reduce the intelligibility of the voice signal.

[0061] In addition, VoLTE technology has deficiencies in data personalization optimization. The noise reduction and echo cancellation algorithms of VoLTE are usually designed based on general models and cannot be adjusted individually according to the specific needs of each user. This means that users cannot customize the noise reduction effect according to their usage habits and preferences, thus restricting the further improvement of the user experience. In a complex noise environment, the noise reduction effect of VoLTE technology may not fully meet the user's needs.

[0062] Although VoLTE adopts a variety of noise reduction technologies, in practical applications, these technologies may not be able to completely eliminate background noise. For example, in a noisy street or industrial environment, the noise reduction technology of VoLTE may not be able to effectively filter out all background noise, thus affecting the clarity of the call.

[0063] In view of this, in the embodiments of the present application, a user voice enhancement method, device, equipment and medium based on federated learning are provided, which can effectively improve the quality of voice communication and the user experience effect.

[0064] The following specifically elaborates on the embodiments of the present application in conjunction with the accompanying drawings:

[0065] The user voice enhancement method based on federated learning provided by the embodiments of the present application relates to the field of communication technologies. The user voice enhancement method based on federated learning provided by the embodiments of the present application can be applied to a terminal, can also be applied to a server, or can be software running on a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto; the server side can be configured as an independent physical server, can also be configured as a server cluster or distributed system composed of multiple physical servers, or can be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application implementing the user voice enhancement method based on federated learning, etc., but is not limited to the above forms.

[0066] This application can be used in numerous general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0067] It should be noted that in each specific embodiment of this application, when it comes to relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of this application need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or redirecting to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of this application will be obtained.

[0068] Figure 1 is an optional flowchart of the user voice enhancement method based on federated learning provided by the embodiments of this application, Figure 1 The method in may include but is not limited to steps S110 to S130:

[0069] Step S110, receiving call request information forwarded by the IP multimedia system, where the IP multimedia system interacts with the user terminal corresponding to the user mobile number;

[0070] Step S120, retrieving user data corresponding to the user mobile number according to the call request information;

[0071] Step S130, determining that the audio enhancement service has been enabled for the user mobile number according to the user data, sending a first status code to the IP multimedia system and executing a warm-up process, where the first status code is used to indicate that the call request information has been received and processed.

[0072] The warm-up process includes:

[0073] Obtain the parameters of the personalized audio enhancement neural network model from the local network parameter database to form a personalized audio enhancement neural network model, and establish a user audio enhancement channel. The parameters of the personalized audio enhancement neural network model are dynamically trained through a federated learning algorithm;

[0074] Receive the currently to-be-processed audio data packet corresponding to the user's mobile number forwarded by the IP multimedia system through the user audio enhancement channel, so as to perform enhancement processing on the currently to-be-processed audio data packet through the personalized audio enhancement neural network model.

[0075] In the embodiment of the present application, it is determined according to user data that the user's mobile number has not subscribed to the audio enhancement service, and a second status code is sent to the IP multimedia system, where the second status code is used to indicate that the call request information has been received but refused to be executed.

[0076] As Figure 2 shown in the interaction scenario, when the user initiates a call request through the user terminal, the IP multimedia subsystem (IMS) forwards the user's call request to the voice enhancement application server (VEAS). After receiving the call request information, VEAS retrieves user data based on the user's international mobile subscriber identification number (IMSI) or mobile subscriber ISDN number (MSISDN). If it is determined according to user data that the mobile number corresponding to the user has subscribed to the voice enhancement service, VEAS will return a 200OK response (the first status code) to IMS and start the warm-up process to prepare for audio enhancement for the upcoming call; otherwise, if the user has not subscribed to this service, a 403Forbidden response (the second status code) will be returned.

[0077] In the warm-up process, VEAS obtains the parameters of the user's personalized audio enhancement neural network model from the local network parameter database to form a personalized audio enhancement neural network model, and establishes a user audio enhancement channel to receive audio data packets from IMS. When IMS receives the 200OK response and plays audio at the media level, it will forward the audio data in the audio data packet to VEAS for enhancement processing. If IMS receives a 403Forbidden response, it will not enhance the data of this audio data packet, but play the audio to the calling and called users according to the normal process.

[0078] In the embodiment of the present application, the process of dynamically training the parameters of the personalized audio enhancement neural network model through a federated learning algorithm includes, but is not limited to, the following steps:

[0079] Perform homomorphic encryption on the currently to-be-processed audio data packet to obtain the data to be trained;

[0080] Locally train the personalized audio enhancement neural network model with the data to be trained to obtain the model parameters to be aggregated;

[0081] Send all the model parameters to be aggregated to the federated learning central server for aggregated training to obtain the optimized model parameters;

[0082] Update the personalized audio enhancement neural network model parameters in the local network parameter database according to the optimized model parameters in the federated learning central server.

[0083] In the embodiments of the present application, the loss function in the federated learning training process is as follows:

[0084] Loss=|M·X (1) -S|2;

[0085] In the formula, S is the target clean speech.

[0086] The process of homomorphic encryption for the current audio data packet to be processed in the federated training process is as follows:

[0087] X'=[X[n]];

[0088] Among them, homomorphic encryption satisfies [f(x,y)]=f([x],[y]), and f is any operation.

[0089] Federated learning training objective:

[0090]

[0091] Among them, loss i is the loss function of a single user.

[0092] The process of aggregated training according to all the model parameters to be aggregated is as follows:

[0093]

[0094] In the formula, θ t+1 represents the optimized model parameters after the (t + 1)-th training aggregation; D represents the data volume of all models under the current federated learning; D i represents the data volume of the i-th model to be aggregated participating in the aggregation; represents the optimized model parameters corresponding to the i-th model to be aggregated after the t-th aggregation training; K represents the total number of models to be aggregated participating in the aggregation; i represents the serial number of the model to be aggregated; t represents the number of training times.

[0095] Exemplarily, the federated learning central server interacts with 20 core networks deployed in different locations. Each core network corresponds to a model to be aggregated, and the value of i is 1, 2, 3, 4, 5,..., 20; Di represents the data volume of the i-th core network itself; represents the parameters of the i-th core network after the t-th training. The value of K is 20, indicating that the models of 20 core networks deployed in different locations participate in aggregation.

[0096] After completing the aggregation of model parameters, the embodiments of the present application update the parameters of each personalized audio enhancement neural network model through the following formula:

[0097]

[0098] In the formula, λ represents a preset parameter weight, which is a constant.

[0099] It can be understood that as Figure 3 shown, VEAS includes a proxy (PROXY), a central controller, a homomorphic encryption module, an encrypted data database, a local training module, a model aggregation module, and a local network parameter database. The federated learning center server includes a proxy (PROXY), a model update module, a central aggregation module, and a model initialization module. The data interaction between VEAS and the federated learning center server is shown in Table 1:

[0100] Table 1

[0101]

[0102]

[0103] In the embodiments of the present application, as Figure 4 shown, when the current audio data packet to be processed has not been received for more than the first preset duration to execute the pre-release process, and a query request is sent to the IP multimedia system to query whether the current call has ended;

[0104] When the reply information sent by the IP multimedia system based on the query request has not been received for more than the second preset duration, release the user audio enhancement channel;

[0105] When the call end information sent by the IP multimedia system based on the query request is received within the second preset duration, release the user audio enhancement channel.

[0106] In the embodiments of the present application, when the current audio data packet to be processed is received within the first preset duration, it is determined whether the current audio data packet to be processed carries a call cancellation instruction (Cancel message);

[0107] If it is determined that the call cancellation instruction is carried, release the user audio enhancement channel;

[0108] It is determined that there is no call cancellation instruction, and the current audio data packet to be processed is enhanced through a personalized audio enhancement neural network model.

[0109] It can be understood that the process of enhancing the current audio data packet to be processed by the personalized audio enhancement neural network model includes but is not limited to the following steps:

[0110] Step 1: Preprocess the current audio data packet to be processed: VEAS integrates the received audio data packets into audio data to obtain an audio input signal x[n];

[0111] Step 2: Perform Fourier transform on the audio input signal x[n]: VEAS performs Fourier transform on the input signal to obtain a frequency-domain signal and gets a complex spectrum D; specifically, the following formula:

[0112]

[0113] In the formula, w[n] is a window function; N is the window length; m is the time frame index; k is the frequency index;

[0114] Step 3: Perform enhanced processing of neural network forward propagation on the complex spectrum:

[0115] Step 3.1: Complex LSTM:

[0116]

[0117] F out =(F rr -F ii )+j(F ri +F ir );

[0118] In the formula, LSTM i ,LSTM r respectively represent the LSTM units for processing the imaginary part and the real part;

[0119] Step 3.2: Complex convolution:

[0120]

[0121] In the formula, W r ,W i are the real part and the imaginary part of the complex convolution kernel of the neural network.

[0122] Step 3.3: Forward propagation calculation:

[0123] X r 1 =D r [m,k],X i 1= D i [m, k] is the real and imaginary parts of the input signal after Fourier transform.

[0124]

[0125] In the formula, prelu is the activation function, and the calculation method is a is the learnable parameter; g is the normalization function, and the calculation method is where u and σ are the mean and standard deviation of X respectively, and γ and β are the learnable parameters; W (n) is the complex convolution kernel of the neural network;

[0126] H 1 = LSTM(X (N ));

[0127]

[0128] M = H (N) ;

[0129] In the formula, N is the number of encoding layers;

[0130] Y = M · X (1) ;

[0131]

[0132] In the formula, y[n] is the result after speech enhancement.

[0133] Based on the above content, it can be seen that the complete implementation process of the method in the embodiments of this application is as Figure 5 shown, and specifically includes the following processes:

[0134] First, the verification process:

[0135] 1.1 Receive the user call notification: The IMS (IP Multimedia Subsystem) forwards the user call notification to the VeAS (Voice Enhancement Application Server).

[0136] 1.2 Query the user service: The VeAS queries whether the user has subscribed to the voice enhancement service.

[0137] 1.3 Service activation check: The VeAS queries whether the user service is activated; if the user has not subscribed to the voice enhancement service, the VeAS replies with 403 Forbidden and exits. If the user has subscribed to the service, the VeAS replies with 200 OK and enters the warm-up process.

[0138] Second, the warm-up process:

[0139] 2.1 Obtain the user's personalized audio enhancement neural network model: When the user has subscribed to the service, VEAS attempts to obtain the user's personalized audio enhancement neural network model obtained through federated learning training. If the acquisition fails, the system returns 403 Forbidden and exits. If the acquisition is successful, proceed to the next step.

[0140] 2.2 Set up the audio enhancement channel: After successfully obtaining the user's personalized audio enhancement neural network model, VEAS sets up the audio enhancement channel to prepare for processing the current audio data packet to be processed.

[0141] 2.3 Wait for the current audio data packet to be processed: VEAS enters a waiting state and waits to receive the current audio data packet to be processed before entering the workflow.

[0142] Third, Workflow:

[0143] 3.1 Timeout check: VEAS starts a timer to detect whether the reception of the user's audio packet times out. If no audio packet is received after the timeout, enter the pre-release process.

[0144] 3.2 Exit detection: The VEAS central controller analyzes the received signaling stream. If a CANCEL signaling is received, enter the release process.

[0145] 3.2 Data preprocessing: The VEAS PROXY module receives the current audio data packet to be processed forwarded by IMS, integrates the received current audio data packet to be processed, and preprocesses the integrated voice data.

[0146] 3.3 Audio enhancement processing: VEAS calls the user's personalized audio enhancement neural network model in the audio enhancement channel to perform personalized enhancement on the preprocessed and integrated voice data.

[0147] 3.4 Forward the enhanced audio: VEAS splits the enhanced audio, restores it to a voice packet, and forwards it to IMS.

[0148] 3.5 Wait for the user's audio packet: VEAS enters a waiting state, waits to receive the current audio data packet to be processed, and returns to 3.1.

[0149] Fourth, Pre-release process:

[0150] 4.1 Query the call status: VeAS sends a query request to IMS to check whether the user's call has ended.

[0151] 4.2 Call status processing: If the call has ended, VEAS enters the release process and exits the service. If the call has not ended, VEAS returns to the workflow and continues with the current audio data packet to be processed.

[0152] 4.3 Timeout Detection: If VEAS does not receive the user call status response replied by IMS within the specified time, the call is considered abnormal and enters the release process.

[0153] Fifth, Release Process:

[0154] Release Channel Resources: When the user call ends / abnormal, VEAS releases the audio enhancement channel and the process ends.

[0155] As can be seen from the above content, the method of the embodiment of the present application, the user voice enhancement method based on federated learning, can effectively reduce the interference of background noise on noisy streets or in noisy meetings, thereby significantly improving the call quality and enabling users to clearly hear the voice content of the other party. At the same time, the method of the present application can also automatically adjust the voice enhancement parameters according to the user's usage habits to provide a more comfortable call experience for users. In addition, since the method of the present application uses the homomorphic encryption technology, the user's voice data is always in an encrypted state during the transmission and processing process, greatly reducing the risk of data leakage, thereby providing more reliable privacy protection for users.

[0156] Refer to Figure 6 , the embodiment of the present application provides a user voice enhancement device based on federated learning, and the device includes:

[0157] The first module 610 is configured to receive the call request information forwarded by the IP multimedia system, and the IP multimedia system interacts with the user terminal corresponding to the user mobile number;

[0158] The second module 620 is configured to retrieve the user data corresponding to the user mobile number according to the call request information;

[0159] The third module 630 is configured to determine that the audio enhancement service has been activated for the user mobile number according to the user data, send the first status code to the IP multimedia system and execute the warm-up process, and the first status code is used to indicate that the call request information has been received and processed;

[0160] The warm-up process includes:

[0161] Obtain the parameters of the personalized audio enhancement neural network model from the local network parameter database to form a personalized audio enhancement neural network model, and establish a user audio enhancement channel, and the parameters of the personalized audio enhancement neural network model are dynamically trained through the federated learning algorithm;

[0162] Receive the current audio data packet corresponding to the user mobile number forwarded by the IP multimedia system through the user audio enhancement channel, so as to enhance the current audio data packet to be processed through the personalized audio enhancement neural network model.

[0163] It can be understood that the content in the above method embodiments is applicable to the device embodiments. The functions specifically implemented by the device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0164] An embodiment of the present application further provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above method is implemented. The electronic device can be any intelligent terminal including a tablet computer, a vehicle-mounted computer, etc.

[0165] It can be understood that the content in the above method embodiments is applicable to the device embodiments. The functions specifically implemented by the device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0166] Please refer to Figure 7 , Figure 7 which shows the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0167] A processor 710, which can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;

[0168] A memory 720, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 720 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 720 and are called by the processor 710 to execute the above method of the embodiments of the present application;

[0169] An input / output interface 730, which is used to implement information input and output;

[0170] A communication interface 740, which is used to implement communication interaction between this device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as mobile network, WIFI, Bluetooth, etc.);

[0171] The bus 750 transmits information among various components of the device (such as the processor 710, the memory 720, the input / output interface 730, and the communication interface 740).

[0172] Among them, the processor 710, the memory 720, the input / output interface 730, and the communication interface 740 achieve communication connections with each other inside the device through the bus 750.

[0173] The embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above method is implemented.

[0174] It can be understood that the content in the above method embodiments is applicable to the embodiments of this storage medium. The functions specifically implemented by the embodiments of this storage medium are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0175] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory can optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and their combinations.

[0176] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0177] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.

[0178] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0179] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.

[0180] As used in the specification of this application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0181] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Here, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or similar expressions refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0182] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned unit division is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.

[0183] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0184] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0185] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store programs.

[0186] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present application. Any modification, equivalent replacement, and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.

Claims

1. A method for user speech enhancement based on federated learning, characterized in that: The method comprises the following steps: Receiving the call request information forwarded by the IP multimedia system, the IP multimedia system interacts with the user terminal corresponding to the user mobile number; Retrieving user data corresponding to the user mobile number according to the call request information; Determining, according to the user data, that the user mobile number has activated the audio enhancement service, sending a first status code to the IP multimedia system and executing a warm-up process, wherein the first status code is used to indicate that the call request information has been received and processed; The preheating process includes: Acquire the parameters of the personalized audio enhancement neural network model from the local network parameter database to form a personalized audio enhancement neural network model, and establish a user audio enhancement channel, wherein the parameters of the personalized audio enhancement neural network model are dynamically trained through a federated learning algorithm; The current audio data packet to be processed corresponding to the user mobile number forwarded by the IP multimedia system is received through the user audio enhancement channel, so as to enhance the current audio data packet to be processed through the personalized audio enhancement neural network model.

2. The method according to claim 1, characterized in that The step of performing enhancement processing on the current audio data packet to be processed by using the personalized audio enhancement neural network model includes: Preprocessing the current audio data packet to be processed, an audio input signal; Performing Fourier transform on the audio input signal to obtain a complex spectrum; The complex spectrum is subjected to a neural network forward propagation enhancement process.

3. The method according to claim 1, characterized in that The method further comprises the following steps: It is determined according to the user data that the audio enhancement service is not activated for the user mobile number, and a second status code is sent to the IP multimedia system, where the second status code is used to indicate that the call request information has been received but refused to be executed.

4. The method according to claim 1, characterized in that: The process of dynamically training the parameters of the personalized audio enhancement neural network model through a federated learning algorithm includes: Performing homomorphic encryption on the current audio data packet to be processed to obtain data to be trained; Locally training the personalized audio enhancement neural network model using the data to be trained to obtain model parameters to be aggregated; Sending all the model parameters to be aggregated to the federated learning center server for aggregation training to obtain optimized model parameters; The personalized audio enhancement neural network model parameters in the local network parameter database are updated according to the optimized model parameters in the federated learning center server.

5. The method according to claim 4, characterized in that The process of performing aggregation training according to all the parameters of the models to be aggregated is as follows: In the formula, θ t+1 represents the optimized model parameters after the t+1th training aggregation; D represents the data volume of all models under the current federated learning; D i Indicates the data volume of the i-th model to be aggregated that participates in the aggregation; represents the optimized model parameters corresponding to the i-th model to be aggregated after the t-th aggregation training; K represents the total number of all models to be aggregated participating in the aggregation; i represents the sequence number of the model to be aggregated; t represents the number of training times.

6. The method according to claim 1, characterized in that The method further comprises the following steps: When the current audio data packet to be processed is not received for more than a first preset time, a query request is sent to the IP multimedia system, wherein the query request is used to query whether the current call is ended; When no reply information sent by the IP multimedia system based on the query request is received for more than a second preset time period, releasing the user audio enhancement channel; When the call end information sent by the IP multimedia system based on the query request is received within the second preset time period, the user audio enhancement channel is released.

7. The method according to claim 6, characterized in that The step of receiving the current audio data packet to be processed corresponding to the user mobile number forwarded by the IP multimedia system through the user audio enhancement channel, and performing enhancement processing on the current audio data packet to be processed through the personalized audio enhancement neural network model, includes: When the current audio data packet to be processed is received during the first preset time period, determining whether the current audio data packet to be processed carries a call cancellation instruction; Determine to carry the call cancellation instruction and release the user audio enhancement channel; It is determined that the call cancellation instruction is not carried, and the current audio data packet to be processed is enhanced by using the personalized audio enhancement neural network model.

8. A user speech enhancement device based on federated learning, characterized in that: The device comprises: The first module is used to receive the call request information forwarded by the IP multimedia system, and the IP multimedia system interacts with the user terminal corresponding to the user mobile number; The second module is used to retrieve the user data corresponding to the user mobile number according to the call request information; A third module is used to determine, according to the user data, that the user mobile number has activated the audio enhancement service, send a first status code to the IP multimedia system and execute a warm-up process, wherein the first status code is used to indicate that the call request information has been received and processed; The preheating process includes: Acquire the parameters of the personalized audio enhancement neural network model from the local network parameter database to form a personalized audio enhancement neural network model, and establish a user audio enhancement channel, wherein the parameters of the personalized audio enhancement neural network model are dynamically trained through a federated learning algorithm; The current audio data packet to be processed corresponding to the user mobile number forwarded by the IP multimedia system is received through the user audio enhancement channel, so as to enhance the current audio data packet to be processed through the personalized audio enhancement neural network model.

9. An electronic device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • WebRTC speech enhancement system and method based on multi-modal large model

    CN121011196A