Voiceprint model optimization method and device based on reinforcement learning, equipment and medium
Through the voiceprint model optimization method based on reinforcement learning, the policy network of the voiceprint model is optimized using historical recognition cases and KL divergence constraints, and the existing voiceprint model's problem of high difficulty in optimization and low recognition accuracy is solved, achieving a more efficient recognition effect.
Patent Information
- Application Number
- CN202510493314.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The existing voiceprint model is difficult to optimize in practical applications, and lacks effective means for misjudgment of cases, resulting in a decrease in recognition accuracy and being unable to adapt to new voice features or complex environments.
Using a method based on reinforcement learning, an optimized data set is constructed by collecting historical recognition cases of voiceprint model, randomly initializing the policy network, calculating reward values and updating parameters, optimizing the objective function using the advantage function and KL divergence term, loop iteration until the optimal policy network is obtained, and the voiceprint recognition results are optimized.
It improves the recognition accuracy and stability of the voiceprint model, can adapt to new voice features and complex environments, reduces misjudgment, and improves the reliability and efficiency of recognition.
Smart Images

Figure CN120299461A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a method, device, equipment and medium for optimizing a voiceprint model based on reinforcement learning. Background Art
[0002] A voiceprint is the voice feature contained in speech that can characterize and identify a speaker, and voiceprint recognition is the process of identifying the speaker corresponding to a segment of speech according to the voiceprint features of the speech to be recognized. Due to the particularity of sound, compared with other behavioral characteristics, voiceprint recognition has unique physiological characteristics. Therefore, theoretically, just like fingerprints, it is very rare for two people to have the same voiceprint features. This characteristic makes voiceprint recognition technology show great potential as a new means of identity verification in risk prevention and control and customer identity verification in multiple fields.
[0003] In the field of medical and health, voiceprint recognition can be applied in remote medical consultations. For example, patients communicate with doctors about their conditions through video or voice calls. To ensure the authenticity of the patient's identity and information security, the patient can be authenticated through a voiceprint recognition model; it can also be applied in, for example, the operation permission verification of high-precision medical equipment, that is, the identity of the operator can be verified through voiceprint recognition to ensure that the operator has the corresponding qualifications and permissions.
[0004] In the financial field, voiceprint recognition can be applied to telephone service scenarios such as insurance business or banking business. When customers handle business by phone, the customer's identity is verified through a voiceprint recognition model to ensure transaction security; or during the insurance claim reporting process, the insurance company can more accurately identify the customer's identity through the voiceprint model, reducing misjudgment and fraud.
[0005] However, the existing voiceprint models still have some deficiencies in actual applications. Once the voiceprint model is launched, it is difficult to optimize. There is a lack of effective optimization means for misjudgment cases. Retraining the voiceprint model may cause the model to forget the knowledge learned before, thus affecting the cases that were originally judged correctly. Therefore, when the voiceprint recognition model faces new voice features or complex voice environments, it may not be able to adjust in time, thereby affecting the recognition accuracy. Summary of the Invention
[0006] In view of the above deficiencies of the prior art, the purpose of the present invention is to provide a method, device, equipment and medium for optimizing a voiceprint model based on reinforcement learning that can be applied to the medical field, fintech or other related fields. Its main purpose is to improve the recognition accuracy of the voiceprint model by performing efficient and reliable optimization on the voiceprint model.
[0007] The technical solution of the present invention is as follows:
[0008] The first aspect of the present invention provides a method for optimizing a voiceprint model based on reinforcement learning, including:
[0009] Collect historical recognition cases of the voiceprint model, and construct an optimization data set according to the historical recognition cases;
[0010] Randomly initialize the policy network, sample audio data from the optimization data set and input it into the voiceprint model, control the voiceprint model to perform voiceprint recognition according to the initialized policy, and obtain the recognition result and corresponding reward value of each sample;
[0011] Calculate the advantage function according to the recognition result and reward value of each sample, and update the parameters of the policy network based on the advantage function;
[0012] Optimize the pre-constructed objective function according to the parameters before and after the update of the policy network. The objective function includes a KL divergence term for constraining the policy update amplitude;
[0013] Loop through the process of sampling audio data for voiceprint recognition and policy update until the objective function meets the preset conditions to obtain the optimal policy network;
[0014] Based on the optimal policy network, control the voiceprint model to perform voiceprint recognition according to the optimal policy to optimize the recognition result.
[0015] The second aspect of the present invention provides a device for optimizing a voiceprint model based on reinforcement learning, including:
[0016] A data acquisition module for collecting historical recognition cases of the voiceprint model and constructing an optimization data set according to the historical recognition cases;
[0017] An initialization module for randomly initializing the policy network, sampling audio data from the optimization data set and inputting it into the voiceprint model, and controlling the voiceprint model to perform voiceprint recognition according to the initialized policy to obtain the recognition result and corresponding reward value of each sample;
[0018] A policy update module for calculating the advantage function according to the recognition result and reward value of each sample, and updating the parameters of the policy network based on the advantage function;
[0019] A function optimization module for optimizing the pre-constructed objective function according to the parameters before and after the update of the policy network. The objective function includes a KL divergence term for constraining the policy update amplitude;
[0020] An iteration control module for looping through the process of sampling audio data for voiceprint recognition and policy update until the objective function meets the preset conditions to obtain the optimal policy network;
[0021] A model optimization module, configured to control the voiceprint model to perform voiceprint recognition according to an optimal policy based on the optimal policy network, so as to optimize the recognition result.
[0022] The third aspect of the present invention provides a computer device, including at least one processor; and,
[0023] a memory communicatively connected to the at least one processor; wherein,
[0024] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor, so that the at least one processor can execute the above-mentioned voiceprint model optimization method based on reinforcement learning.
[0025] The fourth aspect of the present invention provides a non-volatile computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by one or more processors, the one or more processors can execute the above-mentioned voiceprint model optimization method based on reinforcement learning.
[0026] Beneficial effects: The present invention discloses a voiceprint model optimization method, device, equipment and medium based on reinforcement learning. Compared with the prior art, the embodiments of the present invention include: collecting historical recognition cases of the voiceprint model, constructing an optimization data set according to the historical recognition cases; randomly initializing a policy network, sampling audio data from the optimization data set and inputting it into the voiceprint model, controlling the voiceprint model to perform voiceprint recognition according to the initialized policy, and obtaining the recognition result of each sample and the corresponding reward value; calculating an advantage function according to the recognition result and reward value of each sample, and updating the parameters of the policy network based on the advantage function; optimizing a pre-constructed objective function according to the parameters before and after the update of the policy network, where the objective function includes a KL divergence term for constraining the policy update amplitude; repeatedly executing the process of sampling audio data for voiceprint recognition and policy update until the objective function meets a preset condition to obtain an optimal policy network; controlling the voiceprint model to perform voiceprint recognition according to the optimal policy network to optimize the recognition result. By optimizing the voiceprint model through reinforcement learning and KL divergence constraint, the stability and efficiency of policy update are improved, reliable optimization of the voiceprint model is achieved, and the recognition accuracy of the voiceprint model is improved. Description of the Drawings
[0027] To more clearly illustrate the solutions in the present invention, the following provides a brief introduction to the drawings required for the description of the embodiments of the present invention. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0028] Figure 1 Schematic diagram of an application environment for the method for optimizing a voiceprint model based on reinforcement learning provided in an embodiment of the present invention;
[0029] Figure 2 Flowchart of a method for optimizing a voiceprint model based on reinforcement learning provided in an embodiment of the present invention;
[0030] Figure 3 Flowchart of step S201 in the method for optimizing a voiceprint model based on reinforcement learning provided in an embodiment of the present invention;
[0031] Figure 4 Flowchart of step S202 in the method for optimizing a voiceprint model based on reinforcement learning provided in an embodiment of the present invention;
[0032] Figure 5 Flowchart of step S203 in the method for optimizing a voiceprint model based on reinforcement learning provided in an embodiment of the present invention;
[0033] Figure 6 Flowchart of step S204 in the method for optimizing a voiceprint model based on reinforcement learning provided in an embodiment of the present invention;
[0034] Figure 7 Schematic diagram of functional modules of a device for optimizing a voiceprint model based on reinforcement learning provided in an embodiment of the present invention;
[0035] Figure 8 Schematic diagram of the hardware structure of a computer device provided in an embodiment of the present invention. Detailed implementation manners
[0036] To make the objectives, technical solutions and effects of the present invention clearer and more definite, the present invention is further described in detail below. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. The embodiments of the present invention are introduced below with reference to the accompanying drawings.
[0037] The method for optimizing a voiceprint model based on reinforcement learning provided in an embodiment of the present invention can be applied in, for example Figure 1In the application environment, it includes a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 can include various connection types, such as wired and / or wireless communication links, etc.
[0038] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (only for example).
[0039] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, laptop portable computers, and desktop computers, etc.
[0040] The server 105 can be a server that provides various services, such as a background server that provides support for the content browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (only for example). The background server can analyze and process data such as received user requests, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device. The server 105 can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability in traditional physical hosts and VPS services (″Virtual Private Server″, or simply ″VPS″). The server 105 can also be a server of a distributed system, or a server combined with a blockchain.
[0041] It should be noted that the method for optimizing the voiceprint model based on reinforcement learning provided in the embodiments of the present application can generally be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103. Correspondingly, the device for optimizing the voiceprint model based on reinforcement learning provided in the embodiments of the present invention can also be set in the first terminal device 101, the second terminal device 102, or the third terminal device 103. Alternatively, the method for optimizing the voiceprint model based on reinforcement learning provided in the embodiments of the present invention can generally also be executed by the server 105. Correspondingly, the device for optimizing the voiceprint model based on reinforcement learning provided in the embodiments of the present invention can generally be set in the server 105.
[0042] It should be understood that the numbers of the above terminal devices, networks, and servers are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.
[0043] As Figure 2 shown, the method for optimizing the voiceprint model based on reinforcement learning provided in the embodiments of the present invention specifically includes the following steps:
[0044] S201. Collect historical recognition cases of the voiceprint model, and construct an optimization data set according to the historical recognition cases.
[0045] In this embodiment, to optimize the existing voiceprint model, historical recognition cases of the voiceprint model are collected from actual business scenarios (such as medical consultation scenarios, financial insurance scenarios, etc.). Based on the audio data and corresponding recognition results that the voiceprint model has processed in actual business applications, the corresponding historical recognition cases are collected. These historical recognition cases include correctly recognized cases and misrecognized cases. By collecting these historical recognition cases, a case data set including correctly recognized and misrecognized cases is constructed for subsequent model optimization training, providing a basis for the continuous improvement of the model. Preferably, in addition to directly using historical recognition cases, more diverse training data can be generated through data augmentation processing (such as adding noise, adjusting pitch, speed change, etc.) to comprehensively cover various situations that the model may encounter in actual applications and enhance the generalization ability of the model.
[0046] Exemplarily, in the medical and health scenario, for example, in remote medical consultation, the voiceprint recognition results when patients talk to doctors are collected, including cases of correctly recognizing the patient's identity and misjudgment cases, etc., to construct an optimization data set. In the financial scenario, for example, in insurance claim settlement services, the voiceprint recognition results when customers handle claim settlement services are collected, including cases of correctly recognizing the customer's identity and misjudgment cases, etc., to construct an optimization data set. At the same time, in various scenarios, more diverse training data can be generated through data augmentation methods such as adding background noise.
[0047] S202. Randomly initialize the policy network, sample audio data from the optimization dataset and input it into the voiceprint model, and control the voiceprint model to perform voiceprint recognition according to the initialized policy to obtain the recognition result and the corresponding reward value of each sample.
[0048] In this embodiment, the policy network is a deep learning model, such as a convolutional neural network (CNN), a recurrent neural network (RNN), or other network structures suitable for processing audio data, etc. It is used to output the decision-making policy of the voiceprint model. The parameters of the policy network determine how the voiceprint model identifies and judges the input audio data, that is, to identify whether the input audio data and the reference audio (such as the audio pre-recorded by the customer) are of the same person. When optimizing the voiceprint model, first randomly initialize the parameters of the policy network, randomly sample a batch of audio data from the optimization dataset as samples and input them into the voiceprint model. For each input sample, the voiceprint model performs voiceprint recognition with the initialized policy corresponding to the current policy network, so as to obtain the recognition result and the corresponding reward value of each sample, providing a corresponding feedback signal for subsequent policy updates.
[0049] The reward value is obtained through a reward function defined in advance based on business requirements. For example, a positive reward (such as +1) is given when the voiceprint model correctly recognizes, and a negative reward (such as -1) is given for incorrect recognition, etc. By designing positive and negative reward functions, the model can learn from correct and incorrect cases and feedback to update the policy, thereby effectively improving the recognition accuracy. Of course, in other embodiments, other reward functions can also be sampled to obtain the reward value, such as giving different degrees of negative rewards according to the severity of misjudgment, or giving rewards according to the recognition confidence, etc. This embodiment does not limit this.
[0050] For example, in the scenario of remote medical consultation, after randomly initializing the policy network, sample the audio data of the patient's call with the doctor from the optimization dataset and input it into the voiceprint model, and calculate the reward value according to the recognition result. For example, when the voiceprint model correctly recognizes the patient's identity through the input audio data, a reward of +1 is given, and a reward of -1 is given for misjudgment. Another example is in scenarios such as claim reporting or telephone banking in the financial field. After randomly initializing the policy network, sample the audio data of the customer handling business by phone from the optimization dataset and input it into the voiceprint model. If the voiceprint model can accurately recognize the customer's identity, a reward of +1 is given, and a reward of -1 is given for misjudgment.
[0051] S203. Calculate the advantage function according to the recognition result and the reward value of each sample, and update the parameters of the policy network based on the advantage function.
[0052] In this embodiment, the advantage function refers to the advantage of the return of taking a certain action relative to the average return under a given state. Specifically, the advantage function can be estimated by means such as Generalized Advantage Estimation (GAE), Monte Carlo Tree Search (MCTS), and TD (Temporal Difference) learning. This embodiment does not limit this. Since the advantage function can more accurately evaluate the pros and cons of each action, in this embodiment, the advantage function is calculated based on the recognition result and the reward value of each sample, and the parameters of the policy network are adjusted according to the value of the advantage function, that is, the policy improvement is guided by the advantage function, and the sampling data is fully utilized to make the model more inclined to take actions with greater advantages, so as to more effectively update the policy network to optimize the recognition effect of the voiceprint model.
[0053] S204. Optimize the pre-constructed objective function according to the parameters before and after the update of the policy network. The objective function includes a KL divergence term for constraining the policy update amplitude.
[0054] In this embodiment, an objective function is pre-constructed to calculate the expected value of the cumulative reward to evaluate the performance of the policy network. The values of the corresponding objective functions are calculated based on the parameters before and after the update of the policy network, and the performance of the policy network is evaluated based on the values of the objective functions to determine whether it is improving. Thus, the entire optimization process is guided by maximizing the expected value of the cumulative reward to ensure that the learning direction of the policy network is towards maximizing the cumulative reward, and the policy network is ensured to learn the optimal policy by optimizing the objective function, thereby improving the performance of the voiceprint recognition model. At the same time, in order to prevent the policy update amplitude from being too large, resulting in the model forgetting the knowledge learned previously and affecting the cases that were originally judged correctly, the pre-constructed objective function in this embodiment includes a KL divergence term. The KL divergence specifically refers to the Kullback-Leibler divergence, which is an asymmetric measure of the difference between two probability distributions and is used to measure the difference between two probability distributions. By restricting the KL divergence between the old and new policies, the policy update amplitude is ensured to be appropriate, and the stability of the model during the optimization process is improved.
[0055] S205. Repeat the process of sampling audio data for voiceprint recognition and policy update until the optimal policy network is obtained when the objective function meets the preset conditions.
[0056] Repeat the above processes of sampling, recognition, calculating the advantage function, updating the policy network, and calculating the objective function until the objective function meets certain convergence conditions, such as the parameter update amplitude of the policy network is less than a certain threshold, or the value of the objective function reaches a certain stable value, or the preset number of iterations is reached, etc., so as to obtain the optimal policy network.
[0057] That is, in each iteration, the policy network selects actions according to the latest parameters to guide the voiceprint recognition model to recognize audio data. Through iterative loops, the policy network is gradually optimized to achieve optimal performance after multiple updates. Specifically, in the initial stage, the parameters of the policy network are randomly initialized, so the actions it selects may not be ideal, and the performance of the voiceprint model is poor at this time. In the intermediate stage, as the policy network is continuously updated, more reasonable actions will be selected after each update, gradually improving the performance of the voiceprint recognition model. In the convergence stage, when the parameters of the policy network converge to the optimal state, the optimal actions will be selected, and the behavior of the voiceprint recognition model will be dynamically adjusted based on the optimization method of reinforcement learning to maximize the performance of the voiceprint recognition model.
[0058] S206. Based on the optimal policy network, control the voiceprint model to perform voiceprint recognition according to the optimal policy to optimize the recognition result.
[0059] In this embodiment, after multiple iterative optimizations, a policy network with optimal performance is finally obtained. The optimized policy network is deployed in practical applications for voiceprint recognition tasks, so as to use the optimal policy to guide the optimization of the voiceprint recognition model, thereby optimizing the recognition result of the model and improving the accuracy and reliability of voiceprint recognition. The optimization method of the voiceprint model based on reinforcement learning and KL divergence constraint is particularly applicable to remote authentication scenarios in the insurance business, such as claim reporting, telemarketing, etc., which can quickly identify customer identities, reduce misjudgment and fraud behaviors, and improve the customer experience and service quality.
[0060] In the above embodiment, the present invention discloses a method for optimizing a voiceprint model based on reinforcement learning. By collecting historical recognition cases of the voiceprint model, an optimization data set is constructed according to the historical recognition cases; the policy network is randomly initialized, audio data is sampled from the optimization data set and input into the voiceprint model, and the voiceprint model is controlled to perform voiceprint recognition according to the initialized policy to obtain the recognition result and the corresponding reward value of each sample; the advantage function is calculated according to the recognition result and the reward value of each sample, and the parameters of the policy network are updated based on the advantage function; the pre-constructed objective function is optimized according to the parameters before and after the update of the policy network, and the KL divergence term for constraining the policy update amplitude is included in the objective function; the process of sampling audio data for voiceprint recognition and policy update is repeatedly executed until the optimal policy network is obtained when the objective function meets the preset conditions; based on the optimal policy network, control the voiceprint model to perform voiceprint recognition according to the optimal policy to optimize the recognition result. The voiceprint model is optimized through reinforcement learning and KL divergence constraint, improving the stability and efficiency of policy update, realizing reliable optimization of the voiceprint model, and improving the recognition accuracy of the voiceprint model.
[0061] In one embodiment, asFigure 3 As shown in Figure 3 , step S201 includes:
[0062] S301. Collect historical recognition cases of the voiceprint model from the business scenario, where the historical recognition cases include a set of successful recognition cases and a set of failed recognition cases;
[0063] S302. Preprocess the audio data in each historical recognition case, and use the preprocessed audio data and the corresponding historical output of the model as an audio pair to construct an optimized database.
[0064] In this embodiment, when constructing the data set, data is collected from the environment where the voiceprint model is actually applied, such as business scenarios like remote medical consultations and telephone banking services. A large amount of voiceprint recognition data has been accumulated in these scenarios. Extract the recognition records of the voiceprint model from the actual business system, including audio files and the output results of the model. Label each case to clarify whether it belongs to a successful case or a failed case. Divide the historical recognition cases into two categories, a set of successful recognition cases and a set of failed recognition cases, that is, cases where the model correctly recognizes whether an audio pair belongs to the same person, and cases where the model incorrectly recognizes whether an audio pair belongs to the same person. By collecting both successful and failed cases, it is possible to comprehensively cover various situations that the model may encounter in actual applications, providing rich data support for subsequent optimization.
[0065] After that, corresponding preprocessing is performed on the audio data in the historical recognition cases to improve the data quality and enhance the training effect of the model. Specifically, preprocessing methods such as sampling quantization, pre-emphasis, windowing, and noise removal can be used. Among them, sampling quantization is to convert the audio signal into a digital signal and adjust the sampling rate to adapt to the model input; pre-emphasis is to enhance high-frequency signals and reduce the loss of high-frequency signals during transmission; windowing is to divide the audio signal into multiple short-time windows for easy analysis; noise removal is to remove background noise through filtering and other methods to improve the clarity of the audio signal. Combine the preprocessed audio data with the historical output of the model (that is, the previous recognition result of the model) into a data pair. For example, for an audio pair, the historical output of the model may be "the same person" or "different people", and this historical output may be correctly or incorrectly judged, that is, this case belongs to the set of successful recognition cases or the set of failed recognition cases, providing high-quality training data for the optimization of the voiceprint model, while ensuring the diversity and pertinence of the data, laying a solid foundation for subsequent model optimization.
[0066] In one embodiment, as Figure 4 shown in Figure 4 , step S202 includes:
[0067] S401. Sample audio pairs from the optimized dataset, input the audio data in the audio pairs into the voiceprint model, and control the voiceprint model to identify whether the audio data is of the same person as the corresponding reference audio according to the initialization strategy, so as to obtain the recognition result of each sample;
[0068] S402. Confirm whether the current voiceprint model is correctly recognized according to the recognition result of each sample, the corresponding model historical output, and the case set where the sample is located;
[0069] S403. Give a positive reward when the voiceprint model is correctly recognized, and give a negative reward when the voiceprint model is incorrectly recognized, to obtain the corresponding reward value.
[0070] In this embodiment, audio pairs (s, a) are sampled from the optimized dataset as training samples, where s is the input audio data (including the audio to be recognized and the corresponding reference audio), and a is the model historical output (i.e., the recognition result of the voiceprint model in previous business applications). Specifically, methods such as random sampling and stratified sampling can be used to ensure the diversity and representativeness of the samples. The sampled audio data is input into the voiceprint model, and the voiceprint model identifies the input audio data according to the initialization strategy, and outputs a judgment result on whether the audio pair belongs to the same person, that is, the recognition result of each sample is obtained, so as to obtain the initial performance of the model, providing a basis for subsequent reward calculation and policy update. Then, by comparing the current recognition result with the model historical output, combined with the case set where the sample is located, it is judged whether the current recognition is correct, and the corresponding reward value is given based on the pre-set reward function. The reward value is associated with the recognition result of each sample. Specifically, when the voiceprint model correctly recognizes whether the audio pair belongs to the same person, a positive reward value (such as +1) is given, indicating that the model's decision is correct; when the voiceprint model incorrectly recognizes whether the audio pair belongs to the same person, a negative reward value (such as -1) is given, indicating that the model's decision is wrong. Through positive and negative reward values, a clear feedback signal is provided for the model to help the model learn the correct decision-making mode, and the reward value is used as the optimization guidance to guide the model to be more inclined to make correct decisions in subsequent training.
[0071] Specifically, when determining whether the current voiceprint model recognition is correct, if the current recognition result is consistent with the historical output and the sample belongs to the set of successful recognition cases (i.e., the model correctly recognizes whether two audio files belong to the same person), it is determined that the recognition is correct; if the current recognition result is inconsistent with the historical output and the sample belongs to the set of successful recognition cases (i.e., the model incorrectly recognizes whether two audio files belong to the same person), it is determined that the recognition is incorrect; if the current recognition result is inconsistent with the historical output and the sample belongs to the set of failed recognition cases (i.e., the model corrects a previous incorrect recognition), it is determined that the recognition is correct; if the current recognition result is consistent with the historical output and the sample belongs to the set of failed recognition cases (i.e., the model still incorrectly recognizes whether the audio pair belongs to the same person), it is determined that the recognition is incorrect. Through this judgment logic, it is possible to clearly feedback whether the current recognition result of the model is correct, which is conducive to targeted optimization, helping the model learn how to correct previous mistakes while avoiding forgetting previously correctly recognized cases, thereby improving the overall performance of the model.
[0072] In one embodiment, as Figure 5 shown, step S203 includes:
[0073] S501. Calculate the corresponding advantage function through Generalized Advantage Estimation according to the recognition result and reward value of each sample;
[0074] S502. Calculate the policy gradient according to the advantage function, and update the parameters of the policy network according to the policy gradient.
[0075] In this embodiment, after the voiceprint model recognizes each sampled audio pair and obtains the corresponding reward value (positive reward or negative reward) according to whether the recognition result is correct, the corresponding advantage function is calculated through Generalized Advantage Estimation. Generalized Advantage Estimation (GAE) is a method for calculating the advantage function, which balances bias and variance by combining multi-step temporal difference errors, thereby improving the accuracy and stability of the advantage function estimation. Specifically, the advantage function A(s,a) is calculated through the following formula:
[0076]
[0077] where γ is the discount factor, λ is the GAE parameter, V(s t ) is the estimated value of the value function for processing the t-th sample, V(s t+1 ) is the estimated value of the value function for processing the t+1-th sample, T is the total number of samples for calculating the advantage function, R t+1It is the reward value obtained by processing the (t + 1)-th sample. By adjusting the GAE parameter λ, the bias and variance trade-off of GAE can be controlled. When λ = 0, GAE degenerates into a one-step advantage estimation. When λ = 1, GAE approaches the Monte Carlo estimation, thereby effectively reducing the variance in the policy update process and improving the stability of training.
[0078] Based on the calculated advantage function, the parameters of the policy network are updated according to the policy gradient method. The policy gradient method updates the parameters of the policy network by calculating the policy gradient. Specifically, it first calculates the derivative of the policy network output (action probability) with respect to the parameters, then multiplies the derivative of the action probability by the advantage function, and then takes the expectation over all states and actions to obtain the policy gradient. According to the calculated policy gradient, the parameters of the policy network are updated according to the gradient ascent method, and a learning rate is added during the update to control the step size of the parameter update, so that the policy network gradually tends to take actions with high rewards, and thus tends to take actions with greater advantages in subsequent decisions, so as to improve the performance of the policy network and enable it to select better actions in the voiceprint recognition task, thereby improving the performance of the voiceprint recognition model.
[0079] In one embodiment, as Figure 6 shown, step S204 includes:
[0080] S601. Call a pre-constructed objective function, where the objective function includes a KL divergence term for constraining the policy update amplitude;
[0081] S602. Calculate the probability ratio of the actions selected by the old and new policies and the KL divergence value according to the parameters of the policy network before the update and the parameters of the policy network after the update;
[0082] S603. Update the value of the objective function according to the probability ratio, the advantage function, and the KL divergence value.
[0083] In this embodiment, an objective function including a KL divergence term is predefined. After each update of the policy network based on the advantage function, the objective function is called to calculate the value of the objective function corresponding to the current policy adjustment, so as to evaluate each update of the policy network and ensure that the policy network is moving in the direction of maximizing the cumulative reward. By updating and optimizing the objective function, it is ensured that the policy network can learn the optimal policy. Specifically, the predefined objective function is:
[0084]
[0085] where is the probability ratio, ε is a truncation parameter, is the KL divergence, β is the weight parameter of the KL divergence, are the old parameters of the policy network, πθ is the new parameter of the policy network, is the output of the policy network under the old parameters, that is, the probability of taking action a in state s based on the old policy, π θ (a|s) is the output of the policy network under the new parameters, that is, the probability of taking action a in state s based on the new policy, and clip is the clipping function.
[0086] That is, when optimizing the objective function, according to the parameters of the policy network before and after the update, calculate the probability ratio of the old and new policies to select actions and the KL divergence value. The probability ratio refers to the probability ratio of the old and new policies to select the same action, which is used to measure the change in preference for the same action before and after the policy update. The KL divergence value measures the difference between the old and new policy distributions and is used to evaluate the magnitude of the policy update. By calculating the probability ratio, the change in preference for the same action before and after the policy update can be intuitively evaluated. By calculating the KL divergence value, the magnitude of the policy update can be quantified, providing a basis for the subsequent update of the objective function, limiting the magnitude of the policy update, and avoiding excessive model updates leading to forgetting. Then, according to the probability ratio, the advantage function, and the KL divergence value, update the value of the objective function according to the above formula. Evaluate the performance of the policy network through the objective function that includes the expectation of the reward value and the KL divergence term, providing a guiding basis for the optimization of the policy network and ensuring that the parameters of the policy network will be adjusted in the optimization direction.
[0087] In one embodiment, the method further includes:
[0088] Obtain pre-constructed validation set data;
[0089] Evaluate the performance of the voiceprint model using the validation set data according to a preset evaluation strategy;
[0090] Dynamically adjust the reward function used to calculate the reward value according to the performance evaluation result.
[0091] In this embodiment, after optimizing the voiceprint model based on reinforcement learning and KL divergence constraint, the performance of the voiceprint model is evaluated. For example, the performance of the voiceprint model is evaluated at fixed time intervals, or according to the amount of processed data of the voiceprint model, and the performance is evaluated when the processed data reaches a preset quantity, so as to ensure that the performance of the voiceprint model always meets the business requirements. Specifically, first collect validation set data from the actual business scenario or generate it through methods such as data augmentation. The validation set is a set of samples independent of the training data and is used to evaluate the performance of the model on unseen data. By obtaining the pre-constructed validation set data to provide an independent evaluation environment, the performance of the model can be evaluated more objectively. Based on the validation set data, the performance of the voiceprint model is evaluated according to a preset evaluation strategy. For example, input the validation set data into the voiceprint model for recognition, and calculate corresponding evaluation metrics such as accuracy, recall, F1 score, etc. after obtaining the recognition result, and confirm whether the evaluation metrics meet the preset metric thresholds. The performance of the model on the validation set is comprehensively evaluated through multiple evaluation metrics to ensure the reliability and effectiveness of the model in actual applications, and changes in the model performance can also be detected in a timely manner to further optimize the model performance.
[0092] Furthermore, with the optimization and performance evaluation of the voiceprint model, the reward function used to calculate the reward value is dynamically adjusted according to the performance evaluation results at different model stages to optimize the training direction of the model. Specifically, when the performance of the voiceprint model improves, the amplitude of the positive reward is correspondingly reduced and the amplitude of the negative reward is increased; when the performance of the voiceprint model decreases, the amplitude of the positive reward is correspondingly increased and the amplitude of the negative reward is reduced. For example, when it is evaluated that the accuracy or recall of the model is low, the amplitude of the positive reward can be increased and the amplitude of the negative reward can be reduced, etc. Since the performance of the model is different at different stages, by dynamically adjusting the reward function, the size and distribution of the reward value can be adjusted according to the current performance of the model to make it more adaptable to the training needs of the model, enabling the model to focus on different optimization goals at different stages, better balancing the accuracy and recall of the model, and thus improving the overall performance of the model.
[0093] It should be noted that there is not necessarily a certain order among the above steps. Those of ordinary skill in the art can understand according to the description of the embodiments of the present invention that in different embodiments, the above steps can have different execution orders, that is, they can be executed in parallel or exchanged, etc.
[0094] Further reference Figure 7 to Figure 2 As an implementation of the method shown above, the present invention provides an embodiment of a voiceprint model optimization device based on reinforcement learning. This device embodiment corresponds to the method embodiment shown in Figure 2 and this device can be specifically applied to various electronic devices.
[0095] As Figure 7 shown, the reinforcement learning-based voiceprint model optimization device 70 described in this embodiment includes:
[0096] A data acquisition module 701, configured to collect historical recognition cases of the voiceprint model, and construct an optimization data set according to the historical recognition cases;
[0097] An initialization module 702, configured to randomly initialize a policy network, sample audio data from the optimization data set and input it into the voiceprint model, and control the voiceprint model to perform voiceprint recognition according to the initialized policy to obtain the recognition result of each sample and the corresponding reward value;
[0098] A policy update module 703, configured to calculate an advantage function according to the recognition result and reward value of each sample, and update the parameters of the policy network based on the advantage function;
[0099] A function optimization module 704, configured to optimize a pre-constructed objective function according to the parameters before and after the update of the policy network, and the objective function includes a KL divergence term for constraining the policy update amplitude;
[0100] An iteration control module 705, configured to repeatedly execute the process of sampling audio data for voiceprint recognition and policy update until the objective function meets a preset condition to obtain an optimal policy network;
[0101] A model optimization module 706, configured to control the voiceprint model to perform voiceprint recognition according to the optimal policy network to optimize the recognition result.
[0102] The module referred to in the present invention refers to a series of computer program instruction segments that can complete specific functions, which is more suitable for describing the execution process of the reinforcement learning-based voiceprint model optimization than a program. For the specific implementation manners of each module, please refer to the corresponding method embodiments above and will not be elaborated here.
[0103] In one embodiment, the data acquisition module 701 includes:
[0104] A case acquisition unit, configured to collect historical recognition cases of the voiceprint model from a business scenario, and the historical recognition cases include a set of successful recognition cases and a set of failed recognition cases;
[0105] A preprocessing unit, configured to preprocess the audio data in each historical recognition case, and use the preprocessed audio data and the corresponding model historical output as an audio pair to construct an optimization database.
[0106] In one embodiment, the initialization module 702 includes:
[0107] A sampling input unit, configured to sample audio pairs from the optimized dataset, input the audio data in the audio pairs into the voiceprint model, control the voiceprint model to identify whether the audio data is the same person as the corresponding reference audio according to the initialization strategy, and obtain the recognition result of each sample;
[0108] A result judgment unit, configured to confirm whether the current voiceprint model is correctly recognized according to the recognition result of each sample, the corresponding model historical output, and the case set where the sample is located;
[0109] A reward unit, configured to give a positive reward when the voiceprint model is correctly recognized, and give a negative reward when the voiceprint model is incorrectly recognized, to obtain corresponding reward values.
[0110] In one embodiment, the policy update module 703 includes:
[0111] A first calculation unit, configured to calculate a corresponding advantage function by generalized advantage estimation according to the recognition result and reward value of each sample;
[0112] A policy parameter update unit, configured to calculate a policy gradient according to the advantage function, and update the parameters of the policy network according to the policy gradient.
[0113] In one embodiment, the function optimization module 704 includes:
[0114] A call unit, configured to call a pre-constructed objective function, where the objective function includes a KL divergence term for constraining the policy update amplitude;
[0115] A second calculation unit, configured to calculate the probability ratio of the old and new policies to select actions and the KL divergence value according to the parameters of the policy network before update and the parameters of the policy network after update;
[0116] An objective function update unit, configured to update the value of the objective function according to the probability ratio, the advantage function, and the KL divergence value.
[0117] In one embodiment, the apparatus 70 further includes:
[0118] An acquisition module, configured to acquire pre-constructed validation set data;
[0119] A performance evaluation module, configured to evaluate the performance of the voiceprint model using the validation set data according to a preset evaluation strategy;
[0120] A dynamic adjustment module, configured to dynamically adjust the reward function used to calculate the reward value according to the performance evaluation result.
[0121] In one embodiment, the dynamic adjustment module is specifically configured to:
[0122] When the performance of the voiceprint model improves, correspondingly reduce the amplitude of the positive reward and increase the amplitude of the negative reward;
[0123] When the performance of the voiceprint model decreases, correspondingly increase the amplitude of the positive reward and reduce the amplitude of the negative reward.
[0124] In the above embodiment, the present invention discloses a voiceprint model optimization device based on reinforcement learning. By collecting historical recognition cases of the voiceprint model, an optimization data set is constructed according to the historical recognition cases; the policy network is randomly initialized, and audio data is sampled from the optimization data set and input into the voiceprint model, and the voiceprint model is controlled to perform voiceprint recognition according to the initialized policy to obtain the recognition result of each sample and the corresponding reward value; the advantage function is calculated according to the recognition result and the reward value of each sample, and the parameters of the policy network are updated based on the advantage function; the pre-constructed objective function is optimized according to the parameters before and after the update of the policy network, and the objective function includes a KL divergence term for constraining the amplitude of policy update; the above process of sampling audio data for voiceprint recognition and policy update is cyclically executed until the optimal policy network is obtained when the objective function meets the preset conditions; based on the optimal policy network, the voiceprint model is controlled to perform voiceprint recognition according to the optimal policy to optimize the recognition result. By optimizing the voiceprint model through reinforcement learning and KL divergence constraint, the stability and efficiency of policy update are improved, reliable optimization of the voiceprint model is realized, and the recognition accuracy of the voiceprint model is improved.
[0125] Another embodiment of the present invention provides a computer device, as Figure 8 shown, the computer device 80 includes:
[0126] One or more processors 801 and a memory 802. Figure 8 Taking one processor 801 as an example for introduction, the processor 801 and the memory 802 can be connected by a bus or other means. Figure 8 Taking the connection by bus as an example.
[0127] The processor 801 is used to complete various control logics of the computer device 80, and it can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a single-chip microcomputer, an ARM (Acorn RISC Machine), or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination of these components. Additionally, the processor 801 can also be any conventional processor, microprocessor, or state machine. The processor 801 can also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in combination with a DSP and / or any other such configuration.
[0128] The memory 802, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions corresponding to the method for optimizing the voiceprint model based on reinforcement learning in the embodiments of the present invention. The processor 801 executes various functional applications and data processing of the computer device 80 by running the non-volatile software programs, instructions, and units stored in the memory 802, that is, implements the method for optimizing the voiceprint model based on reinforcement learning in the above method embodiments.
[0129] The memory 802 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device 80. In addition, the memory 802 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 802 optionally includes a memory remotely set relative to the processor 801, and these remote memories can be connected to the computer device 80 through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof. One or more units are stored in the memory 802 and, when executed by one or more processors 801, execute the steps of the method for optimizing the voiceprint model based on reinforcement learning in any of the above method embodiments.
[0130] In the above embodiments, the present invention discloses a computer device, which collects historical recognition cases of a voiceprint model, constructs an optimized data set according to the historical recognition cases; randomly initializes a policy network, samples audio data from the optimized data set and inputs it into the voiceprint model, controls the voiceprint model to perform voiceprint recognition according to the initialized policy, and obtains the recognition result and the corresponding reward value of each sample; calculates an advantage function according to the recognition result and the reward value of each sample, and updates the parameters of the policy network based on the advantage function; optimizes a pre-constructed objective function according to the parameters before and after the update of the policy network, and the objective function includes a KL divergence term for constraining the policy update amplitude; repeatedly executes the processes of sampling audio data for voiceprint recognition and policy update until the objective function meets the preset conditions to obtain an optimal policy network; based on the optimal policy network, controls the voiceprint model to perform voiceprint recognition according to the optimal policy to optimize the recognition result. By optimizing the voiceprint model through reinforcement learning and KL divergence constraint, the stability and efficiency of policy update are improved, reliable optimization of the voiceprint model is achieved, and the recognition accuracy of the voiceprint model is improved.
[0131] An embodiment of the present invention provides a non-volatile computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by one or more processors, the steps of the voiceprint model optimization method based on reinforcement learning in any of the above method embodiments are executed.
[0132] In the above embodiments, the present invention discloses a non-volatile computer-readable storage medium, which collects historical recognition cases of a voiceprint model, constructs an optimized data set according to the historical recognition cases; randomly initializes a policy network, samples audio data from the optimized data set and inputs it into the voiceprint model, controls the voiceprint model to perform voiceprint recognition according to the initialized policy, and obtains the recognition result and the corresponding reward value of each sample; calculates an advantage function according to the recognition result and the reward value of each sample, and updates the parameters of the policy network based on the advantage function; optimizes a pre-constructed objective function according to the parameters before and after the update of the policy network, and the objective function includes a KL divergence term for constraining the policy update amplitude; repeatedly executes the processes of sampling audio data for voiceprint recognition and policy update until the objective function meets the preset conditions to obtain an optimal policy network; based on the optimal policy network, controls the voiceprint model to perform voiceprint recognition according to the optimal policy to optimize the recognition result. By optimizing the voiceprint model through reinforcement learning and KL divergence constraint, the stability and efficiency of policy update are improved, reliable optimization of the voiceprint model is achieved, and the recognition accuracy of the voiceprint model is improved.
[0133] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0134] The present invention can be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0135] In summary, in the method, device, equipment, and medium for optimizing a voiceprint model based on reinforcement learning disclosed in the present invention, the method includes: collecting historical recognition cases of the voiceprint model, and constructing an optimization data set according to the historical recognition cases; randomly initializing a policy network, sampling audio data from the optimization data set and inputting it into the voiceprint model, controlling the voiceprint model to perform voiceprint recognition according to the initialized policy, and obtaining the recognition result and the corresponding reward value of each sample; calculating an advantage function according to the recognition result and reward value of each sample, and updating the parameters of the policy network based on the advantage function; optimizing a pre-constructed objective function according to the parameters before and after the update of the policy network, and the objective function includes a KL divergence term for constraining the policy update amplitude; repeatedly executing the process of sampling audio data for voiceprint recognition and policy update until the objective function meets a preset condition to obtain an optimal policy network; based on the optimal policy network, controlling the voiceprint model to perform voiceprint recognition according to the optimal policy to optimize the recognition result. By optimizing the voiceprint model through reinforcement learning and KL divergence constraint, the stability and efficiency of policy update are improved, reliable optimization of the voiceprint model is achieved, and the recognition accuracy of the voiceprint model is improved.
[0136] Of course, those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium, and when the computer program is executed, it can include the processes of the above method embodiments. The storage medium can be a memory, a magnetic disk, a floppy disk, a flash memory, an optical memory, etc.
[0137] It should be noted that if there are software tools or components of other companies in the embodiments of the present application, they are only used for illustrative introduction and do not represent actual use. It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.
Claims
1. An optimization method for a voiceprint model based on reinforcement learning, characterized in that Including: Collecting historical recognition cases of the voiceprint model, and constructing an optimized dataset according to the historical recognition cases; Randomly initializing a policy network, sampling audio data from the optimized dataset and inputting it into the voiceprint model, controlling the voiceprint model to perform voiceprint recognition according to the initialized policy, and obtaining the recognition result and corresponding reward value of each sample; Calculating an advantage function according to the recognition result and reward value of each sample, and updating the parameters of the policy network based on the advantage function; Optimizing a pre-constructed objective function according to the parameters of the policy network before and after the update, where the objective function includes a KL divergence term for constraining the policy update amplitude; Repeatedly execute the process of sampling audio data for voiceprint recognition and policy update until the objective function meets the preset conditions to obtain the optimal policy network; Based on the optimal policy network, controlling the voiceprint model to perform voiceprint recognition according to the optimal policy to optimize the recognition result.
2. The method for optimizing a voiceprint model based on reinforcement learning according to claim 1, wherein The collecting historical recognition cases of the voiceprint model and constructing an optimized dataset according to the historical recognition cases includes: Collecting historical recognition cases of the voiceprint model from the business scenario, where the historical recognition cases include a set of successful recognition cases and a set of failed recognition cases; Preprocessing the audio data in each historical recognition case, and using the preprocessed audio data and the corresponding model historical output as an audio pair to construct an optimized database.
3. The method for optimizing a voiceprint model based on reinforcement learning according to claim 2, wherein The sampling audio data from the optimized dataset and inputting it into the voiceprint model, controlling the voiceprint model to perform voiceprint recognition according to the initialized policy, and obtaining the recognition result and corresponding reward value of each sample includes: Sampling audio pairs from the optimized dataset, inputting the audio data in the audio pairs into the voiceprint model, and controlling the voiceprint model to recognize whether the audio data is the same person as the corresponding reference audio according to the initialized policy to obtain the recognition result of each sample; Confirming whether the current voiceprint model is recognized correctly according to the recognition result of each sample, the corresponding model historical output, and the case set where the sample is located; Giving a positive reward when the voiceprint model is recognized correctly, and giving a negative reward when the voiceprint model is recognized incorrectly to obtain the corresponding reward value.
4. The method for optimizing a voiceprint model based on reinforcement learning according to claim 1, characterized in that The calculating an advantage function according to the recognition result and reward value of each sample, and updating the parameters of the policy network based on the advantage function includes: Calculating the corresponding advantage function through generalized advantage estimation according to the recognition result and reward value of each sample; Calculating the policy gradient according to the advantage function, and updating the parameters of the policy network according to the policy gradient.
5. The method for optimizing a voiceprint model based on reinforcement learning according to claim 1, wherein The optimizing a pre-constructed objective function according to the parameters of the policy network before and after the update, where the objective function includes a KL divergence term for constraining the policy update amplitude includes: Invoking a pre-constructed objective function, where the objective function includes a KL divergence term for constraining the policy update amplitude; Calculating the probability ratio of the old and new policy selection actions and the KL divergence value according to the parameters of the policy network before the update and the parameters of the policy network after the update; Update the value of the objective function according to the probability ratio, the advantage function, and the KL divergence value.
6. The method for optimizing a voiceprint model based on reinforcement learning according to claim 1, wherein Further comprising: Obtain pre-constructed validation set data; Evaluate the performance of the voiceprint model using the validation set data according to a preset evaluation strategy; Dynamically adjust the reward function used to calculate the reward value according to the performance evaluation result.
7. The method for optimizing a voiceprint model based on reinforcement learning according to claim 6, characterized in that, The dynamically adjusting the reward function used to calculate the reward value according to the performance evaluation result specifically includes: When the performance of the voiceprint model improves, correspondingly reduce the amplitude of the positive reward and increase the amplitude of the negative reward; When the performance of the voiceprint model decreases, correspondingly increase the amplitude of the positive reward and reduce the amplitude of the negative reward.
8. An acoustic model optimization device based on reinforcement learning, characterized in that, Comprising: A data acquisition module, configured to acquire historical recognition cases of a voiceprint model, and construct an optimization data set according to the historical recognition cases; An initialization module, configured to randomly initialize a policy network, sample audio data from the optimization data set and input it into the voiceprint model, control the voiceprint model to perform voiceprint recognition according to the initialization policy, and obtain the recognition result of each sample and the corresponding reward value; A policy update module, configured to calculate an advantage function according to the recognition result and the reward value of each sample, and update the parameters of the policy network based on the advantage function; A function optimization module, configured to optimize a pre-constructed objective function according to the parameters before and after the update of the policy network, where the objective function includes a KL divergence term for constraining the policy update amplitude; An iteration control module, configured to repeatedly execute the process of performing voiceprint recognition on the sampled audio data and policy update until the objective function meets a preset condition to obtain an optimal policy network; A model optimization module, configured to control the voiceprint model to perform voiceprint recognition according to the optimal policy network based on the optimal policy network to optimize the recognition result.
9. A computer device, characterized in that, Comprising at least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method for optimizing a voiceprint model based on reinforcement learning according to any one of claims 1-7.
10. A non-volatile computer-readable storage medium, characterized in that, The non-volatile computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by one or more processors, the one or more processors can execute the method for optimizing a voiceprint model based on reinforcement learning according to any one of claims 1-7.
Citation Information
Patent Citations
Method for training voiceprint representation model and related device
CN110491393A
Improved reinforcement learning AGV path planning method
CN117826713A
Code generation model fine tuning method and device based on clustering and natural language strategy optimization algorithm
CN118468982A
Method and apparatus for service allocation based on reinforcement learning
WO2021208720A1