Voiceprint model optimization method and device based on reinforcement learning, equipment and medium
By using a reinforcement learning-based voiceprint model optimization method and leveraging historical recognition cases and KL divergence constraints, the policy network of the voiceprint model is optimized, solving the problems of high optimization difficulty and low recognition accuracy in existing technologies, and achieving more efficient recognition results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2026-04-07
AI Technical Summary
Existing voiceprint models are difficult to optimize in practical applications, lack effective means to deal with misjudgment cases, resulting in a decrease in recognition accuracy and an inability to adapt to new speech features or complex environments.
A reinforcement learning-based approach is adopted. An optimized dataset is constructed by collecting historical recognition cases of the voiceprint model. The policy network is randomly initialized for voiceprint recognition, the reward value is calculated and the policy network parameters are updated, and the policy update amplitude is constrained by the KL divergence term. The optimization is iterated until the preset conditions are met to obtain the optimal policy network.
It improves the recognition accuracy and stability of the voiceprint model, adapts to different speech environments, reduces misjudgments, and enhances the reliability and efficiency of recognition.
Smart Images

Figure CN120299461B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and medium for optimizing voiceprint models based on reinforcement learning. Background Technology
[0002] Voiceprints are the phonological features contained in speech that represent and identify the speaker. Voiceprint recognition is the process of identifying the speaker corresponding to a segment of speech based on its voiceprint features. Due to the unique nature of sound, voiceprint recognition, compared to other behavioral characteristics, possesses unique physiological characteristics. Therefore, theoretically, like fingerprints, very few two people have the same voiceprint features. This characteristic makes voiceprint recognition technology a promising emerging means of identity verification in risk control and customer authentication across multiple fields.
[0003] In the healthcare field, voiceprint recognition can be applied to remote medical consultations, such as when patients communicate with doctors about their condition via video or voice calls. To ensure the authenticity of the patient's identity and information security, a voiceprint recognition model can be used to verify the patient's identity. It can also be applied to, for example, the verification of operating permissions for high-precision medical equipment. Voiceprint recognition can be used to verify the identity of the operator to ensure that the operator has the corresponding qualifications and permissions.
[0004] In the financial sector, voiceprint recognition can be applied to telephone service scenarios such as insurance or banking. When customers conduct business over the phone, the voiceprint recognition model verifies their identity to ensure transaction security. Alternatively, during the insurance claims process, the voiceprint model enables insurance companies to more accurately identify customers and reduce misjudgments and fraud.
[0005] However, current voiceprint models still have some shortcomings in practical applications. Once a voiceprint model is deployed, optimization becomes difficult, and there is a lack of effective optimization methods for misjudgments. Retraining the voiceprint model may cause it to forget previously learned knowledge, thus affecting previously correctly judged cases. Therefore, this means that voiceprint recognition models may not be able to adjust in time when faced with new speech features or complex speech environments, thereby affecting recognition accuracy. Summary of the Invention
[0006] In view of the shortcomings of the prior art, the purpose of this invention is to provide a method, apparatus, device and medium for optimizing voiceprint models based on reinforcement learning that can be applied to the medical field, financial technology or other related fields. Its main purpose is to improve the recognition accuracy of voiceprint models by optimizing them efficiently and reliably.
[0007] The technical solution of the present invention is as follows:
[0008] The first aspect of this invention provides a method for optimizing a voiceprint model based on reinforcement learning, comprising:
[0009] Collect historical recognition cases of the voiceprint model, and construct an optimized dataset based on the historical recognition cases;
[0010] A random initialization policy network is used to sample audio data from the optimized dataset and input it into the voiceprint model. The voiceprint model is then controlled to perform voiceprint recognition according to the initialization policy to obtain the recognition result and corresponding reward value for each sample.
[0011] The advantage function is calculated based on the recognition result and reward value of each sample, and the parameters of the policy network are updated based on the advantage function.
[0012] The pre-constructed objective function is optimized based on the parameters before and after the policy network update. The objective function includes a KL divergence term used to constrain the magnitude of the policy update.
[0013] The process of sampling audio data and performing voiceprint recognition and policy update is repeated until the objective function meets the preset conditions to obtain the optimal policy network.
[0014] Based on the optimal policy network, the voiceprint model is controlled to perform voiceprint recognition according to the optimal policy in order to optimize the recognition results.
[0015] A second aspect of the present invention provides a reinforcement learning-based speaker model optimization device, comprising:
[0016] The data acquisition module is used to collect historical recognition cases of the voiceprint model and construct an optimized dataset based on the historical recognition cases;
[0017] An initialization module is used to randomly initialize the policy network, sample audio data from the optimization dataset and input it into the voiceprint model, control the voiceprint model to perform voiceprint recognition according to the initialization policy, and obtain the recognition result and corresponding reward value for each sample.
[0018] The policy update module is used to calculate the advantage function based on the recognition result and reward value of each sample, and update the parameters of the policy network based on the advantage function;
[0019] The function optimization module is used to optimize a pre-constructed objective function based on the parameters before and after the policy network update. The objective function includes a KL divergence term used to constrain the magnitude of the policy update.
[0020] The iterative control module is used to repeatedly execute the above-mentioned sampled audio data for voiceprint recognition and policy update until the objective function meets the preset conditions to obtain the optimal policy network.
[0021] The model optimization module is used to control the voiceprint model to perform voiceprint recognition according to the optimal strategy based on the optimal strategy network, so as to optimize the recognition results.
[0022] A third aspect of the present invention provides a computer device including at least one processor; and,
[0023] A memory communicatively connected to the at least one processor; wherein,
[0024] The memory stores instructions that can be executed by the at least one processor, which enables the at least one processor to perform the aforementioned reinforcement learning-based voiceprint model optimization method.
[0025] A fourth aspect of the present invention provides a non-volatile computer-readable storage medium storing computer-executable instructions, which, when executed by one or more processors, cause the one or more processors to perform the above-described reinforcement learning-based voiceprint model optimization method.
[0026] Beneficial Effects: This invention discloses a method, apparatus, device, and medium for optimizing a voiceprint model based on reinforcement learning. Compared with existing technologies, the embodiments of this invention, through a method, apparatus, device, and medium for optimizing a voiceprint model based on reinforcement learning, include: collecting historical recognition cases of the voiceprint model and constructing an optimization dataset based on the historical recognition cases; randomly initializing a policy network, sampling audio data from the optimization dataset and inputting it into the voiceprint model, controlling the voiceprint model to perform voiceprint recognition according to the initialization policy, and obtaining the recognition result and corresponding reward value for each sample; calculating an advantage function based on the recognition result and reward value of each sample, and updating the parameters of the policy network based on the advantage function; optimizing a pre-constructed objective function based on the parameters of the policy network before and after the update, wherein the objective function includes a KL divergence term used to constrain the magnitude of the policy update; cyclically executing the above process of sampling audio data for voiceprint recognition and policy update until the objective function meets preset conditions to obtain the optimal policy network; and controlling the voiceprint model to perform voiceprint recognition according to the optimal policy based on the optimal policy network to optimize the recognition result. By optimizing the voiceprint model using reinforcement learning and KL divergence constraints, the stability and efficiency of policy updates are improved, enabling reliable optimization of the voiceprint model and enhancing its recognition accuracy. Attached Figure Description
[0027] To more clearly illustrate the solutions in this invention, the accompanying drawings used in the description of the embodiments of this invention will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0028] Figure 1 A schematic diagram of an application environment for the voiceprint model optimization method based on reinforcement learning provided in an embodiment of the present invention;
[0029] Figure 2 A flowchart of a voiceprint model optimization method based on reinforcement learning provided in an embodiment of the present invention;
[0030] Figure 3 This is a flowchart of step S201 in the voiceprint model optimization method based on reinforcement learning provided in an embodiment of the present invention.
[0031] Figure 4 This is a flowchart of step S202 in the voiceprint model optimization method based on reinforcement learning provided in an embodiment of the present invention.
[0032] Figure 5 This is a flowchart of step S203 in the voiceprint model optimization method based on reinforcement learning provided in the embodiments of the present invention.
[0033] Figure 6 This is a flowchart of step S204 in the voiceprint model optimization method based on reinforcement learning provided in an embodiment of the present invention.
[0034] Figure 7 A schematic diagram of the functional modules of the voiceprint model optimization device based on reinforcement learning provided in an embodiment of the present invention;
[0035] Figure 8 A schematic diagram of the hardware structure of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0036] To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the invention is further described in detail below. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. The embodiments of the invention are described below with reference to the accompanying drawings.
[0037] The reinforcement learning-based voiceprint model optimization method provided in this invention can be applied to, for example... Figure 1In the application environment, it includes a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0038] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (for example only).
[0039] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0040] Server 105 can be a server providing various services, such as a backend server supporting the content browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend server can analyze and process received user requests and other data, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices. Server 105 can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the shortcomings of traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"), such as high management difficulty and weak business scalability. Server 105 can also be a server for a distributed system or a server combined with blockchain.
[0041] It should be noted that the reinforcement learning-based voiceprint model optimization method provided in this application embodiment can generally be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103. Correspondingly, the reinforcement learning-based voiceprint model optimization device provided in this embodiment can also be located in the first terminal device 101, the second terminal device 102, or the third terminal device 103. Alternatively, the reinforcement learning-based voiceprint model optimization method provided in this embodiment can generally be executed by the server 105. Correspondingly, the reinforcement learning-based voiceprint model optimization device provided in this embodiment can generally be located in the server 105.
[0042] It should be understood that the number of terminal devices, networks, and servers listed above is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be used.
[0043] like Figure 2 As shown, the voiceprint model optimization method based on reinforcement learning provided in this embodiment of the invention specifically includes the following steps:
[0044] S201. Collect historical recognition cases of the voiceprint model and construct an optimized dataset based on the historical recognition cases.
[0045] In this embodiment, to optimize the existing voiceprint model, historical recognition cases of the voiceprint model are collected from real-world business scenarios (such as medical consultation scenarios, financial insurance scenarios, etc.). Based on the audio data and corresponding recognition results already processed by the voiceprint model in actual business applications, corresponding historical recognition cases are collected. These historical recognition cases include correctly recognized cases and incorrectly recognized cases. By collecting these historical recognition cases, a dataset containing correctly and incorrectly recognized cases is constructed for subsequent model optimization training, providing a foundation for continuous model improvement. Preferably, in addition to directly using historical recognition cases, more diverse training data can be generated through data augmentation processing (such as adding noise, adjusting pitch, changing speed, etc.) to comprehensively cover various situations that the model may encounter in practical applications, enhancing the model's generalization ability.
[0046] For example, in healthcare scenarios, such as telemedicine consultations, voiceprint recognition results from patient-doctor conversations can be collected, including cases of correct patient identification and cases of misidentification, to build an optimized dataset. In financial scenarios, such as insurance claims services, voiceprint recognition results from customers processing claims can be collected, including cases of correct customer identification and cases of misidentification, to build an optimized dataset. Furthermore, in various scenarios, data augmentation methods such as adding background noise can be used to generate more diverse training data.
[0047] S202. Randomly initialize the policy network, sample audio data from the optimized dataset and input it into the voiceprint model, control the voiceprint model to perform voiceprint recognition according to the initialization policy, and obtain the recognition result and corresponding reward value for each sample.
[0048] In this embodiment, the policy network is a deep learning model, such as a convolutional neural network (CNN), a recurrent neural network (RNN), or other network structures suitable for processing audio data. It is used to output the decision strategy for the voiceprint model. The parameters of the policy network determine how the voiceprint model identifies and judges the input audio data, that is, whether the input audio data and the reference audio (such as audio pre-recorded by the customer) belong to the same person. When optimizing the voiceprint model, the parameters of the policy network are first randomly initialized. A batch of audio data is randomly sampled from the optimization dataset as samples and input into the voiceprint model. For each input sample, the voiceprint model performs voiceprint recognition using the initialization strategy corresponding to the current policy network, thereby obtaining the recognition result and corresponding reward value for each sample, providing corresponding feedback signals for subsequent policy updates.
[0049] The reward value is obtained through a pre-defined reward function based on business requirements. For example, a positive reward (e.g., +1) is given when the voiceprint model correctly identifies the voiceprint, while a negative reward (e.g., -1) is given for incorrect identification. By designing positive and negative reward functions, the model can learn from correct and incorrect cases and update its strategy accordingly, thereby effectively improving the recognition accuracy. Of course, in other embodiments, other reward functions can be sampled to obtain reward values, such as giving different levels of negative rewards based on the severity of misjudgment, or giving rewards based on the confidence level of recognition, etc. This embodiment does not limit this.
[0050] For example, in telemedicine consultations, after randomly initializing the policy network, audio data from patient-doctor conversations is sampled from the optimized dataset and input into the voiceprint model. A reward value is calculated based on the recognition results; for instance, if the voiceprint model correctly identifies the patient using the input audio data, the reward is +1, and if it misidentifies, the reward is -1. Similarly, in financial scenarios such as claims reporting or telephone banking, after randomly initializing the policy network, audio data from customers conducting business over the phone is sampled from the optimized dataset and input into the voiceprint model. If the voiceprint model accurately identifies the customer, the reward is +1, and if it misidentifies, the reward is -1.
[0051] S203. Calculate the advantage function based on the recognition result and reward value of each sample, and update the parameters of the policy network based on the advantage function.
[0052] In this embodiment, the advantage function refers to the advantage of taking a certain action relative to the average reward in a given state. Specifically, the advantage function can be estimated through methods such as Generalized Advantage Estimation (GAE), Monte Carlo Tree Search (MCTS), and Temporal Difference (TD) learning; this embodiment does not limit the specific method used. Since the advantage function can more accurately evaluate the merits of each action, this embodiment calculates the advantage function based on the recognition result and reward value of each sample, and adjusts the parameters of the policy network according to the value of the advantage function. That is, the advantage function guides policy improvement, fully utilizing the sampled data to make the model more inclined to take actions with a greater advantage, thereby more effectively updating the policy network to optimize the recognition performance of the voiceprint model.
[0053] S204. Optimize the pre-constructed objective function based on the parameters before and after the policy network update. The objective function includes a KL divergence term used to constrain the policy update magnitude.
[0054] In this embodiment, a pre-constructed objective function is used to calculate the expected value of the cumulative reward to evaluate the performance of the policy network. The objective function value is calculated based on the parameters of the policy network before and after the update. The performance of the policy network is then evaluated based on the objective function value. This maximizes the expected value of the cumulative reward to guide the entire optimization process, ensuring that the policy network's learning direction is towards maximizing the cumulative reward. Optimizing the objective function ensures that the policy network can learn the optimal policy, thereby improving the performance of the voiceprint recognition model. Simultaneously, to prevent excessively large policy updates from causing the model to forget previously learned knowledge and thus affecting previously correctly judged cases, the pre-constructed objective function in this embodiment includes a KL divergence term. KL divergence specifically refers to Kullback-Leibler divergence, an asymmetric measure of the difference between two probability distributions. By limiting the KL divergence between the old and new policies, the magnitude of policy updates is ensured to be moderate, improving the model's stability during the optimization process.
[0055] S205. Repeat the above process of sampling audio data for voiceprint recognition and policy update until the objective function meets the preset conditions to obtain the optimal policy network.
[0056] Repeat the above process of sampling, identification, calculation of the advantage function, updating the policy network, and calculation of the objective function until the objective function meets certain convergence conditions, such as the parameter update magnitude of the policy network being less than a certain threshold, or the value of the objective function reaching a certain stable value, or reaching a preset number of iterations, etc., thereby obtaining the optimal policy network.
[0057] In each iteration, the policy network selects actions based on the latest parameters to guide the voiceprint recognition model in recognizing audio data. Through iterative iteration, the policy network is gradually optimized to achieve optimal performance after multiple updates. Specifically, in the initial stage, the parameters of the policy network are randomly initialized, so the actions it selects may not be ideal, resulting in poor performance of the voiceprint model. In the intermediate stage, as the policy network is continuously updated, it selects more reasonable actions after each update, gradually improving the performance of the voiceprint recognition model. In the convergence stage, when the parameters of the policy network converge to the optimal state, it selects the optimal action, dynamically adjusting the behavior of the voiceprint recognition model based on reinforcement learning optimization methods to maximize its performance.
[0058] S206. Based on the optimal policy network, control the voiceprint model to perform voiceprint recognition according to the optimal policy in order to optimize the recognition results.
[0059] In this embodiment, after multiple iterations and optimizations, the optimal policy network is finally obtained. This optimized policy network is then deployed in a practical application for voiceprint recognition tasks. The optimal policy guides the optimization of the voiceprint recognition model, thereby improving the model's recognition results and enhancing the accuracy and reliability of voiceprint recognition. The voiceprint model optimization method based on reinforcement learning and KL divergence constraints is particularly suitable for remote identity verification scenarios in insurance business, such as claims reporting and telephone sales. It can quickly identify customer identities, reduce misjudgments and fraudulent behavior, and improve customer experience and service quality.
[0060] In the above embodiments, this invention discloses a voiceprint model optimization method based on reinforcement learning. The method involves collecting historical recognition cases from the voiceprint model and constructing an optimization dataset based on these cases; randomly initializing a policy network; sampling audio data from the optimization dataset and inputting it into the voiceprint model; controlling the voiceprint model to perform voiceprint recognition according to the initialization policy; obtaining the recognition result and corresponding reward value for each sample; calculating an advantage function based on the recognition result and reward value for each sample; updating the parameters of the policy network based on the advantage function; optimizing a pre-constructed objective function based on the parameters before and after the policy network update, where the objective function includes a KL divergence term to constrain the policy update magnitude; iteratively executing the above process of sampling audio data for voiceprint recognition and policy update until the objective function meets preset conditions to obtain the optimal policy network; and controlling the voiceprint model to perform voiceprint recognition according to the optimal policy based on the optimal policy network to optimize the recognition results. By optimizing the voiceprint model through reinforcement learning and KL divergence constraints, the stability and efficiency of policy updates are improved, achieving reliable optimization of the voiceprint model and improving the recognition accuracy of the voiceprint model.
[0061] In one embodiment, such as Figure 3 As shown, step S201 includes:
[0062] S301. Collect historical recognition cases of the voiceprint model from business scenarios, wherein the historical recognition cases include a set of successful recognition cases and a set of failed recognition cases;
[0063] S302. Preprocess the audio data in each historical recognition case, and use the preprocessed audio data and the corresponding historical output of the model as audio pairs to construct an optimized database.
[0064] In this embodiment, when constructing the dataset, data is collected from the actual application environment of the voiceprint model, such as business scenarios like telemedicine consultation and telephone banking services. These scenarios have accumulated a large amount of voiceprint recognition data. The recognition records of the voiceprint model are extracted from the actual business system, including audio files and model output results. Each case is labeled to clarify whether it is a successful case or a failed case. The historical recognition cases are divided into two categories: a set of successful recognition cases and a set of failed recognition cases. That is, cases where the model correctly identifies whether the audio pair belongs to the same person, and cases where the model incorrectly identifies whether the audio pair belongs to the same person. By collecting successful and failed cases, we can comprehensively cover various situations that the model may encounter in actual applications, providing rich data support for subsequent optimization.
[0065] Next, the audio data from historical recognition cases undergoes preprocessing to improve data quality and enhance model training performance. Specifically, this preprocessing can be achieved through sampling quantization, pre-emphasis, windowing, and noise removal. Sampling quantization converts the audio signal into a digital signal and adjusts the sampling rate to suit the model input; pre-emphasis enhances high-frequency signals and reduces transmission loss; windowing divides the audio signal into multiple short-time windows for easier analysis; and noise removal uses filtering and other methods to remove background noise and improve audio signal clarity. The preprocessed audio data is then combined with the model's historical outputs (i.e., previous recognition results) to form a data pair. For example, for an audio pair, the model's historical output might be "same person" or "different person," indicating whether the judgment was correct or incorrect. This means the case belongs to either the set of successful or failed recognition cases, providing high-quality training data for voiceprint model optimization while ensuring data diversity and relevance, laying a solid foundation for subsequent model optimization.
[0066] In one embodiment, such as Figure 4 As shown, step S202 includes:
[0067] S401. Sample audio pairs from the optimized dataset, input the audio data of the audio pairs into the voiceprint model, control the voiceprint model to identify whether the audio data and the corresponding reference audio belong to the same person according to the initialization strategy, and obtain the recognition result of each sample;
[0068] S402. Based on the recognition result of each sample, the corresponding historical output of the model, and the case set in which the sample is located, confirm whether the current voiceprint model has correctly identified the voiceprint.
[0069] S403. When the voiceprint model is correctly identified, a positive reward is given; when the voiceprint model is incorrectly identified, a negative reward is given, and a corresponding reward value is obtained.
[0070] In this embodiment, audio pairs (s, a) are sampled from the optimized dataset as training samples, where s is the input audio data (including the audio to be identified and the corresponding reference audio), and a is the model's historical output (i.e., the recognition results of the voiceprint model in previous business applications). Specifically, random sampling, stratified sampling, or other methods can be used to ensure the diversity and representativeness of the samples. The sampled audio data is input into the voiceprint model, which identifies the input audio data according to the initialization strategy and outputs a judgment result on whether the audio pair belongs to the same person, thus obtaining the recognition result for each sample. This yields the initial performance of the model, providing a foundation for subsequent reward calculation and strategy updates. Then, by comparing the current recognition result with the model's historical output and considering the case set in which the sample belongs, it is determined whether the current recognition is correct. Based on the pre-set reward function, a corresponding reward value is given, and the reward value is associated with the recognition result of each sample. Specifically, when the voiceprint model correctly identifies whether the audio pair belongs to the same person, a positive reward value (such as +1) is given, indicating that the model's decision is correct; when the voiceprint model incorrectly identifies whether the audio pair belongs to the same person, a negative reward value (such as -1) is given, indicating that the model's decision is incorrect. Through positive and negative reward values, clear feedback signals are provided to the model to help the model learn the correct decision-making pattern, and the reward value serves as an optimization guide, leading the model to be more inclined to make correct decisions in subsequent training.
[0071] Specifically, when determining whether the current voiceprint model's recognition is correct, the following logic applies: if the current recognition result is consistent with historical outputs and the sample belongs to the set of successful recognition cases (i.e., the model correctly identified whether two audio recordings belong to the same person), then the recognition is considered correct. If the current recognition result is inconsistent with historical outputs and the sample belongs to the set of successful recognition cases (i.e., the model incorrectly identified whether two audio recordings belong to the same person), then the recognition is considered incorrect. If the current recognition result is inconsistent with historical outputs and the sample belongs to the set of failed recognition cases (i.e., the model corrected its previous incorrect recognition), then the recognition is considered correct. If the current recognition result is consistent with historical outputs and the sample belongs to the set of failed recognition cases (i.e., the model still incorrectly identified whether the audio pair belongs to the same person), then the recognition is considered incorrect. This judgment logic clearly reflects whether the model's current recognition result is correct and facilitates targeted optimization. It helps the model learn how to correct previous errors while avoiding forgetting previously correctly recognized cases, thereby improving the overall performance of the model.
[0072] In one embodiment, such as Figure 5 As shown, step S203 includes:
[0073] S501. Based on the identification result and reward value of each sample, calculate the corresponding advantage function through generalized advantage estimation;
[0074] S502. Calculate the policy gradient based on the advantage function, and update the parameters of the policy network based on the policy gradient.
[0075] In this embodiment, after the voiceprint model identifies each sampled audio pair and obtains the corresponding reward value (positive or negative reward) based on whether the identification result is correct, the corresponding advantage function is calculated through generalized advantage estimation. Generalized advantage estimation (GAE) is a method for calculating the advantage function. It balances bias and variance by combining multi-step temporal difference errors, thereby improving the accuracy and stability of the advantage function estimation. Specifically, the advantage function A(s,a) is calculated using the following formula:
[0076]
[0077] Where γ is the discount factor, λ is the GAE parameter, and V(s) t V(s) is the estimated value of the value function for processing the t-th sample. t+1 R is the estimated value of the value function for the (t+1)th sample, where T is the total number of samples used to calculate the dominance function, and R is the sum of the values of the samples used to calculate the dominance function. t+1This is the reward value obtained from processing the (t+1)th sample. By adjusting the GAE parameter λ, the bias and variance tradeoff of GAE are controlled. When λ = 0, GAE degenerates into a one-step advantage estimate, and when λ = 1, GAE approaches Monte Carlo estimation, thereby effectively reducing the variance in the policy update process and improving the stability of training.
[0078] The parameters of the policy network are updated using the policy gradient method based on the calculated advantage function. This method updates the network parameters by calculating the policy gradient. Specifically, it first calculates the derivative of the policy network output (action probability) with respect to the parameters, then multiplies the derivative of the action probability with the advantage function, and finally calculates the expectation over all states and actions to obtain the policy gradient. The parameters are then updated using the gradient ascent method based on this gradient. A learning rate is incorporated during the update to control the step size of the parameter updates, gradually making the policy network more inclined to take high-reward actions. This leads to a greater preference for actions with a higher advantage in subsequent decisions, improving the network's performance and enabling it to select better actions in the voiceprint recognition task, thereby enhancing the performance of the voiceprint recognition model.
[0079] In one embodiment, such as Figure 6 As shown, step S204 includes:
[0080] S601. Call the pre-built objective function, which includes a KL divergence term for constraining the policy update magnitude;
[0081] S602. Based on the parameters of the policy network before the update and the parameters of the policy network after the update, calculate the probability ratio of the new and old policies for selecting actions and the KL divergence value.
[0082] S603. Update the value of the objective function based on the probability ratio, the dominance function, and the KL divergence value.
[0083] In this embodiment, a predefined objective function including a KL divergence term is used. After each update of the policy network based on the dominance function, this objective function is called to calculate the value of the objective function corresponding to this policy adjustment, thereby evaluating each policy network update and ensuring that the policy network moves in the direction of maximizing cumulative reward. By updating and optimizing the objective function, the policy network is ensured to learn the optimal policy. Specifically, the predefined objective function is:
[0084]
[0085] in, It is the probability ratio, and ε is the cutoff parameter. β is the KL divergence, and β is the weighting parameter of the KL divergence. These are the old parameters of the policy network, π.θ These are new parameters for the policy network. This is the output of the policy network under the old parameters, that is, the probability of taking action a in state s based on the old policy, π. θ (a|s) is the output of the policy network under the new parameters, that is, the probability of taking action a in state s based on the new policy, and clip is the clipping function.
[0086] In optimizing the objective function, based on the parameters of the policy network before and after the update, the probability ratio of the new and old policies in selecting the same action and the KL divergence value are calculated. The probability ratio represents the ratio of the probability of selecting the same action under the new and old policies, measuring the change in preference for the same action before and after the policy update. The KL divergence value measures the difference between the distributions of the new and old policies, evaluating the magnitude of the policy update. Calculating the probability ratio allows for a direct assessment of the change in preference for the same action before and after the policy update, while calculating the KL divergence value quantifies the magnitude of the policy update, providing a basis for subsequent objective function updates, limiting the magnitude of policy updates, and preventing excessive model updates that could lead to forgetting. Then, the objective function is updated according to the above formula based on the probability ratio, dominance function, and KL divergence value. The performance of the policy network is evaluated using an objective function that includes the expected reward value and the KL divergence term, providing guidance for policy network optimization and ensuring that the policy network parameters are adjusted in the optimization direction.
[0087] In one embodiment, the method further includes:
[0088] Obtain pre-built validation set data;
[0089] The performance of the voiceprint model is evaluated using the validation set data according to a preset evaluation strategy.
[0090] The reward function used to calculate the reward value is dynamically adjusted based on the performance evaluation results.
[0091] In this embodiment, after optimizing the voiceprint model based on reinforcement learning and KL divergence constraints, the performance of the voiceprint model is evaluated. For example, the performance can be evaluated at fixed time periods or based on the amount of data processed, with performance evaluation performed when a preset amount of data is processed, ensuring that the voiceprint model's performance always meets business requirements. Specifically, validation set data is first collected from actual business scenarios or generated through data augmentation methods. The validation set is a set of samples independent of the training data, used to evaluate the model's performance on unseen data. Obtaining pre-constructed validation set data provides an independent evaluation environment, enabling a more objective evaluation of the model's performance. Based on the validation set data, the performance of the voiceprint model is evaluated according to a preset evaluation strategy. For example, the validation set data is input into the voiceprint model for recognition. After obtaining the recognition results, corresponding evaluation metrics are calculated, such as accuracy, recall, and F1 score. It is confirmed whether the evaluation metrics meet preset thresholds. By comprehensively evaluating the model's performance on the validation set using multiple evaluation metrics, the reliability and effectiveness of the model in practical applications are ensured, and changes in model performance can be detected in a timely manner for further optimization.
[0092] Furthermore, as the voiceprint model is optimized and its performance evaluated, the reward function used to calculate the reward value is dynamically adjusted based on the performance evaluation results at different model stages to optimize the model's training direction. Specifically, when the performance of the voiceprint model improves, the magnitude of positive rewards is reduced and the magnitude of negative rewards is increased; when the performance of the voiceprint model decreases, the magnitude of positive rewards is increased and the magnitude of negative rewards is decreased. For example, when the accuracy or recall of the model is low, the magnitude of positive rewards can be increased and the magnitude of negative rewards reduced, and so on. Since the model's performance varies at different stages, dynamically adjusting the reward function can adjust the magnitude and distribution of reward values according to the model's current performance to better adapt to the model's training needs. This allows the model to focus on different optimization objectives at different stages, better balancing the model's accuracy and recall, thereby improving the overall performance of the model.
[0093] It should be noted that there is no necessary order between the above steps. Those skilled in the art will understand from the description of the embodiments of the present invention that the above steps may have different execution orders in different embodiments, that is, they may be executed in parallel or in turn, etc.
[0094] Further reference Figure 7 As a response to the above Figure 2 The present invention provides an embodiment of a reinforcement learning-based voiceprint model optimization device, which implements the method shown. Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0095] like Figure 7 As shown, the reinforcement learning-based voiceprint model optimization device 70 described in this embodiment includes:
[0096] Data acquisition module 701 is used to collect historical recognition cases of voiceprint models and construct an optimized dataset based on the historical recognition cases;
[0097] The initialization module 702 is used to randomly initialize the policy network, sample audio data from the optimized dataset and input it into the voiceprint model, control the voiceprint model to perform voiceprint recognition according to the initialization policy, and obtain the recognition result and corresponding reward value for each sample.
[0098] The strategy update module 703 is used to calculate the advantage function based on the recognition result and reward value of each sample, and update the parameters of the strategy network based on the advantage function;
[0099] The function optimization module 704 is used to optimize the pre-constructed objective function based on the parameters before and after the policy network update. The objective function includes a KL divergence term used to constrain the policy update magnitude.
[0100] The iterative control module 705 is used to repeatedly execute the above-mentioned process of sampling audio data for voiceprint recognition and policy update until the objective function meets the preset conditions to obtain the optimal policy network.
[0101] The model optimization module 706 is used to control the voiceprint model to perform voiceprint recognition according to the optimal strategy based on the optimal strategy network, so as to optimize the recognition results.
[0102] The module referred to in this invention is a series of computer program instruction segments that can perform specific functions. It is more suitable than a program for describing the optimization execution process of the voiceprint model based on reinforcement learning. For the specific implementation of each module, please refer to the corresponding method embodiments above, which will not be repeated here.
[0103] In one embodiment, the data acquisition module 701 includes:
[0104] The case collection unit is used to collect historical recognition cases of the voiceprint model from business scenarios. The historical recognition cases include a set of successful recognition cases and a set of failed recognition cases.
[0105] The preprocessing unit is used to preprocess the audio data in each historical recognition case, and to construct an optimized database by using the preprocessed audio data and the corresponding historical output of the model as audio pairs.
[0106] In one embodiment, the initialization module 702 includes:
[0107] The sampling input unit is used to sample audio pairs from the optimized dataset, input the audio data of the audio pairs into the voiceprint model, control the voiceprint model to identify whether the audio data and the corresponding reference audio belong to the same person according to the initialization strategy, and obtain the recognition result of each sample;
[0108] The result judgment unit is used to confirm whether the current voiceprint model is correctly identified based on the recognition result of each sample, the corresponding historical output of the model, and the case set to which the sample belongs.
[0109] The reward unit is used to give a positive reward when the voiceprint model is correctly identified and a negative reward when the voiceprint model is incorrectly identified, thereby obtaining the corresponding reward value.
[0110] In one embodiment, the policy update module 703 includes:
[0111] The first calculation unit is used to calculate the corresponding advantage function based on the recognition result and reward value of each sample through generalized advantage estimation.
[0112] The policy parameter update unit is used to calculate the policy gradient based on the advantage function and update the parameters of the policy network based on the policy gradient.
[0113] In one embodiment, the function optimization module 704 includes:
[0114] The calling unit is used to call a pre-built objective function, which includes a KL divergence term for constraining the policy update magnitude;
[0115] The second calculation unit is used to calculate the probability ratio of the new and old policy selection actions and the KL divergence value based on the parameters of the policy network before the update and the parameters of the policy network after the update.
[0116] The objective function update unit is used to update the value of the objective function based on the probability ratio, the dominance function, and the KL divergence value.
[0117] In one embodiment, the device 70 further includes:
[0118] The acquisition module is used to acquire pre-built validation set data;
[0119] The performance evaluation module is used to evaluate the performance of the voiceprint model using the validation set data according to a preset evaluation strategy.
[0120] A dynamic adjustment module is used to dynamically adjust the reward function used to calculate the reward value based on the performance evaluation results.
[0121] In one embodiment, the dynamic adjustment module is specifically used for:
[0122] When the performance of the voiceprint model improves, the magnitude of positive rewards is reduced and the magnitude of negative rewards is increased accordingly.
[0123] When the performance of the voiceprint model decreases, the magnitude of positive rewards is increased and the magnitude of negative rewards is decreased accordingly.
[0124] In the above embodiments, this invention discloses a voiceprint model optimization device based on reinforcement learning. The device collects historical recognition cases of the voiceprint model and constructs an optimization dataset based on these cases. A policy network is randomly initialized, and audio data is sampled from the optimization dataset and input into the voiceprint model. The voiceprint model is controlled to perform voiceprint recognition according to the initialized policy, obtaining the recognition result and corresponding reward value for each sample. An advantage function is calculated based on the recognition result and reward value of each sample, and the policy network parameters are updated based on the advantage function. A pre-constructed objective function is optimized based on the parameters of the policy network before and after the update. The objective function includes a KL divergence term to constrain the policy update amplitude. The process of sampling audio data for voiceprint recognition and policy update is repeated until the objective function meets preset conditions to obtain the optimal policy network. Based on the optimal policy network, the voiceprint model is controlled to perform voiceprint recognition according to the optimal policy to optimize the recognition result. By optimizing the voiceprint model through reinforcement learning and KL divergence constraints, the stability and efficiency of policy updates are improved, achieving reliable optimization of the voiceprint model and improving the recognition accuracy of the voiceprint model.
[0125] Another embodiment of the present invention provides a computer device, such as... Figure 8 As shown, the computer device 80 includes:
[0126] One or more processors 801 and memory 802, Figure 8 The following section uses a processor 801 as an example. The processor 801 and the memory 802 can be connected via a bus or other means. Figure 8 Taking the example of a connection between China and Israel via a bus.
[0127] Processor 801 is used to perform various control logics of computer device 80, and can be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), microcontroller, ARM (Acorn RISC Machine) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of these components. Furthermore, processor 801 can also be any conventional processor, microprocessor, or state machine. Processor 801 can also be implemented as a combination of computing devices, such as a combination of DSP and microprocessor, multiple microprocessors, one or more microprocessors combined with DSP and / or any other such configuration.
[0128] The memory 802, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions corresponding to the reinforcement learning-based voiceprint model optimization method in this embodiment of the invention. The processor 801 executes various functional applications and data processing of the computer device 80 by running the non-volatile software programs, instructions, and units stored in the memory 802, thereby implementing the reinforcement learning-based voiceprint model optimization method in the above method embodiment.
[0129] The memory 802 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device 80. Furthermore, the memory 802 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 802 may optionally include memory remotely located relative to the processor 801, and these remote memories may be connected to the computer device 80 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. One or more units stored in the memory 802, when executed by one or more processors 801, perform the steps of the reinforcement learning-based voiceprint model optimization method in any of the above method embodiments.
[0130] In the above embodiments, the present invention discloses a computer device that collects historical recognition cases of a voiceprint model and constructs an optimized dataset based on these cases; randomly initializes a policy network, samples audio data from the optimized dataset and inputs it into the voiceprint model, controls the voiceprint model to perform voiceprint recognition according to the initialized policy, and obtains the recognition result and corresponding reward value for each sample; calculates an advantage function based on the recognition result and reward value of each sample, and updates the parameters of the policy network based on the advantage function; optimizes a pre-constructed objective function based on the parameters of the policy network before and after the update, the objective function including a KL divergence term to constrain the policy update magnitude; iteratively executes the above process of sampling audio data for voiceprint recognition and policy update until the objective function meets preset conditions to obtain the optimal policy network; based on the optimal policy network, controls the voiceprint model to perform voiceprint recognition according to the optimal policy to optimize the recognition result. By optimizing the voiceprint model through reinforcement learning and KL divergence constraints, the stability and efficiency of policy updates are improved, reliable optimization of the voiceprint model is achieved, and the recognition accuracy of the voiceprint model is improved.
[0131] This invention provides a non-volatile computer-readable storage medium storing computer-executable instructions. When these computer-executable instructions are executed by one or more processors, they perform the steps of the reinforcement learning-based voiceprint model optimization method described in any of the above method embodiments.
[0132] In the above embodiments, the present invention discloses a non-volatile computer-readable storage medium. It collects historical recognition cases from a voiceprint model and constructs an optimized dataset based on these cases. A policy network is randomly initialized, and audio data is sampled from the optimized dataset and input into the voiceprint model. The voiceprint model is controlled to perform voiceprint recognition according to the initialized policy, obtaining the recognition result and corresponding reward value for each sample. An advantage function is calculated based on the recognition result and reward value of each sample, and the policy network parameters are updated based on the advantage function. A pre-constructed objective function is optimized based on the parameters of the policy network before and after the update. The objective function includes a KL divergence term to constrain the policy update magnitude. The process of sampling audio data for voiceprint recognition and policy update is repeated until the objective function meets preset conditions to obtain the optimal policy network. Based on the optimal policy network, the voiceprint model is controlled to perform voiceprint recognition according to the optimal policy to optimize the recognition result. By optimizing the voiceprint model through reinforcement learning and KL divergence constraints, the stability and efficiency of policy updates are improved, achieving reliable optimization of the voiceprint model and improving the recognition accuracy of the voiceprint model.
[0133] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0134] This invention can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This invention can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0135] In summary, the method, apparatus, device, and medium for optimizing a voiceprint model based on reinforcement learning disclosed in this invention include: collecting historical recognition cases of the voiceprint model and constructing an optimization dataset based on the historical recognition cases; randomly initializing a policy network, sampling audio data from the optimization dataset and inputting it into the voiceprint model, controlling the voiceprint model to perform voiceprint recognition according to the initialization policy, and obtaining the recognition result and corresponding reward value for each sample; calculating an advantage function based on the recognition result and reward value for each sample, and updating the parameters of the policy network based on the advantage function; optimizing a pre-constructed objective function based on the parameters of the policy network before and after the update, wherein the objective function includes a KL divergence term used to constrain the magnitude of the policy update; cyclically executing the above process of sampling audio data for voiceprint recognition and policy update until the objective function meets preset conditions to obtain an optimal policy network; and controlling the voiceprint model to perform voiceprint recognition according to the optimal policy based on the optimal policy network to optimize the recognition result. By optimizing the voiceprint model through reinforcement learning and KL divergence constraints, the stability and efficiency of policy updates are improved, reliable optimization of the voiceprint model is achieved, and the recognition accuracy of the voiceprint model is improved.
[0136] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.). The computer program can be stored in a non-volatile, computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The storage medium can be a memory, magnetic disk, floppy disk, flash memory, optical storage, etc.
[0137] It should be noted that any software tools or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. It should be understood that the application of this invention is not limited to the examples described above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.
Claims
1. A method for optimizing a voiceprint model based on reinforcement learning, characterized in that, include: Collect historical recognition cases of the voiceprint model, and construct an optimized dataset based on the historical recognition cases; A random initialization policy network is used to sample audio data from the optimized dataset and input it into the voiceprint model. The voiceprint model is then controlled to perform voiceprint recognition according to the initialization policy to obtain the recognition result and corresponding reward value for each sample. The advantage function is calculated based on the recognition result and reward value of each sample, and the parameters of the policy network are updated based on the advantage function. The pre-constructed objective function is optimized based on the parameters before and after the policy network update. The objective function includes a KL divergence term used to constrain the magnitude of the policy update. The process of sampling audio data and performing voiceprint recognition and policy update is repeated until the objective function meets the preset conditions to obtain the optimal policy network. Based on the optimal policy network, the voiceprint model is controlled to perform voiceprint recognition according to the optimal policy in order to optimize the recognition results.
2. The voiceprint model optimization method based on reinforcement learning according to claim 1, characterized in that, The historical recognition cases collected by the voiceprint model are used to construct an optimized dataset, including: Collect historical recognition cases of the voiceprint model from business scenarios. The historical recognition cases include a set of successful recognition cases and a set of failed recognition cases. The audio data in each historical recognition case is preprocessed, and the preprocessed audio data and the corresponding historical output of the model are used as audio pairs to construct an optimized database.
3. The voiceprint model optimization method based on reinforcement learning according to claim 2, characterized in that, The process of sampling audio data from the optimized dataset and inputting it into the voiceprint model, controlling the voiceprint model to perform voiceprint recognition according to the initialization strategy, and obtaining the recognition result and corresponding reward value for each sample includes: Audio pairs are sampled from the optimized dataset, and the audio data in the audio pairs is input into the voiceprint model. The voiceprint model is controlled to identify whether the audio data and the corresponding reference audio belong to the same person according to the initialization strategy, so as to obtain the recognition result of each sample. Based on the recognition result of each sample, the corresponding historical output of the model, and the case set to which the sample belongs, confirm whether the current voiceprint model has correctly identified the voiceprint. A positive reward is given when the voiceprint model is correctly identified, and a negative reward is given when the voiceprint model is incorrectly identified, and a corresponding reward value is obtained.
4. The voiceprint model optimization method based on reinforcement learning according to claim 1, characterized in that, The step of calculating an advantage function based on the recognition result and reward value of each sample, and updating the parameters of the policy network based on the advantage function, includes: Based on the identification result and reward value of each sample, the corresponding advantage function is calculated through generalized advantage estimation; The policy gradient is calculated based on the advantage function, and the parameters of the policy network are updated based on the policy gradient.
5. The voiceprint model optimization method based on reinforcement learning according to claim 1, characterized in that, The optimization of the pre-constructed objective function based on the parameters before and after the policy network update, wherein the objective function includes a KL divergence term used to constrain the magnitude of the policy update, including: Invoke a pre-built objective function, which includes a KL divergence term to constrain the policy update magnitude; Based on the parameters of the policy network before and after the update, calculate the probability ratio of the new and old policies for selecting actions and the KL divergence value. The objective function is updated based on the probability ratio, the dominance function, and the KL divergence value.
6. The voiceprint model optimization method based on reinforcement learning according to claim 1, characterized in that, Also includes: Obtain pre-built validation set data; The performance of the voiceprint model is evaluated using the validation set data according to a preset evaluation strategy. The reward function used to calculate the reward value is dynamically adjusted based on the performance evaluation results.
7. The voiceprint model optimization method based on reinforcement learning according to claim 6, characterized in that, The step of dynamically adjusting the reward function used to calculate the reward value based on the performance evaluation results specifically includes: When the performance of the voiceprint model improves, the magnitude of positive rewards is reduced and the magnitude of negative rewards is increased accordingly. When the performance of the voiceprint model decreases, the magnitude of positive rewards is increased and the magnitude of negative rewards is decreased accordingly.
8. A device for optimizing a voiceprint model based on reinforcement learning, characterized in that, include: The data acquisition module is used to collect historical recognition cases of the voiceprint model and construct an optimized dataset based on the historical recognition cases; An initialization module is used to randomly initialize the policy network, sample audio data from the optimization dataset and input it into the voiceprint model, control the voiceprint model to perform voiceprint recognition according to the initialization policy, and obtain the recognition result and corresponding reward value for each sample. The policy update module is used to calculate the advantage function based on the recognition result and reward value of each sample, and update the parameters of the policy network based on the advantage function; The function optimization module is used to optimize a pre-constructed objective function based on the parameters before and after the policy network update. The objective function includes a KL divergence term used to constrain the magnitude of the policy update. The iterative control module is used to repeatedly execute the above-mentioned sampled audio data for voiceprint recognition and policy update until the objective function meets the preset conditions to obtain the optimal policy network. The model optimization module is used to control the voiceprint model to perform voiceprint recognition according to the optimal strategy based on the optimal strategy network, so as to optimize the recognition results.
9. A computer device, characterized in that, Includes at least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the reinforcement learning-based voiceprint model optimization method according to any one of claims 1-7.
10. A non-volatile computer-readable storage medium, characterized in that, The non-volatile computer-readable storage medium stores computer-executable instructions that, when executed by one or more processors, cause the one or more processors to perform the reinforcement learning-based voiceprint model optimization method according to any one of claims 1-7.
Citation Information
Patent Citations
Method for training voiceprint representation model and related device
CN110491393A
Improved reinforcement learning AGV path planning method
CN117826713A