Voice anonymization method and device, computer equipment and storage medium

By extracting speech features for K nearest neighbor matching and fusion, and combining speech synthesis model to generate anonymous voice, it solves the balance between speech quality and privacy protection in the fields of financial technology and medical health, and achieves high-quality and natural anonymous voice generation.

CN120296796APending Publication Date: 2025-07-11PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510493297.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the fields of financial technology and medical health, it is difficult to maintain the accuracy of speech recognition and the clinical value of speech data while protecting user privacy, resulting in distortion of speech signals or removal of key features.

Method used

By obtaining the source voice of the target speaker, using the speech processing model to extract speech features, perform K-nearest neighbor matching and feature fusion, and generate anonymous voice using the speech synthesis model to ensure speech quality and privacy protection.

Benefits of technology

Generate high-quality and natural anonymous voice, protect the privacy of the speaker, maintain the accuracy of the voice recognition system and the effectiveness of medical diagnosis, and improve the safety and quality of financial services and medical services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296796A_ABST
    Figure CN120296796A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and the field of financial science and technology and medical health, and discloses a voice anonymization method and device, computer equipment and a storage medium, and the method comprises the steps: obtaining a source voice of a target speaker, and extracting voice features from a voice signal of the source voice through a voice processing model; performing K-nearest neighbor matching on the extracted voice features and voice features in a voice feature library to obtain the first K feature points which are closest to each other; performing average fusion on the speech features of the K feature points to obtain a target speech feature; converting the target voice feature into a voice signal by using a voice synthesis model to obtain an anonymized target voice; therefore, the anonymized voice which is high in quality, natural and capable of better protecting the privacy of the speaker can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence, fintech, and healthcare, and particularly relates to a voice anonymization method, apparatus, computer device, and computer-readable storage medium. Background Art

[0002] Currently, with the wide application of voice technology, the privacy protection of voice data has become an urgent problem to be solved. Voice anonymization technology aims to change the speaker identity of voice while protecting the privacy of the speaker and preventing identity theft or being used for other malicious purposes. This technology has important application values in multiple fields, such as fintech and healthcare. However, the current voice anonymization technology also faces different problems in fintech and healthcare respectively.

[0003] In the fintech field, voice anonymization needs to ensure the accuracy and reliability of the voice recognition system while protecting user privacy. However, existing anonymization technologies may cause voice signal distortion and reduce the accuracy rate of voice recognition, thus affecting the efficiency and security of financial services. For example, in voice transaction or identity verification scenarios, if the anonymized voice cannot be accurately recognized, it may lead to transaction failures or user identity verification errors, thereby affecting the user experience and the normal operation of financial services.

[0004] In the healthcare field, voice anonymization technology needs to ensure the availability of voice data for medical diagnosis and research while protecting patient privacy. However, excessive anonymization may remove key features in the voice, resulting in the loss of clinical value of voice data. For example, in mental health support or voice medical record keeping, the emotional expression and intonation changes of the voice are important diagnostic bases. If the anonymization process weakens these features, it may affect doctors' judgment of patients' conditions and reduce the effectiveness of voice data in medical scenarios.

[0005] That is, there is a certain balance problem between voice quality and privacy protection in existing voice anonymization technologies. Therefore, how to provide a voice anonymization method, apparatus, computer device, and computer-readable storage medium that can generate high-quality, natural, and better protect the privacy of speakers is an urgent problem for those skilled in the art currently. Summary of the Invention

[0006] In view of the deficiencies of the above-mentioned prior art, the purpose of the present invention is to provide a voice anonymization method, apparatus, computer device, and computer-readable storage medium, aiming to solve the problem of how to generate high-quality, natural, and better protect the privacy of speakers in anonymized voices.

[0007] To achieve the above purpose, the present invention adopts the following technical solutions:

[0008] In a first aspect, the present invention provides a voice anonymization method, which includes:

[0009] Obtain the source voice of the target speaker, and use a voice processing model to extract voice features from the voice signal of the source voice;

[0010] Perform K-nearest neighbor matching on the extracted voice features and the voice features in the voice feature library to obtain the top K feature points with the closest distances;

[0011] Average and fuse the voice features of the K feature points to obtain target voice features;

[0012] Use a voice synthesis model to convert the target voice features into a voice signal to obtain the anonymized target voice.

[0013] In a second aspect, the present invention provides a voice anonymization device, which includes:

[0014] A feature extraction module, configured to obtain the source voice of the target speaker, and use a voice processing model to extract voice features from the voice signal of the source voice;

[0015] A K-nearest neighbor matching module, configured to perform K-nearest neighbor matching on the extracted voice features and the voice features in the voice feature library to obtain the top K feature points with the closest distances;

[0016] An average fusion module, configured to average and fuse the voice features of the K feature points to obtain target voice features;

[0017] A feature conversion module, configured to use a voice synthesis model to convert the target voice features into a voice signal to obtain the anonymized target voice.

[0018] In a third aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned voice anonymization method is implemented.

[0019] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the above-mentioned voice anonymization method is implemented.

[0020] Compared with the prior art, the present invention provides a voice anonymization method, apparatus, computer device, and computer-readable storage medium. Among them, by obtaining the source voice of the target speaker, a voice processing model is used to extract voice features from the voice signal of the source voice; the extracted voice features are subjected to K-nearest neighbor matching with the voice features in the voice feature library to obtain the top K feature points with the closest distance; the voice features of the K feature points are averaged and fused to obtain target voice features; a voice synthesis model is used to convert the target voice features into a voice signal to obtain the anonymized target voice; thus, the present invention can generate anonymized voices with high quality, naturalness, and better protection of the speaker's privacy. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0022] Figure 1 It is a schematic diagram of the application environment of a voice anonymization method provided by an embodiment of the present invention.

[0023] Figure 2 It is a schematic flowchart of a voice anonymization method provided by an embodiment of the present invention.

[0024] Figure 3 It is a schematic diagram of the program modules of a voice anonymization apparatus provided by an embodiment of the present invention.

[0025] Figure 4 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present invention.

[0026] Figure 5 It is another schematic diagram of the structure of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0028] It should be understood that when used in the specification of the present invention and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.

[0029] It should also be understood that the term "and / or" used in the specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0030] As used in the specification of the present invention and the appended claims, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrases "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" depending on the context.

[0031] In addition, in the description of the specification of the present invention and the appended claims, the terms "first", "second", "third", etc. are only used for differential description and cannot be understood as indicating or implying relative importance.

[0032] The reference to "one embodiment" or "some embodiments" etc. described in the specification of the present invention means that a specific feature, structure or characteristic described in connection with that embodiment is included in one or more embodiments of the present invention. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0033] It should be understood that the magnitudes of the sequence numbers of the steps in the following embodiments do not mean the order of execution is prior or subsequent. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0034] In order to illustrate the technical solutions of the present invention, the following specific embodiments are used for illustration.

[0035] A voice anonymization method provided by an embodiment of the present invention can be applied, for example, in Figure 1In the application environment shown, the client communicates with the server via a network. Among them, the client includes, but is not limited to, computer devices such as a personal digital assistant (PDA), a palm computer, a desktop computer, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a cloud computer device, etc. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0036] Please refer to Figure 2 , an embodiment of the present invention provides a voice anonymization method, where the method includes the following steps:

[0037] S100. Obtain the source voice of the target speaker, and use a voice processing model to extract voice features from the voice signal of the source voice;

[0038] S200. Perform K-nearest neighbor matching on the extracted voice features and the voice features in the voice feature library to obtain the top K feature points with the closest distance;

[0039] S300. Average and fuse the voice features of the K feature points to obtain target voice features;

[0040] S400. Use a voice synthesis model to convert the target voice features into a voice signal to obtain the anonymized target voice.

[0041] In specific implementation, the voice anonymization method of this embodiment achieves the technical effect of generating anonymized voices with high quality, naturalness, and better protection of the speaker's privacy through a series of carefully designed steps. The specific analysis is as follows:

[0042] 1. High-quality voice generation:

[0043] In step S100, by using an advanced voice processing model to extract voice features from the source voice of the target speaker, the accuracy and integrity of voice feature extraction are ensured. High-quality voice feature extraction is the basis for generating natural voices.

[0044] In step S400, a pre-trained voice synthesis model is used to convert the target voice features into a voice signal. The pre-trained voice synthesis model can generate high-quality and natural voice signals, thus ensuring the naturalness and intelligibility of the finally anonymized target voice.

[0045] 2. Naturalness Enhancement:

[0046] Step S200 finds the K feature points closest to the speech feature of the source speech from the speech feature library of the target speaker through K-Nearest Neighbor Matching (KNN Matching); this similarity-based matching method can preserve the speech style and intonation features of the target speaker, making the subsequent generated target speech more natural in style;

[0047] Specifically, the following explains the meaning of "the top K feature points with the closest distance" in the present invention:

[0048] Meaning: Among all the feature points (i.e., speech features) in the speech feature library, find the top K feature points with the smallest distance from the target feature point (i.e., the speech feature of the source speech); the distance values of the top K feature points may be different, but they are the top K with the smallest distance among all the feature points, and K is a positive integer;

[0049] Example: Suppose there are 10 feature points in the speech feature library. Through calculation, it is found that among the distances from the target feature point, the distance of the 1st feature point is 0.1, the distance of the 2nd feature point is 0.2, the distance of the 3rd feature point is 0.3,..., and the distance of the 10th feature point is 1.0; if K = 3 is selected, then the top 3 feature points with the closest distance are the 1st, 2nd, and 3rd feature points, and their distances are 0.1, 0.2, and 0.3 respectively.

[0050] Step S300 performs average fusion on the K feature points, further smoothing the fluctuations of the speech features and avoiding the unnatural feeling that a single feature point may bring. This fusion method can generate more stable and natural target speech features, thereby enhancing the overall naturalness of the subsequent generated target speech.

[0051] 3. Enhanced Privacy Protection:

[0052] The method of this embodiment changes the speech feature of the source speech by matching and fusing the speech feature of the source speech with the speech feature library of the target speaker, thereby obscuring the identity information of the speaker. This anonymization technology based on feature replacement and fusion can effectively protect the privacy of the speaker and prevent the identity from being recognized.

[0053] Through K-Nearest Neighbor Matching and feature fusion, the generated speech signal no longer directly corresponds to the speech feature of any single speaker, but combines the information of multiple similar feature points. This processing method further reduces the risk of the target speech being restored or recognized as a specific speaker, thereby better protecting the privacy of the speaker.

[0054] In summary, through the above method in this embodiment, voice anonymization can not only generate high-quality and natural voice signals, but also effectively protect the privacy of the speaker. This technology has important application value in fields with high privacy requirements such as fintech and healthcare, and can meet the dual needs of users for voice quality and privacy protection.

[0055] Understandably, the voice anonymization method provided by the embodiments of the present invention can be applied to voice anonymization scenarios related to the healthcare field. The following are some specific examples:

[0056] Application Scenario 1: Remote Medical Consultation

[0057] In remote medical consultations, patients communicate with doctors via voice, describing sensitive information such as symptoms and medical histories. To protect the privacy of patients and prevent the leakage of voice data, the voice anonymization method of the present invention can play a role in the following aspects:

[0058] Privacy Protection: During the voice communication between patients and doctors, the system anonymizes the patients' voices in real time, changing the speaker characteristics of the voices to ensure that even if the voice data is intercepted or stored, the identity of the patients cannot be recognized.

[0059] Accuracy of Medical Diagnosis: The anonymized voice signals maintain high quality and naturalness, ensuring that doctors can accurately understand the patients' descriptions, including important information such as intonation and emotional expression. This not only protects the privacy of patients but also ensures the accuracy of medical diagnosis.

[0060] Data Storage and Analysis: The anonymized voice data can be securely stored and used for subsequent medical research without revealing the personal privacy of patients. This is of great significance for improving the quality of medical services and conducting medical research.

[0061] Application Scenario 2: Mental Health Support

[0062] In mental health support services, patients communicate with psychiatrists via voice, sharing emotions and mental states. This information is very sensitive and the privacy of patients needs to be strictly protected. The voice anonymization method of the present invention can play a role in the following aspects:

[0063] Privacy Protection: During the voice communication between patients and psychiatrists, the system anonymizes the patients' voices, obscuring the identity characteristics of the speakers to ensure that the privacy of patients is not leaked.

[0064] Naturalness of Emotional Expression: The anonymized voice signals maintain high quality and naturalness, ensuring that psychiatrists can accurately capture the emotional expressions and intonation changes of patients. This is crucial for psychiatrists to accurately judge the emotional state and psychological problems of patients.

[0065] Remote monitoring and recording: The anonymized voice data can be used for remotely monitoring the changes in a patient's condition or stored as medical record. This not only protects the patient's privacy but also provides valuable data support for subsequent treatment and research.

[0066] From the above two specific application scenarios, it can be seen that the voice anonymization method provided by the embodiments of the present invention has important application value in the field of healthcare. It can not only effectively protect the patient's privacy but also ensure the availability and accuracy of voice data in medical diagnosis and mental health support, thereby improving the quality and safety of medical services.

[0067] It can be understood that the voice anonymization method provided by the embodiments of the present invention can also be applied to voice anonymization scenarios related to the fintech field. The following are some specific examples:

[0068] Application scenario 1: Voice transaction system

[0069] In the fintech field, a voice transaction system allows users to perform financial transaction operations, such as transferring funds, making payments, and querying account balances, through voice commands. To protect the privacy of users and prevent voice data from being maliciously exploited, the voice anonymization method of the present invention can play a role in the following aspects:

[0070] Privacy protection: After a user issues a voice command, the system first processes the voice through the voice anonymization method to change the identity characteristics of the speaker, so that even if the voice data is intercepted, the true identity of the user cannot be recognized.

[0071] Transaction security: The voice signal after anonymization processing still maintains high quality and naturalness, ensuring that the voice recognition system can accurately understand the user's command and smoothly complete the transaction operation. This not only protects the user's privacy but also does not affect the security of the transaction and the user experience.

[0072] Application scenario 2: Intelligent voice customer service system

[0073] Intelligent voice customer service systems are widely used in the fintech field for customer service and support, such as answering account questions, product consultations, and handling complaints. To protect customer privacy and prevent voice data leakage, the voice anonymization method of the present invention can play a role in the following aspects:

[0074] Privacy protection: When a customer interacts with an intelligent customer service system through voice, the system can perform real-time anonymization processing on the customer's voice to blur the customer's identity characteristics and ensure that the customer's privacy is not leaked.

[0075] Service quality: The anonymized voice signal maintains high quality and naturalness, ensuring that the intelligent customer service system can accurately understand customer needs and provide corresponding services and support. This protects customer privacy without affecting the service quality and customer satisfaction of the customer service system.

[0076] Through the above two specific application scenarios, it can be seen that the voice anonymization method provided by the embodiment of the present invention has important application value in the field of financial technology. It can not only effectively protect the privacy of users and customers, but also ensure the accuracy and reliability of voice recognition and intelligent customer service systems, thereby improving the security and user experience of financial technology services.

[0077] Furthermore, in one embodiment, the speech anonymization method, wherein the step S100, obtaining the source speech of the target speaker, and extracting speech features from the speech signal of the source speech using a speech processing model, specifically comprises the steps of:

[0078] Acquiring the initial speech of the target speaker through a recording device;

[0079] After performing noise reduction, sampling rate adjustment, volume normalization and silent segment removal processing on the initial speech, the source speech is obtained;

[0080] The speech signal of the source speech is input into the pre-trained speech processing model to generate speech features of the source speech.

[0081] In specific implementation, this embodiment obtains the initial speech of the target speaker and performs noise reduction, sampling rate adjustment, volume normalization and silent segment removal, and then uses the pre-trained speech processing model to extract speech features, which can effectively ensure the high quality and consistency of speech features and lay a solid foundation for subsequent anonymization processing. Specifically, the preprocessing step can remove background noise and silent segments in the initial speech, adjust the sampling rate and normalize the volume, etc., thereby improving the quality of the speech signal and making it more suitable for subsequent feature extraction. By extracting speech features through a pre-trained speech processing model, the advanced nature and generalization ability of the model can be fully utilized to ensure that the extracted features contain rich semantic and speech information, and can effectively remove the personal characteristics of the speaker, thereby protecting privacy while providing strong support for generating high-quality, natural anonymized speech. This process not only improves the overall effect of speech anonymization, but also enhances the robustness and reliability of the system.

[0082] The specific implementation process of the steps in this embodiment is roughly as follows:

[0083] 1) Get the target speaker's initial speech

[0084] First, obtain the initial speech samples of the target speaker through a recording device or a voice acquisition system. These samples can be the speech recorded by the user in a specific scenario (such as voice assistant interaction, remote medical consultation, etc.), or the samples extracted from an existing speech database.

[0085] Ensure that the sampling rate, bit depth, and channel settings of the recording device meet the requirements of subsequent processing. For example, use a mono recording with a sampling rate of 16 kHz and a bit depth of 16 bits.

[0086] 2) Preprocess the initial speech

[0087] Preprocess the obtained initial speech to improve the quality and consistency of the speech signal.

[0088] Noise reduction processing: Use a noise reduction algorithm (such as Wiener filtering or a deep learning noise reduction model) to remove background noise and ensure the clarity of the speech signal.

[0089] Sampling rate adjustment: If the sampling rate of the initial speech does not meet the requirements, adjust it to the target sampling rate (such as 16 kHz) through resampling technology.

[0090] Volume normalization: Normalize the speech signal to keep its volume at a consistent level and avoid affecting subsequent processing due to volume differences.

[0091] Remove silent segments: Detect and remove the silent segments in the speech signal to reduce unnecessary data processing.

[0092] 3) Use the preprocessed speech as the source speech

[0093] Define the preprocessed speech signal as the source speech, which serves as the input for subsequent feature extraction.

[0094] Ensure that the quality of the source speech meets the requirements of subsequent processing, including the integrity and consistency of the speech signal.

[0095] 4) Select a pre-trained speech processing model

[0096] Select a pre-trained speech processing model for extracting speech features from the source speech. Commonly used models include deep learning-based feature extraction models.

[0097] 5) Extract speech features

[0098] Input the speech signal of the source speech into the pre-trained speech processing model to extract speech features.

[0099] Feature extraction: The model will decompose the speech signal into a series of feature vectors, which contain the acoustic information, semantic information, and speaker identity information of the speech.

[0100] Feature selection: According to the anonymization requirements, select features that are irrelevant to the speaker's identity (such as semantic and prosodic features), and remove or obscure identity-related features.

[0101] Feature output: Use the extracted speech features as the input for subsequent anonymization processing, ensuring the quality and consistency of the features.

[0102] Through the above process, this embodiment can not only obtain high-quality source speech, but also extract speech features suitable for anonymization processing through a pre-trained speech processing model. This process ensures the high quality and consistency of the speech signal, provides a solid foundation for subsequent anonymization processing, and protects the privacy of the speaker.

[0103] Further, in one embodiment, in the speech anonymization method, in step S200, the extracted speech features are matched with the speech features in the speech feature library by K-nearest neighbor to obtain the top K feature points with the closest distance, which specifically includes the steps of:

[0104] Construct the speech feature library of the target speaker, where the speech feature library contains multiple speech features of the target speaker;

[0105] Calculate the vector distance between the speech features of the extracted source speech and each speech feature in the speech feature library to obtain the calculation result;

[0106] According to the calculation result, select the top K feature points in the speech feature library with the smallest vector distance from the speech features of the source speech.

[0107] In specific implementation, through constructing the speech feature library of the target speaker and performing K-nearest neighbor matching on the speech features of the source speech with the speech features in the speech feature library, this embodiment can achieve accurate feature matching and effective privacy protection. Specifically, constructing the speech feature library provides rich reference samples for the matching process, ensuring the diversity and accuracy of the matching. By calculating the vector distance between the speech features of the source speech and each speech feature in the speech feature library, the similarity between the two can be quantified, so as to accurately find the K feature points most similar to the source speech. This similarity-based matching method can not only retain the speech style and intonation features of the target speaker, but also further obscure the speaker's identity information by fusing multiple similar feature points, so as to effectively protect the privacy of the speaker while generating high-quality and natural speech. This process provides an accurate and diverse basis for subsequent feature fusion and speech synthesis, significantly improving the overall effect of speech anonymization.

[0108] Among them, the specific implementation process of the steps in this embodiment is roughly as follows:

[0109] 1) Construct a voice feature library of the target speaker

[0110] Collect multiple voice samples of the target speaker, extract voice features from them, and construct a voice feature library.

[0111] Voice sample collection: Collect multiple voice samples in different contexts from the voice data of the target speaker to ensure the diversity and representativeness of the voice feature library.

[0112] Feature extraction: Extract voice features from each voice sample using the same voice processing model as the source voice.

[0113] Feature storage: Store the extracted feature vectors in the voice feature library to form a set containing multiple feature vectors.

[0114] 2) Calculate the vector distance

[0115] Calculate the vector distance between the voice features of the extracted source voice and each voice feature in the voice feature library to obtain the calculation result.

[0116] Distance metric selection: Select a suitable distance metric method, such as Euclidean distance, cosine similarity, or Manhattan distance, etc.

[0117] Point-by-point calculation: Calculate the distance between each feature vector in the feature library and the source voice feature vector.

[0118] Store the calculation result: Store the calculated distance values as a list or array for subsequent feature point selection.

[0119] 3) Select K feature points

[0120] According to the calculated distance results, select the top K feature points with the smallest distance as the nearest neighbor feature points, where K is a positive integer.

[0121] Sorting: Sort the calculated distance values to find the top K values with the smallest distance.

[0122] Select feature points: According to the sorting results, select the corresponding top K feature vectors as the nearest neighbor feature points.

[0123] Verification: Check whether the selected feature points meet the anonymization requirements, such as whether they are diverse enough and do not contain identifiable personal features.

[0124] 4) Diversity verification of feature points (optional)

[0125] Verify the diversity of the selected K feature points to ensure that they can represent various voice styles of the target speaker.

[0126] Diversity assessment: Calculate the distances or similarities between K feature points and evaluate their diversity.

[0127] Adjustment selection: If the diversity is insufficient, the value of K can be adjusted or the feature points can be reselected to ensure that the subsequent generated target voice has sufficient naturalness and anonymity.

[0128] 5) Privacy protection verification of feature points (optional)

[0129] Verify whether the selected K feature points can effectively protect the privacy of the speaker and prevent identity leakage.

[0130] Privacy assessment: Check whether the K feature points remove identifiable personal characteristics, such as specific timbre or intonation.

[0131] Adjustment processing: If it is found that some feature points may disclose privacy, these feature points can be further processed, such as blurred processing or replaced with other feature points.

[0132] Through the above process, this embodiment can not only accurately select K feature points that are most similar to the source voice, but also ensure the diversity and privacy protection ability of these feature points. This process provides high-quality input for subsequent feature fusion and speech synthesis, significantly improving the overall effect of voice anonymization, while ensuring that the generated voice is both natural and difficult to be recognized as a specific speaker.

[0133] Further, in one embodiment, in the voice anonymization method, wherein, in step S300, the voice features of the K feature points are averaged and fused to obtain target voice features, which specifically includes the steps of:

[0134] Extract the voice features corresponding to each of the K feature points;

[0135] Align the dimensions and structures of the voice features corresponding to each of the K feature points extracted;

[0136] Calculate the element-wise average value of the voice features corresponding to each of the K feature points after dimension and structure alignment to obtain the target voice features.

[0137] Further, in the voice anonymization method, wherein, in calculating the element-wise average value of the voice features corresponding to each of the K feature points after dimension and structure alignment to obtain the target voice features, it specifically includes the steps of:

[0138] Calculate the element-wise average value of the voice features corresponding to each of the K feature points after dimension and structure alignment to obtain the first preliminary voice features;

[0139] Perform feature smoothing on the first preliminary speech feature through a smoothing filter;

[0140] Perform feature enhancement on the first preliminary speech feature after feature smoothing to obtain the target speech feature.

[0141] Further, in the speech anonymization method, where performing feature enhancement on the first preliminary speech feature after feature smoothing to obtain the target speech feature specifically includes the steps of:

[0142] Perform feature enhancement on the first preliminary speech feature after feature smoothing to obtain a second preliminary speech feature;

[0143] Evaluate the quality and privacy of the second preliminary speech feature to generate an evaluation result;

[0144] Analyze the evaluation result, and when the evaluation result meets the preset quality and privacy requirements, use the second preliminary speech feature as the target speech feature.

[0145] In specific implementation, this embodiment realizes the generation of target speech features with high quality, naturalness, and better privacy protection through a series of fine feature processing steps. First, by aligning the dimensions and structures of the speech features of K feature points, the basic consistency of feature fusion is ensured, providing a unified framework for subsequent processing. Then, the first preliminary speech feature is obtained through element-wise average calculation. This process not only retains the style features of the target speaker but also further obscures the identity information of the original speaker by fusing multiple similar features. Subsequently, feature smoothing and enhancement are performed on the first preliminary speech feature to further optimize the quality and naturalness of the speech feature, making it more suitable for the requirements of speech synthesis. Finally, by evaluating whether the second preliminary speech feature meets the preset quality and privacy requirements, the reliability of the generated target speech feature in terms of quality and privacy protection is ensured. This series of steps not only improves the overall effect of speech anonymization but also ensures the high quality and naturalness of the subsequent generated target speech through a multi-layer optimization and evaluation mechanism, while maximizing the protection of the speaker's privacy.

[0146] Among them, the specific implementation process of the steps in this embodiment is roughly as follows:

[0147] 1) Feature alignment

[0148] Align the dimensions and structures of the speech features of K feature points.

[0149] Feature extraction: Extract speech feature vectors from K feature points.

[0150] Dimension matching: Ensure that the dimensions of all feature vectors are consistent. If the dimensions are inconsistent, adjustments can be made through interpolation or cropping.

[0151] Structural alignment: Normalize the feature vectors to make them consistent in structure. For example, through standardization, make the mean of all feature vectors 0 and the standard deviation 1.

[0152] 2) Element-wise average calculation

[0153] Calculate the element-wise average of the speech features corresponding to each of the K feature points after dimension and structural alignment to obtain the first preliminary speech feature.

[0154] Element-wise calculation: Sum the values in each feature dimension and then divide by K to obtain the average value of each dimension.

[0155] Generate preliminary features: Combine the calculated average values into a new feature vector, which is the first preliminary speech feature.

[0156] 3) Feature smoothing and enhancement

[0157] Perform feature smoothing and feature enhancement on the first preliminary speech feature to obtain the second preliminary speech feature.

[0158] Feature smoothing: Use a smoothing filter (such as a Gaussian filter or a moving average filter) to smooth the first preliminary speech feature and reduce the sudden changes and discontinuities of the feature values.

[0159] Feature enhancement: Improve the naturalness and intelligibility of the subsequent target speech by adjusting the dynamic range of the feature values or enhancing certain key features (such as prosodic features or intonation features).

[0160] Generate optimized features: Use the smoothed and enhanced feature vector as the second preliminary speech feature.

[0161] 4) Feature evaluation

[0162] Evaluate the second preliminary speech feature. When the evaluation result meets the preset requirements, use the second preliminary speech feature as the target speech feature.

[0163] Quality evaluation: Evaluate the second preliminary speech feature using objective evaluation metrics (such as signal-to-noise ratio, harmonic-to-noise ratio, etc.) and subjective evaluation methods (such as auditory tests).

[0164] Privacy evaluation: Check whether the second preliminary speech feature removes identifiable personal features to ensure privacy protection.

[0165] Meet the requirements: If the evaluation result meets the preset quality and privacy requirements, use the second preliminary voice feature as the target voice feature; otherwise, return to step 3) for further optimization.

[0166] 5) Feature adjustment (optional)

[0167] If the evaluation result does not meet the preset quality and privacy requirements, adjust the second preliminary voice feature.

[0168] Adjustment strategy: According to the evaluation result, adjust the degree of feature smoothing or enhancement, or reselect K feature points.

[0169] Repeat evaluation: Repeat the evaluation process in step 4) until the generated feature meets the preset quality and privacy requirements.

[0170] Final determination: Use the feature that meets the preset quality and privacy requirements as the target voice feature for subsequent speech synthesis.

[0171] Through the above process, this embodiment can not only generate high-quality target voice features through feature fusion, but also further optimize the quality and privacy protection ability of the features through feature smoothing, enhancement and evaluation processes. This series of steps ensures the reliability of the generated target voice features in terms of naturalness, intelligibility and privacy protection, providing a solid foundation for subsequent speech synthesis.

[0172] Furthermore, in one embodiment, for the voice anonymization method, in step S400, using a speech synthesis model to convert the target voice feature into a voice signal to obtain the anonymized target voice, specifically includes the steps:

[0173] Load a pre-trained voice style transfer model and the speech synthesis model;

[0174] Use the target voice feature as the input, and through the speech synthesis model, convert the target voice feature into a voice signal to obtain an intermediate voice;

[0175] Use the intermediate voice as the input, and through the voice style transfer model, convert the intermediate voice into the target style to obtain the target voice.

[0176] In specific implementation, through the collaborative action of the speech synthesis model and the speech style transfer model, this embodiment achieves the generation of high-quality, natural, and style-consistent target speech. First, the target speech features are input into the pre-trained speech synthesis model to generate intermediate speech. This process utilizes advanced speech synthesis technology to ensure that the generated speech signal reaches a high level in terms of sound quality and naturalness. Subsequently, the intermediate speech is subjected to style conversion through the pre-trained speech style transfer model to further adjust the timbre, rhythm, and speech rate of the speech to conform to the target style. This dual processing not only improves the naturalness and intelligibility of the speech but also ensures the consistency and anonymity of the speech style, effectively protecting the privacy of the speaker. The finally generated target speech not only retains the style characteristics of the target speaker but is also difficult to trace back to the original speaker, significantly enhancing the overall effect of speech anonymization and being applicable to scenarios with high privacy protection requirements, such as the fintech and healthcare fields.

[0177] Among them, the specific implementation process of the steps in this embodiment is roughly as follows:

[0178] 1) Load the pre-trained speech synthesis model

[0179] Load the pre-trained speech synthesis model, which can convert speech features into speech signals.

[0180] Model selection: Select a suitable pre-trained speech synthesis model.

[0181] Model loading: Load the speech synthesis model into the system to ensure that the speech synthesis model can receive the target speech features as input.

[0182] Model verification: Verify the performance of the speech synthesis model to ensure that it can generate high-quality speech signals.

[0183] 2) Input the target speech features into the speech synthesis model

[0184] Input the optimized target speech features into the speech synthesis model to generate an intermediate speech signal.

[0185] Feature input: Pass the target speech features as input to the speech synthesis model.

[0186] Speech generation: The speech synthesis model generates the corresponding speech signal according to the input target speech features to obtain the intermediate speech.

[0187] Intermediate speech saving: Save the generated intermediate speech as an audio file or temporarily store it in memory for subsequent processing.

[0188] 3) Load the pre-trained speech style transfer model

[0189] Load a pre-trained voice style transfer model that can convert the style of the intermediate voice to the target style.

[0190] Model selection: Select a suitable pre-trained voice style transfer model.

[0191] Model loading: Load the voice style transfer model into the system to ensure that the voice style transfer model can receive the intermediate voice as input.

[0192] Model verification: Verify the performance of the voice style transfer model to ensure that it can effectively convert the style of the intermediate voice to the target style.

[0193] 4) Input the intermediate voice into the voice style transfer model

[0194] Input the generated intermediate voice into the voice style transfer model for style conversion.

[0195] Voice input: Pass the intermediate voice as input to the voice style transfer model.

[0196] Style conversion: The voice style transfer model processes the intermediate voice, adjusting its timbre, prosody, and speaking rate to match the target style.

[0197] Target voice generation: Generate the target voice after style conversion.

[0198] 5) Verify and optimize the target voice (optional)

[0199] Verify and optimize the generated target voice to ensure its quality and style consistency.

[0200] Quality assessment: Use objective assessment metrics (such as PESQ, STOI) and subjective assessment methods (such as auditory tests) to assess the quality of the target voice.

[0201] Style verification: Check whether the target voice conforms to the preset target style to ensure the accuracy of style conversion.

[0202] Optimization adjustment: If the evaluation results do not meet the requirements, fine-tune the parameters of the voice style transfer model and regenerate the target voice until the quality requirements are met.

[0203] Final output: Use the verified target voice as the final output for subsequent applications.

[0204] Through the above process, this embodiment can not only generate high-quality intermediate speech through a speech synthesis model, but also further optimize the style consistency of the speech through a speech style transfer model. This process ensures high-quality performance of the generated target speech in terms of naturalness, intelligibility, and style consistency, while effectively protecting the privacy of the speaker, and is applicable to scenarios with high privacy protection requirements such as fintech and healthcare.

[0205] As can be seen from the above method embodiments, the speech anonymization method provided by the present invention includes: obtaining the source speech of the target speaker, and extracting speech features from the speech signal of the source speech by using a speech processing model; performing K-nearest neighbor matching on the extracted speech features and the speech features in the speech feature library to obtain the top K feature points with the closest distance; averaging and fusing the speech features of the K feature points to obtain target speech features; and using a speech synthesis model to convert the target speech features into a speech signal to obtain the anonymized target speech. In this way, the anonymized speech with high quality, naturalness, and better protection of the speaker's privacy can be generated by the method of the present invention.

[0206] It should be understood that although the method operation steps as described in the embodiments or flowcharts are provided in this application, based on routine or non-creative labor, there may be more or fewer operation steps, and these operation steps are not necessarily executed in the order of the embodiments or flowcharts. The step order listed in the embodiments or flowcharts is only one way among many step execution orders and does not represent the only execution order. It should be noted that there is not necessarily a certain order between the above steps. Those of ordinary skill in the art can understand according to the description of the embodiments of the present invention that in different embodiments, the above steps can have different execution orders, that is, they can be executed in parallel, or exchanged, etc. Moreover, at least a part of the steps in the embodiments or flowcharts may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately, alternately, or synchronously with at least a part of other steps or sub-steps or stages of other steps.

[0207] Based on the above method embodiments, please refer to Figure 3 , another embodiment of the present invention further provides a speech anonymization device, where the device includes:

[0208] A feature extraction module 11, configured to obtain the source speech of the target speaker, and extract speech features from the speech signal of the source speech by using a speech processing model;

[0209] The K-nearest neighbor matching module 12 is used to perform K-nearest neighbor matching on the extracted speech features and the speech features in the speech feature library to obtain the top K feature points with the closest distances.

[0210] The average fusion module 13 is used to perform average fusion on the speech features of the K feature points to obtain the target speech features.

[0211] The feature conversion module 14 is used to convert the target speech features into speech signals by using a speech synthesis model to obtain the anonymized target speech.

[0212] Further, in one embodiment, for the speech anonymization device, wherein the feature extraction module 11 is specifically used for:

[0213] Obtain the initial speech of the target speaker through a recording device;

[0214] After performing noise reduction, sampling rate adjustment, volume normalization, and silent segment removal processing on the initial speech, obtain the source speech;

[0215] Input the speech signal of the source speech into the pre-trained speech processing model to generate the speech features of the source speech.

[0216] Further, in one embodiment, for the speech anonymization device, wherein the K-nearest neighbor matching module 12 is specifically used for:

[0217] Construct the speech feature library of the target speaker, and the speech feature library contains multiple speech features of the target speaker;

[0218] Calculate the vector distance between the extracted speech features of the source speech and each speech feature in the speech feature library to obtain the calculation result;

[0219] According to the calculation result, select the top K feature points in the speech feature library with the smallest vector distance from the speech features of the source speech.

[0220] Further, in one embodiment, for the speech anonymization device, wherein the average fusion module 13 is specifically used for:

[0221] Extract the speech features corresponding to each of the K feature points;

[0222] Align the dimensions and structures of the speech features corresponding to each of the K feature points;

[0223] Perform element-wise average calculation on the speech features corresponding to each of the K feature points after dimension and structure alignment to obtain the target speech features.

[0224] Further, for the voice anonymization device, wherein calculating the element-wise average of the voice features corresponding to each of the K feature points after dimension and structure alignment to obtain the target voice feature specifically includes:

[0225] Calculating the element-wise average of the voice features corresponding to each of the K feature points after dimension and structure alignment to obtain a first preliminary voice feature;

[0226] Performing feature smoothing on the first preliminary voice feature through a smoothing filter;

[0227] Performing feature enhancement on the first preliminary voice feature after feature smoothing to obtain the target voice feature.

[0228] Further, for the voice anonymization device, wherein performing feature enhancement on the first preliminary voice feature after feature smoothing to obtain the target voice feature specifically includes:

[0229] Performing feature enhancement on the first preliminary voice feature after feature smoothing to obtain a second preliminary voice feature;

[0230] Evaluating the quality and privacy of the second preliminary voice feature to generate an evaluation result;

[0231] Analyzing the evaluation result, and when the evaluation result meets the preset quality and privacy requirements, using the second preliminary voice feature as the target voice feature.

[0232] Further, in one embodiment, for the voice anonymization device, wherein the feature conversion module 14 is specifically used for:

[0233] Loading a pre-trained voice style transfer model and the voice synthesis model;

[0234] Using the target voice feature as input, and converting the target voice feature into a voice signal through the voice synthesis model to obtain an intermediate voice;

[0235] Using the intermediate voice as input, and converting the intermediate voice into a target style through the voice style transfer model to obtain the target voice.

[0236] It should be noted that in the device embodiment of the present invention, for the information interaction, execution process, etc. between the above modules, since they are based on the same concept as the method embodiment of the present invention, their specific functions and the technical effects brought about can be specifically referred to the method embodiment part above, and will not be elaborated here.

[0237] Based on the above method embodiments, another embodiment of the present invention further provides a computer device, which can be a server, and its internal structure diagram can be as Figure 4 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements the functions or steps of the voice anonymization method on the server side in any of the above method embodiments.

[0238] Based on the above method embodiments, another embodiment of the present invention further provides a computer device, which can be a client, and its internal structure diagram can be as Figure 5 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements the functions or steps of the voice anonymization method on the client side in any of the above method embodiments.

[0239] Those skilled in the art can understand that Figure 4 the Figure 5 structural schematic diagram shown is only a schematic diagram of a part of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more components than those shown in the figure, or combine some components, or have different component arrangements.

[0240] Among them, the so-called processor can be a CPU, and the processor can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0241] Among them, the memory includes a readable storage medium, internal memory, etc. Among them, the internal memory can be the memory of a computer device, and the internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The readable storage medium can be the hard disk of a computer device, and in some other embodiments, it can also be an external storage device of the computer device. For example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the memory can also include both the internal storage unit of the computer device and the external storage device. The memory is used to store the operating system, application programs, boot loaders, data, and other programs, such as the program code of a computer program. The memory can also be used to temporarily store the data that has been output or will be output.

[0242] Based on the above method embodiments, another embodiment of the present invention also provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the voice anonymization method in any one of the above method embodiments. The computer-readable storage medium can be non-volatile or volatile.

[0243] It should be noted that for the functions or steps that the above computer-readable storage medium or computer device can achieve, and the technical effects brought by the functions / steps, reference can be made to the relevant descriptions in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0244] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc. The memory components or memories of the operating environment disclosed herein are intended to include one or more of these and / or any other suitable types of memories.

[0245] Those skilled in the art can clearly understand that for the convenience and simplicity of description, in the device embodiments of the present invention, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present invention. The specific working processes of the units and modules in the above device can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0246] Those of ordinary skill in the art will realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0247] In the embodiments provided by the present invention, it should be understood that the disclosed device / computer equipment and method can be implemented in other ways. For example, the device / computer equipment embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be electrical, mechanical or other forms.

[0248] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0249] It should be noted that if non-company software tools or components appear in the embodiments of this application, they are only for illustrative introduction and do not represent actual use. The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A voice anonymization method, characterized in that, Including: Obtain the source speech of the target speaker, and use a speech processing model to extract speech features from the speech signal of the source speech; Perform K-nearest neighbor matching on the extracted speech features and the speech features in the speech feature library to obtain the top K feature points with the closest distance; Average and fuse the speech features of the K feature points to obtain target speech features; Use a speech synthesis model to convert the target speech features into a speech signal to obtain the anonymized target speech.

2. The voice anonymization method according to claim 1, characterized in that, The obtaining the source speech of the target speaker and using a speech processing model to extract speech features from the speech signal of the source speech includes: Obtain the initial speech of the target speaker through a recording device; After performing noise reduction, sampling rate adjustment, volume normalization, and silent segment removal on the initial speech, obtain the source speech; Input the speech signal of the source speech into the pre-trained speech processing model to generate the speech features of the source speech.

3. The voice anonymization method according to claim 1, characterized in that The performing K-nearest neighbor matching on the extracted speech features and the speech features in the speech feature library to obtain the top K feature points with the closest distance includes: Construct the speech feature library of the target speaker, where the speech feature library contains multiple speech features of the target speaker; Calculate the vector distance between the extracted speech features of the source speech and each speech feature in the speech feature library to obtain a calculation result; According to the calculation result, select the top K feature points in the speech feature library with the smallest vector distance from the speech features of the source speech.

4. The voice anonymization method according to claim 1, characterized in that The averaging and fusing the speech features of the K feature points to obtain target speech features includes: Extract the speech features corresponding to each of the K feature points; Align the dimensions and structures of the speech features corresponding to the K feature points extracted; Perform element-wise average calculation on the speech features corresponding to the K feature points after dimension and structure alignment to obtain the target speech features.

5. The voice anonymization method according to claim 4, characterized in that, The performing element-wise average calculation on the speech features corresponding to the K feature points after dimension and structure alignment to obtain the target speech features includes: Perform element-wise average calculation on the speech features corresponding to the K feature points after dimension and structure alignment to obtain the first preliminary speech features; Perform feature smoothing processing on the first preliminary speech features through a smoothing filter; Perform feature enhancement processing on the first preliminary speech features after feature smoothing processing to obtain the target speech features.

6. The voice anonymization method according to claim 5, characterized in that, The performing feature enhancement processing on the first preliminary speech features after feature smoothing processing to obtain the target speech features includes: Perform feature enhancement processing on the first preliminary speech features after feature smoothing processing to obtain the second preliminary speech features; Evaluate the quality and privacy of the second preliminary speech features to generate an evaluation result; Analyze the evaluation result, and when the evaluation result meets the preset quality and privacy requirements, use the second preliminary speech features as the target speech features.

7. The voice anonymization method according to any one of claims 1-6, characterized in that, The using a speech synthesis model to convert the target speech features into a speech signal to obtain the anonymized target speech includes: Load the pre-trained voice style transfer model and the voice synthesis model; Use the target voice feature as the input, and convert the target voice feature into a voice signal through the voice synthesis model to obtain an intermediate voice; Use the intermediate voice as the input, and convert the intermediate voice into the target style through the voice style transfer model to obtain the target voice.

8. A voice anonymization device, characterized in that, It includes: A feature extraction module, which is used to obtain the source voice of the target speaker and extract voice features from the voice signal of the source voice by using a voice processing model; A K-nearest neighbor matching module, which is used to perform K-nearest neighbor matching on the extracted voice features and the voice features in the voice feature library to obtain the top K feature points with the closest distance; An average fusion module, which is used to perform average fusion on the voice features of the K feature points to obtain target voice features; A feature conversion module, which is used to convert the target voice features into voice signals by using a voice synthesis model to obtain the anonymized target voice.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the voice anonymization method according to any one of claims 1-7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the voice anonymization method according to any one of claims 1-7.