Generative audio anonymization reconstruction method and device based on sound source separation and semantic preservation, equipment and program product
By employing an audio anonymization reconstruction method that separates sound sources and preserves semantics, the problem of stream disruption caused by unauthorized human audio track processing in audio data privacy protection is solved, thereby improving the scene coherence and comprehensibility of audio.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-13
- Publication Date
- 2026-03-24
AI Technical Summary
Existing audio data privacy protection technologies are prone to causing audio stream breaks when processing unauthorized human audio tracks, affecting scene coherence and comprehensibility.
The audio is separated into a speaker's audio track and an ambient background audio track by using sound source separation technology. After authentication, the unauthorized speaker's audio track is anonymized and then resynthesized with the ambient background audio track to generate an anonymized scene audio track.
It effectively avoids audio stream interruptions, improves the continuity and comprehensibility of audio scenes, and takes into account the privacy of unauthorized human audio tracks.
Smart Images

Figure CN121725795A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of speech privacy protection and audio signal processing, and particularly relates to a generative audio anonymization reconstruction method, device, equipment and program product based on sound source separation and semantic reservation. BACKGROUND
[0002] The current mainstream technical means of audio data privacy protection mainly include signal layer mute processing, content layer information shielding coding and data segment deletion, etc. Although these methods can effectively reduce the risk of sensitive information leakage, they all have significant limitations. Among them, the signal layer mute processing directly cuts off the time domain continuity of the audio signal, resulting in a sudden interruption in the sense of hearing; the content layer information shielding coding preserves the audio time length feature, but destroys the integrity of the language logic chain; and the data segment deletion method may cause context semantic fault due to excessive pruning.
[0003] The above technical intervention will cause different degrees of audio stream rupture phenomenon, which is difficult to guarantee the scene coherence of the audio, and ultimately leads to the structural decline of the intelligibility and overall quality of the audio. SUMMARY
[0004] The present application provides a generative audio anonymization reconstruction method, device, equipment and program product based on sound source separation and semantic reservation, which can avoid the audio stream rupture phenomenon, improve the scene coherence of the audio, and further improve the intelligibility and overall quality of the audio.
[0005] According to an aspect of the present application, a generative audio anonymization reconstruction method based on sound source separation and semantic reservation is provided, which comprises: Obtaining an original audio and performing sound source separation on the original audio to obtain at least one speaker track and an environmental background track; Authenticating each speaker track to determine the authentication result of each speaker track; wherein the authentication result of the speaker track includes an authorized person track and an unauthorized person track; Performing anonymization processing on the unauthorized person track to obtain an anonymized person track; Performing re-synthesis on the anonymized person track and the environmental background track to obtain an anonymized scene track.
[0006] According to another aspect of the present application, a generative audio anonymization reconstruction device based on sound source separation and semantic reservation is provided, which comprises: An original audio sound source separation module for obtaining an original audio and performing sound source separation on the original audio to obtain at least one speaker track and an environmental background track; The speaker track authentication module is configured to authenticate each of the speaker tracks and determine an authentication result of each of the speaker tracks, wherein the authentication result of the speaker track includes an authorized person track and an unauthorized person track. The unauthorized person track anonymization module is configured to anonymize the unauthorized person track to obtain an anonymized person track. The anonymized scene track generation module is configured to synthesize the anonymized person track and the environmental background track to obtain an anonymized scene track.
[0007] According to another aspect of the present application, an electronic device is provided, which comprises: at least one processor; and a memory connected to the at least one processor in communication; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the method for generative audio anonymization reconstruction based on sound source separation and semantic preservation according to any one of the embodiments of the present application.
[0008] According to another aspect of the present application, a computer readable storage medium is provided, which stores computer instructions for enabling a processor to implement the method for generative audio anonymization reconstruction based on sound source separation and semantic preservation according to any one of the embodiments of the present application when executed by the processor.
[0009] According to another aspect of the present application, a computer program product is provided, which comprises a computer program for implementing the method for generative audio anonymization reconstruction based on sound source separation and semantic preservation according to any one of the embodiments of the present application when executed by a processor.
[0010] The technical solution of the embodiments of the present application can maximize the semantic information of the speech content corresponding to the unauthorized person track and the environmental information of the scene by performing sound source separation on the original audio to obtain at least one speaker track and an environmental background track, authenticating each speaker track to obtain an authentication result of each speaker track, anonymizing the unauthorized person track to obtain an anonymized person track, and synthesizing the anonymized person track and the environmental background track to obtain an anonymized scene track. In addition, the privacy of the unauthorized person track can also be taken into account, thereby avoiding the audio stream breaking phenomenon, improving the scene coherence of the audio, and further improving the intelligibility and overall quality of the audio.
[0011] It is to be understood that the details set forth herein do not limit the scope of the embodiments of the application to the specific embodiments described. The foregoing detailed description has set forth various embodiments of the devices and / or processes via the use of specific terminology. However, embodiments of the application should not be construed as limited to the foregoing aspects, since the same can be modified, changed and / or combined in various ways. BRIEF DESCRIPTION OF DRAWINGS
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.
[0013] Figure 1 is a flow chart of a generative audio anonymization reconstruction method based on sound source separation and semantic reservation provided by the first embodiment of the present application; Figure 2 is a flow chart of a generative audio anonymization reconstruction method based on sound source separation and semantic reservation provided by the second embodiment of the present application; Figure 3 is a structural schematic diagram of a generative audio anonymization reconstruction device based on sound source separation and semantic reservation provided by the third embodiment of the present application; Figure 4 is a structural schematic diagram of an electronic device implementing the generative audio anonymization reconstruction method based on sound source separation and semantic reservation of the embodiments of the present application. DETAILED DESCRIPTION
[0014] In order for those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.
[0015] It should be noted that the terms "first", "second", and the like in the description and in the claims of the present application and the above-described accompanying drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or device that includes a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products, or devices.
[0016] Embodiment one Figure 1 A flowchart of a generative audio anonymization reconstruction method based on sound source separation and semantic preservation is provided for the first embodiment of the present application. The embodiments of the present application can be applicable to the anonymization of audio collected in a conference scene or a live scene. The method can be executed by a generative audio anonymization reconstruction device based on sound source separation and semantic preservation. The generative audio anonymization reconstruction device based on sound source separation and semantic preservation can be realized in the form of hardware and / or software. The generative audio anonymization reconstruction device based on sound source separation and semantic preservation can be configured in an electronic device that carries the function of generative audio anonymization reconstruction based on sound source separation and semantic preservation, such as a recording device, a mobile terminal, or a server.
[0017] Referring to Figure 1 The generative audio anonymization reconstruction method based on sound source separation and semantic preservation includes: S101, obtaining original audio and performing sound source separation on the original audio to obtain at least one speaker track and an environmental background track.
[0018] The original audio is audio collected by a recording device. Optionally, in a conference scene or a live scene, the original audio can include a speaker track of at least one speaker and an environmental background track of the scene. The speaker track is a sound track corresponding to the sound source of the speaker in the original audio. The environmental background track is a scene track of the scene where the original audio is located. Illustratively, the recording device can be a recording pen or a Bluetooth headset, etc.
[0019] Specifically, a recording device is used to obtain original audio in a conference scene or a live scene. An independent component analysis algorithm, a non-negative matrix factorization algorithm, or a computational auditory scene analysis algorithm, etc. is used to perform sound source separation on the original audio to obtain at least one speaker track and an environmental background track.
[0020] In an optional embodiment of the present invention, before acquiring the original audio, the method further includes: when a pointing adjustment operation is detected or the ambient noise exceeds a preset ambient noise threshold, adjusting the pointing of the recording device according to the device orientation, holding posture, or historical sound source direction of the recording device.
[0021] The pointing adjustment operation is the trigger operation for adjusting the pointing of the recording device. For example, the pointing adjustment operation may include clicking a button on the recording device a first preset number of times within a preset duration. The preset duration is a pre-set detection duration for the recording device buttons. The first preset number of clicks is a pre-set number of button clicks corresponding to the pointing adjustment operation. For example, the preset duration could be 5 seconds; the first preset number of clicks could be 3 times. For example, the pointing adjustment operation may also include a preset click time interval between each click. Referring to the above example, the preset click time interval between the first two clicks could be less than 1 second; the preset click time interval between the second and third clicks could be greater than 1 second and less than or equal to 3 seconds. Ambient noise is the sum of all sounds in the scene except for the speaker's voice. Ambient noise is used to characterize the interference of other sounds in the scene on the speaker's voice. The preset ambient noise threshold is a pre-set minimum value of ambient noise required when the pointing of the recording device needs to be adjusted.
[0022] Compared to other scenarios where recording devices are fixed in a speaker's position, such as the ear or neck, the position of recording devices in meeting or live streaming scenarios is not fixed. For example, the recording device may be worn on the speaker's chest, held by the speaker, or held by another speaker. In this case, the relative position between the recording device and the sound source is more variable, and the same recording device will also capture audio from different speakers in the scene, making it more difficult to guarantee the quality of the original audio.
[0023] Device orientation refers to the orientation of the recording device. Device orientation reflects the direction or angle of the sensitive components of the recording device relative to the sound source. Device orientation can affect recording quality, sound clarity, background noise suppression, and the capture of specific sound effects. Holding posture refers to the speaker's holding posture of the recording device. Holding posture can affect recording quality and recording stability. Historical sound source direction reflects the distribution of sound source locations over a historical period prior to the current moment. Historical sound source direction is used to characterize the frequency of different sound source directions in a scene. Historical sound source direction can be used to predict the sound source direction at the current moment and can be used to pre-adjust the directivity of the sound source direction at the current moment based on historical sound source direction. Recording device pointing is used to characterize the recording device's sensitivity to sounds from different directions. For example, recording device pointing can include the recording device's directivity pattern or beam pointing. For example, directivity patterns include cardioid, supercardioid, or figure-eight patterns. By adjusting the recording device pointing, the sound source corresponding to the recording device's pointing can be enhanced. This can improve the acquisition quality of the obtained raw audio.
[0024] Optionally, when a speaker's directional adjustment operation on the recording device is detected, or when environmental noise in the scene exceeds a preset environmental noise threshold, the relative angle between the recording device's orientation and the sound source is detected. When the recording device has a directional microphone, the directional mode of the recording device is adjusted to the directional mode closest to the relative angle; when the recording device has a multi-microphone array, the beam direction of the recording device is adjusted to the beam direction corresponding to the relative angle.
[0025] Optionally, when a speaker's directional adjustment operation on the recording device is detected, or when environmental noise in the scene exceeds a preset environmental noise threshold, the holding direction of the recording device is identified based on the user's grip posture. When the recording device has a directional microphone, the directional mode of the recording device is adjusted to the directional mode closest to the holding direction; when the recording device has a multi-microphone array, the beam direction of the recording device is adjusted to the holding direction.
[0026] Optionally, when a speaker's directional adjustment operation on the recording device is detected, or when environmental noise in the scene exceeds a preset environmental noise threshold, the direction of the high-frequency sound source is identified based on the historical sound source direction of the recording device. When the recording device has a directional microphone, the directional mode of the recording device is adjusted to the directional mode closest to the direction of the high-frequency sound source; when the recording device has a multi-microphone array, the beam direction of the recording device is adjusted to the direction of the high-frequency sound source.
[0027] Optionally, when a speaker's directional adjustment operation on the recording device is detected, or when environmental noise in the scene exceeds a preset environmental noise threshold, the historical sound source direction change trajectory can be determined based on the historical sound source direction of the recording device. When the recording device has a directional microphone, the directional mode of the recording device is adjusted according to the historical sound source direction change trajectory; when the recording device has a multi-microphone array, the beam direction of the recording device is adjusted according to the historical sound source direction change trajectory.
[0028] When this solution detects a directional adjustment operation or ambient noise exceeding a preset ambient noise threshold, it pre-adjusts the directional of the recording device based on the device's orientation, holding posture, or historical sound source direction. Through directional pickup or beamforming, it enhances the original audio, improving the signal-to-noise ratio of human voice to background separation. This also improves the sound source separation effect of the original audio, thereby enhancing the anonymization reconstruction effect and stability of the audio.
[0029] S102. Authenticate each speaker's audio track and determine the authentication result of each speaker's audio track; wherein, the authentication result of the speaker's audio track includes the authorized speaker's audio track and the unauthorized speaker's audio track.
[0030] Before a meeting or live stream, recording devices can be used to authorize the voices of speakers who will participate in the meeting or live stream beforehand. The authentication result of the speaker's audio track is used to characterize the speaker's participation authority in the scenario. Specifically, the authorized speaker's audio track indicates that the corresponding speaker has the corresponding participation authority. This can be understood as the speaker corresponding to the authorized speaker's audio track having pre-authorized their participation before joining the meeting or live stream. The unauthorized speaker's audio track indicates that the corresponding speaker does not have the corresponding participation authority. This can be understood as the speaker corresponding to the unauthorized speaker's audio track not having pre-authorized their participation before joining the meeting or live stream.
[0031] Optionally, voiceprints of speakers participating in the meeting or live stream can be pre-collected to generate voiceprint templates. These templates are then entered into a voiceprint database, which is subsequently sent to the recording device. Optionally, at least one recording device can be used.
[0032] Specifically, the speaker's voiceprint corresponding to each speaker's voice track can be matched against each voiceprint template in the voiceprint library on the recording device. If a match is found between the speaker's voiceprint corresponding to a voice track and a specific voiceprint template in the recording device's voiceprint library (e.g., if a similarity template with the corresponding voiceprint is greater than or equal to a preset similarity threshold), the authentication result for that voice track is determined to be an authorized voice track. If no match is found between the speaker's voiceprint corresponding to a voice track and any voiceprint template in the recording device's voiceprint library (e.g., if no similarity template with the corresponding voiceprint is greater than or equal to the preset similarity threshold), the authentication result for that voice track is determined to be an unauthorized voice track. The preset similarity threshold is a pre-defined minimum similarity value between voiceprints during voiceprint matching. For example, the preset similarity threshold could be 80%.
[0033] S103. Anonymize the unauthorized voice track to obtain an anonymized voice track.
[0034] In meeting or live streaming scenarios, the speaker corresponding to an unauthorized audio track may be a temporary or confidential participant. Although the speaker in the unauthorized audio track has not obtained prior authorization, directly applying mainstream audio data privacy protection methods—such as muting, censoring, and / or deleting data segments from the unauthorized audio track—can cause varying degrees of audio interruption, making it difficult to guarantee the continuity of the audio context. This ultimately leads to a structural decline in the comprehensibility and overall quality of the audio. Furthermore, an unauthorized audio track merely indicates that the speaker has not obtained prior authorization; it doesn't necessarily indicate the importance of the content being spoken. In some meeting or live streaming scenarios, the content in the unauthorized audio track is actually crucial. Therefore, for meeting or live streaming scenarios, balancing the integrity of the meeting or live streaming content with the privacy of the unauthorized audio track is of paramount importance.
[0035] Anonymous audio tracks are the result of anonymizing unauthorized audio tracks. In other words, anonymized audio tracks retain the spoken content of the unauthorized audio track, but remove the speaker's identity.
[0036] Specifically, spectral distortion or speech synthesis methods can be used to hide the voiceprint features in the unauthorized speaker's voice track while preserving its semantic, emotional, speech rate, and tone content. Furthermore, metadata anonymization can be applied to the semantic content of the unauthorized speaker's voice track. This involves removing or replacing metadata associated with the speaker's identity within the semantic content of the unauthorized speaker's voice track. For example, metadata associated with the speaker's identity may include the speaker's name and identifier.
[0037] S104. Resynthesize the anonymized human voice track and the ambient background sound track to obtain the anonymized scene sound track.
[0038] The anonymized scene audio track is the result of recombining an anonymized human voice track and an ambient background audio track. Distinguishing the anonymized scene audio track from the authorized human voice track can be understood as effectively differentiating the content of authorized audio sources from other audio sources in meeting or live streaming scenarios. This allows for a more complete and intuitive representation of the spoken content corresponding to the authorized human voice track, while preserving the content corresponding to the unauthorized human voice track and the ambient background audio track to the greatest extent possible. This facilitates subsequent tasks such as voice transcription and content retrieval based on the authorized human voice track and the anonymized scene audio track, and also ensures the privacy of the unauthorized human voice track.
[0039] Specifically, basic mixing or layered mixing can be used to resynthesize the anonymized vocal track and the ambient background audio track to obtain an anonymized scene audio track.
[0040] The technical solution of this invention separates the original audio source to obtain at least one speaker audio track and an ambient background audio track. Each speaker audio track is authenticated to obtain the authentication result. Unauthorized speaker audio tracks are anonymized to obtain an anonymized audio track. The anonymized audio track and the ambient background audio track are then resynthesized to obtain an anonymized scene audio track. This method maximizes the preservation of semantic information of the speech content corresponding to the unauthorized speaker audio track and the environmental information of the scene. Furthermore, it also ensures the privacy of the unauthorized speaker audio track, thereby avoiding audio stream interruptions, improving the scene coherence of the audio, and ultimately enhancing the comprehensibility and overall quality of the audio.
[0041] In an optional embodiment of the present invention, after recombining the anonymized human voice track and the ambient background audio track to obtain an anonymized scene audio track, the method further includes: detecting an audio track output operation; and outputting a target audio track according to the audio track output operation; wherein the target audio track includes an authorized human voice track or an anonymized scene audio track.
[0042] The audio track output operation is used to select between authorized human audio tracks and anonymized scene audio tracks. The target audio track is the track selected by the audio track output operation. For example, the target audio track includes either an authorized human audio track or an anonymized scene audio track. Optionally, during the playback, export, or transcription stage of audio in a meeting or live broadcast scenario, the corresponding target audio track can be selected for output based on the audio track output operation. This improves the targeting of the target audio during playback, export, or transcription, thereby better adapting to the needs of subsequent applications.
[0043] Optionally, the system detects user input to the audio track output on the mobile terminal's front-end page. Upon detecting an audio track output operation, the system outputs the target audio track based on the selected audio track. Optionally, the mobile terminal's front-end page may have input elements corresponding to authorized human voice tracks or anonymized scene audio tracks. Correspondingly, the audio track output operation can be a click operation on the input element corresponding to the authorized human voice track or anonymized scene audio track on the front-end page.
[0044] Optionally, the recording device can detect if the user clicks the recording device's buttons a second preset number of times. This second preset number is a pre-defined number of button presses corresponding to a specific audio output operation. For example, the second preset number of presses for an authorized person's audio track could be 4 times; the second preset number of presses for an anonymous scene audio track could be 5 times.
[0045] This solution outputs either an authorized speaker's audio track or an anonymous scene audio track through audio track output operations. On the one hand, it can more completely and intuitively reflect the speech content corresponding to the authorized speaker's audio track. On the other hand, it can also retain the content corresponding to the unauthorized speaker's audio track and the environmental background audio track to the greatest extent, which facilitates subsequent tasks such as voice writing and content retrieval based on the authorized speaker's audio track and the anonymous scene audio track, realizing the subsequent application of the authorized speaker's audio track and the anonymous scene audio track.
[0046] Example 2 Figure 2 This is a flowchart of a generative audio anonymization reconstruction method based on sound source separation and semantic preservation, provided in Embodiment 2 of the present invention. Based on the above embodiments, this embodiment of the present invention specifies "anonymizing unauthorized human voice tracks to obtain anonymized human voice tracks" as "extracting features from the unauthorized human voice tracks to obtain original speaker identity features and original semantic features; anonymizing the original speaker identity features to obtain anonymous speaker identity features; and recombining the anonymous speaker identity features and original semantic features to obtain anonymized human voice tracks." This method preserves the semantic information of the speech content corresponding to the unauthorized human voice tracks, maintains the scene coherence of the audio, and improves the efficiency of anonymizing the unauthorized human voice tracks. It should be noted that parts not detailed in this embodiment of the present invention can be found in the descriptions of other embodiments.
[0047] See Figure 2 The generative audio anonymization reconstruction method based on sound source separation and semantic preservation, as shown, includes: S201. Obtain the original audio and perform sound source separation on the original audio to obtain at least one speaker audio track and an ambient background audio track.
[0048] S202. Authenticate each speaker's audio track and determine the authentication result of each speaker's audio track; wherein, the authentication result of the speaker's audio track includes the authorized speaker's audio track and the unauthorized speaker's audio track.
[0049] S203. Extract features from the unauthorized speaker's audio track to obtain the original speaker's identity features and original semantic features.
[0050] The original speaker identity features are used to characterize the identity of the speaker corresponding to the unauthorized speaker's audio track. For example, the original speaker identity features may include the acoustic and linguistic features of the speaker corresponding to the unauthorized speaker's audio track. Specifically, the acoustic features of the speaker corresponding to the unauthorized speaker's audio track describe the physical properties of the speaker's voice. For example, the acoustic features of the speaker corresponding to the unauthorized speaker's audio track include timbre and / or accent. The linguistic features of the speaker corresponding to the unauthorized speaker's audio track are used to characterize the speaker's vocabulary habits and language style. For example, the linguistic features of the speaker corresponding to the unauthorized speaker's audio track include sentence structure and / or repetition rate. For example, sentence structure includes "long sentences" and "short sentences." Repetition rate includes "high repetition rate" and "low repetition rate."
[0051] The original semantic features are used to characterize the semantic content corresponding to the unauthorized speaker's audio track. Unlike the original speaker identity features, the original semantic features focus more on the semantic content of the speaker corresponding to the unauthorized speaker's audio and / or the implicit meaning of the semantic content.
[0052] In an optional embodiment of the present invention, the original semantic features include at least one of the following: original transcribed text, original semantic vector, and original phoneme sequence representation. The original transcribed text is the textual content corresponding to the speech of the unauthorized human voice track. The original transcribed text directly reflects the semantic content corresponding to the unauthorized human voice track. The original semantic vector is the low-dimensional vector encoding result of the speech or corresponding textual content of the unauthorized human voice track. The original semantic vector can capture the semantic similarity and relationships between the unauthorized human voice tracks. The original phoneme sequence representation is the phoneme decomposition result of the speech signal of the unauthorized human voice track. The original phoneme sequence representation can reflect the pronunciation details and structure of the speech of the unauthorized human voice track. This solution, by concretizing the original semantic features into at least one of the original transcribed text, original semantic vector, and original phoneme sequence representation, and by selecting typical features from the original semantic features, improves the feature extraction efficiency of the original semantic features, thereby improving the efficiency of generative audio anonymization reconstruction based on source separation and semantic preservation.
[0053] Specifically, acoustic processing technology is used to extract the identity features of the unauthorized speaker's audio track, obtaining the original speaker's identity features. Speech recognition technology is used to extract the semantic features of the unauthorized speaker's audio track, obtaining the original semantic features.
[0054] S204. Anonymize the original speaker's identity features to obtain anonymous speaker identity features.
[0055] Anonymous speaker identity features can be the result of anonymizing the original speaker identity features.
[0056] Specifically, the original speaker identity features can be adjusted so that their identifiability is less than or equal to a preset identifiability threshold, resulting in anonymous speaker identity features. Identifiability is used to characterize the degree to which the original speaker identity features are identifiable. The preset identifiability threshold is the maximum identifiability value after anonymizing the original speaker features. The preset identifiability threshold is used to measure the degree of anonymization of the original speaker identity features.
[0057] In an optional embodiment of the present invention, the original speaker identity features include original audio features, original timbre features, and original speaker embedding features; correspondingly, the original speaker identity features are anonymized to obtain anonymous speaker identity features, including: performing frequency shifting or pitch shifting and / or formant shifting on the original audio features to obtain anonymous audio features; performing timbre reconstruction on the original timbre features to obtain anonymous timbre features; and performing feature permutation on the original speaker embedding features to obtain anonymous speaker embedding features.
[0058] The original audio features characterize the frequency of the audio signal corresponding to the unauthorized speaker's audio track. The original timbre features characterize the timbre of the speaker corresponding to the unauthorized speaker's audio track. The original speaker embedding features are low-dimensional vectors representing the speaker's identity corresponding to the unauthorized speaker's audio track. The anonymized audio features are the anonymized result of the original audio features. The anonymized timbre features are the anonymized result of the original timbre features. The anonymized speaker embedding features are the anonymized result of the original speaker embedding features.
[0059] Specifically, time-domain methods or parametric methods can be used to shift or modulate the original audio features to obtain anonymous audio features. For example, time-domain methods can involve resampling, such as upsampling or downsampling to change the pitch period, thereby achieving pitch shifting. Parametric methods can adjust the frequency of a sinusoidal signal to achieve pitch shifting.
[0060] Specifically, either direct or indirect methods can be used to perform formant shifting on the original audio features to obtain anonymous audio features. For example, the direct method first calculates the formant frequency values and then modifies them. The indirect method indirectly modifies the formant frequency values by changing the pole positions or the frequency of the line spectrum.
[0061] Specifically, a vocoder can be used to resynthesize the original timbre features into a time-domain waveform, and the timbre can be changed by adjusting the parameters of the vocoder during the synthesis process to obtain anonymous timbre features.
[0062] Specifically, a preset speaker embedding feature is obtained, and the original speaker embedding feature is replaced using the preset speaker embedding feature to obtain the anonymous speaker embedding feature. The preset speaker embedding feature is a pre-defined speaker embedding feature used to replace the original speaker embedding feature.
[0063] This scheme obtains anonymous audio features by performing frequency shifting or pitch shifting and / or formant shifting on the original audio features, obtains anonymous timbre features by performing timbre reconstruction on the original timbre features, and obtains anonymous speaker embedding features by performing feature substitution on the original speaker embedding features. From the three dimensions of audio, timbre and speaker embedding features, it realizes the anonymization of the original speaker identity features, improves the efficiency of the anonymization of the original speaker identity features, and thus improves the anonymization efficiency of unauthorized human audio tracks.
[0064] S205. The anonymous speaker's identity features and the original semantic features are resynthesized to obtain an anonymized human voice track.
[0065] Specifically, a feature fusion algorithm can be used to resynthesize the anonymous speaker's identity features and the original semantic features to obtain an anonymized audio track.
[0066] In an optional embodiment of the present invention, while extracting features from the unauthorized human voice to obtain the original speaker identity features and the original semantic features, the method further includes: extracting features from the unauthorized human voice to obtain the original prosodic features; and recombining the anonymous speaker identity features and the original semantic features to obtain an anonymized human voice track, including: recombining the anonymous speaker identity features, the original semantic features, and the original prosodic features to obtain an anonymized human voice track.
[0067] Primitive prosodic features are the characteristics in the speech corresponding to the unauthorized speaker's voice track used to express language rhythm, intonation, stress, and pauses. Primitive prosodic features are used to reflect the speaker's emotional state, linguistic intention, and discourse structure corresponding to the unauthorized speaker's voice track.
[0068] In an optional embodiment of the present invention, the original prosodic features include at least one of the original fundamental frequency profile, original speech rate, and original energy envelope. The original fundamental frequency profile is a curve showing the change of the fundamental frequency of the speech signal from the unauthorized human voice track over time. The original fundamental frequency profile reflects the pitch fluctuation pattern in the speech of the unauthorized human voice track. The original speech rate is the number of syllables or words pronounced by the speaker corresponding to the unauthorized human voice track per unit time. The original speech rate reflects the tempo of the speech of the unauthorized human voice track. The original energy envelope is a curve showing the energy change of the speech signal from the unauthorized human voice track within a syllable or phoneme period. The original energy envelope reflects the change in intensity or loudness of the speech of the unauthorized human voice track over time. This solution, by specifying the original prosodic features as at least one of the original fundamental frequency profile, original speech rate, and original energy envelope, and by selecting typical features from the original prosodic features, improves the feature extraction efficiency of the original prosodic features, thereby improving the efficiency of generative audio anonymization reconstruction based on source separation and semantic preservation.
[0069] Specifically, an autocorrelation algorithm can be used to extract features from unauthorized human voices to obtain the original prosodic features. A feature fusion algorithm can be used to resynthesize the anonymous speaker's identity features, the original semantic features, and the original prosodic features to obtain an anonymized voice track.
[0070] This scheme resynthesizes anonymous speaker identity features, original semantic features, and original prosodic features to obtain anonymized human voice tracks. Based on the original semantic features, it further supplements the original prosodic features, thereby improving the completeness of semantic information in unauthorized human voice tracks.
[0071] S206. Resynthesize the anonymized human voice track and the ambient background sound track to obtain the anonymized scene sound track.
[0072] The technical solution of this invention extracts features from unauthorized human audio tracks to obtain original speaker identity features and original semantic features. The original speaker identity features are then anonymized to obtain anonymous speaker identity features. Finally, the anonymous speaker identity features and the original semantic features are resynthesized to obtain an anonymized human audio track. This method improves the anonymization efficiency of unauthorized human audio tracks while preserving the semantic information of the scene and maintaining the scene coherence of the audio.
[0073] Example 3 Figure 3 This is a schematic diagram of a generative audio anonymization reconstruction device based on source separation and semantic preservation, provided in Embodiment 3 of the present invention. This embodiment of the invention is applicable to the anonymization of audio collected in conference or live streaming scenarios. The device can execute a generative audio anonymization reconstruction method based on source separation and semantic preservation. The device can be implemented in hardware and / or software and can be configured in an electronic device that carries the generative audio anonymization reconstruction function based on source separation and semantic preservation, such as a recording device, a mobile terminal, or a server.
[0074] See Figure 3 The generative audio anonymization reconstruction device based on sound source separation and semantic preservation, as shown, includes: an original audio sound source separation module 301, a speaker audio track authentication module 302, an unauthorized speaker audio track anonymization module 303, and an anonymized scene audio track generation module 304. The original audio sound source separation module 301 is used to acquire the original audio and perform sound source separation on the original audio to obtain at least one speaker audio track and an ambient background audio track; the speaker audio track authentication module 302 is used to authenticate each of the speaker audio tracks and determine the authentication result of each speaker audio track; wherein the authentication result of the speaker audio track includes authorized speaker audio tracks and unauthorized speaker audio tracks; the unauthorized speaker audio track anonymization module 303 is used to anonymize the unauthorized speaker audio tracks to obtain an anonymized speaker audio track; and the anonymized scene audio track generation module 304 is used to resynthesize the anonymized speaker audio track and the ambient background audio track to obtain an anonymized scene audio track.
[0075] The technical solution of this invention separates the original audio source to obtain at least one speaker audio track and an ambient background audio track. Each speaker audio track is authenticated to obtain the authentication result. Unauthorized speaker audio tracks are anonymized to obtain an anonymized audio track. The anonymized audio track and the ambient background audio track are then resynthesized to obtain an anonymized scene audio track. This method maximizes the preservation of semantic information of the speech content corresponding to the unauthorized speaker audio track and the environmental information of the scene. Furthermore, it also ensures the privacy of the unauthorized speaker audio track, thereby avoiding audio stream interruptions, improving the scene coherence of the audio, and ultimately enhancing the comprehensibility and overall quality of the audio.
[0076] In an optional embodiment of the present invention, the unauthorized human voice track anonymization module 303 includes: an unauthorized human voice feature extraction unit, used to extract features from the unauthorized human voice track to obtain original speaker identity features and original semantic features; a speaker identity feature anonymization unit, used to anonymize the original speaker identity features to obtain anonymous speaker identity features; and an anonymized human voice track generation unit, used to resynthesize the anonymous speaker identity features and the original semantic features to obtain an anonymized human voice track.
[0077] In an optional embodiment of the present invention, the original speaker identity features include original audio features, original timbre features, and original speaker embedding features; correspondingly, the speaker identity feature anonymization unit includes: an anonymous audio feature generation subunit, used to perform frequency shifting or pitch shifting processing and / or formant shifting processing on the original audio features to obtain anonymous audio features; an anonymous timbre feature generation subunit, used to perform timbre reconstruction on the original timbre features to obtain anonymous timbre features; and an anonymous speaker embedding feature generation subunit, used to perform feature permutation on the original speaker embedding features to obtain anonymous speaker embedding features.
[0078] In an optional embodiment of the present invention, the unauthorized human voice track anonymization module 303 further includes: an original prosodic feature acquisition unit, used to extract features from the unauthorized human voice while extracting features from the unauthorized human voice track to obtain original speaker identity features and original semantic features, to obtain original prosodic features; the anonymized human voice track generation unit includes: an anonymized human voice track generation subunit, used to resynthesize the anonymous speaker identity features, the original semantic features and the original prosodic features to obtain anonymized human voice track.
[0079] In an optional embodiment of the present invention, the original semantic features include at least one of the original transcribed text, the original semantic vector, and the original phoneme sequence representation; the original prosodic features include at least one of the original fundamental frequency profile, the original speech rate, and the original energy envelope.
[0080] In an optional embodiment of the present invention, the apparatus further includes: an audio track output operation detection module, configured to detect an audio track output operation after the anonymized human voice track and the ambient background audio track are resynthesized to obtain an anonymized scene audio track; and a target audio track output module, configured to output a target audio track according to the audio track output operation; wherein the target audio track includes an authorized human voice track or an anonymized scene audio track.
[0081] In an optional embodiment of the present invention, the device further includes: a recording device pointing adjustment module, configured to adjust the pointing of the recording device according to the device orientation, holding posture or historical sound source direction when a pointing adjustment operation is detected or the ambient noise exceeds a preset ambient noise threshold before the original audio is acquired.
[0082] The generative audio anonymization reconstruction apparatus based on sound source separation and semantic preservation provided in this embodiment of the invention can execute the generative audio anonymization reconstruction method based on sound source separation and semantic preservation provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.
[0083] The acquisition, storage, and application of original audio and other materials involved in the technical solutions of this invention comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0084] Example 4 Figure 4 A schematic diagram of an electronic device 400 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0085] like Figure 4As shown, the electronic device 400 includes at least one processor 401 and a memory, such as a read-only memory (ROM) 402 or a random access memory (RAM) 403, communicatively connected to the at least one processor 401. The memory stores computer programs executable by the at least one processor. The processor 401 can perform various appropriate actions and processes based on the computer program stored in the ROM 402 or loaded into the RAM 403 from storage unit 408. The RAM 403 may also store various programs and data required for the operation of the electronic device 400. The processor 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0086] Multiple components in electronic device 400 are connected to I / O interface 405, including: input unit 406, such as keyboard, mouse, etc.; output unit 407, such as various types of displays, speakers, etc.; storage unit 408, such as disk, optical disk, etc.; and communication unit 409, such as network card, modem, wireless transceiver, etc. Communication unit 409 allows electronic device 400 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0087] Processor 401 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 401 performs the various methods and processes described above, such as generative audio anonymization reconstruction methods based on source separation and semantic preservation.
[0088] In some embodiments, the generative audio anonymization reconstruction method based on source separation and semantic preservation can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 400 via ROM 402 and / or communication unit 409. When the computer program is loaded into RAM 403 and executed by processor 401, one or more steps of the generative audio anonymization reconstruction method based on source separation and semantic preservation described above can be performed. Alternatively, in other embodiments, processor 401 can be configured to perform the generative audio anonymization reconstruction method based on source separation and semantic preservation by any other suitable means (e.g., by means of firmware).
[0089] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include: implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0090] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0091] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0092] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0093] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0094] A computing system can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system. It addresses the shortcomings of traditional physical hosts and VPS (Virtual Private Server) services, such as high management difficulty and weak business scalability.
[0095] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0096] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A generative audio anonymization reconstruction method based on sound source separation and semantic preservation, characterized in that, The method includes: The original audio is acquired, and the original audio is separated into sound sources to obtain at least one speaker audio track and an ambient background audio track. The authentication results of each speaker's audio track are determined; wherein, the authentication results of each speaker's audio track include authorized speaker audio tracks and unauthorized speaker audio tracks. The unauthorized voice track is anonymized to obtain an anonymized voice track; The anonymized human voice track and the environmental background audio track are resynthesized to obtain an anonymized scene audio track.
2. The method according to claim 1, characterized in that, The process of anonymizing the unauthorized voice track to obtain an anonymized voice track includes: Feature extraction is performed on the unauthorized speaker's audio track to obtain the original speaker identity features and original semantic features; The original speaker identity features are anonymized to obtain anonymous speaker identity features; The anonymous speaker's identity features and the original semantic features are resynthesized to obtain an anonymous human voice track.
3. The method according to claim 2, characterized in that, The original speaker identity features include original audio features, original timbre features, and original speaker embedding features; Accordingly, the anonymization process of the original speaker identity features to obtain anonymous speaker identity features includes: Perform frequency shifting or pitch shifting and / or formant shifting on the original audio features to obtain anonymous audio features; Phonological reconstruction is performed on the original timbre features to obtain anonymous timbre features; The original speaker embedding features are subjected to feature permutation to obtain anonymous speaker embedding features.
4. The method according to claim 2, characterized in that, In addition to extracting features from the unauthorized speaker's audio track to obtain the original speaker's identity features and original semantic features, the process also includes: Feature extraction is performed on the unauthorized human voice to obtain the original prosodic features; The process of resynthesizing the anonymous speaker's identity features and the original semantic features to obtain an anonymized audio track includes: The anonymous speaker's identity features, the original semantic features, and the original prosodic features are resynthesized to obtain an anonymous human voice track.
5. The method according to claim 4, characterized in that, The original semantic features include at least one of the original transcribed text, the original semantic vector, and the original phoneme sequence representation; the original prosodic features include at least one of the original fundamental frequency profile, the original speech rate, and the original energy envelope.
6. The method according to claim 1, characterized in that, After recombining the anonymized human voice track and the ambient background sound track to obtain the anonymized scene sound track, the method further includes: Detect audio track output operation; According to the audio track output operation, the target audio track is output; wherein, the target audio track includes authorized human voice tracks or anonymized scene audio tracks.
7. The method according to claim 1, characterized in that, Before acquiring the original audio, the following is also included: When a pointing adjustment operation is detected or the ambient noise exceeds a preset ambient noise threshold, the pointing of the recording device is adjusted according to the device orientation, holding posture, or historical sound source direction.
8. A generative audio anonymization reconstruction device based on sound source separation and semantic preservation, characterized in that, The device includes: The original audio source separation module is used to acquire the original audio and perform source separation on the original audio to obtain at least one speaker audio track and an ambient background audio track. The speaker audio track authentication module is used to authenticate each of the speaker audio tracks and determine the authentication result of each speaker audio track; wherein, the authentication result of the speaker audio track includes authorized speaker audio tracks and unauthorized speaker audio tracks; An unauthorized human voice track anonymization module is used to anonymize the unauthorized human voice track to obtain an anonymized human voice track; An anonymized scene audio track generation module is used to resynthesize the anonymized human voice track and the environmental background audio track to obtain an anonymized scene audio track.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the generative audio anonymization reconstruction method based on sound source separation and semantic preservation as described in any one of claims 1-7.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the generative audio anonymization reconstruction method based on sound source separation and semantic preservation according to any one of claims 1-7.
Citation Information
Patent Citations
Voice privacy for far field voice control devices using remote voice services
CN118411995A
End-to-end background reservation voice conversion method based on context learning
CN120299452A
Voice anonymization and model training method, device, equipment, medium and product
CN121214967A
Anonymization privacy protection method and system for voice information retention
CN121459789A
System and Method for Multi-Channel Speech Privacy Processing
US20240296826A1