User voice data security encryption method and system

By performing short-time Fourier transform and random frequency segmentation encryption on the user's voice signal, combined with amplitude encryption, the problem of imperfect voiceprint information encryption in the existing technology is solved, and more efficient voice data security protection is achieved.

CN120260580AActive Publication Date: 2025-07-04GUANGZHOU JIUSI INTELLIGENT TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510702783.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-04
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

In the existing voice data security protection technology, the encryption protection measures for voiceprint information are not perfect enough, and traditional encryption algorithms cannot effectively cover up the voiceprint characteristics in voiceprint information, resulting in the risk of illegal acquisition of voiceprint data during transmission and storage.

Method used

By performing short-time Fourier transform on the user's voice signal to obtain the spectral graph, calculate the degree of voiceprint representation of the data points, and set the frequency segmentation ratio according to the degree of voiceprint representation, randomly arrange the frequency interval for segmentation, randomly select encrypted data, and combine amplitude encryption to realize the encryption of the user's voice signal.

Benefits of technology

It improves the encryption effect of voice signals, enhances the ability to cover voiceprint features, increases the randomness of encrypted data and the difficulty of decryption, and enhances the security of voice data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260580A_ABST
    Figure CN120260580A_ABST
Patent Text Reader

Abstract

The invention relates to the field of data encryption processing, in particular to a user voice data security encryption method and system. The method comprises the following steps: performing short-time Fourier transform on a user voice signal to obtain a spectrogram; calculating the voiceprint representation degree of each data point in the spectrogram; obtaining a pre-obtained frequency interval; setting a frequency segmentation proportion of each data point according to the voiceprint representation degree; randomly arranging the frequency segmentation proportions of all the data points at one moment, and according to the arrangement sequence of any arrangement result, segmenting the frequency interval to obtain the segmentation result of the arrangement; analyzing the frequency deviation degree of the segmentation result of each arrangement; selecting the segmentation result of the arrangement corresponding to the maximum value of the frequency deviation degree as a target segmentation result; randomly selecting an integer from each interval of the target segmentation result as frequency encryption data of the corresponding data point; amplitude encryption of each data point is completed; therefore, user voice signal encryption is realized. And the voice signal encryption effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data encryption processing, and particularly to a method and system for securely encrypting user voice data. Background Art

[0002] With the rapid development of information technology, voice interaction technology has been widely used in people's life and work, such as voice control, voice payment, etc. As a biometric recognition technology, voiceprint recognition has played an increasingly important role in fields such as identity authentication and access control in recent years. Due to the different physiological characteristics of each person, such as the shape of the vocal cords, the length and shape of the vocal tract, the shape of the oral cavity, the position of the teeth and tongue, etc., the voice emitted has unique voiceprint characteristics, which makes voiceprint recognition have high accuracy and reliability. However, precisely because voiceprint information can uniquely identify a person's identity, once the voiceprint data is leaked, it may bring serious security risks to users, such as identity theft, privacy infringement, property loss, etc. In the existing voice data security protection technologies, the voice content itself is usually encrypted and transmitted and stored using encryption algorithms to prevent information from being stolen or tampered with. However, the encryption protection measures for voiceprint information are not yet perfect. Many systems may only use simple encryption methods or even store voiceprint information in plain text when collecting and storing it, which makes the voiceprint data face the risk of being illegally obtained during the transmission process and when stored in the database.

[0003] Traditional encryption algorithms do not consider the voiceprint characteristics in voice information and adopt a unified encryption method for all voice data, resulting in the fact that the voiceprint characteristics in voice information cannot be well masked after encryption, so the voiceprint information in the voice can be extracted through simple analysis and processing. Therefore, how to better mask the voiceprint information in the voice through encryption has become the research focus of the present invention. Summary of the Invention

[0004] In order to solve the problem of how to better mask the voiceprint information in the voice through encryption, the present invention provides a method and system for securely encrypting user voice data.

[0005] In a first aspect, the present invention provides a method for securely encrypting user voice data, adopting the following technical solution: A method for securely encrypting user voice data includes the steps of: Obtain a user voice signal; Perform a short-time Fourier transform on the user voice signal to obtain a spectrogram; calculate the voiceprint representation degree of each data point in the spectrogram, where the voiceprint representation degree is negatively correlated with the distance between the data point and the formant, and is also negatively correlated with the frequency difference and variation law difference between the data point and the harmonic. Obtain the pre-obtained frequency interval; set the frequency division ratio of each data point according to the degree of voiceprint representation, and the frequency division ratio of the data point is positively correlated with the degree of voiceprint representation; randomly arrange the frequency division ratios of all data points at a moment, and according to the arrangement order of any arrangement result, divide the frequency interval to obtain the division result of this arrangement; analyze the frequency offset degree of the division result of each arrangement, and the frequency offset degree reflects the difference between the interval center corresponding to the frequency division ratio of each data point and the original frequency value of the data point; select the division result of the arrangement corresponding to the maximum frequency offset degree as the target division result; randomly select an integer in each interval of the target division result as the frequency encryption data corresponding to the data point; Complete the amplitude encryption of each data point; to achieve the encryption of the user voice signal.

[0006] According to the correlation between each piece of information in the user voice signal and the voiceprint characteristics, the present invention encrypts each piece of information differently, so that the information with a stronger correlation with the voiceprint characteristics is encrypted more strongly, thereby improving the encryption effect; further, when analyzing the correlation between each piece of information and the voiceprint characteristics, information such as formants and harmonic frequencies that can better reflect the voiceprint characteristics is used to accurately reflect the situation of each piece of information having voiceprint characteristics, providing a basis for subsequent precise encryption; further, when encrypting the user voice signal, the encryption data is set by randomly selecting data in the interval, increasing the randomness of the encrypted data, better masking the regular information of the data, and improving the encryption effect; further, when encrypting the user voice signal, the division ratio is set according to the degree of voiceprint representation, so that the selection interval of the encrypted data of the information with a relatively large voiceprint characteristic correlation is larger, increasing the randomness of the encrypted data of the information with a large voiceprint characteristic correlation, and improving the encryption effect of the information with a large voiceprint characteristic correlation. Further, in order to make the frequency of the data point have a large difference from the encrypted data, the frequency offset degree is introduced to accurately measure the frequency difference between the data point and the corresponding interval, thereby laying a foundation for subsequent improvement of the encryption effect.

[0007] Preferably, calculating the degree of voiceprint representation of each data point in the spectrogram includes: Obtain the formant region, harmonic frequency, and variation model of each data point in the spectrogram; if the data point belongs to the formant region of the spectrogram, set the peak distance of the data point to a preset first anti-zero value, otherwise, connect the data point with the geometric center of the formant region, and use the distance between the data point and the intersection of the connection line as the peak distance; add the difference between the frequency of the data point and the harmonic frequency with the smallest frequency difference to a preset second anti-zero value to obtain the harmonic difference; obtain the variation law reference model of the data point, and add the difference between the variation model of the data point and the variation law reference model to a preset third anti-zero value to obtain the variation law difference; calculate the voiceprint characterization degree according to the peak distance, harmonic difference, and variation law difference of the data point, and the voiceprint characterization degree is negatively correlated with the peak distance, harmonic difference, and variation law difference.

[0008] The present invention takes into account that the frequencies of formants and harmonics can strongly reflect voiceprint characteristics, so by introducing the peak distance between the data point and the formant, the difference between the frequency of the data point and the harmonic frequency, and the variation law difference, it can accurately reflect the correlation between the data point and the voiceprint characteristics.

[0009] Preferably, the method for obtaining the variation model includes: Obtain a preset region centered on each data point, construct a Gaussian mixture model based on all data points within the preset region, and use the Gaussian mixture model of the preset region as the variation model of each data point.

[0010] Preferably, the method for obtaining the variation law reference model includes: Obtain the similarity of the variation models of every two data points and record it as the variation similarity, and cluster the variation similarities of all two data points into two categories; record the data points corresponding to the variation similarities in the category with the largest average variation similarity as the law reference data points; cluster the law reference data points according to the variation similarity, and record the law reference data points in one category as the same-category law data points; round up the average value of the corresponding parameters of the variation models of the same-category law data points to obtain the comprehensive parameters, and use the variation model composed of the comprehensive parameters as the law model; Obtain the law model with the smallest difference from the variation model of the data point as the variation law reference model of the data point.

[0011] There are various laws representing voiceprint characteristics in the voice signal of the present invention. By introducing a classification method, data with different laws are separated, thereby increasing the accuracy of fitting the law model.

[0012] Preferably, the method for setting the frequency division ratio of each data point according to the voiceprint characterization degree includes: Divide the voiceprint characterization degree of the data point by the sum of the voiceprint characterization degrees of all data points at a moment, and use the obtained quotient value as the frequency division ratio of the data point.

[0013] The present invention reflects the frequency division ratio through the proportion of the voiceprint characterization degree, so that the data points with a larger voiceprint characterization degree have a larger frequency division ratio, providing a basis for improving the encryption effect of the information with a larger voiceprint characterization degree subsequently.

[0014] Preferably, the process of dividing the frequency interval according to the arrangement order of any one of the arrangement results to obtain the division result of this arrangement includes: According to the arrangement order of any one of the arrangement results, successively intercept corresponding proportion intervals on the frequency interval based on the frequency division ratio to obtain a number of intervals.

[0015] Preferably, the process of analyzing the frequency deviation degree of the division result of each arrangement includes: Obtain the frequency values at the center points of each interval in the division result of any one of the arrangements, and record the absolute value of the difference between the frequency value at the center point of the interval corresponding to the data point and the frequency division ratio as the frequency deviation value of the data point; use the voiceprint characterization degree of the data point as the weight, and perform weighted summation on the frequency deviation values of all data points at a moment to obtain the frequency deviation degree of the division result of this arrangement.

[0016] The present invention introduces the voiceprint characterization degree as the weight to weight the frequency deviation value, so that the data points with a larger voiceprint characterization degree have a greater impact on the frequency deviation degree, providing a basis for improving the encryption effect of the data points with a larger voiceprint characterization degree subsequently.

[0017] Preferably, the process of completing the amplitude encryption of each data point includes: Obtain the amplitude interval obtained in advance; According to the method of obtaining the target division result of the frequency interval, obtain the target division result of the amplitude interval, and randomly select an integer in each interval of the target division result of the amplitude interval as the amplitude encryption data corresponding to the data point.

[0018] Preferably, the process of realizing the encryption of the user voice signal includes: Take the spectrogram composed of the frequency encryption data and the amplitude encryption data as the encrypted spectrogram, and perform inverse transformation on the encrypted spectrogram to obtain the encrypted signal of the user voice signal.

[0019] In a second aspect, the present invention provides a user voice data security encryption system, adopting the following technical solution: A user voice data security encryption system includes: a processor and a memory, and the memory stores computer program instructions, which implement the above-mentioned user voice data security encryption method when the computer program instructions are executed by the processor.

[0020] By adopting the above technical solution, a computer program is generated from the above user voice data security encryption method and stored in a memory to be loaded and executed by a processor, so as to manufacture a terminal device based on the memory and the processor, which is convenient to use.

[0021] The present invention has the following technical effects: According to the correlation between each piece of information in the user voice signal and the voiceprint feature, different encryption is performed on each piece of information, so that the information with a stronger correlation with the voiceprint feature is encrypted more strongly, thereby improving the encryption effect; Further, when analyzing the correlation between each piece of information and the voiceprint feature, information such as formants and harmonic frequencies that can better reflect the voiceprint feature is used to accurately reflect the situation of each piece of information having a voiceprint feature, providing a basis for subsequent precise encryption; Further, when encrypting the user voice signal, the encryption data is set by randomly selecting data in the interval, which increases the randomness of the encrypted data, better masks the regular information of the data, and improves the encryption effect; Further, when encrypting the user voice signal, the segmentation ratio is set according to the degree of voiceprint representation, so that the selection interval of the encrypted data of the information with a relatively large correlation of voiceprint features is larger, increasing the randomness of the encrypted data of the information with a relatively large correlation of voiceprint features and improving the encryption effect of the information with a large correlation of voiceprint features.

[0022] Further, in order to make the frequency of the data points have a large difference from the encrypted data, the degree of frequency offset is introduced to accurately measure the frequency difference between the data points and the corresponding interval, thereby laying a foundation for subsequent improvement of the encryption effect. Description of the Drawings

[0023] Figure 1 is a flowchart of the method in a user voice data security encryption method according to an embodiment of the present invention. Detailed Embodiment

[0024] An embodiment of the present invention discloses a user voice data security encryption method, referring to Figure 1 , including step S1-step S5: S1: Obtain a user voice signal.

[0025] Specifically, collect the user voice signal.

[0026] S2: Perform a short-time Fourier transform on the user voice signal to obtain a spectrogram; calculate the degree of voiceprint representation of each data point in the spectrogram, and the degree of voiceprint representation is negatively correlated with the distance between the data point and the formant, and is also negatively correlated with the difference and variation law difference between the frequency of the data point and the harmonic.

[0027] It should be noted that in order to better conceal the voiceprint features in the voice signal, it is necessary to encrypt the data information with strong voiceprint feature description ability. Therefore, it is necessary to analyze the ability of each data information in the voice signal to describe the voiceprint features.

[0028] S20: Perform short-time Fourier transform on the user voice signal to obtain a spectrogram.

[0029] It should be noted that the spectrogram obtained by short-time Fourier transform can reflect both the frequency-domain information and the time-sequence information of the signal. In order to better analyze the ability of each information in the voice signal to describe the voiceprint features, it is necessary to perform short-time Fourier transform on the voice signal to obtain a spectrogram. Performing short-time Fourier transform on the voice signal is a prior art and will not be elaborated here.

[0030] S21: Calculate the voiceprint characterization degree of each data point in the spectrogram.

[0031] It should be noted that formants, harmonics, and amplitude variation features can better reflect voiceprint features. Therefore, the ability of each information in the voice signal to reflect voiceprint features can be analyzed through formants, harmonics, and amplitude variation features.

[0032] Preferably, as an example, calculating the voiceprint characterization degree of each data point in the spectrogram includes: Obtain the formant region, harmonic frequencies, and variation models of each data point in the spectrogram; It should be noted that there are many existing methods for obtaining the formant region and harmonic frequencies in the spectrogram. Any one of the existing methods can be selected to obtain the formant region and harmonic frequencies. Here, an example method is given: obtain the formant region through image processing methods such as threshold segmentation, obtain the fundamental frequency through the autocorrelation method, and multiply the fundamental frequency by an integer multiple to obtain the harmonic frequencies.

[0033] Calculate the peak distance, harmonic difference, and variation law difference according to the formant region, harmonic frequencies, and variation model in the spectrogram.

[0034] Take the reciprocal of the product of the normalized value of the peak distance, the normalized value of the harmonic difference, and the normalized value of the variation law difference of the data point as the voiceprint characterization degree of the data point.

[0035] In the above embodiments, the peak distance, harmonic difference, and variation law difference are involved. Next, the determination methods of the peak distance, harmonic difference, and variation law difference need to be described.

[0036] Among them, the method for obtaining the peak distance is introduced first.

[0037] If a data point belongs to the resonance peak region of the spectrogram, set the peak distance of the data point to a preset first anti-zero value; otherwise, connect the data point to the geometric center of the resonance peak region, and use the distance between the data point and the intersection of the connection line as the peak distance.

[0038] It can be understood that the larger the peak distance, the farther the data point is from the resonance peak region, so the data point participates less in the information described by the resonance peak. The resonance peak is an important piece of information for describing voiceprint characteristics, so the data point reflects less voiceprint characteristic information.

[0039] Then, introduce the method for obtaining the harmonic difference.

[0040] Obtain the harmonic difference by adding the difference between the frequency of the data point and the harmonic frequency with the smallest frequency difference to a preset second anti-zero value. Exemplarily, the difference between the frequency of the data point and the harmonic frequency with the smallest frequency difference can be calculated by calculating the absolute value of the difference or the standard deviation.

[0041] It can be understood that the harmonic frequency is also important information for reflecting voiceprint characteristics. If the difference between the frequency of the data point and the harmonic frequency is large, the data point describes less voiceprint characteristic information.

[0042] Finally, introduce the method for obtaining the variation law difference.

[0043] Obtain a preset region centered on each data point, construct a Gaussian mixture model based on all the data points within the preset region, and use the Gaussian mixture model of the preset region as the variation model of each data point. Obtain the similarity of the variation models of every two data points, which is denoted as the variation similarity. Exemplarily, the similarity of the variation models of two data points can be reflected by the cosine similarity of the parameters between the variation models of the two data points. The parameters of the variation model can be the mean variances of each single Gaussian model in the Gaussian mixture model.

[0044] Cluster the variation similarities of all pairs of data points into two categories; denote the data points corresponding to the variation similarities in the category with the largest mean of variation similarities as the regular reference data points; cluster the regular reference data points according to the variation similarities, and denote the regular reference data points in one category as the same-category regular data points; denote the ceiling value of the mean of the corresponding parameters of the variation models of the same-category regular data points as the comprehensive parameter, and use the variation model composed of the comprehensive parameters as the regular model; obtain the regular model with the smallest difference from the variation model of the data point as the variation law reference model of the data point; obtain the variation law difference by adding the difference between the variation model of the data point and the variation law reference model to a preset third anti-zero value. Exemplarily, the difference between the variation model of the data point and the variation law reference model can be reflected by the parameter difference between the models, and the difference between the parameters of the models can be reflected by the Euclidean distance between the parameters of the models.

[0045] It can be understood that there are various variation law information in the voice signal that reflects the voiceprint characteristics. By decomposing the law model of each law for description and comparing the data points with the law model with the smallest difference, the compliance of the data points with the variation law of the voiceprint is judged.

[0046] S3: Obtain the pre-obtained frequency range; set the frequency division ratio of each data point according to the voiceprint characterization degree, and the frequency division ratio of the data point is positively correlated with the voiceprint characterization degree; randomly arrange the frequency division ratios of all data points at a moment, and according to the arrangement order of any arrangement result, divide the frequency range to obtain the division result of this arrangement; analyze the frequency offset degree of the division result of each arrangement, and the frequency offset degree reflects the difference between the interval center corresponding to the frequency division ratio of each data point and the original frequency value of the data point; select the division result of the arrangement corresponding to the maximum frequency offset degree as the target division result; randomly select an integer in each interval of the target division result as the frequency encryption data corresponding to the data point.

[0047] It should be noted that in order to mask the voiceprint characteristic information in the voice signal, other frequency values can be used for replacement. However, the fixed value replacement is easily decrypted through the regularity of the encrypted data, thus exposing the voiceprint characteristic information. Therefore, the replacement can be carried out by randomly selecting replacement values within the interval. There are more optional values within the interval, so there are more types of values of the encrypted frequency data, and it is difficult to decode the data through the data law. In order to better mask the information with a large voiceprint characterization degree, the selection interval of the encrypted frequency of the information with a large voiceprint characterization degree in the voice signal can be set larger.

[0048] S30: Obtain the pre-obtained frequency range.

[0049] It should be noted that people's voices generally have a certain frequency range, so the common frequency range of people's voices can be used as the frequency range, and then the selection interval of the encrypted data of each data is set by dividing the frequency range.

[0050] Optionally, obtaining the pre-obtained frequency range includes: Taking the frequency range of people's voices as the pre-obtained frequency range. For example, the frequency range of people's voices is 200–5000Hz.

[0051] S31: Set the frequency division ratio of each data point according to the voiceprint characterization degree.

[0052] It should be noted that, in order to better conceal information with a large voiceprint characterization degree, the selection interval of the encrypted frequency of the information with a large voiceprint characterization degree in the voice signal can be set to be larger, and based on this, the selection interval of the encrypted frequency of each data point is intercepted in the frequency interval.

[0053] Preferably, as an example, the frequency division ratio of each data point is set according to the voiceprint characterization degree, including: Dividing the voiceprint characterization degree of the data point by the sum of the voiceprint characterization degrees of all data points at a moment, and taking the obtained quotient value as the frequency division ratio of the data point.

[0054] It can be understood that the frequency division ratio is reflected by the proportion of the voiceprint characterization degree of the data point, so that the frequency division ratio of the data point with a large voiceprint characterization degree is larger, providing a basis for stronger encryption of the data point with a large voiceprint characterization in the subsequent process.

[0055] S32: Randomly arrange the frequency division ratios of all data points at a moment. According to the arrangement order of any one arrangement result, divide the frequency interval to obtain the division result of this arrangement; analyze the frequency offset degree of the division result of each arrangement, where the frequency offset degree reflects the difference between the interval center corresponding to the frequency division ratio of each data point and the original frequency value of this data point; select the division result of the arrangement corresponding to the maximum frequency offset degree as the target division result.

[0056] It should be noted that, in order to make the difference between the encrypted frequency and the pre-encrypted frequency of each data point relatively large, when dividing the interval, the division interval of the data point should be far from the pre-encrypted frequency. Therefore, based on this, interval division control is performed to obtain the interval of the data point.

[0057] Preferably, as an example, randomly arrange the frequency division ratios of all data points at a moment. According to the arrangement order of any one arrangement result, divide the frequency interval to obtain the division result of this arrangement; analyze the frequency offset degree of the division result of each arrangement, where the frequency offset degree reflects the difference between the interval center corresponding to the frequency division ratio of each data point and the original frequency value of this data point; select the division result of the arrangement corresponding to the maximum frequency offset degree as the target division result, including: Randomly arrange the frequency division ratios of all data points at a moment. According to the arrangement order of any one arrangement result, successively intercept corresponding proportional intervals on the frequency interval based on the frequency division ratio to obtain several intervals.

[0058] Obtain the frequency values at the center points of each interval in the segmentation results of any permutation, and record the absolute value of the difference between the data point and the frequency value at the center point of the interval corresponding to the frequency segmentation ratio as the frequency deviation value of the data point; use the voiceprint characterization degree of the data point as the weight, and perform weighted summation on the frequency deviation values of all data points at a moment to obtain the frequency deviation degree of the segmentation result of this permutation combination.

[0059] Select the segmentation result of the permutation corresponding to the maximum frequency deviation degree as the target segmentation result.

[0060] It can be understood that by selecting the segmentation result with the largest frequency deviation degree as the target segmentation result. This can make the difference between the encrypted data and the original data larger, thereby increasing the decryption difficulty.

[0061] S33: Randomly select an integer in each interval of the target segmentation result as the frequency encrypted data corresponding to the data point.

[0062] It can be understood that by encrypting in the way of randomly selecting data instead of using the mapping relationship rule for encryption, the problem of the mapping relationship rule exposing the voiceprint information can be prevented, and the encryption accuracy can be improved.

[0063] S4: Complete the amplitude encryption of each data point to realize the encryption of the user voice signal.

[0064] It should be noted that in order to prevent the amplitude information of the data point from exposing the voiceprint information, it is necessary to encrypt the amplitude of each data point.

[0065] Preferably, as an example, to complete the amplitude encryption of each data point to realize the encryption of the user voice signal, including: It should be noted that there is generally a certain amplitude range in people's voices, so the common amplitude range of people's voices can be used as the amplitude interval.

[0066] Optionally, obtaining the pre-obtained amplitude interval includes: Taking the amplitude range of people's voices as the pre-obtained amplitude interval. For example, the amplitude range of people's voices is 20–85 Pa.

[0067] Divide the voiceprint characterization degree of the data point by the sum of the voiceprint characterization degrees of all data points at a moment, and use the obtained quotient value as the amplitude segmentation ratio of the data point. Based on the amplitude segmentation ratio and the amplitude interval, in the same way as obtaining the target segmentation result of the frequency interval, obtain the target segmentation result of the amplitude interval, and randomly select an integer in each interval of the target segmentation result of the amplitude interval as the amplitude encrypted data corresponding to the data point. Use the spectrogram composed of the frequency encrypted data and the amplitude encrypted data as the encrypted spectrogram, and perform inverse transformation on the encrypted spectrogram to obtain the encrypted signal of the user voice signal.

[0068] It should be noted that the method of obtaining the signal by inverse-transforming the encrypted spectrogram is a prior art and will not be elaborated here.

[0069] It can be understood that the above encryption sets different interval ranges according to the situation of the voiceprint features contained in the data points, so that the encrypted data of the data points with more voiceprint feature information has more data selection types, improving the randomness of its encrypted data and strengthening its encryption effect.

[0070] It should be added that the method for obtaining the key includes: using the difference between the frequency of the data point and the frequency at the center point of the interval corresponding to the frequency division ratio as the frequency offset value.

[0071] Obtain the amplitude deviation value of the data point in the same way.

[0072] Sort the voiceprint characterization degrees of all data points at a moment according to the division order of the target division result of the frequency to obtain a voiceprint characterization degree sequence. In the same sorting way, sort the frequency offset values of all data points at a moment to obtain a frequency offset value sequence, and splice the frequency offset value sequence behind the voiceprint characterization degree sequence to obtain the frequency key at a moment; Obtain the amplitude key at a moment in the same way.

[0073] S5: Perform decryption processing.

[0074] Preferably, as an example, performing decryption processing includes: Performing a short-time Fourier transform on the encrypted signal of the user voice signal to obtain an encrypted spectrogram; It should be noted that the time window length of the short-time Fourier transform is a length agreed upon in advance by the encryptor and the decryptor, and the time window lengths of all short-time Fourier transforms for encryption and decryption are the same.

[0075] Obtain the frequency keys at each moment, and decompose the voiceprint characterization degree sequence and the frequency offset value sequence in the frequency key. Calculate the frequency division ratio according to each voiceprint characterization degree in the voiceprint characterization degree sequence. According to the frequency division ratio and the corresponding sorting way of the voiceprint characterization degree, divide the frequency interval; obtain the interval to which the frequency of each data point in the encrypted spectrogram belongs in the division result and record it as the analysis interval of each data point. According to the order of the voiceprint characterization degree corresponding to the analysis interval in the voiceprint characterization degree sequence, obtain the frequency offset value corresponding to the analysis interval in the frequency offset value sequence, add the frequency offset value to the center point of the analysis interval, and use the obtained sum value as the original frequency value of the data point corresponding to the analysis interval. In the same way, decrypt the original amplitude of the data point according to the amplitude key.

[0076] The spectrogram formed by the original amplitude values and original frequency values of all data points is inversely transformed to obtain the original speech signal.

[0077] An embodiment of the present invention also discloses a user voice data security encryption system, including a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, a user voice data security encryption method according to the present invention is implemented.

[0078] The above system also includes other components well known to those skilled in the art, such as a communication bus and a communication interface. Their settings and functions are known in the art, so they will not be elaborated here.

[0079] In the present invention, the aforementioned memory can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or apparatus. For example, a computer-readable storage medium can be any suitable magnetic storage medium or magneto-optical storage medium, such as a resistive random access memory, a dynamic random access memory, a static random access memory, an enhanced dynamic random access memory, a high-bandwidth memory, a hybrid memory cube, etc., or any other medium that can be used to store the required information and can be accessed by an application program, a module, or both. Any such computer storage medium can be part of the device or accessible or connectable to the device.

Claims

1. A method for secure encryption of user voice data, characterized in that, Including the steps of: Obtaining a user voice signal; Performing a short-time Fourier transform on the user voice signal to obtain a spectrogram; calculating the voiceprint characterization degree of each data point in the spectrogram, where the voiceprint characterization degree is negatively correlated with the distance between the data point and the formant, and is also negatively correlated with the frequency difference and the variation law difference between the data point and the harmonic; Obtaining a pre-obtained frequency interval; setting the frequency segmentation ratio of each data point according to the voiceprint characterization degree, where the frequency segmentation ratio of the data point is positively correlated with the voiceprint characterization degree; randomly arranging the frequency segmentation ratios of all data points at a moment, and dividing the frequency interval according to the arrangement order of any arrangement result to obtain the segmentation result of this arrangement; analyzing the frequency offset degree of the segmentation result of each arrangement, where the frequency offset degree reflects the difference between the center of the interval corresponding to the frequency segmentation ratio of each data point and the original frequency value of the data point; Selecting the segmentation result of the arrangement corresponding to the maximum frequency offset degree as the target segmentation result; randomly selecting an integer in each interval of the target segmentation result as the frequency encryption data corresponding to the data point; Completing the amplitude encryption of each data point; to achieve the encryption of the user voice signal.

2. The method for securely encrypting user voice data according to claim 1, characterized in that, The calculating of the voiceprint characterization degree of each data point in the spectrogram includes: Obtaining the formant region, harmonic frequencies, and the variation model of each data point in the spectrogram; if the data point belongs to the formant region of the spectrogram, setting the peak distance of the data point as a preset first anti-zero value, otherwise, connecting the data point with the geometric center of the formant region, and taking the distance between the data point and the intersection of the connection line as the peak distance; adding the difference between the frequency of the data point and the harmonic frequency with the smallest frequency difference to a preset second anti-zero value to obtain the harmonic difference; obtaining the variation law reference model of the data point, and adding the difference between the variation model of the data point and the variation law reference model to a preset third anti-zero value to obtain the variation law difference; calculating the voiceprint characterization degree according to the peak distance, harmonic difference, and variation law difference of the data point, where the voiceprint characterization degree is negatively correlated with the peak distance, harmonic difference, and variation law difference.

3. A method for securely encrypting user voice data according to claim 2, characterized in that, The method for obtaining the variation model includes: Taking each data point as the center to obtain a preset region, constructing a Gaussian mixture model based on all data points in the preset region, and taking the Gaussian mixture model of the preset region as the variation model of each data point.

4. A method for securely encrypting user voice data according to claim 3, characterized in that, The method for obtaining the variation law reference model includes: Obtaining the similarity of the variation models of every two data points and recording it as the variation similarity, clustering all the variation similarities of every two data points into two categories; recording the data points corresponding to the variation similarities in the category with the largest average variation similarity as the regular reference data points; clustering the regular reference data points according to the variation similarity, and recording the regular reference data points in one category as the same-category regular data points; taking the ceiling value of the average value of the corresponding parameters of the variation models of the same-category regular data points as the comprehensive parameter, and taking the variation model composed of the comprehensive parameters as the regular model; Obtaining the regular model with the smallest difference from the variation model of the data point as the variation law reference model of the data point.

5. A method for securely encrypting user voice data according to claim 1, characterized in that The setting of the frequency segmentation ratio of each data point according to the voiceprint characterization degree includes: Divide the voiceprint characterization degree of a data point by the sum of the voiceprint characterization degrees of all data points at a moment, and use the obtained quotient value as the frequency division ratio of the data point.

6. A method for secure encryption of user voice data according to claim 1, characterized in that, The dividing the frequency interval according to the arrangement order of any one of the arrangement results to obtain the division result of this arrangement includes: According to the arrangement order of any one of the arrangement results, successively intercept corresponding proportion intervals on the frequency interval based on the frequency division ratio to obtain a number of intervals.

7. A method for secure encryption of user voice data according to claim 1, characterized in that, The analyzing the frequency deviation degree of the division result of each arrangement includes: Obtain the frequency value at the center point of each interval in the division result of any one arrangement, and record the absolute value of the difference between the frequency value at the center point of the interval corresponding to the data point and the frequency division ratio as the frequency deviation value of the data point; use the voiceprint characterization degree of the data point as the weight, and perform a weighted sum of the frequency deviation values of all data points at a moment to obtain the frequency deviation degree of the division result of this arrangement.

8. A method for securely encrypting user voice data according to claim 1, characterized in that, The completing the amplitude encryption of each data point includes: Obtain the amplitude interval obtained in advance; According to the way of obtaining the target division result of the frequency interval, obtain the target division result of the amplitude interval, and randomly select an integer in each interval of the target division result of the amplitude interval as the amplitude encryption data corresponding to the data point.

9. A method for secure encryption of user voice data according to claim 1, characterized in that The realizing the encryption of the user voice signal includes: Take the spectrogram composed of the frequency encryption data and the amplitude encryption data as the encrypted spectrogram, and perform an inverse transformation on the encrypted spectrogram to obtain the encrypted signal of the user voice signal.

10. A user voice data security encryption system, characterized in that including: A processor and a memory, where the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a user voice data security encryption method according to any one of claims 1-9 is implemented.

Citation Information

Patent Citations

  • Voice encryption and decryption method

    CN102624518A

  • Voice data processing method and terminal

    CN105096937A

  • Voiceprint authentication method and device, electronic equipment and storage medium

    CN116821881A

  • Formant frequency estimation method, apparatus, and medium in speech recognition

    US20070192088A1

  • Voiceprint recognition method and device employing deep learning, and apparatus

    WO2021051608A1