A method for constructing representations based on specific person audio

By building an analysis model that combines human voice and environmental data and generating comprehensive characterization parameters, the problem of inaccurate single feature analysis in traditional audio characterization technology is solved, and more accurate and real-time audio identification is achieved.

CN116913285BActive Publication Date: 2025-09-05CHINA ACADEMY OF INFORMATION & COMM
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310980913.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-04
Publication Date
2025-09-05
Estimated Expiration
2043-08-04

AI Technical Summary

Technical Problem

Traditional audio characterization technology only focuses on a single feature and ignores the impact of ambient sound, resulting in inaccurate analysis.

Method used

Build a human voice analysis model, combine human voice and environmental data, generate comprehensive characterization parameters, and perform audio identification through multi-dimensional data analysis and real-time processing technology.

Benefits of technology

It improves the accuracy and real-time performance of audio identification, comprehensively considers environmental factors, and provides richer audio feature representation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116913285B_ABST
    Figure CN116913285B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for constructing a representation based on specific person audio, which relates to the field of audio analysis technology. The specific steps include: step S100, obtaining human voice data and environmental data in specific person audio data; step S200, constructing a human voice analysis model, analyzing the human voice data and generating a human voice analysis coefficient; step S300, performing combined analysis on the human voice data and environmental data to generate an environmental analysis coefficient; step S400, integrating the human voice analysis coefficient and the environmental analysis coefficient to generate a representation parameter for the specific person audio data; step S500, performing threshold analysis on the representation parameter, and characterizing and marking the specific person audio according to the analysis result. The present invention takes into account the impact of the environment on the audio, thereby more comprehensively analyzing the authenticity of the audio; adopting real-time means to obtain and process data, the specific person audio can be represented and identified in real time, thereby increasing the real-time and accuracy of the identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of audio analysis, and in particular to a method for constructing a representation based on the audio of a specific person. Background Art

[0002] After searching, a comparative document with publication number CN107680601B proposed an identity verification method and device based on spectrogram and phoneme retrieval. The comparative document provides an identity verification method based on spectrogram and phoneme retrieval, including: obtaining a spectrogram corresponding to a sample audio file; obtaining speech feature parameters of the sample audio file; constructing a phoneme recognition model, inputting the speech feature parameters into the phoneme recognition model to perform phoneme retrieval and obtain matching phonemes; marking the matching phonemes on the spectrogram, and matching vowels or vowel groups with the same markings. The identity test is carried out together to determine whether the identity authentication of the person to be identified corresponding to the sample audio file has passed. By constructing a phoneme recognition model, the phonemes that meet the requirements in the sample audio file are retrieved, and the phonemes that meet the requirements are compared with the spectrogram corresponding to the sample audio file to identify the identity of the person to be identified corresponding to the sample audio file. This is more accurate than manual comparison, and multiple phonemes that meet the requirements are retrieved through the phoneme recognition model, which further improves the accuracy of the comparison, solves the technical problems of searching and finding phonemes in actual voiceprint identification, and visualizes the phonemes to improve the identification efficiency of case handlers.

[0003] Combining the comparative documents with the prior art, we find that the prior art still has the following deficiencies:

[0004] 1. Traditional audio characterization techniques often focus on single features, such as phonemes, spectrum, and energy, which can lead to inaccuracies in analysis.

[0005] 2. Ambient sound has a significant impact on the quality and characteristics of audio, which is often overlooked by traditional audio characterization techniques.

[0006] To solve the above-mentioned problems, a representation construction method based on specific person audio is proposed. Summary of the Invention

[0007] The purpose of the present invention is to provide a method for constructing a representation based on the audio of a specific person to address the deficiencies in the background technology.

[0008] In order to achieve the above object, the present invention provides the following technical solutions:

[0009] The method for constructing a representation based on the audio of a specific person comprises the following steps:

[0010] Step S100: Acquire human voice data and environmental data from audio data of a specific person;

[0011] Step S200: constructing a vocal analysis model, analyzing the vocal data and generating a vocal analysis coefficient;

[0012] Step S300: performing combined analysis on the human voice data and the environmental data to generate an environmental analysis coefficient;

[0013] Step S400: integrating the human voice analysis coefficient and the environmental analysis coefficient to generate characterization parameters for the specific person's audio data;

[0014] Step S500: Perform threshold analysis on the characterization parameters, and characterize and mark the audio of the specific person based on the analysis results.

[0015] In a preferred embodiment, the human voice data includes the human voice speech rate change rate and the voice intonation change amplitude. The speech rate change rate reflects the degree of change of the speaker's speech rate in the audio. The greater the speech rate change rate, the lower the corresponding degree of synthesis of the specific person's audio. The voice intonation change amplitude refers to the high and low frequencies of the sound, and also indicates the degree of change of the speaker's voice in the audio. The greater the voice intonation change amplitude, the lower the corresponding degree of synthesis of the specific person's audio.

[0016] In a preferred embodiment, the environmental data includes the ambient sound intensity change rate and the ambient sound frequency change rate. The ambient sound intensity change rate indicates the degree of change in the sound intensity of the ambient sound in the audio. The greater the ambient sound intensity change rate, the greater the influence of the corresponding specific person audio when performing synthetic audio identification. The ambient sound frequency change rate indicates the degree of change in the sound frequency of the ambient sound in the audio. The greater the sound frequency change degree, the greater the influence of the corresponding specific person audio when performing synthetic audio identification.

[0017] In a preferred embodiment, the steps of constructing the human voice analysis model include:

[0018] Use appropriate microphones or recording equipment to collect audio data of specific people and environmental sound data;

[0019] Preprocess audio data and remove background noise using audio signal processing algorithms and tools;

[0020] Extracting speech rate change rate: The speech rate change rate is obtained by calculating the ratio between the audio duration and the speech rate;

[0021] Extracting the amplitude of intonation changes: Using a fundamental frequency estimation algorithm based on the autocorrelation function method to calculate the fundamental frequency changes in the audio;

[0022] Extracting the rate of change of ambient sound intensity: Using audio signal processing methods to estimate the rate of change of ambient sound intensity, the audio is framed and then the energy change rate of each frame is calculated.

[0023] Extracting the rate of change of ambient audio frequency: Using spectrum analysis algorithm to estimate the rate of change of ambient audio frequency;

[0024] Normalize the extracted features of the human voice speech rate change rate and intonation change amplitude to generate the human voice analysis coefficient;

[0025] Normalize and weight the extracted features of the ambient sound intensity change rate and the ambient sound frequency change rate to generate an environmental analysis coefficient;

[0026] Perform combined analysis on the vocal analysis coefficient and the environmental analysis coefficient, and use weighted average or logical operation method to integrate the vocal and environmental analysis results to generate representation parameters;

[0027] Set a threshold or threshold value, compare the characterization parameter with the threshold value, and determine the authenticity of the audio of a specific person.

[0028] In a preferred embodiment, the step of generating the vocal analysis coefficient includes:

[0029] Divide the specific human audio into L identification intervals, obtain the human voice speed change rate v and the voice intonation change amplitude h for each identification interval, add the human voice speed change rate v and the voice intonation change amplitude h for each identification interval, take the root of the sum, take the absolute value of the root-sum result, and divide it by the error correction constant K to obtain the human voice analysis coefficient α;

[0030] Among them, K and α are both greater than 0. The larger the value of the vocal analysis coefficient α, the greater the emotional fluctuation of the vocal expression in the specific person's audio, the more the language expression tends to be artificial, the higher the authenticity, and the lower the probability of synthesis. Conversely, the lower the value of the vocal analysis coefficient α, the smaller the emotional fluctuation of the vocal expression in the specific person's audio, the more mechanical the language expression, the lower the authenticity, and the higher the probability of synthesis.

[0031] In a preferred embodiment, the step of generating the environmental analysis coefficient includes:

[0032] Obtain the ambient sound intensity change rate q and the ambient sound frequency change rate p within the same single identification interval as step S200, perform weighted analysis on the ambient sound intensity change rate q and the ambient sound frequency change rate p within the single identification interval, square the weighted analysis result, and multiply it by the error compensation constant , calculate and obtain the environmental analysis coefficient γ;

[0033] Among them, the weight parameter m1+m2=2.1325, the weight parameter m1 is greater than the weight parameter m2. When the environmental analysis coefficient γ is larger, the influence of the ambient sound is greater, and the influence on the audio identification processing will also be greater; conversely, the smaller the environmental analysis coefficient γ is, the smaller the influence of the ambient sound is, and the influence on the audio identification processing will also be smaller.

[0034] In a preferred embodiment, the step of integrating the vocal analysis coefficient and the environmental analysis coefficient includes:

[0035] Perform parameter analysis on the vocal analysis coefficient α and the environmental analysis coefficient γ within a single interval, and obtain the representation parameter β by dividing the vocal analysis coefficient α by the cube root of the environmental analysis coefficient γ and adding the vocal analysis coefficient α multiplied by the inverse of the environmental analysis coefficient γ;

[0036] The larger the representation parameter β is, the greater the calculation proportion of the human voice analysis coefficient α is, the corresponding authenticity is higher, and the environmental influence effect in this case is weaker than the human voice influence effect in the specific human audio; conversely, the smaller the representation parameter β is, the smaller the calculation proportion of the human voice analysis coefficient α is, the corresponding authenticity is lower, and the environmental influence effect in this case is stronger than the human voice influence effect in the specific human audio.

[0037] In a preferred embodiment, the process of performing threshold analysis on the characterization parameters and performing characterization feature marking is as follows:

[0038] Set characterization parameter comparison thresholds SH1 and SH2, where the characterization parameter comparison threshold SH1 is greater than the characterization parameter comparison threshold SH2, substitute the characterization parameter β into the characterization parameter comparison thresholds SH1 and SH2 for comparison, and if the characterization parameter β is less than the characterization parameter comparison threshold SH2, mark the audio of the specific person as a false characterization parameter;

[0039] Substitute the characterization parameter β into the characterization parameter comparison thresholds SH1 and SH2 for comparison. If the characterization parameter β is less than the characterization parameter comparison threshold SH1 and the characterization parameter β is greater than the characterization parameter comparison threshold SH2, mark the audio of the specific person as suspicious.

[0040] The characterization parameter β is substituted into the characterization parameter comparison thresholds SH1 and SH2 for comparison. If the characterization parameter β is greater than the characterization parameter comparison threshold SH1, the audio of the specific person is marked as a true characterization parameter.

[0041] In the above technical solution, the technical effects and advantages provided by the present invention are:

[0042] 1. The present invention integrates multi-dimensional data such as the rate of change of human voice speed, the amplitude of voice tone change, the rate of change of ambient sound intensity, and the rate of change of ambient sound frequency in the audio of a specific person for analysis, rather than being limited to single feature extraction.

[0043] 2. The present invention introduces environmental sound data for analysis, taking into account the impact of the environment on the audio, thereby more comprehensively analyzing the authenticity of the audio.

[0044] 3. The present invention uses real-time means to acquire and process data, which can characterize and identify the audio of a specific person in real time, thereby increasing the real-time nature and accuracy of the identification.

[0045] 4. The present invention adopts a multi-level combined analysis to comprehensively process the human voice analysis coefficient and the environmental analysis coefficient, thereby more comprehensively analyzing and processing the audio representation of a specific person.

[0046] In summary, the present invention utilizes a human voice analysis model that utilizes comprehensive multidimensional data (including the rate of change of human voice speed, the amplitude of voice intonation change, the rate of change of ambient sound intensity, the rate of change of ambient sound frequency, etc.), combined with advanced real-time processing technology, to comprehensively analyze the audio of a specific individual. By considering environmental factors such as background noise and the degree of ambient sound variability, the model can more accurately characterize the audio characteristics of a specific individual. Furthermore, a combined analysis method is used to integrate human voice analysis coefficients and environmental analysis coefficients to generate more comprehensive characterization parameters, thereby more accurately identifying and labeling the audio of a specific individual. Compared to traditional audio characterization techniques, this model can provide richer and more accurate audio feature representations, improving identification accuracy and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction to the drawings required for use in the embodiments will be given below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0048] Figure 1 This is a flowchart of a method for constructing a representation based on audio of a specific person according to the present invention. DETAILED DESCRIPTION

[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0050] See also Figure 1 As shown, the method for constructing a representation based on the audio of a specific person described in this embodiment includes the following steps:

[0051] Step S100: Acquire human voice data and environmental data from audio data of a specific person;

[0052] Step S200: constructing a vocal analysis model, analyzing the vocal data and generating a vocal analysis coefficient;

[0053] Step S300: performing combined analysis on the human voice data and the environmental data to generate an environmental analysis coefficient;

[0054] Step S400: integrating the human voice analysis coefficient and the environmental analysis coefficient to generate characterization parameters for the specific person's audio data;

[0055] Step S500: Perform threshold analysis on the characterization parameters, and characterize and mark the audio of the specific person based on the analysis results.

[0056] Audio representation construction refers to converting sound data into specific numerical representations so that subsequent audio analysis tasks can proceed smoothly;

[0057] In this solution, the characterization is constructed based on the audio of the person.

[0058] The human voice data includes the human voice speed change rate and the voice intonation change amplitude. The speech speed change rate reflects the degree of change of the speaker's speech speed in the audio. The greater the speech speed change rate, the lower the degree of synthesis of the corresponding specific person's audio. The voice intonation change amplitude refers to the high and low frequencies of the sound, and also indicates the degree of change of the speaker's voice intonation in the audio. The greater the voice intonation change amplitude, the lower the degree of synthesis of the corresponding specific person's audio.

[0059] It should be noted that the speech rate change rate and intonation change rate can be obtained through Praat, OpenSMILE, librosa, or MATLAB. These tools and libraries can be customized and expanded according to different requirements and data formats, and can be easily used to extract features related to speech rate and intonation.

[0060] If the speech rate in an audio file is high, it means that the speaker's speaking speed varies greatly at different time periods, such as fast or slow speech. If the intonation in an audio file is large, it means that the speaker's voice changes significantly, such as fluctuating intonation or switching between high and low pitch.

[0061] The rate of change of speech rate can be used to analyze the speaker's speaking speed performance in different emotional states or contexts. For example, when a speaker is excited, nervous, or sad, the speaking speed may change significantly. The amplitude of intonation change can be used to analyze the speaker's emotional expression or emphasis. For example, when a speaker emphasizes a certain point, the intonation may change significantly.

[0062] The environmental data includes the ambient sound intensity change rate and the ambient sound frequency change rate. The ambient sound intensity change rate indicates the degree of change in the sound intensity of the ambient sound in the audio. The greater the ambient sound intensity change rate, the greater the influence of the corresponding specific person audio when performing synthetic audio identification. The ambient sound frequency change rate indicates the degree of change in the sound frequency of the ambient sound in the audio. The greater the degree of change in the sound frequency, the greater the influence of the corresponding specific person audio when performing synthetic audio identification.

[0063] It should be noted that the ambient sound does not include human voices. The two data, the rate of change of the ambient sound intensity and the degree of change of the ambient sound frequency, can be extracted from the identification audio using audio processing software and signal processing tools, such as Adobe Audition and Audacity.

[0064] The ambient sound intensity change rate can be used as the first indicator to assess the impact of ambient sound. The ambient sound intensity change rate indicates the rate of change of the ambient sound intensity over a period of time. If the ambient sound intensity change rate is large, it means that the ambient sound intensity has changed significantly in a short period of time, which may have a significant impact on the authenticity of the audio. Conversely, if the ambient sound intensity change rate is small, it means that the ambient sound intensity changes slowly and has a relatively small impact on the audio.

[0065] Frequency is the energy distribution of different frequency components in sound. The frequency change rate of ambient sound can be used as a second indicator of the degree of ambient sound impact. The higher the frequency change of ambient sound, the greater the impact of ambient sound. The impact of ambient sound will cause the frequency components of sound to change. If the impact of ambient sound is large, more frequency components may be introduced, making the sound frequency change more obvious. On the contrary, if the impact of ambient sound is small, the sound frequency change may be relatively small.

[0066] The steps of constructing the human voice analysis model include:

[0067] Use appropriate microphones or recording equipment to collect audio data of specific people and environmental sound data;

[0068] Preprocess the audio data and remove background noise using audio signal processing algorithms and tools, such as noise suppression algorithms and filters;

[0069] Extracting speech rate change rate: The speech rate change rate is obtained by calculating the ratio between the audio duration and the speech rate;

[0070] Extracting the amplitude of intonation changes: Using a fundamental frequency estimation algorithm based on the autocorrelation function method to calculate the fundamental frequency changes in the audio;

[0071] Extracting the rate of change of ambient sound intensity: Using audio signal processing methods to estimate the rate of change of ambient sound intensity, the audio is framed and then the energy change rate of each frame is calculated.

[0072] Extracting the rate of change of ambient audio frequency: Using spectrum analysis algorithm to estimate the rate of change of ambient audio frequency;

[0073] Normalize the extracted features of the human voice speech rate change rate and intonation change amplitude to generate the human voice analysis coefficient;

[0074] Normalize and weight the extracted features of the ambient sound intensity change rate and the ambient sound frequency change rate to generate an environmental analysis coefficient;

[0075] Perform combined analysis on the vocal analysis coefficient and the environmental analysis coefficient, and use weighted average or logical operation method to integrate the vocal and environmental analysis results to generate representation parameters;

[0076] Set a threshold or threshold value, compare the characterization parameter with the threshold value, and determine the authenticity of the audio of a specific person.

[0077] It should be noted that the technologies involved in the human voice analysis model are:

[0078] Real-time data collection: Use real-time audio collection equipment, such as microphone arrays, to collect specific human audio and environmental sounds in real time.

[0079] Real-time signal processing: Use efficient signal processing algorithms, such as the Fast Fourier Transform (FFT) algorithm, to process audio data in real time.

[0080] Real-time feature extraction: Use high-performance feature extraction algorithms to extract speech rate and intonation features of audio in real time.

[0081] Real-time analysis and integration: Use real-time data processing and integration algorithms to generate vocal analysis coefficients and environmental analysis coefficients in real time.

[0082] Real-time threshold analysis: Uses real-time data analysis and judgment algorithms to authenticate the authenticity of audio from a specific person in real time.

[0083] The steps to generate vocal analysis coefficients include:

[0084] The specific human audio is divided into L identification intervals. The human voice speed change rate v and the voice intonation change amplitude h of each identification interval are obtained. The human voice speed change rate v and the voice intonation change amplitude h of each identification interval are added and the root is taken. The absolute value of the root is taken and divided by the error correction constant K to obtain the human voice analysis coefficient α. The specific formula of the human voice analysis coefficient α is:

[0085] α= (K>0);

[0086] Among them, K and α are both greater than 0. The larger the value of the vocal analysis coefficient α, the greater the emotional fluctuation of the vocal expression in the specific person's audio, the more the language expression tends to be artificial, the higher the authenticity, and the lower the probability of synthesis. Conversely, the lower the value of the vocal analysis coefficient α, the smaller the emotional fluctuation of the vocal expression in the specific person's audio, the more mechanical the language expression, the lower the authenticity, and the higher the probability of synthesis.

[0087] The steps to generate environmental analysis coefficients include:

[0088] Obtain the ambient sound intensity change rate q and the ambient sound frequency change rate p within the same single identification interval as step S200, perform weighted analysis on the ambient sound intensity change rate q and the ambient sound frequency change rate p within the single identification interval, square the weighted analysis result, and multiply it by the error compensation constant , calculate and obtain the environmental analysis coefficient γ, and the steps for generating the environmental analysis coefficient γ are:

[0089] γ = (m1 q+m2 p) 2 (w>0);

[0090] Among them, the weight parameter m1+m2=2.1325, the weight parameter m1 is greater than the weight parameter m2. When the environmental analysis coefficient γ is larger, the influence of the ambient sound is greater, and the influence on the audio identification processing will also be greater; conversely, the smaller the environmental analysis coefficient γ is, the smaller the influence of the ambient sound is, and the influence on the audio identification processing will also be smaller.

[0091] The step of integrating the vocal analysis coefficient and the environmental analysis coefficient includes:

[0092] The vocal analysis coefficient α and the environmental analysis coefficient γ in a single interval are subjected to parameter analysis. The characterization parameter β is obtained by dividing the vocal analysis coefficient α by the cube root of the environmental analysis coefficient γ and adding the vocal analysis coefficient α multiplied by the inverse of the environmental analysis coefficient γ. The formula for obtaining the characterization parameter β is:

[0093] β= +α (β>0);

[0094] The larger the representation parameter β is, the greater the calculation proportion of the human voice analysis coefficient α is, the corresponding authenticity is higher, and the environmental influence effect in this case is weaker than the human voice influence effect in the specific human audio; conversely, the smaller the representation parameter β is, the smaller the calculation proportion of the human voice analysis coefficient α is, the corresponding authenticity is lower, and the environmental influence effect in this case is stronger than the human voice influence effect in the specific human audio.

[0095] The process of performing threshold analysis on the characterization parameters and performing characterization feature marking is as follows:

[0096] Set characterization parameter comparison thresholds SH1 and SH2, where the characterization parameter comparison threshold SH1 is greater than the characterization parameter comparison threshold SH2, substitute the characterization parameter β into the characterization parameter comparison thresholds SH1 and SH2 for comparison, and if the characterization parameter β is less than the characterization parameter comparison threshold SH2, mark the audio of the specific person as a false characterization parameter;

[0097] Substitute the characterization parameter β into the characterization parameter comparison thresholds SH1 and SH2 for comparison. If the characterization parameter β is less than the characterization parameter comparison threshold SH1 and the characterization parameter β is greater than the characterization parameter comparison threshold SH2, mark the audio of the specific person as suspicious.

[0098] The characterization parameter β is substituted into the characterization parameter comparison thresholds SH1 and SH2 for comparison. If the characterization parameter β is greater than the characterization parameter comparison threshold SH1, the audio of the specific person is marked as a true characterization parameter.

[0099] It should be noted that the above involves data matching technology, and the design steps of this data matching technology are as follows:

[0100] Collect a set of known real and fake person-specific audios and extract the corresponding representation parameter β from each audio;

[0101] This set of known audios is analyzed, the mean and standard deviation of the characterization parameter β are calculated, and the characterization parameter comparison thresholds SH1 and SH2 are determined based on the mean and standard deviation.

[0102] In summary, the present invention utilizes a human voice analysis model that utilizes comprehensive multidimensional data (including the rate of change of human voice speed, the amplitude of voice intonation change, the rate of change of ambient sound intensity, the rate of change of ambient sound frequency, etc.), combined with advanced real-time processing technology, to comprehensively analyze the audio of a specific individual. By considering environmental factors such as background noise and the degree of ambient sound variability, the model can more accurately characterize the audio characteristics of a specific individual. Furthermore, a combined analysis method is used to integrate human voice analysis coefficients and environmental analysis coefficients to generate more comprehensive characterization parameters, thereby more accurately identifying and labeling the audio of a specific individual. Compared to traditional audio characterization techniques, this model can provide richer and more accurate audio feature representations, improving identification accuracy and efficiency.

[0103] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.

[0104] It should be understood that the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. A and B can be singular or plural. Furthermore, the character " / " as used herein generally indicates an "or" relationship between the associated objects, but it may also indicate an "and / or" relationship. For specific understanding, please refer to the context.

[0105] In this application, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, "at least one of a, b, or c" can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.

[0106] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0107] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0108] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, and may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment as needed.

[0109] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0110] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A method for constructing a representation based on audio of a specific person, characterized in that: The method comprises the following steps: Step S100: Acquire human voice data and environmental data from audio data of a specific person; The human voice data includes the rate of change of human voice speed and the amplitude of change of intonation; The environmental data includes the rate of change of the intensity of the ambient sound and the rate of change of the frequency of the ambient sound; Step S200: constructing a vocal analysis model, analyzing the vocal data and generating a vocal analysis coefficient; The steps to generate vocal analysis coefficients include: Divide the specific human audio into L identification intervals, obtain the human voice speed change rate v and the voice intonation change amplitude h for each identification interval, add the human voice speed change rate v and the voice intonation change amplitude h for each identification interval, take the root of the sum, take the absolute value of the root-sum result, and divide it by the error correction constant K to obtain the human voice analysis coefficient α; Step S300: performing combined analysis on the human voice data and the environmental data to generate an environmental analysis coefficient; The steps to generate environmental analysis coefficients include: Obtain the ambient sound intensity change rate q and the ambient sound frequency change rate p within the same single identification interval as step S200, perform weighted analysis on the ambient sound intensity change rate q and the ambient sound frequency change rate p within the single identification interval, square the weighted analysis result, and multiply it by the error compensation constant , calculate and obtain the environmental analysis coefficient γ; Step S400: integrating the human voice analysis coefficient and the environmental analysis coefficient to generate characterization parameters for the specific person's audio data; Step S500: Perform threshold analysis on the characterization parameters, and characterize and mark the audio of the specific person based on the analysis results.

2. The method for constructing a representation based on a specific person's audio according to claim 1, characterized in that: The steps of constructing the human voice analysis model include: Use appropriate microphones or recording equipment to collect audio data of specific people and environmental sound data; Preprocess audio data and remove background noise using audio signal processing algorithms and tools; Extracting speech rate change rate: The speech rate change rate is obtained by calculating the ratio between the audio duration and the speech rate; Extracting the amplitude of intonation changes: Using a fundamental frequency estimation algorithm based on the autocorrelation function method to calculate the fundamental frequency changes in the audio; Extracting the rate of change of ambient sound intensity: Using audio signal processing methods to estimate the rate of change of ambient sound intensity, the audio is framed and then the energy change rate of each frame is calculated. Extracting the rate of change of ambient audio frequency: Using spectrum analysis algorithm to estimate the rate of change of ambient audio frequency; Normalize the extracted features of the human voice speech rate change rate and intonation change amplitude to generate the human voice analysis coefficient; Normalize and weight the extracted features of the ambient sound intensity change rate and the ambient sound frequency change rate to generate an environmental analysis coefficient; Perform combined analysis on the vocal analysis coefficient and the environmental analysis coefficient, and use weighted average or logical operation method to integrate the vocal and environmental analysis results to generate representation parameters; Set a threshold or threshold value, compare the characterization parameter with the threshold value, and determine the authenticity of the audio of a specific person.

3. The method for constructing a representation based on a specific person's audio according to claim 2, characterized in that: The step of integrating the vocal analysis coefficient and the environmental analysis coefficient includes: Perform parameter analysis on the vocal analysis coefficient α and the environmental analysis coefficient γ within a single interval, and obtain the representation parameter β by dividing the vocal analysis coefficient α by the cube root of the environmental analysis coefficient γ and adding the vocal analysis coefficient α multiplied by the inverse of the environmental analysis coefficient γ; The larger the representation parameter β is, the greater the calculation proportion of the human voice analysis coefficient α is, the corresponding authenticity is higher, and the environmental influence effect in this case is weaker than the human voice influence effect in the specific human audio; conversely, the smaller the representation parameter β is, the smaller the calculation proportion of the human voice analysis coefficient α is, the corresponding authenticity is lower, and the environmental influence effect in this case is stronger than the human voice influence effect in the specific human audio.

4. The method for constructing a representation based on a specific person's audio according to claim 3, characterized in that: The process of performing threshold analysis on the characterization parameters and performing characterization feature marking is as follows: Set characterization parameter comparison thresholds SH1 and SH2, where the characterization parameter comparison threshold SH1 is greater than the characterization parameter comparison threshold SH2, substitute the characterization parameter β into the characterization parameter comparison thresholds SH1 and SH2 for comparison, and if the characterization parameter β is less than the characterization parameter comparison threshold SH2, mark the audio of the specific person as a false characterization parameter; Substitute the characterization parameter β into the characterization parameter comparison thresholds SH1 and SH2 for comparison. If the characterization parameter β is less than the characterization parameter comparison threshold SH1 and the characterization parameter β is greater than the characterization parameter comparison threshold SH2, mark the audio of the specific person as suspicious. The characterization parameter β is substituted into the characterization parameter comparison thresholds SH1 and SH2 for comparison. If the characterization parameter β is greater than the characterization parameter comparison threshold SH1, the audio of the specific person is marked as a true characterization parameter.

Citation Information

Patent Citations

  • A method and apparatus for identity verification based on spectrograms and phoneme retrieval

    CN107680601B

  • Voice detection method and device

    CN110931020A

  • Distinguishing user speech from background speech in speech-dense environments

    US20180033454A1