System and method for distinguishing original speech from synthetic speech in IoT environment
By extracting features of user voice and environmental factors in the Internet of Things (IoT) environment, generating and sorting scores, and using dynamic thresholds to distinguish between raw and synthesized speech, the accuracy and security issues of speech recognition in the IoT environment are solved, and the reliability of speech recognition is improved.
Patent Information
- Application Number
- CN202480030015.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-23
- Filing Date
- 2024-04-26
- Publication Date
- 2025-11-28
AI Technical Summary
Existing technologies and methods struggle to distinguish between raw and synthesized speech in Internet of Things (IoT) environments, particularly in smart home environments.
By extracting multiple features from user speech and environmental factors, scores are generated and ranked based on the type of user speech and environmental factors. The ranked scores are then compared with dynamic thresholds to determine whether the user speech is original or synthesized.
It improves the accuracy of distinguishing between raw and synthesized speech in IoT environments, reduces recognition errors caused by environmental noise and changes in user distance, and enhances the reliability and security of user speech recognition.
Smart Images

Figure CN121039734A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates generally to speech detection, and more particularly, to a system and method for distinguishing between original speech and synthetic speech in an Internet of Things (IoT) environment. BACKGROUND
[0002] The advent of Internet of Things (IoT) technology has led to the development of a smart home environment that is controllable through voice commands. Although such a smart home environment provides a user with the convenience of controlling devices using their voice, the related IoT environment can face voice-based spoofing attacks and / or voice mimicry that can deceive the user into divulging personal information and / or gaining access to devices present in the smart home environment. For example, in a voice-based spoofing attack, an attacker can use mimicry and / or recorded playback techniques to deceive a user. That is, these voice-based spoofing attacks can involve an attacker mimicking a user’s voice and / or playing a recording of a user’s voice in order to gain access to sensitive data and / or resources, which can result in serious privacy breaches and / or data loss and potential security risks to users of these systems. As another example, a voice mimicry artist can steal personal data by impersonating a user’s voice and thus make it difficult for a voice assistant (VA) to distinguish between a user’s voice and a synthetic or mimicked voice without another identifying input (e.g., an image of the user, a fingerprint of the user, etc.).
[0003] Accordingly, in order to protect users from potential security risks, there can be a need for a system and / or method that can distinguish between original user speech and synthetic speech in an IoT environment.
[0004] There can be a variety of measures that can be used to distinguish between original speech and synthetic speech, which can include, but are not limited to, integrating biometrics (e.g., facial recognition, fingerprint scanning, etc.). However, due to their complexity and associated costs (e.g., additional sensors, degraded user experience, etc.), it can be difficult to employ such measures.
[0005] Further, for example, in scenarios where a user can have difficulty enunciating due to physiological reasons, there can be instances where a VA can not be able to recognize the user’s speech. In order to mitigate this issue, a sensor can be fixed to the user’s throat. However, this approach can result in usability issues and can compromise accuracy in obtaining information from the sound source. As another example, the presence of environmental acoustic artifacts can interfere with the ability to extract features for user recognition. As another example, the distance of a user from a device (e.g., a VA, a smart speaker, etc.) can affect the accuracy of user recognition.
[0006] Accordingly, there is a need for further improvements in speech detection techniques, as the demand for smart home environments can be constrained by privacy and security risks of IoT environments. Improvements are presented herein. These improvements can also be applicable to other voice biometric, security identification, and identity recognition techniques. SUMMARY
[0007] According to an aspect of the disclosure, a method of controlling an electronic device for distinguishing between original speech and synthetic speech in an Internet of Things (IoT) environment includes obtaining a user speech and environmental factors associated with a user-initiated request; extracting a plurality of features from the user speech and the environmental factors; obtaining scores using the plurality of features and ranking the scores based on types of the user speech and the environmental factors; and determining the user speech as the original speech or the synthetic speech by comparing the ranked scores with a dynamic threshold.
[0008] The operation of determining the user speech as the original speech of the synthetic speech can include determining the user speech as the original speech based on the ranked scores being higher than the dynamic threshold, and determining the user speech as the synthetic speech based on the ranked scores being lower than the dynamic threshold.
[0009] The method can include performing re-verification of the user speech based on the ranked scores being equal to the dynamic threshold. The operation of performing the re-verification can include searching for phrases of spatial features and temporal features similar to the spatial features and the temporal features of the user speech and the environmental factors in a combined database, selecting a phrase for least change in environment from the searched phrases, simulating environmental factors of an environment of the selected phrase using a plurality of IoT devices, prompting a user to speak the selected phrase to re-verify the user speech, counting a number of times that the re-verification of the user speech is performed, and stopping the re-verification based on the number of times reaching a predefined count value.
[0010] The method can include processing the user-initiated request, wherein the user-initiated request includes the user speech and the environmental factors, extracting the plurality of features from the processed user speech and the environmental factors, mapping the plurality of features to features stored in a speaker database, identifying a user who initiated the user-initiated request based on the mapping, separating each channel of the user speech and the environmental factors, and generating a first output and a second output, wherein the first output includes a first channel of the user speech and the second output includes a combination of each channel of the environmental factors.
[0011] The plurality of features can include spatial features and temporal features of the user speech and the environmental factors. In some embodiments, the operation of extracting the plurality of features can include pre-processing the user speech and the environmental factors separately, wherein the pre-processing can include performing normalization, pre-emphasis, and frame blocking; extracting the plurality of features from the pre-processed user speech and the pre-processed environmental factors based on at least one of frequency, energy, zero-crossing rate, or mel-frequency cepstral coefficients (MFCCs); and separating the plurality of features of the user speech and the environmental factors by performing feature separation and dimensionality reduction.
[0012] The operation of extracting the plurality of features from the pre-processed user speech and the pre-processed environmental factors can include performing a continuous wavelet transform on the pre-processed user speech and the pre-processed environmental factors and generating a scalogram that visualizes the transform; extracting the plurality of features, wherein the plurality of features include at least one of periodic variations, non-periodic variations, or temporal variations; and performing separation of the plurality of features separately.
[0013] The operation of obtaining the score can include pre-processing the plurality of features; searching for one or more attributes of each feature of the plurality of features; determining an upper bound and a lower bound of a specification of each feature of the plurality of features; determining a first type of the user speech, wherein the first type includes at least one of regular speech or irregular speech; determining a second type of the environmental factors based on at least one of past patterns or user history stored in a combination database, wherein the second type includes at least one of known environmental factors or unknown environmental factors; selecting a kernel; extracting the plurality of features of the user speech and the environmental factors based on the kernel; and computing the score using the optimized plurality of features, and ranking the score based on the first type of the user speech and the second type of the environmental factors, wherein the optimized plurality of features can include a portion of the plurality of features optimized using a regression function.
[0014] The operation of selecting the kernel includes selecting the kernel from a plurality of kernels by comparing a combined feature vector of each kernel of the plurality of kernels using a cost function and selecting the kernel having a minimum distance based on the cost function.
[0015] The plurality of features of the user speech can include at least one of spatial features of the user speech or temporal features of the user speech. The spatial features of the user speech can include at least one of a fundamental frequency, a formant frequency, a speech variability, an amplitude, or a fall amplitude. The temporal features of the user speech can include at least one of a pause duration, a maximum pause duration, a minimum pause duration, or a zero-crossing rate.
[0016] The plurality of characteristics of the environmental factors can include at least one of a spatial characteristic of the environmental factors or a temporal characteristic of the environmental factors. The spatial characteristic of the environmental factors can include at least one of a spectral band energy, a spectral flux, a spectral maximum, a rise amplitude, or a fall amplitude. The temporal characteristic of the environmental factors can include at least one of a pause duration, a maximum pause duration, a minimum pause duration, periodicity, or aperiodicity.
[0017] The operation of determining the user speech as the original speech or the synthesized speech can include: performing a threshold verification on the ranked scores; determining a dynamic threshold based on the spatial and temporal characteristics of the user speech and the environmental factors stored in the combined database; re-performing the ranking based on the dynamic threshold; determining the user speech as the original speech based on the ranked scores being above the dynamic threshold; and determining the user speech as the synthesized speech based on the ranked scores being below the dynamic threshold.
[0018] According to an aspect of the disclosure, an electronic device for distinguishing original speech from synthesized speech in an IoT environment includes a memory storing instructions and one or more processors operatively coupled to the memory. The one or more processors are configured to execute the instructions to obtain user speech and environmental factors associated with a user-initiated request, extract a plurality of characteristics from the user speech and the environmental factors, obtain scores using the plurality of characteristics and rank the scores based on types of the user speech and the environmental factors, and determine the user speech as the original speech or the synthesized speech by comparing the ranked scores with a dynamic threshold.
[0019] The one or more processors are further configured to execute further instructions to determine the user speech as the original speech based on the ranked scores being above the dynamic threshold, and determine the user speech as the synthesized speech based on the ranked scores being below the dynamic threshold.
[0020] The one or more processors are further configured to execute further instructions to perform re-verification of the user speech based on the ranked scores being equal to the dynamic threshold. In some embodiments, the operation of performing the re-verification of the user speech can include searching for phrases of the spatial and temporal characteristics similar to the user speech and the spatial and temporal characteristics of the environmental factors in the combined database, selecting a phrase for least change in the environment from the searched phrases, simulating the environmental factors of the environment of the selected phrase using a plurality of IoT devices, prompting the user to speak the selected phrase to re-verify the user speech, counting a number of times the re-verification of the user speech is performed, and stopping the re-verification based on the number of times reaching a predefined count value.
[0021] The one or more processors are further configured to execute further instructions to: process the user-initiated request, wherein the user-initiated request comprises user speech and environmental factors; extract a plurality of features from the processed user speech and the processed environmental factors; map the plurality of features to features stored in a speaker database; identify, based on the mapping, a user that initiated the user-initiated request; separate each channel of the user speech and the environmental factors; and generate a first output and a second output, wherein the first output comprises a first channel of the user speech and the second output comprises a combination of each channel of the environmental factors.
[0022] The plurality of features can comprise spatial features and temporal features of the user speech and the environmental factors. In some embodiments, the one or more processors are further configured to execute further instructions to: pre-process the user speech and the environmental factors separately, wherein the pre-processing can comprise performing normalization, pre-emphasis, and frame blocking; extract the plurality of features from the pre-processed user speech and the pre-processed environmental factors based on at least one of frequency, energy, zero-crossing rate, or MFCC; and separate the plurality of features of the user speech and the environmental factors by performing feature separation and dimensionality reduction.
[0023] The one or more processors are further configured to execute further instructions to: perform a continuous wavelet transform on the pre-processed user speech and the pre-processed environmental factors and generate a scalogram that visualizes the transform; extract a plurality of features, wherein the plurality of features comprise at least one of periodic variations, non-periodic variations, or temporal variations; and perform separation of the plurality of features separately.
[0024] The one or more processors are further configured to execute further instructions to: pre-process the plurality of features; search for one or more attributes of each feature in the plurality of features; determine an upper bound and a lower bound of a specification of each feature in the plurality of features; determine a first type of the user speech, wherein the first type comprises at least one of regular speech or irregular speech; determine a second type of the environmental factors based on at least one of past patterns stored in a combination database or a user history, wherein the second type comprises at least one of known environmental factors or unknown environmental factors; select a kernel; extract the plurality of features of the user speech and the environmental factors based on the kernel; and compute a score using the optimized plurality of features, wherein the optimized plurality of features comprise a portion of the plurality of features optimized using a regression function, and rank the score based on the first type of the user speech and the second type of the environmental factors.
[0025] The one or more processors are further configured to execute further instructions to: select the kernel from a plurality of kernels by comparing a combined feature vector of each kernel in the plurality of kernels using a cost function and selecting the kernel with a minimum distance based on the cost function.
[0026] The one or more processors are further configured to execute further instructions to perform the following operations: performing a threshold verification on the ranked scores; determining a dynamic threshold based on the spatial and temporal features of the user speech and environmental factors stored in the combination database; re-performing the ranking based on the dynamic threshold; determining that the user speech is the original speech based on the ranked scores being above the dynamic threshold; and determining that the user speech is the synthetic speech based on the ranked scores being below the dynamic threshold.
[0027] Additional aspects can be partly set forth in the description that follows and might be apparent from the description, and / or can be learnt by practice of the presented embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0028] The above and other aspects, features, and advantages of certain embodiments of the disclosure can be more apparent from the following description taken in conjunction with the accompanying drawings, in which: Figure 1 depicts a flow diagram illustrating a method for distinguishing original speech from synthetic speech in an Internet of Things (IoT) environment, in accordance with one or more embodiments; Figure 2 depicts a block diagram of a system performing a method for distinguishing original speech from synthetic speech in an IoT environment, in accordance with one or more embodiments; Figure 3 depicts a block diagram of an acoustic effect separation module, in accordance with one or more embodiments; Figure 4 depicts a flow diagram illustrating a method of an acoustic effect separation module, in accordance with one or more embodiments; Figure 5 depicts a block diagram of a feature extraction module, in accordance with one or more embodiments; Figure 6 depicts a flow diagram illustrating a method of extracting spatial features of a user speech and environmental factors, in accordance with one or more embodiments; Figure 7 depicts a flow diagram illustrating a method of extracting temporal features of a user speech and environmental factors, in accordance with one or more embodiments; Figure 8 depicts a block diagram of a score determination module, in accordance with one or more embodiments; Figure 9 depicts a flow diagram illustrating a method of generating scores by a score determination module, in accordance with one or more embodiments; Figure 10 depicts a block diagram of a decision module, in accordance with one or more embodiments; Figure 11 depicts a flow diagram illustrating a method of determining whether a user speech is original speech or synthetic speech by a decision module, in accordance with one or more embodiments; Figure 12 a block diagram depicting a simulation module, in accordance with one or more embodiments; Figure 13 a block diagram depicting a sound effect generation sub-module, in accordance with one or more embodiments; Figure 14 a flow diagram illustrating a method of re-verification of a user voice performed by a simulation module, in accordance with one or more embodiments; Figure 15A a first use case of distinguishing between original voice and synthetic voice in an IoT environment, in accordance with one or more embodiments; and Figure 15B a second use case of distinguishing between original voice and synthetic voice in an IoT environment, in accordance with one or more embodiments, in accordance with an embodiment; and Figure 16 a flow diagram illustrating a method of distinguishing between original voice and synthetic voice in an IoT environment, in accordance with one or more embodiments. DETAILED DESCRIPTION
[0029] In the following description, numerous specific details are set forth to provide a thorough understanding of the present disclosure. However, it will be apparent to one skilled in the art that the specific details need not be used to practice the present disclosure. In other instances, well-known methods, procedures, components, and circuits have not been described in detail as not to unnecessarily obscure aspects of the present disclosure. Additionally, it is to be noted that the description and drawings merely preferred embodiments of the present disclosure and that it is not to be limited to the exact construction and arrangements described. Additionally, it is to be noted that the phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting.
[0030] Further, in the present specification, reference to “one embodiment”, “one or more embodiments”, or “an embodiment” means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the disclosure. The appearances of the phrase “in one embodiment” in various places in the specification are not necessarily referring to the same embodiment, nor are separate or alternative embodiments mutually exclusive of other embodiments. Also, the term “one” as used herein does not restrict the quantity to one, but rather to at least one. Additionally, various features are described which can be described by some embodiments but not by other embodiments. Similarly, various requirements are described which can be requirements for some embodiments but not for other embodiments.
[0031] In the description of the figures, like-referenced characters can be used to designate like elements throughout the figures. It will be understood that a singular form of a noun intended to include one or more of the things, unless the relevant context clearly dictates otherwise. As used herein, each of the terms "and / or" where used, means "and / or", one or more of the items in the groups can occur. As used herein, each of the terms "formed on" or "formed with" or "formed from" means "formed on" or "formed with" or "formed from" one or more of the items in the groups. As used herein, each of the terms "coupled with" or "coupled to" or "coupled in" means that the element is "coupled with" or "coupled to" or "coupled in" one or more of the items in the groups.
[0032] It will be understood that the particular order in which the blocks are presented in the processes / flow charts is illustrative. Based upon design choices and other factors, the specific order of blocks can be re-arranged and / or some blocks can be omitted. The accompanying claims set forth elements of the various blocks in sample order, and are not meant to be limited to the particular sequence presented in the claims.
[0033] As shown in the figures, embodiments herein can be described and shown in terms of blocks that perform described functions. These blocks (which can be referred to herein as units or modules, or by names such as devices, logic, circuitry, controllers, counters, comparators, generators, converters, etc.) can be physically implemented by analog and / or digital circuitry, including one or more of logic gates, integrated circuits, microprocessors, memory circuits, passive electronic components, active electronic components, optical components, etc.
[0034] In the present disclosure, the singular forms "a," "an," and "the" are intended to include one or more items, and are used interchangeably with the term "one or more." The term "or" is used in the inclusive sense (i.e., "and / or") unless the context clearly indicates otherwise. The term "based on" is used to describe one or more factor to which determination, action, or result is based, which can include one or more factors selected from one or more past values of a subject, current value(s) of the subject, or expected future value(s) of the subject. The term "processor" can refer to a single processor or multiple processors, and the term "processing unit" can refer to a single processing unit or multiple processing units. When a processor is described as performing an operation and the processor is referred to as performing an additional operation, the multiple operations can be performed by the single processor or any one or a combination of the multiple processors.
[0035] It will be understood that the number and / or arrangement of components, modules, etc. depicted in the drawings is provided as an example. In practice, there can be additional components, fewer components, different components, or differently arranged components than those depicted in the drawings. Furthermore, two or more components depicted in the drawings can be implemented within a single component, or a single component can be implemented as multiple, distributed components. Alternatively or additionally, a set of one or more components can be integrated into one or more other components, and / or the set can be implemented as integrated circuits, software and / or a combination of software and circuitry.
[0036] Various embodiments of the present disclosure are described hereinafter, with reference to the drawings.
[0037] Referring to Figure 1 , a flowchart showing a method 100 for distinguishing between original speech and synthetic speech in an Internet of Things (IoT) environment is depicted, in accordance with an embodiment. The ability to distinguish between original speech and synthetic speech in an IoT environment can be increasingly important. For example, as the number of connected devices can be increasing, the likelihood of malicious users using artificial or synthetic speech to manipulate data and gain access to sensitive information can similarly increase. To potentially prevent these threats, speech detection techniques can be able to accurately identify whether a given input is from a true (original) source or artificially created speech.
[0038] Referring to Figure 2 the system described in conjunction with Figure 1 , a method 100 for distinguishing between original speech and synthetic speech in an IoT environment is described. In the flowchart of the method 100, Figure 1The two blocks shown in succession, or sometimes these blocks can be performed in the reverse order, depending on the functionality involved. Any process descriptions or blocks in the flow diagrams should be understood as representing modules, segments, or portions of code which can include one or more executable instructions for implementing the specified logical function or operation (s) in the process and alternative implementations are included within the scope of the example embodiments in which there can be an order other than the shown or discussed order for executing activities or functions depending on the functionality involved, that are not explicitly shown or discussed. Further, representations of the process or block can be understood as representing a decision made by a hardware structure (such as a state machine). The flow diagram begins with operation 102 and proceeds to operation 108.
[0039] At operation 102, user speech and / or environmental factors associated with a user-initiated request can be isolated. In one embodiment, environmental factors can include background noise from sounds from indoor sources such as, but not limited to, water supply, heating / cooling, service facilities, nearby other people speaking, etc. and / or from outdoor sources such as, but not limited to, traffic, buildings, neighborhood activities, weather, etc. In an embodiment, environmental factors can include, but are not limited to, clock ticking, fan noise, traffic noise, etc.
[0040] At operation 104, a plurality of features can be extracted from the user speech and environmental factors. In one embodiment, the plurality of features can include spatial and / or temporal features of the user speech and / or environmental factors. Spatial features can refer to how sound waves propagate in space, which can include features such as, but not limited to, intensity, frequency range, directionality, or focus in the environment where the sound is detected from multiple sources. Temporal features can relate to time-based aspects such as, but not limited to, duration or timing between words in the user-initiated request, which can be used to identify different speakers based on different patterns of the speakers with respect to time within a given area.
[0041] At operation 106, scores can be generated and the generated scores can be ranked. In one embodiment, the scores can be generated using the extracted features of the user voice and the environmental factors, and the scores can be ranked based on the type of user voice and the type of environmental factors. The type of user voice can include, but is not limited to, regular voice and irregular voice based on past patterns or user history stored in the combination database. In an embodiment, the regular voice type can refer to the typical and / or normal voice of the user. Alternatively or additionally, the irregular voice type can refer to voice patterns that exhibit altered voice quality, tone breaks, or other abnormalities that can be attributed to changes in vocal cord behavior. In some embodiments, the irregular voice can be indicative of specific vocal cord disorders such as, but not limited to, nodules, polyps, cysts, etc. Further, the type of environmental factors can include, but is not limited to, known environmental factors and / or unknown environmental factors based on past patterns or user history stored in the combination database. In an embodiment, the known environmental factors can include, but are not limited to, environmental acoustics that can be present in the room at the same time of day on different days. Alternatively or additionally, the unknown environmental factors can include acoustics that can be new and / or different and / or first occurrence.
[0042] At operation 108, the user voice can be determined to be original voice or synthetic voice by comparing the generated (and ranked) scores with a dynamic threshold. In one embodiment, the user voice can be determined to be original voice if the generated (and ranked) scores are higher than the dynamic threshold, and / or the user voice can be determined to be synthetic voice if the generated (and ranked) scores are lower than the dynamic threshold.
[0043] Referring to Figure 2 , a block diagram for distinguishing original voice from synthetic voice in an IoT environment is depicted in accordance with one or more embodiments. The system 200 can include an acoustic separation module 202 that can be configured to separate a user voice and environmental factors associated with a user initiated request, and the acoustic separation module 202 is described with reference to Figure 3 and Figure 4 In one embodiment, the acoustic separation module 202 can separate the user voice and environmental factors to enable the system 200 to distinguish original voice from synthetic voice and understand and interpret the user initiated request, which can result in a relatively more accurate response when compared to related voice detection devices.
[0044] Referring to Figure 3 , a block diagram of the acoustic separation module 202 is depicted in accordance with an embodiment. To perform the separation of the user voice and environmental factors associated with the user initiated request, the acoustic separation module 202 can include a signal processing sub-module 302.
[0045] The signal processing sub-module 302 can be configured to process the user-initiated request, which can include encoding the user-initiated request and separating the user speech and environmental factors associated with the user-initiated request. In an embodiment, the signal processing sub-module 302 can perform a series of operations on the user-initiated request, including but not limited to a Fast Fourier Transform (FFT), a log amplitude spectrum, a mel-scale, a Discrete Cosine Transform (DCT), and the like.
[0046] In an embodiment, the user-initiated request can be and / or can include an audio signal, and the FFT performed by the signal processing sub-module 302 can be used to analyze the audio signal by converting the audio signal from a time domain to a frequency domain. For example, the conversion of the audio signal can allow for the identification of different frequency components that can be present in the audio signal. To apply the FFT to the audio signal, the audio data can be converted from an analog signal to a digital format. For example, an FFT algorithm can be applied to the digital audio data to obtain a frequency analysis of the audio signal. The signal processing sub-module 302 can use at least one of the known FFT algorithms, such as but not limited to the Cooley-Tukey algorithm, the Bluestein algorithm, the Winograd Fourier Transform Algorithm, and the like, to perform the FFT of the audio signal.
[0047] In an embodiment, the log amplitude spectrum can be performed by calculating the log of the amplitude spectrum of the Fourier transform of the audio signal. The log amplitude spectrum can be used in audio signal processing to represent the frequency content of the signal.
[0048] In an embodiment, the mel-scale can refer to a process of mapping the frequency spectrum of an audio signal to a mel-scale, which can be based on the perception of sound by the human ear. That is, the mel-scale can be based on a non-linear transformation of frequency that can reflect the way in which the human ear can perceive different frequencies. In an embodiment, the mel-scale can be performed to provide a relatively more accurate representation of the frequency content of the audio signal.
[0049] The mel-scale can be used in audio signal processing applications and can be combined with other techniques, such as Fourier transforms, to enable effective analysis and manipulation of audio signals.
[0050] In an embodiment, the DCT can be performed to convert a time domain representation of an audio signal to a frequency domain representation. Similar to other types of Fourier transforms, the DCT can provide for the identification of different frequency components present in the audio signal.
[0051] The acoustic effect separation module 202 can also include a speaker verification sub-module 304, which can be configured to extract features from the processed user speech and environmental factors, and can also be configured to map the extracted features to features stored in a speaker database. In an embodiment, the mapping can be used to determine (e.g., identify) the user that initiated the request.
[0052] Further, the acoustic effect separation module 202 can include a separator and chunker sub-module 306 that can be configured to separate each channel of the user speech and the environmental factors and combine each channel of the environmental factors together to generate two outputs, the two outputs including one output for the user speech and another output for the environmental factors. In one embodiment, the acoustic effect separation module 202 can perform encoding of the received user speech and the environmental factors, separate each channel of the user speech and the environmental factors, and then decode. That is, the acoustic effect separation module 202 can perform meaningful separation of the channels for the user speech and the environmental factors. For example, when a user initiates a request that includes a combination of the user speech and the environmental factors (e.g., environmental noise from a fan, street traffic, or outdoor activities), the acoustic effect separation module 202 can process the received request and can separate each channel of the user speech and each environmental factor using the separator and chunker sub-module 306 and can provide two channels at the output.
[0053] Referring to Figure 4 , a flowchart showing a method 400 of acoustic effect separation module is depicted, in accordance with an embodiment. At operation 402, a user initiated request can be processed. In one embodiment, the user initiated request can include user speech and environmental factors.
[0054] At operation 404, features can be extracted from the processed user speech and the environmental factors and the extracted features can be mapped with features stored in a speaker database to determine the user who initiated the request. At operation 406, each channel of the user speech and the environmental factors can be separated.
[0055] At operation 408, each channel of the environmental factors can be combined together to generate two outputs, the two outputs including one output for the user speech and another output for the environmental factors.
[0056] Returning to Figure 2 , the system 200 can further include a feature extraction module 204. The feature extraction module 204 can be configured to extract a plurality of features from the user speech and the environmental factors, as described with reference to Figure 5 , Figure 6 and Figure 7 . In one embodiment, the plurality of features can include spatial and temporal features of the user speech and the environmental factors.
[0057] Referring to Figure 5 , a block diagram of the feature extraction module 204 is depicted, in accordance with one or more embodiments.
[0058] The feature extraction module 204 can include a speaker speech analysis sub-module 502 and an environmental factors analysis sub-module 504 that can be configured to perform extraction of spatial and temporal features of the user speech and the environmental factors, respectively.
[0059] In an embodiment, the speaker voice analysis sub-module 502 can perform pre-processing, feature extraction, and feature separation of the user voice to extract spatial features of the user voice. The pre-processing can include, but is not limited to, performing normalization, pre-emphasis, and frame blocking to potentially improve the quality of the user voice and facilitate further processing. The pre-processed user voice can then be subjected to feature extraction based on at least one of frequency, energy, zero-crossing rate, Mel-frequency cepstral coefficients (MFCCs), and the like. In an embodiment, the MFCCs can be extracted by performing operations that can include, but are not limited to, performing an FFT, applying a mel-scale filter, generating a log of the filter output, and subsequently applying a DCT and deriving derivatives from the resulting coefficients.
[0060] In an embodiment, the extracted features can be separated by performing feature separation and dimensionality reduction.
[0061] The speaker voice analysis sub-module 502 can also perform pre-processing, continuous wavelet transform, feature extraction, and feature separation of the extracted features of the user voice to extract temporal features of the user voice. The pre-processing can include, but is not limited to, performing normalization, pre-emphasis, and frame blocking to potentially improve the quality of the user voice and facilitate further processing. The continuous wavelet transform can be performed on the pre-processed user voice, and a scalogram can be generated to visualize the transform.
[0062] In an embodiment, the continuous wavelet transform can be performed in an operation that can include, but is not limited to, pre-processing the user voice using a Morlet continuous wavelet transform followed by normalization. The resulting output can be fed to a feature learning layer that can receive input from layers such as, but not limited to, a convolutional layer, a max-pooling layer, and a rectified linear unit (ReLU) function. The output from the feature learning layer can be provided to a classification layer that can also receive input from layers such as, but not limited to, a soft-max layer, a flatten layer, and a fully connected layer, and can display the output.
[0063] In an embodiment, features (e.g., periodic, aperiodic, temporal variations) can be extracted and separated from the user voice.
[0064] In one embodiment, the environmental factors analysis submodule 504 can preprocess, extract features, and separate features from environmental factors to extract spatial features of the environmental factors. Preprocessing may include, but is not limited to, performing normalization, pre-emphasis, and frame segmentation to potentially improve the quality of the environmental factors and facilitate further processing. Feature extraction can be performed on the preprocessed environmental factors based on at least one of frequency, energy, zero-crossing rate, MFCC, etc. In one embodiment, MFCC can be extracted by performing operations, including but not limited to performing FFT, applying Mel-scale filtering, and generating the logarithm of the filter output, and subsequently applying DCT and deriving derivatives from the obtained coefficients.
[0065] In this embodiment, the extracted features can be separated by performing feature separation and / or dimensionality reduction.
[0066] The environmental factors analysis submodule 504 can also perform preprocessing, continuous wavelet transform, feature extraction, and separation of features of the extracted environmental factors. Preprocessing may include performing normalization, pre-emphasis, and frame segmentation to potentially improve the quality of the user's speech and facilitate further processing. Continuous wavelet transform can be performed on the preprocessed environmental factors, and scale maps can be generated to visualize the transform.
[0067] In one embodiment, a continuous wavelet transform can be performed by executing operations that may include preprocessing environmental factors using the Morlet continuous wavelet transform followed by normalization. The resulting output can be fed into a feature learning layer, which may receive input from layers such as, but not limited to, convolutional layers, max-pooling layers, and ReLU functions. The output from the feature learning layer can be provided to a classification layer, which may also receive input from layers such as, but not limited to, soft-max layers, flattening layers, fully connected layers, etc., and the output can be displayed.
[0068] In the embodiments, features (e.g., periodicity, non-periodicity, time variation) can be extracted and separated from environmental factors.
[0069] Reference Figure 6 According to an embodiment, a flowchart illustrating a method 600 for extracting spatial features of user speech and environmental factors is depicted. In operation 602, the user speech and environmental factors may be processed separately. In one embodiment, preprocessing may include performing at least one of normalization, pre-emphasis, or frame segmentation.
[0070] In operation 604, features can be extracted from preprocessed user speech and environmental factors based on at least one of frequency, energy, zero-crossing rate, MFCC, etc. In operation 606, features of user speech and environmental factors can be separated by performing feature separation and dimensionality reduction.
[0071] Reference Figure 7According to an embodiment, a flowchart illustrating a method 700 for extracting temporal features of user speech and environmental factors is depicted. In operation 702, the user speech and environmental factors may be processed separately. In one embodiment, preprocessing may include performing at least one of normalization, pre-emphasis, frame segmentation, etc.
[0072] In operation 704, a continuous wavelet transform can be performed on the preprocessed user speech and environmental factors, and a scaling plot can be generated to visualize the transform. In operation 706, features of the user speech and environmental factors can be extracted and separated. In one embodiment, the features may include, but are not limited to, periodic, non-periodic, or time-varying characteristics.
[0073] return Figure 2 The system 200 may also include a score determination module 206. The score determination module 206 can be configured to generate scores for multiple features of the extracted user speech and environmental factors, and to sort the scores based on the type of user speech and environmental factors, as shown in the reference... Figure 8 and Figure 9 As stated above.
[0074] Reference Figure 8 A block diagram of a fraction determination module 206 is depicted according to one or more embodiments.
[0075] The score determination module 206 may include a preprocessing submodule 802, which can be configured to preprocess multiple features of user speech and environmental factors extracted by the feature extraction module 204.
[0076] The score determination module 206 may further include an attribute search submodule 804, which can be configured to perform a search for one or more attributes of each of the multiple features. The score determination module 206 may also include a peak-based clustering submodule 806 and a combination submodule 808. The peak-based clustering submodule 806 can be configured to determine an upper and / or lower specification limit for each of the multiple features. The combination submodule 808 can be configured to perform multiple functions, including but not limited to: determining the type of user speech (which may include, but is not limited to, regular or irregular speech) based on past patterns and / or user history stored in a combination database; determining the type of environmental factors (which may include, but is not limited to, known or unknown environmental factors); selecting an appropriate kernel; and extracting multiple features of user speech and environmental factors for the selected kernel.
[0077] In one embodiment, an appropriate kernel can be selected from multiple kernels by comparing the combined feature vector of each kernel with a cost function and selecting the kernel with the shortest distance (e.g., minimum distance) based on the cost function.
[0078] Multiple features of user speech may include spatial features (such as, but not limited to, fundamental frequency, formant frequency, speech variability, amplitude, drop amplitude, etc.). Multiple features of user speech may include temporal features (such as, but not limited to, pause duration, maximum pause duration, minimum pause duration, zero crossover rate, etc.).
[0079] Multiple characteristics of environmental factors may include spatial characteristics (such as, but not limited to, spectral band energy, spectral flux, spectral maximum value, rise rate, fall rate, etc.). Multiple characteristics of environmental factors may include temporal characteristics (such as pause duration, maximum pause duration, minimum pause duration, periodicity, non-periodicity, etc.).
[0080] The score determination module 206 may include an optimization submodule 810, which can be configured to calculate a score using multiple optimized features and to sort the scores based on the determined types of user speech and environmental factors, wherein a regression function is used to optimize the multiple features.
[0081] Reference Figure 9 According to an embodiment, a flowchart illustrating a method 900 for generating scores by a score determination module 206 is depicted.
[0082] In operation 902, multiple features of the extracted user speech and environmental factors can be preprocessed. In operation 904, one or more attributes of each of the multiple features can be searched. In operation 906, an upper and lower specification limit for each of the multiple features can be determined. In operation 908, multiple functions can be performed, which may include at least one of the following: determining the type of user speech (e.g., regular speech, irregular speech) and the type of environmental factors (e.g., known environmental factors, unknown environmental factors) based on past patterns and / or user history stored in a combined database, selecting an appropriate kernel, and extracting multiple features of user speech and environmental factors for the selected kernel.
[0083] In an embodiment, the multiple features may include spatial features. and time characteristics Feature matrix As shown in Equation 1.
[0084]
[0085] Referring to Equation 1, and It can be a positive integer greater than one (1).
[0086] In this embodiment, kernel functions can be used. From the feature matrix Select the kernel.
[0087] For example, kernel functions It can be used to transform nonlinear decision surfaces into linear equations in a higher number of dimensional spaces.
[0088] In an embodiment, a formula that can be expressed as an equation similar to Equation 2 can be used to select a linear kernel.
[0089]
[0090] In an embodiment, a formula that can be represented as an equation similar to Equation 3 can be used to select the polynomial kernel.
[0091]
[0092] Referring to Equation 3, It can represent the degree of a polynomial.
[0093] In an embodiment, this can be done for each kernel (e.g., ) Calculate kernel functions The independent variable that minimizes the value of the equation can be expressed as an equation similar to Equation 4.
[0094]
[0095] Kernel functions can The minimum value of each independent variable and the cost function For comparison, cost function It can be represented as an equation similar to Equation 5.
[0096]
[0097] Referring to Equation 5, It can represent spatial features. It can represent time features, and It can represent the deviation coefficient that can be continuously evaluated based on optimization.
[0098] In operation 910, multiple optimized features can be used to calculate scores, and the calculated scores can be ranked based on the determined types of user speech and environmental factors. In one embodiment, a regression function can be used to optimize multiple features.
[0099] In the embodiment, the regression function is applied to the extracted feature matrix. It can be represented as an equation similar to Equation 6.
[0100]
[0101] Referring to Equation 6, This represents the Y-intercept and can be a constant term. , , and Indicates the slope coefficient. Indicates the spatial characteristics of environmental factors. Indicates the temporal characteristics of environmental factors. Representing the spatial features of user speech, and This represents the temporal characteristics of the user's speech.
[0102] In an embodiment, Spatial features that can represent and be equal to user speech Temporal characteristics of user voice (For example, ),and Spatial characteristics that can represent environmental features and are equivalent to environmental factors Temporal characteristics of environmental factors (For example, ).
[0103] In some embodiments, another user's voice may be presented. In such embodiments, it may be provided by... and This indicates another user's voice.
[0104] In an embodiment, the following can be used: Optimize the regression function , It can represent residuals or model errors, and scores can be calculated using equations similar to Equation 7.
[0105]
[0106] Referring to Equation 7, and It is a positive number (e.g., >0 and >0), Greater than (For example, > ),and and It is used to determine the weights for authentication of user voice.
[0107] In this embodiment, considering the consistency of the temporal characteristics of environmental factors and the dependence of user voice on environmental factors, the maximum weight can be assigned to the temporal characteristics.
[0108]
[0109] Referring to Equation 8, the standard deviation of the user's speech temporal characteristics can be minimized. Standard deviation of time characteristics of environmental factors Method of selection and .
[0110]
[0111] Referring to Equation 9, the standard deviation of the spatial features of the user's speech can be used as the criterion. Standard deviation of spatial characteristics of environmental factors Minimizeable way to choose and .
[0112] return Figure 2 The system 200 may also include a decision module 208. The decision module 208 may be configured to determine whether the user's speech is raw or synthesized speech by comparing the generated score with a dynamic threshold, as shown in reference [reference needed]. Figure 10 and Figure 11 As stated above.
[0113] Reference Figure 10 A block diagram of decision module 208 is depicted according to one or more embodiments.
[0114] The decision module 208 may include a threshold verification submodule 1002, a ranking submodule 1004, and a decision submodule 1006. In one embodiment, the threshold verification submodule 1002 may be configured to perform threshold verification of the ranking received from the score determination module 206. In the event of discrepancies, the threshold verification submodule 1002 may also be configured to determine a dynamic threshold based on spatial and temporal characteristics of user speech and environmental factors stored in a combined database.
[0115] The sorting submodule 1004 can be configured to re-perform sorting based on a determined dynamic threshold, and the decision submodule 1006 can be configured to determine whether the user speech is original speech by comparing the generated score with the dynamic threshold. For example, if the generated score is higher than the dynamic threshold, the user speech can be identified as original speech, and if the generated score is lower than the dynamic threshold, the user speech can be identified as synthesized speech.
[0116] Reference Figure 11According to an embodiment, a flowchart illustrating a method 1100 for determining whether user speech is original speech or synthesized speech by a decision module is depicted. In operation 1102, a threshold verification may be performed on the sorting received from the score determination module 206, and in the event of any discrepancies, a dynamic threshold may be determined based on the spatial and temporal characteristics of user speech and environmental factors stored in a combined database. In operation 1104, the sorting may be re-executed based on the determined dynamic threshold. In operation 1106, the user speech may be determined. If the generated score is higher than the dynamic threshold, the user speech may be determined as original speech, and if the generated score is lower than the dynamic threshold, the user speech may be determined as synthesized speech.
[0117] Decision module 208 can activate simulation module 210 to perform re-verification of the user's voice if the generated score approaches a dynamic threshold, as shown in reference. Figure 12 As stated above.
[0118] Reference Figure 12 According to an embodiment, a block diagram 1200 of the simulation module 210 is depicted. For example... Figure 12 As shown, the simulation module 210 can be activated by the decision module 208 to perform re-verification of the user's voice. In one embodiment, the simulation module 210 may include a search submodule 1202, a selection submodule 1204, a sound effect generation submodule 1206, and a counter submodule 1208.
[0119] Search submodule 1202 can be configured to search a combined database for phrases with spatial and temporal features that may be similar to the spatial and temporal features of the received user speech and environmental factors. In one embodiment, search submodule 1202 can be configured to create a combined cost function for the received user speech and environmental factors, search the combined database, and determine the phrase whose cost function is closest to the created combined cost function. Selection submodule 1204 can be configured to select a phrase from the searched phrases based on the least variation (minimum) for a specific environment. Sound effect generation submodule 1206 can be configured to use multiple IoT devices to simulate environmental factors identical to the environment of the selected phrase and prompt the user to speak the selected phrase to re-verify the user speech, as shown in the reference. Figure 13 to Figure 1 As described in 5. The counter submodule 1208 can be configured to count the number of times user voice re-authentication is performed and stop re-authentication based on reaching a predefined count. In one embodiment, the counter submodule 1208 can be configured to select a re-authentication question, count the re-authentication questions prompted to the user, and stop re-authentication when a predefined (maximum) count is reached.
[0120] Reference Figure 13 A block diagram of a sound effect generation submodule 1206 is depicted according to one or more embodiments.Figure 13 As shown, the sound effects generation submodule 1206 may include a parameterization execution submodule 1302. In one embodiment, the parameterization execution submodule 1302 may be and / or may include a Mel-based parameterization execution submodule. The parameterization execution submodule 1302 may be configured to perform the inverse operation of the natural logarithm function of the selected phrase. For example, the natural logarithm can be used as a transform function to linearize data, which facilitates data analysis, and the inverse operation of the natural logarithm can be used to obtain the original value from the transformed data. The parameterization execution submodule 1302 may also be configured to perform an inverse fast Fourier transform (IFFT) calculation to transform the frequency domain signal back to the corresponding time domain signal. In an embodiment, the parameterization execution submodule 1302 may be configured to perform framing of the output. Framing in parameterization may refer to the process of dividing the output of the IFFT calculation into a series of relatively short overlapping frames.
[0121] The sound effects generation submodule 1206 may further include a density function calculation submodule 1304, which can be configured to calculate a probability density function. The sound effects generation submodule 1206 may also include an error optimization submodule 1306, which can be configured to minimize the difference and / or error between the predicted speech and the user's actual speech. For example, the sound effects generation submodule 1206 may use a loss function, a cost function, etc., to measure the difference between the predicted speech and the user's actual speech, and apply optimization techniques to determine the speech that minimizes the loss function. The sound effects generation submodule 1206 may also include a vocoder submodule 1308, which can be configured to synthesize speech from the user in real time.
[0122] Reference Figure 14 According to an embodiment, a flowchart of a method 1400 for re-verifying user voice by an analog module is depicted. In operation 1402, a search can be performed in a combined database to find phrases with spatial and temporal characteristics that may be similar to the spatial and temporal characteristics of the received user voice and environmental factors. In operation 1404, a phrase can be selected from the searched phrases based on the least variation (minimum) for a specific environment. In operation 1406, multiple IoT devices can be used to simulate the same environmental factors as for each environment of the selected phrase, and the user can be prompted to say the selected phrase to re-verify the user voice. In operation 1408, the number of times the user voice re-verification is performed is counted, and in operation 1408, re-verification can be stopped based on reaching a predefined count. In one embodiment, the predefined count may be a user-defined value.
[0123] Reference Figure 15A According to one or more embodiments, a first use case is described for distinguishing between raw speech and synthesized speech in an IoT environment. Figure 15AAs shown, a user can request a voice assistant (e.g., Bixby) to read unread emails. Upon receiving the request, the voice assistant (e.g., Bixby), which may have already implemented this disclosure, can process user voice and environmental factors, question the user's voice (e.g., it may be unable to determine whether the received voice is actual user voice or synthesized voice), and can perform re-verification by prompting the user to turn off the smart light bulb located on the left side of the bed. In an embodiment, the user may realize that there is no smart light bulb on the left side of the bed and can respond accordingly to the voice assistant (e.g., Bixby). The voice assistant (e.g., Bixby) can determine based on the user's response that the received voice is original voice and can continue reading the unread emails. Optionally, if the user responds by instructing the user to turn off the light and / or performing an action accordingly, the voice assistant (e.g., Bixby) can determine that the received voice is synthesized voice and can refuse the request to read the unread emails.
[0124] Reference Figure 15B According to one or more embodiments, a second use case is described for distinguishing between raw speech and synthesized speech in an IoT environment. Figure 15B As shown, a user can request a voice assistant (VA) to play music based on the environment. Upon receiving the request, the VA, which may have been implemented using this disclosure, can process the user's voice and environmental factors, including ambient sounds such as rain. The VA can select and play music that may be suitable for the current environment (e.g., rain), potentially enhancing the user's overall listening experience.
[0125] Figure 16 A flowchart illustrating a method 1600 for distinguishing between raw speech and synthesized speech in an IoT environment according to one or more embodiments is depicted.
[0126] According to an embodiment, a method 1600 for controlling an electronic device for distinguishing between raw speech and synthesized speech in an IoT environment includes: in operation S1605, a sound effect separation module 202 obtains user speech and environmental factors associated with a user-initiated request; in operation S1610, a feature extraction module 204 extracts multiple features from the user speech and environmental factors; in operation S1615, a score determination module 206 uses the extracted multiple features of the user speech and environmental factors to obtain a score and sorts the scores based on the type of the user speech and environmental factors; and in operation S1620, a decision module 208 determines the user speech as raw speech or synthesized speech by comparing the generated score with a dynamic threshold.
[0127] The method may include: obtaining a user-initiated request. User-initiated requests may include requests for filtering user speech.
[0128] The method may include: obtaining input data, which includes user voice and environmental factors associated with a user-initiated request. The input data may include data that is a combination (or mixture or synthesis) of user voice and environmental factors.
[0129] Environmental factors can be described as environmental information, environmental audio, environmental noise, or ambient ambient audio.
[0130] The acquisition of operation S1615 can be described as separation, division, filtering, or identification.
[0131] The extraction in operation S1610 can be described as acquisition or identification.
[0132] Multiple features can be described as feature information.
[0133] The acquisition of operation S1605 can be described as generation or identification.
[0134] A score can be described as score information. A score can be described as an indication of the degree of a score's result or result information.
[0135] Dynamic thresholds can be described as predetermined values.
[0136] If the generated score is higher than the dynamic threshold, the user's speech can be identified as original speech; if the generated score is lower than the dynamic threshold, the user's speech can be identified as synthesized speech.
[0137] The method may include: if the generated score is greater than a dynamic threshold, then the decision module 208 determines the user's speech as the original speech.
[0138] The method may include: if the generated score is equal to or less than a dynamic threshold, then the decision module 208 determines the user's speech as synthesized speech.
[0139] The method may include: performing a re-verification of the user's voice by the simulation module 1200 based on the generated score being the same as the dynamic threshold.
[0140] The re-verification process may include: the search submodule 1202 searching the combined database for phrases with spatial and temporal features that may be similar to the spatial and temporal features of the received user speech and environmental factors; the selection submodule 1204 selecting the phrase with the least variation (minimum) for a specific environment from the searched phrases; the sound effect generation submodule 1206 using multiple IoT devices to simulate the same environmental factors for each environment of the selected phrase and prompting the user to say the selected phrase to re-verify the user speech; and the counter submodule 1208 counting the number of times the user speech re-verification is performed and stopping the re-verification when a predefined count is reached.
[0141] The method may include: identifying whether the generated score is a dynamic threshold. For example, the method may include: obtaining the difference between the generated score and the dynamic threshold. The method may include: identifying whether the difference is within the threshold range. The method may include: performing a re-verification of the user's speech based on the fact that the difference is included within the threshold range.
[0142] The operation for re-verification can be performed by at least one of the search submodule 1202, the selection submodule 1204, the sound effect generation submodule 1206, or the counter submodule 1208.
[0143] User voice and environmental factors associated with a user-initiated request can be separated by performing operations, which may include: processing the user-initiated request by signal processing submodule 302, the user-initiated request including user voice and environmental factors; extracting features from the processed user voice and environmental factors by speaker verification submodule 304, and mapping the extracted features to features stored in a speaker database to determine the user initiating the request; and separating each channel of user voice and environmental factors by separator and grouping submodule 306 and combining each channel of environmental factors together to generate two outputs, including one output for user voice and another output for environmental factors.
[0144] In an embodiment, the method may include: a first channel for recognizing user voice and a second channel for recognizing environmental factors.
[0145] In this embodiment, environmental factors may include multiple environmental factors. For example, environmental factors may include a first environmental factor and a second environmental factor. The method may include: recognizing a first channel of user speech, a second channel of the first environmental factor, and a third channel of the second environmental factor. Furthermore, the method may include: combining (or synthesizing) the second channel of the first environmental factor with the third channel of the second environmental factor to form a new channel (a fourth channel).
[0146] Multiple features may include spatial and temporal features of user speech and environmental factors. Spatial features of user speech and environmental factors can be extracted separately by performing operations, which may include: preprocessing user speech and environmental factors separately, wherein the preprocessing includes performing normalization, pre-emphasis and frame segmentation; extracting features from the preprocessed user speech and environmental factors based on, but not limited to, frequency, energy, zero crossover rate and MFCC; and separating the features of user speech and environmental factors by performing feature separation and dimensionality reduction.
[0147] Temporal features of user speech and environmental factors can be extracted separately by performing operations, which may include: preprocessing user speech and environmental factors separately, wherein the preprocessing includes performing normalization, pre-emphasis and frame segmentation; performing continuous wavelet transform on the preprocessed user speech and environmental factors and generating scale maps to visualize the transform; and extracting features (such as periodicity, non-periodicity, and temporal variation) and separating the extracted user speech and environmental factors features separately.
[0148] Scores can be generated for multiple features of extracted user speech and environmental factors by performing operations, which may include: preprocessing multiple features of user speech and environmental factors extracted by feature extraction module 204 by preprocessing submodule 802; searching one or more attributes for each of the multiple features by attribute search submodule 804; determining upper and lower bounds of specifications for each of the multiple features by peak-based clustering submodule 806; performing multiple functions by combination submodule 808, including: determining the type of user speech (including regular and non-regular speech) and the type of environmental factors (including known and unknown environmental factors) based on past patterns or user history stored in the combination database, selecting an appropriate kernel, and extracting multiple features of user speech and environmental factors for the selected kernel; and calculating scores using the optimized multiple features by optimization submodule 810, and ranking the scores based on the determined types of user speech and environmental factors, wherein a regression function may be used to optimize the multiple features.
[0149] By comparing the combined feature vector of each kernel with the cost function and selecting the kernel that has the shortest distance to the cost function, an appropriate kernel is chosen from multiple kernels.
[0150] Multiple features of user speech may include spatial features of user speech (such as, but not limited to, fundamental frequency, formant frequency, speech variability, amplitude, and drop amplitude) and temporal features of user speech (including, but not limited to, pause duration, maximum pause duration, minimum pause duration, and zero crossover rate).
[0151] Multiple characteristics of environmental factors may include spatial characteristics of environmental factors (such as, but not limited to, spectral band energy, spectral flux, spectral maximum value, rise rate, and fall rate) and temporal characteristics of environmental factors (including, but not limited to, pause duration, maximum pause duration, minimum pause duration, periodicity, and non-periodicity).
[0152] The determination of user speech may include performing operations such as, but not limited to: performing threshold verification by threshold verification submodule 1002 on the sorting received from score determination module 206, and determining a dynamic threshold based on the spatial and temporal characteristics of user speech and environmental factors stored in the combined database in the presence of any discrepancies; re-performing sorting by sorting submodule 1004 based on the determined dynamic threshold; and determining user speech as original speech by decision submodule 1006 if the generated score is higher than the dynamic threshold, and determining user speech as synthesized speech by decision submodule 1006 if the generated score is lower than the dynamic threshold.
[0153] According to an embodiment, an electronic device for distinguishing between raw speech and synthesized speech in an IoT environment includes: at least one processor configured to: separate (operation 102) user speech and environmental factors associated with a user-initiated request via a sound effect separation module 202; extract (operation 104) multiple features from the user speech and environmental factors via a feature extraction module 204; generate (operation 106) a score using the extracted multiple features of the user speech and environmental factors via a score determination module 206 and sort the score based on the type of the user speech and environmental factors; and determine (operation 108) the user speech as raw speech or synthesized speech by a decision module 208 by comparing the generated score with a dynamic threshold.
[0154] If the generated score is higher than the dynamic threshold, the user's speech can be identified as original speech; if the generated score is lower than the dynamic threshold, the user's speech can be identified as synthesized speech.
[0155] At least one processor may also be configured to perform a re-verification of the user's voice via the simulation module 1200 when the generated score is substantially similar to or equal to a dynamic threshold. The re-verification process may include: the search submodule 1202 searching a combined database for phrases with spatial and temporal features similar to the received user's voice and environmental factors; the selection submodule 1204 selecting the phrase with the least variation for a specific environment from the searched phrases; the sound effect generation submodule 1206 using multiple IoT devices to simulate environmental factors identical to those of the selected phrase for each environment and prompting the user to speak the selected phrase to re-verify the user's voice; and the counter submodule 1208 counting the number of times the user's voice re-verification is performed and stopping the re-verification when a predefined count is reached.
[0156] User voice and environmental factors associated with a user-initiated request can be separated by performing operations, which may include: processing the user-initiated request by signal processing submodule 302, the user-initiated request including user voice and environmental factors; extracting features from the processed user voice and environmental factors by speaker verification submodule 304, and mapping the extracted features to features stored in a speaker database to determine the user initiating the request; and separating each channel of user voice and environmental factors by separator and grouping submodule 306 and combining each channel of environmental factors together to generate two outputs, including one output for user voice and another output for environmental factors.
[0157] It is readily apparent that various aspects of this disclosure provide a system and method for distinguishing between raw and synthesized speech in an IoT environment. Such a system and method may be subject to numerous modifications and variations, all of which are encompassed by the same innovative concept, and all details may be replaced by technically equivalent elements. Therefore, the scope of this disclosure is defined by the appended claims.
Claims
1. A method for controlling an electronic device for distinguishing between raw speech and synthesized speech in an Internet of Things (IoT) environment, the method comprising: Obtain user voice and environmental factors associated with the user-initiated request; Extract multiple features from the user's voice and the environmental factors; The scores are obtained using the multiple features and ranked based on the type of the user's voice and the environmental factors. as well as The user's speech is determined to be either the original speech or the synthesized speech by comparing the sorted score with a dynamic threshold.
2. The method as described in claim 1, wherein, The operation of identifying the user's speech as the original speech of the synthesized speech includes: Based on the ranking score being higher than the dynamic threshold, the user's speech is identified as the original speech; and The user's speech is identified as the synthesized speech if the ranking score is lower than the dynamic threshold.
3. The method of claim 1, further comprising: The user's voice is re-verified based on a ranking score equal to the dynamic threshold. The re-verification process includes: Search the combined database for phrases with spatial and temporal features similar to the spatial and temporal features of the user's speech and the environmental factors; Choose the phrase that least relates to changes in the environment from the search results; Using multiple IoT devices to simulate environmental factors of the selected phrase's context; The user is prompted to say the selected phrase to re-verify the user's voice; The number of times the user's voice re-verification is performed is counted; and Revalidation stops based on the number of times a predefined count value is reached.
4. The method of claim 1, further comprising: Process the user-initiated request, wherein the user-initiated request includes the user's voice and the environmental factors; Extract the multiple features from the processed user voice and environmental factors; Map the multiple features to features stored in the speaker database; Based on the mapping, identify the user who initiated the request. Separate each channel of the user's voice and the environmental factors; and Generate a first output and a second output, wherein the first output includes a first channel of the user's voice, and the second output includes a combination of each channel of the environmental factors.
5. The method of claim 1, wherein, The multiple features include the spatial and temporal features of the user's voice and the environmental factors, and The operations for extracting the multiple features include: The user's voice and the environmental factors are preprocessed respectively, wherein the preprocessing includes normalization, pre-emphasis and frame segmentation; The multiple features are extracted from preprocessed user speech and preprocessed environmental factors based on at least one of frequency, energy, zero-crossing rate, or Mel-frequency cepstral coefficients (MFCCs); and The user's speech and the environmental factors are separated by performing feature separation and dimensionality reduction.
6. The method of claim 5, wherein, The operation of extracting the multiple features from preprocessed user speech and preprocessed environmental factors includes: Perform continuous wavelet transform on preprocessed user speech and preprocessed environmental factors and generate a scale map to visualize the transform; Extract the plurality of features, wherein the plurality of features includes at least one of periodic variation, non-periodic variation, or time variation; and The separation of the multiple features is performed separately.
7. The method of claim 1, wherein, The operations for obtaining the score include: The aforementioned features are preprocessed; Search for one or more attributes of each of the plurality of features; Determine the upper and lower limits of the specifications for each of the plurality of features; Determine a first type of the user's voice, wherein the first type includes at least one of regular voice or unconventional voice; The second type of the environmental factor is determined based on at least one of past patterns or user history stored in the combined database, wherein the second type includes at least one of known environmental factors or unknown environmental factors; Select kernel; Based on the kernel, the multiple features of the user's voice and the environmental factors are extracted; and The score is calculated using multiple optimized features and ranked based on the first type of the user's speech and the second type of the environmental factors, wherein the multiple optimized features include a portion of the multiple features optimized using a regression function.
8. The method of claim 7, wherein, The operation of selecting the kernel includes selecting the kernel from the plurality of kernels in the following manner: The combined feature vectors of each kernel in a plurality of kernels are compared using a cost function, and the kernel with the minimum distance is selected based on the cost function.
9. The method of claim 7, wherein, The plurality of features of the user's speech include at least one of the spatial features or the temporal features of the user's speech. The spatial features of the user's speech include at least one of the following: fundamental frequency, formant frequency, speech variability, amplitude, or descent amplitude. The time characteristics of the user's voice include at least one of the following: pause duration, maximum pause duration, minimum pause duration, or zero crossover rate.
10. The method of claim 7, wherein, The plurality of characteristics of the environmental factors include at least one of the spatial characteristics or the temporal characteristics of the environmental factors. The spatial characteristics of the environmental factors include at least one of spectral band energy, spectral flux, spectral maximum value, and rise or fall amplitude. The time characteristics of the environmental factors include at least one of the following: pause duration, maximum pause duration, minimum pause duration, periodicity, or non-periodicity.
11. The method of claim 1, wherein, The operation of identifying the user's speech as either the original speech or the synthesized speech includes: Perform threshold validation on the sorted scores; The dynamic threshold is determined based on the spatial and temporal characteristics of the user's voice and the environmental factors stored in the combined database; Re-sorting is performed based on the dynamic threshold; Based on a ranking score higher than the dynamic threshold, the user's speech is determined to be the original speech; and Based on the ranking score being lower than the dynamic threshold, the user's speech is determined to be the synthesized speech.
12. An electronic device for distinguishing between raw speech and synthesized speech in an Internet of Things (IoT) environment, the electronic device comprising: Memory, storing instructions; as well as One or more processors are operatively incorporated into the memory, wherein the one or more processors are configured to execute the instructions to perform the following operations: Obtain user voice and environmental factors associated with the user-initiated request; Extract multiple features from the user's voice and the environmental factors; Scores are obtained using the multiple features, and the scores are ranked based on the type of the user's voice and the environmental factors; and The user's speech is determined to be either the original speech or the synthesized speech by comparing the sorted score with a dynamic threshold.
13. The electronic device of claim 12, wherein, The one or more processors are also configured to execute further instructions to perform the following operations: Based on the ranking score being higher than the dynamic threshold, the user's speech is identified as the original speech; and The user's speech is identified as the synthesized speech if the ranking score is lower than the dynamic threshold.
14. The electronic device of claim 12, wherein, The one or more processors are also configured to execute further instructions to perform the following operations: The user's voice is re-verified based on a ranking score equal to the dynamic threshold. The operation of re-verifying the user's voice includes: Search the combined database for phrases with spatial and temporal features similar to the spatial and temporal features of the user's speech and the environmental factors; Choose the phrase that least relates to changes in the environment from the search results; Using multiple IoT devices to simulate environmental factors of the selected phrase's context; The user is prompted to say the selected phrase to re-verify the user's voice; The number of times the user's voice re-verification is performed is counted; and Revalidation stops based on the number of times a predefined count value is reached.
15. The electronic device of claim 12, wherein, The one or more processors are also configured to execute further instructions to perform the following operations: Process the user-initiated request, wherein the user-initiated request includes the user's voice and the environmental factors; The multiple features are extracted from the processed user voice and the processed environmental factors; Map the multiple features to features stored in the speaker database; Based on the mapping, identify the user who initiated the request. Separate each channel of the user's voice and the environmental factors; and Generate a first output and a second output, wherein the first output includes a first channel of the user's voice, and the second output includes a combination of each channel of the environmental factors.