A Forged Speech Intelligent Detection System Based on Multidimensional Fusion

Through the gradient enhancement tree model, the feature combination is optimized based on the deep learning model, and the problem of overfitting and computing overhead is solved, and the efficiency and accuracy of forged speech detection is improved.

CN119479697BActive Publication Date: 2025-07-01HEFEI KEXIN ZHILIAN TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411605592.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-12
Publication Date
2025-07-01
Estimated Expiration
2044-11-12

AI Technical Summary

Technical Problem

Too much feature input will increase the complexity of the model, resulting in overfitting or excessive computational overhead, affecting the efficiency of forged speech detection.

Method used

The importance score of the target feature is determined through the gradient enhancement tree model, and 40% of the features are selected as the features to be selected. A fake speech detection model is established based on the deep learning model. The feature combination is optimized through performance indicators, and the combination of the best performance is gradually screened.

Benefits of technology

The dimension of the feature set is reduced, the efficiency and accuracy of the detection model are improved, the problems of overfitting and excessive computational overhead are avoided, and the reliability of forged speech detection is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119479697B_ABST
    Figure CN119479697B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and specifically discloses a forged speech intelligent detection system based on multi-dimensional fusion. Acquisition module: Determine the importance score of target features, and screen candidate features from the target features according to the importance score; Feature optimization module: Establish a forged speech detection model, and determine the performance score of the forged speech detection model after training; Sort the performance scores, and determine the target performance score and the target set according to the performance score difference; Feature selection module: Determine a new training set, repeat the above steps to obtain the sorting position interval and the target set; Iterate until the performance score meets the preset requirements, determine the standard features and use them to detect forged speech. The present invention can reduce the number of features required for forged speech detection and improve the detection efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to an intelligent forged speech detection system based on multi-dimensional fusion. Background Art

[0002] Forged speech refers to audio that is generated by artificial synthesis or modification and imitates real human voices, usually used to impersonate others' voices or create false speech content. Forged speech technology is widely used in fields such as voice assistants, but it may also bring security risks such as fraud and identity impersonation.

[0003] With the rapid development and wide application of artificial intelligence technology, the technology of generating forged speech is also constantly advancing, gradually approaching real speech in terms of timbre, emotion, and even subtle pronunciation details. This poses a huge challenge to traditional speech detection methods. In some high-risk scenarios, such as identity verification and voice transmission of important information, forged speech may cause serious security problems.

[0004] The detection methods of forged speech usually start from multi-dimensional information such as the spectrum, acoustic features, and biometric features of speech, and construct a forged speech recognition model through multi-dimensional features to identify forged traces and distinguish between real and false speech content. However, too many feature inputs often increase the complexity of the model, resulting in overfitting of the model or excessive computational overhead, thereby affecting the detection efficiency in actual applications. Therefore, how to select effective features and construct an efficient and accurate forged speech detection model is one of the key issues in current research. Summary of the Invention

[0005] The purpose of the present invention is to provide an intelligent forged speech detection system based on multi-dimensional fusion to solve the following technical problems:

[0006] Too many feature inputs often increase the complexity of the model, resulting in overfitting of the model or excessive computational overhead, thereby affecting the detection efficiency in actual applications.

[0007] The purpose of the present invention can be achieved by the following technical solutions:

[0008] An intelligent forged speech detection system based on multi-dimensional fusion, comprising:

[0009] A collection module: preprocess the features for detecting forged speech to obtain target features, determine the importance score P of the target features based on the gradient boosting tree model, sort the target features in ascending order of the importance score, and use the last 40% of the target features in the sorting as candidate features;

[0010] Feature Optimization Module: Establish a forged speech detection model based on a deep learning model, train the forged speech detection model based on a training set, where the training set is obtained based on candidate features, determine the performance metrics of the trained forged speech detection model, and the performance metrics include accuracy A1, precision A2, and recall A3. Determine the performance score A = A1 + A2 + A3;

[0011] Sort the performance scores in descending order. Starting from the first performance score in the sorting, calculate the performance score difference Ai - 1 = Ai - 1 - Ai, where Ai represents the performance score at the i-th position in the sorting. When the performance score difference Ai - 1 ≥ Ays, use the performance scores within the sorting position range [1, i - 1] as the target performance scores, and Ays represents a preset performance score difference threshold;

[0012] Determine the union of the candidate features in the training set corresponding to the target performance scores, and use it as the target set B;

[0013] Feature Selection Module: Determine a training set based on the candidate features in the target set B, determine the performance scores, remove the performance scores less than or equal to Ai - 1, and determine the sorting position range and the target set B;

[0014] Repeat the above steps until the performance scores corresponding to the target set B are all less than the performance score at the end point of the sorting position range in the previous iteration. Determine the training set and performance scores corresponding to the target set B in the previous iteration, use the candidate features in the training set with the maximum performance score as the standard features, and detect forged speech based on the standard features.

[0015] As a further solution of the present invention: In the acquisition module, the process of determining the importance score of the initial features specifically includes:

[0016] Establish a database that stores a preset number of samples. Each sample includes initial features and detection labels, and the detection labels include real and forged;

[0017] Establish a gradient boosting tree model based on the gradient boosting tree framework, train the gradient boosting tree model based on the database, determine the number of times C that a single initial feature serves as a splitting node during the training process, and calculate the importance score P of the initial feature as P = C / Ctot, where Ctot represents the total number of splitting nodes.

[0018] As a further solution of the present invention: In the acquisition module, the process of determining the initial features specifically includes:

[0019] Use the features for detecting forged speech as the first features;

[0020] Standardize the first numerical feature to eliminate the dimension.

[0021] Encode the first non - numerical feature through one - hot encoding.

[0022] Remove the outliers in the first feature through the Isolation Forest algorithm.

[0023] As a further solution of the present invention: in the feature optimization module, performance scores of the same size are in the same position in the sorting.

[0024] As a further solution of the present invention: in the feature optimization module, when there is no performance score difference greater than or equal to the performance score difference threshold, starting from the performance score at the first place in the sorting, calculate the first performance score difference ADYi - 1 = A1 - Ai - 1. When the first performance score difference ADYi - 1≥η*Ays, take [1, i - 1] as the sorting position interval, where η is a preset correction coefficient and η > 1.

[0025] As a further solution of the present invention: in the feature selection module, when there are two or more performance scores that are the same and are the maximum performance scores, take the candidate feature in the training set with the smallest number of candidate features as the standard feature.

[0026] As a further solution of the present invention: in the feature optimization module, the process of determining the training set specifically includes:

[0027] Determine the total number n of the candidate features, and randomly select f candidate features for combination to obtain a training set, where f∈[1, n] and f is an integer.

[0028] As a further solution of the present invention: copy a single training set, and the number of copies is a preset value.

[0029] Advantages of the present invention: First, the target features are sorted in ascending order, and the features ranked in the last 40% are selected as candidate features. Through preliminary screening, the dimension of the feature set is reduced, making subsequent feature optimization more focused on relatively influential features, and improving the efficiency of the detection model; the feature combinations are scored based on performance indicators, and the feature combinations with the optimal performance are gradually screened. By selecting the performance mutation points, it is possible to converge to the feature combinations with higher performance faster, saving computational costs; the union of the candidate features in the training set corresponding to the target performance score is obtained as the target set B, and the optimization process is repeated. The feature sets with performance lower than the performance end point of the previous iteration are screened out, and high-performance features are retained; finally, when the performance scores corresponding to the target set B are all lower than the performance of the previous iteration, the target set B of the previous iteration is taken as the final feature set, where determining the new sorting position interval and the target set is regarded as one iteration; if the number of candidate features in two adjacent iterations is the same, a warning message is sent; through multiple rounds of iteration, it gradually converges to the optimal feature set, ensuring that the selected features can effectively detect forged speech and improve the performance of the model, avoiding the influence of redundant features on the model effect; finally, after multiple rounds of optimization of the candidate features, the feature combinations with high importance and strong correlation with performance indicators are selected to form standard features for the construction of forged speech detection. This method not only considers the importance score of the features, but also integrates multiple performance indicators, making the model effect better and the detection of forged speech more reliable. Description of the Drawings

[0030] The present invention will be further described below with reference to the accompanying drawings.

[0031] Figure 1 It is a schematic flowchart of an intelligent forged speech detection system based on multi-dimensional fusion according to the present invention. Detailed Embodiments

[0032] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0033] Please refer to Figure 1 As shown, the present invention is an intelligent forged speech detection system based on multi-dimensional fusion, including:

[0034] Acquisition module: preprocess the features for detecting forged speech to obtain target features, determine the importance score P of the target features based on the gradient boosting tree model, sort the target features in ascending order of importance score, and take the target features in the last 40% of the sorting as candidate features;

[0035] Feature optimization module: Establish a forged speech detection model based on a deep learning model, train the forged speech detection model based on a training set, where the training set is obtained based on candidate features, determine the performance metrics of the forged speech detection model after training, and the performance metrics include accuracy A1, precision A2, and recall A3. Determine the performance score A = A1 + A2 + A3;

[0036] Sort the performance scores in descending order. Starting from the top performance score in the sorting, calculate the performance score difference Ai-1 = Ai-1 - Ai, where Ai represents the performance score at the i-th position in the sorting. When the performance score difference Ai-1 ≥ Ays, use the performance scores within the sorting position range [1, i - 1] as the target performance scores, and Ays represents a preset performance score difference threshold;

[0037] Determine the union of the candidate features in the training set corresponding to the target performance scores and use it as the target set B;

[0038] Feature selection module: Determine a training set based on the candidate features in the target set B, determine the performance scores, remove the performance scores less than or equal to Ai-1, and determine the sorting position range and the target set B;

[0039] Repeat the above steps until the performance scores corresponding to the target set B are all less than the performance score at the end point of the sorting position range in the previous iteration. Determine the training set and performance scores corresponding to the target set B in the previous iteration. Use the candidate features in the training set corresponding to the maximum performance score as the standard features, and detect the forged speech based on the standard features.

[0040] It should be noted that the target features are sorted in ascending order, and the features ranked in the last 40% are selected as candidate features. Through preliminary screening, the dimension of the feature set is reduced, so that subsequent feature optimization is more concentrated on relatively influential features, improving the efficiency of the detection model; the feature combinations are scored based on performance indicators, and the feature combinations with the optimal performance are gradually screened. By selecting the performance mutation points, it is possible to converge faster to the feature combinations with higher performance, saving computational costs; the union of the candidate features in the training set corresponding to the target performance score is obtained as the target set B, and the optimization process is repeated. The feature sets with performance lower than the performance end point of the previous iteration are screened out, and the high-performance features are retained; finally, when the performance scores corresponding to the target set B are all lower than the performance of the previous iteration, the target set B of the previous iteration is taken as the final feature set, where determining the new sorting position interval and the target set is regarded as one iteration; through multiple rounds of iteration, it gradually converges to the optimal feature set, ensuring that the selected features can not only effectively detect forged speech, but also improve the performance of the model, avoiding the influence of redundant features on the model effect;

[0041] In another preferred embodiment of the present invention, in the acquisition module, the process of determining the importance score of the initial features specifically includes:

[0042] A database is established, and a preset number of samples are stored in the database. Each sample includes initial features and detection labels, and the detection labels include real and forged;

[0043] A gradient boosting tree model is established based on the gradient boosting tree framework, and the gradient boosting tree model is trained based on the database. The number of times C that a single initial feature is used as a splitting node during the training process is determined, and the importance score P of the initial feature is calculated as P = C / Ctot, where Ctot represents the total number of splitting nodes.

[0044] It is worth noting that in the gradient boosting tree model, the feature importance score is usually measured according to the frequency of the splitting nodes. The more splitting nodes a feature has, the greater its impact on the model's prediction, and usually the more important it is; by calculating the ratio of the number of splits to the total number of splits for each feature's importance score, this method can quantify the contribution of each feature to the final classification result (i.e., determining whether the speech is real or forged).

[0045] In another preferred embodiment of the present invention, in the acquisition module, the process of determining the initial features specifically includes:

[0046] The features for detecting forged speech are used as the first features;

[0047] The numerical first features are standardized to eliminate the dimension;

[0048] Encode the non-numerical first feature through one-hot encoding;

[0049] Remove outliers in the first feature through the Isolation Forest algorithm.

[0050] In another preferred embodiment of the present invention, in the feature optimization module, performance scores of the same size are in the same position in the sorting.

[0051] In another preferred embodiment of the present invention, in the feature optimization module, when there is no performance score difference greater than or equal to the performance score difference threshold, starting from the performance score at the first position in the sorting, calculate the first performance score difference ADYi-1 = A1 - Ai-1. When the first performance score difference ADYi-1 ≥ η * Ays, take [1, i - 1] as the sorting position interval, where η is a preset correction coefficient and η > 1.

[0052] It should be noted that in the original feature optimization process, if the difference between the sorted performance scores (i.e., the difference between adjacent performance scores) does not reach the preset difference threshold Ays, a suitable performance score position interval cannot be found according to the standard steps; through the above method, it can be ensured that in the case of small differences in performance scores between features, a relatively large performance change interval can still be found by relaxing the standard, without falling into an infinite loop or being unable to complete feature selection due to strict difference limitations.

[0053] In another preferred embodiment of the present invention, in the feature selection module, when there are two or more performance scores that are the same and are the maximum performance scores, the candidate features in the training set with the smallest number of candidate features are used as the standard features.

[0054] It can be understood that on the premise of ensuring the best performance, a combination with the smallest number of features is selected to make the model more concise, while reducing possible redundant features and improving the efficiency and generalization ability of the model; in the case of consistent performance, selecting a combination with the smallest number of features can reduce the model complexity, avoid overfitting, and improve the running efficiency of the model. This strategy conforms to the "simplicity first" principle of feature selection.

[0055] In another preferred embodiment of the present invention, in the feature optimization module, the process of determining the training set specifically includes:

[0056] Determine the total number n of the candidate features, and randomly select f candidate features for combination to obtain the training set, where f ∈ [1, n] and f is an integer.

[0057] It is worth noting that among the n candidate features, f features are randomly selected for combination to generate different feature subsets, and corresponding training sets are generated using these subsets. By generating multiple different feature combinations, the model is trained and tested on the training sets of different feature combinations to ensure that the best feature combination can be found.

[0058] In another preferred embodiment of the present invention, a single said training set is replicated, and the replication multiple is a preset value.

[0059] It is worth noting that in the training sets of some feature combinations, there may be a situation where the number of samples is insufficient, especially for small-scale feature combinations. By replication, the number of samples in this training set can be increased, thereby balancing the data, ensuring that the model does not suffer from overfitting or underfitting problems due to insufficient data volume, and helping the model to have a deeper learning of specific feature combinations during training, thus improving the accuracy and robustness of the model.

[0060] The above has described in detail an embodiment of the present invention, but the content described is only a preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made according to the scope of the present invention application should still fall within the scope covered by the patent of the present invention.

Claims

1. A forged voice intelligent detection system based on multi-dimensional fusion, characterized in that: include: Acquisition module: pre-processing the features used to detect forged speech to obtain target features, determining the importance score P of the target features based on the gradient boosting tree model, sorting the target features in order of importance score from small to large, and taking the last 40% of the target features in the sorting as candidate features; Feature optimization module: establish a forged voice detection model based on a deep learning model, train the forged voice detection model based on a training set, the training set is obtained based on the selected features, determine the performance indicators of the forged voice detection model after training, the performance indicators include accuracy A1, precision A2 and recall A3, and determine the performance score A=A1+A2+A3; The performance scores are sorted in descending order, starting from the first performance score in the sorting, and the performance score difference Ai-1=Ai-1-Ai is calculated, where Ai represents the performance score of the i-th position in the sorting. When the performance score difference Ai-1≥Ays, the performance score in the sorting position interval [1, i-1] is used as the target performance score, and Ays represents the preset performance score difference threshold; Determine the union of the features to be selected in the training set corresponding to the target performance score, and use it as the target set B; Feature selection module: determine the training set based on the selected features in the target set B, determine the performance score, remove the performance score less than or equal to Ai-1, and determine the sorting position interval and the target set B; Repeat the above steps until the performance scores corresponding to the target set B are all smaller than the performance scores at the end of the sorting position interval in the previous iteration, determine the training set and performance scores corresponding to the target set and B in the previous iteration, use the selected features in the training set corresponding to the maximum performance score as standard features, and detect forged speech based on the standard features.

2. According to claim 1, a forged voice intelligent detection system based on multi-dimensional fusion is characterized in that: In the acquisition module, the process of determining the importance score of the target feature specifically includes: Establishing a database, wherein a preset number of samples are stored in the database, wherein each sample includes a target feature and a detection label, wherein the detection label includes real and fake ones; A gradient boosting tree model is established based on the gradient boosting tree framework, and the gradient boosting tree model is trained based on the database. The number of times C that a single target feature is used as a split node during the training process is determined, and the importance score P = C / Ctot of the target feature is calculated, where Ctot represents the total number of split nodes.

3. According to the multi-dimensional fusion-based forged voice intelligent detection system of claim 1, it is characterized in that: In the acquisition module, the process of determining the target features specifically includes: Using a feature for detecting fake speech as a first feature; Standardize the first feature of the numerical type to eliminate the dimension; Encode the non-numeric first feature by one-hot encoding; The outliers in the first feature are removed by using the isolation forest algorithm.

4. According to the multi-dimensional fusion-based forged voice intelligent detection system of claim 1, it is characterized in that: In the feature optimization module, performance scores of the same size are in the same position in the ranking.

5. According to the multi-dimensional fusion-based forged voice intelligent detection system of claim 1, it is characterized in that: In the feature optimization module, when there is no performance score difference greater than or equal to the performance score difference threshold, starting from the first performance score in the sorting, calculate the first performance score difference ADYi-1=A1-Ai-1, and when the first performance score difference ADYi-1≥η*Ays, use [1, i-1] as the sorting position interval, η is the preset correction coefficient and η>1.

6. The forged voice intelligent detection system based on multi-dimensional fusion according to claim 1 is characterized in that: In the feature selection module, when there are two or more performance scores that are the same and are the maximum performance scores, the candidate features in the training set with the smallest number of candidate features are used as standard features.

7. The forged voice intelligent detection system based on multi-dimensional fusion according to claim 1 is characterized in that: In the feature optimization module, the process of determining the training set specifically includes: The total number n of the features to be selected is determined, and f features to be selected are arbitrarily selected and combined to obtain a training set, where f∈[1,n] and f is an integer.

8. The forged voice intelligent detection system based on multi-dimensional fusion according to claim 7 is characterized in that: A single training set is replicated, and the replication multiple is a preset value.

Citation Information

Patent Citations

  • Forged voice detection method based on multi-feature fusion and device thereof

    CN113488073A

  • Counterfeit voice detection method, system and device fused with large language model, and medium

    CN117577119A