Anti-Fraud Early Warning Method, Device, Electronic Device and Medium Based on Voice Conversion Recognition

By obtaining reference audio data from the preset sound model library and combining reverb, equalization and matching scores, the target audio data is scored, which solves the problem of low sound change recognition accuracy in the prior art, and improves the accuracy and information security of voice change recognition.

CN115910098BActive Publication Date: 2025-06-24WEIKUN (SHANGHAI) TECH SERVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211364571.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-02
Publication Date
2025-06-24
Estimated Expiration
2042-11-02

AI Technical Summary

Technical Problem

The existing anti-fraud warning methods based on voice change recognition have low recognition accuracy and cannot be accurately identified in various scenarios, resulting in low accuracy of voice change recognition.

Method used

By obtaining reference audio data from the preset sound model library, extracting reference audio parameters, and combining reverb, equalization and matching scores, the target audio data is scored from three dimensions to improve the accuracy of voice change recognition.

Benefits of technology

It improves the accuracy of voice change recognition, enhances the recognition ability of voice change processing, and ensures the security of users' information during the call.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115910098B_ABST
    Figure CN115910098B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence, and discloses an anti-fraud warning method, device, electronic device and storage medium based on voice conversion recognition. The method includes: obtaining reference audio data, and extracting reference audio parameters from the reference audio data; collecting and analyzing target audio data of the initiator of an unfamiliar call to obtain target audio parameters, and calculating a reverberation score of the target audio data according to the target audio parameters; when the reverberation score is not greater than a first preset score, performing frequency division on the target audio data to obtain target audio frequency division segments, and calculating an equalization score of the target audio data; when the equalization score is not greater than a second preset score, matching the target audio parameters with the reference audio parameters to obtain a matching score; weighting the reverberation score, the equalization score and the matching score to obtain an accumulated score; when the accumulated score is not greater than a third preset score, determining that the target audio data has not been voice-converted. The present invention can improve the accuracy of target audio voice conversion recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to an anti-fraud warning method, device, electronic device and readable storage medium based on voice conversion recognition. Background Art

[0002] Voice conversion refers to the voice after using voice change technology to change the voice. For example, the voice after being processed by a voice changer.

[0003] With the development of Internet technology, telecom fraud gangs often change their voices before implementing fraud in order to deceive users and achieve the purpose of fraud. Currently, common anti-fraud warning methods based on voice conversion recognition mostly use common voice recognition software on the Internet to intelligently distinguish simple voice processing traces. In addition, current anti-fraud warning methods based on voice conversion recognition lack data volume and cannot accurately identify in each scenario. Therefore, the recognition accuracy is not high and the voice conversion recognition accuracy is relatively low. Summary of the Invention

[0004] The present invention provides an anti-fraud warning method, device, electronic device and readable storage medium based on voice conversion recognition, aiming to improve the accuracy of voice conversion recognition during a call.

[0005] To achieve the above object, an anti-fraud warning method based on voice conversion recognition provided by the present invention includes:

[0006] Obtain reference audio data from a preset voice model library, and extract reference audio parameters from the reference audio data;

[0007] When a user receives an unfamiliar call, collect target audio data of the originator of the unfamiliar call, analyze the target audio data to obtain target audio parameters, and calculate a reverberation score of the target audio data according to the target audio parameters by using a voice audio data analysis algorithm;

[0008] When the reverberation score is greater than a first preset score, determine that the target audio data has been subjected to reverberation processing, and give a warning prompt to the user;

[0009] When the reverberation score is not greater than the first preset score, perform frequency division processing on the target audio data to obtain target audio frequency division segments, and calculate an equalization score of the target audio data according to the target audio frequency division segments, the target audio parameters and the reference audio parameters;

[0010] When the equalization score is greater than a second preset score, determine that the target audio data has been subjected to equalization processing, and give a warning prompt to the user;

[0011] When the equalization score is not greater than the second preset score, match the target audio parameters with the reference audio parameters to obtain a matching score, and weight the reverberation score, the equalization score, and the matching score to obtain an accumulated score;

[0012] When the accumulated score is greater than the third preset score, determine that the target audio data has been pitch-shifted, and give a warning prompt to the user;

[0013] When the accumulated score is not greater than the third preset score, determine that the target audio data has not been pitch-shifted.

[0014] Optionally, calculating the reverberation score of the target audio data by using a sound audio data analysis algorithm according to the target audio parameters includes:

[0015] According to the sound pressure in the target audio parameters, use a sound audio data analysis algorithm to calculate the real-time sound pressure level of the target audio data;

[0016] Convert the real-time sound pressure level into a real-time decibel value;

[0017] Generate an audio decibel diagram according to the real-time decibel value;

[0018] Determine whether the duration of the high decibel value in the audio decibel diagram exceeds a preset threshold;

[0019] When the duration of the high decibel value in the audio decibel diagram does not exceed the preset threshold, determine that no reverberation is added to the target audio data;

[0020] When the duration of the high decibel value in the audio decibel diagram exceeds the preset threshold, determine that reverberation is added to the target audio data, and obtain the duration and decibel value of the high decibel value in the audio decibel diagram;

[0021] According to a preset rule, calculate the difference between the duration and decibel value of the high decibel value in the audio decibel diagram and the duration and decibel value of the high decibel value in the reference audio data to obtain the reverberation score of the target audio data.

[0022] Optionally, calculating the real-time sound pressure level of the target audio data by using a sound audio data analysis algorithm according to the sound pressure in the target audio parameters includes:

[0023] Use the following formula to calculate the real-time sound pressure level SPL of the target audio data:

[0024]

[0025] where P eis the sound pressure among the target audio parameters, and P is a constant reference sound pressure value, generally 2×10 -5 Pa.

[0026] Optionally, calculating the equalization score of the target audio data according to the target audio frequency division segment, the target audio parameters, and the reference audio parameters includes:

[0027] Constructing an audio high-frequency decibel coordinate system and an audio low-frequency decibel coordinate system according to the target audio frequency division segment and the target audio parameters respectively;

[0028] Constructing a reference audio decibel coordinate system according to the reference audio parameters;

[0029] Judging whether the coordinate curve in the audio high-frequency decibel coordinate system is raised and whether the coordinate curve in the audio low-frequency decibel coordinate system is lowered according to the reference audio decibel coordinate system;

[0030] When the coordinate curve in the audio high-frequency decibel coordinate system is not raised and the coordinate curve in the audio low-frequency decibel coordinate system is not lowered, it is determined that the target audio data has not been equalized;

[0031] When the coordinate curve in the audio high-frequency decibel coordinate system is raised and the coordinate curve in the audio low-frequency decibel coordinate system is lowered, respectively intercept the raised part in the audio high-frequency decibel coordinate system and the lowered part in the audio low-frequency decibel coordinate system;

[0032] Calculating the time length and stretching amplitude of the raised part in the audio high-frequency decibel coordinate system and the lowered part in the audio low-frequency decibel coordinate system;

[0033] Calculating the equalization score of the target audio data according to the time length and the stretching amplitude.

[0034] Optionally, matching the target audio parameters with the reference audio parameters to obtain a matching score includes:

[0035] Calculating the parameter similarity between each parameter in the target audio parameters and the corresponding reference audio parameter one by one;

[0036] Calculating the audio similarity between the reference audio data corresponding to the reference audio parameters and the target audio data according to the parameter similarity, and using the audio similarity as the matching score of the target audio data.

[0037] Optionally, extracting the reference audio parameters from the reference audio data includes:

[0038] The reference audio data is converted from analog to digital using pulse code modulation to obtain a reference audio digital signal;

[0039] The reference audio digital signal is analyzed to obtain analyzed audio data;

[0040] The analyzed audio data is used to draw an audio waveform diagram of the reference audio data;

[0041] Based on the audio waveform diagram, various parameters of the reference audio data are calculated using a preset acoustic algorithm to obtain reference audio parameters.

[0042] Optionally, the converting the reference audio data from analog to digital using pulse code modulation to obtain a reference audio digital signal includes:

[0043] The reference audio data is scanned periodically to obtain a reference audio discrete signal;

[0044] The instantaneous values in the reference audio discrete signal are divided into different levels according to the magnitudes of the instantaneous values in the reference audio discrete signal;

[0045] The levels are converted into binary codes to obtain a reference audio digital signal.

[0046] To solve the above problems, the present invention also provides an anti-fraud warning device based on voice conversion recognition, and the device includes:

[0047] A reverberation scoring module, configured to obtain reference audio data from a preset sound model library, extract reference audio parameters from the reference audio data, when a user receives a strange call, collect target audio data of the originator of the strange call, analyze the target audio data to obtain target audio parameters, and calculate a reverberation score of the target audio data using a sound audio data analysis algorithm according to the target audio parameters;

[0048] An equalization scoring module, configured to determine that the target audio data has been reverberated and give a warning prompt to the user when the reverberation score is greater than a first preset score, and perform frequency division processing on the target audio data to obtain target audio frequency division segments when the reverberation score is not greater than the first preset score, and calculate an equalization score of the target audio data according to the target audio frequency division segments, the target audio parameters, and the reference audio parameters;

[0049] An accumulated score judgment module is used to determine that the target audio data has been equalized and give a warning prompt to the user when the equalization score is greater than a second preset score. When the equalization score is not greater than the second preset score, the target audio parameters are matched with the reference audio parameters to obtain a matching score, and the reverberation score, the equalization score, and the matching score are weighted to obtain an accumulated score. When the accumulated score is greater than a third preset score, it is determined that the target audio data has been voice-transformed and a warning prompt is given to the user. When the accumulated score is not greater than the third preset score, it is determined that the target audio data has not been voice-transformed.

[0050] To solve the above problems, the present invention also provides an electronic device, which includes:

[0051] A memory that stores at least one computer program; and

[0052] A processor that executes the computer program stored in the memory to implement the above-mentioned anti-fraud warning method based on voice transformation recognition.

[0053] To solve the above problems, the present invention also provides a computer-readable storage medium, in which at least one computer program is stored, and the at least one computer program is executed by a processor in an electronic device to implement the above-mentioned anti-fraud warning method based on voice transformation recognition.

[0054] In an embodiment of the present invention, by obtaining reference audio data from a preset sound model library and extracting reference audio parameters from the reference audio data, the rich sound model library improves the breadth of the reference audio data, thereby improving the accuracy of voice conversion recognition. Further, when the user receives an unfamiliar call, the target audio data of the originator of the unfamiliar call is collected, the target audio data is parsed to obtain target audio parameters, according to the target audio parameters, the reverberation score of the target audio data is calculated using a sound audio data analysis algorithm, the target audio data is equalized to obtain audio equalization parameters, the target audio data is frequency-divided to obtain target audio frequency-divided segments, according to the target audio frequency-divided segments, the target audio parameters and the reference audio parameters, the equalization score of the target audio data is calculated, the target audio parameters are matched with the reference audio parameters to obtain a matching score, the target audio data is scored from three dimensions of reverberation, timbre and sound parameters to improve the accuracy of voice conversion recognition. Finally, the reverberation score, the equalization score and the matching score are weighted to obtain an accumulated score, and the user is warned according to the accumulated score, so that the criterion for judging whether the target audio data is voice-converted is more comprehensive and the judgment result is more accurate. Therefore, a fraud prevention and warning method, device, equipment and storage medium based on voice conversion recognition provided by the present invention can improve the accuracy of target audio voice conversion recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 It is a schematic flowchart of a fraud prevention and warning method based on voice conversion recognition provided by an embodiment of the present invention;

[0056] Figures 2 to 3 It is a detailed implementation flowchart of one of the steps in the fraud prevention and warning method based on voice conversion recognition provided by an embodiment of the present invention;

[0057] Figure 4 It is a module schematic diagram of a fraud prevention and warning device based on voice conversion recognition provided by an embodiment of the present invention;

[0058] Figure 5 It is an internal structure schematic diagram of an electronic device for implementing a fraud prevention and warning method based on voice conversion recognition provided by an embodiment of the present invention;

[0059] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0061] An embodiment of the present invention provides an anti-fraud warning method based on voice conversion recognition. The execution subject of the anti-fraud warning method based on voice conversion recognition includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided in the embodiments of the present application. In other words, the anti-fraud warning method based on voice conversion recognition can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server can include an independent server, or can include a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0062] Referring to Figure 1 As shown in the flowchart of the anti-fraud warning method based on voice conversion recognition provided by an embodiment of the present invention, in the embodiment of the present invention, the anti-fraud warning method based on voice conversion recognition includes the following steps S1 - S11:

[0063] S1. Obtain reference audio data from a preset voice model library, and extract reference audio parameters from the reference audio data.

[0064] In the embodiment of the present invention, the preset voice model library can be an audio database containing a large amount of audio data. The reference audio data can be unaltered audio data used to provide the reference audio parameters. The reference audio parameters include male children's voice parameters, male adult voice parameters, male elderly voice parameters, female children's voice parameters, female adult voice parameters, female elderly voice parameters, etc.

[0065] In an alternative embodiment of the present invention, to ensure that the obtained reference audio data is unaltered audio data, the reference audio data can be obtained from the voice model library in local storage, enriching the voice model library and improving the breadth of the reference audio data, thereby improving the accuracy of voice conversion recognition of audio data.

[0066] By extracting the reference audio parameters from the reference audio data in the embodiment of the present invention, a basis is provided for determining the category to which the target audio data belongs, and it is possible to determine whether the target audio data has been subjected to voice conversion processing according to the reference audio parameters, improving the accuracy and efficiency of voice conversion recognition of audio data.

[0067] Further, as an alternative embodiment of the present invention, the extracting the reference audio parameters from the reference audio data includes:

[0068] Perform analog-to-digital conversion on the reference audio data using pulse code modulation to obtain a reference audio digital signal;

[0069] Analyze the reference audio digital signal to obtain analyzed audio data;

[0070] Use the analyzed audio data to draw an audio waveform diagram of the reference audio data;

[0071] According to the audio waveform diagram, use a preset acoustic algorithm to calculate various parameters of the reference audio data to obtain reference audio parameters.

[0072] In an embodiment of the present invention, the pulse code modulation may be a sampling technique for digitizing an analog signal, a coding method for transforming an analog voice signal into a digital signal. Various parameters of the reference audio data include parameters such as frequency, bandwidth, and gain.

[0073] In an alternative embodiment of the present invention, by performing analog-to-digital conversion on the reference audio data to obtain a digital signal that can be interpreted by a computer, and then using the FFmeng open-source technology to analyze the reference audio digital signal to obtain analyzed audio data, thereby according to the audio data, using acoustic-related algorithms to calculate each parameter of the reference audio data. Among them, the FFmeng open-source technology may be a set of open-source computer programs that can be used to record, convert digital audio and video, and convert them into streams. The acoustic-related algorithms include an audio frequency calculation formula, an audio gain calculation formula, an audio bandwidth calculation formula, etc.

[0074] Further, in an embodiment of the present invention, the use of the pulse code modulation method to perform analog-to-digital conversion on the reference audio data to obtain a reference audio digital signal includes:

[0075] Perform periodic scanning on the reference audio data to obtain a reference audio discrete signal;

[0076] Divide the instantaneous values in the reference audio discrete signal into different levels according to the magnitudes of the instantaneous values in the reference audio discrete signal;

[0077] Convert the levels into binary codes to obtain a reference audio digital signal.

[0078] In an alternative embodiment of the present invention, by continuously collecting the instantaneous voltage value of the reference audio data at a fixed time interval through high-frequency detection, and then dividing the instantaneous voltage value into different levels, using the rounding principle, dividing the collected instantaneous voltage value into different levels to reduce the workload. Finally, each sampled voltage level needs to be converted into a binary code recognizable by a computer, thereby completing the conversion of the reference audio data to the reference audio digital signal.

[0079] S2. When the user receives an unfamiliar call, collect the target audio data of the originator of the unfamiliar call, analyze the target audio data to obtain target audio parameters, and calculate the reverberation score of the target audio data according to the target audio parameters by using a sound audio data analysis algorithm.

[0080] In the embodiment of the present invention, the unfamiliar call may be a call request sent from a telephone number not in the user's phone contact list. The target audio data may be the call recording data of the originator of the unfamiliar call. The sound audio data analysis algorithm may be a common algorithm for calculating the decibel value of audio data. The reverberation score may be scored by comparing the high-decibel duration in the sound curve of the target audio data with the high-decibel duration in the sound curve of normal audio data.

[0081] In an alternative embodiment of the present invention, when the user receives an unfamiliar call, start collecting the voice of the originator of the unfamiliar call, so as to ensure timely identification of whether the audio of the originator of the unfamiliar call has been voice-transformed, and protect the information and property security of the user.

[0082] Further, as an alternative embodiment of the present invention, the analysis of the target audio data to obtain target audio parameters is the same as the extraction of reference audio parameters from the reference audio data, so it will not be elaborated here.

[0083] In an alternative embodiment of the present invention, when appropriate reverberation is added to the sound, the sound can become loud, full and infectious, making it easier for the listener to be deceived during the call. Therefore, it is first necessary to identify the reverberation of the target audio data to determine whether the target audio data has reverberation added, so as to protect the property information security of the listener during the call.

[0084] In the embodiment of the present invention, according to the target audio parameters, use a sound audio data analysis algorithm to calculate the reverberation score of the target audio data, and identify whether the target audio data has been voice-transformed from the dimension of reverberation, so as to protect the property information security of the listener during the call.

[0085] Further, as an alternative embodiment of the present invention, refer to Figure 2 As shown, the calculation of the reverberation score of the target audio data according to the target audio parameters by using a sound audio data analysis algorithm includes S21 - S27:

[0086] S21. Calculate the real-time sound pressure level of the target audio data according to the sound pressure in the target audio parameters by using a sound audio data analysis algorithm;

[0087] S22. Convert the real-time sound pressure level into a real-time decibel value;

[0088] S23. Generate an audio decibel graph based on the real-time decibel value;

[0089] S24. Determine whether the duration of the high decibel value in the audio decibel graph exceeds a preset threshold;

[0090] S25. When the duration of the high decibel value in the audio decibel graph does not exceed the preset threshold, determine that the target audio data does not have reverb added;

[0091] S26. When the duration of the high decibel value in the audio decibel graph exceeds the preset threshold, determine that the target audio data has reverb added, and obtain the duration and decibel value of the high decibel value in the audio decibel graph;

[0092] S27. According to a preset rule, calculate the difference between the duration and decibel value of the high decibel value in the audio decibel graph and the duration and decibel value of the high decibel value in the reference audio data, to obtain the reverb score of the target audio data.

[0093] In the embodiment of the present invention, the sound pressure may be the root mean square value of the excess instantaneous pressure generated by a sound wave at a certain point. The real-time decibel value may be the decibel value of each time node of the target audio data. The audio decibel graph may be a coordinate system with time as the abscissa and decibel magnitude as the ordinate. The preset threshold may be the average duration of the high decibel value of a normal human voice when there is no reverb added. The preset rule may be a pre-set scoring criterion.

[0094] In an alternative embodiment of the present invention, first, by calculating the real-time decibel value of the target audio data, a decibel-time coordinate system is constructed, so as to facilitate calculating the duration and trend magnitude of the high decibel value of the target audio data, and further complete the calculation of the reverb score of the target audio data.

[0095] Further, as an alternative embodiment of the present invention, the calculating the real-time sound pressure level of the target audio data by using a sound audio data analysis algorithm according to the sound pressure in the target audio parameters includes:

[0096] Calculate the real-time sound pressure level SPL of the target audio data by using the following formula:

[0097]

[0098] Where P e is the sound pressure in the target audio parameters, and P is a constant reference sound pressure value, generally 2×10 -5 Pa.

[0099] In an alternative embodiment of the present invention, to ensure the accuracy of the reverberation score, the time interval for calculating the real-time sound pressure level should be minimized as much as possible, thereby improving the accuracy of the real-time decibel value.

[0100] S3. Determine whether the reverberation score is greater than a first preset score.

[0101] In an embodiment of the present invention, the first preset score may be the maximum score of the reverberation score in the reference audio data.

[0102] In an alternative embodiment of the present invention, by comparing the numerical magnitudes of the reverberation score and the first preset score, it is thus possible to determine whether the reverberation score is greater than the first preset score.

[0103] S4. When the reverberation score is greater than the first preset score, determine that the target audio data has been subjected to reverberation processing, and give a warning prompt to the user.

[0104] In an alternative embodiment of the present invention, when the reverberation score is greater than the first preset score, it proves that the duration and trend magnitude of the high decibel values in the target audio data are greater than those of the high decibel values in the reference audio data. Therefore, it can be determined that the target audio data has been subjected to reverberation processing. Furthermore, a warning prompt is given to the user to ensure the security of the user's property information.

[0105] S5. When the reverberation score is not greater than the first preset score, perform frequency division processing on the target audio data to obtain target audio frequency division segments, and calculate the equalization score of the target audio data according to the target audio frequency division segments, the target audio parameters, and the reference audio parameters.

[0106] In an embodiment of the present invention, the target audio frequency division segments include a target audio high-frequency segment and a target audio low-frequency segment.

[0107] In an alternative embodiment of the present invention, when sound is equalized, the quality and tone of the voice of the unfamiliar call initiator can be improved, making it easier for the call recipient to be deceived during the call. Therefore, it is first necessary to perform frequency division processing on the target audio data, and determine whether the target audio data has been equalized according to the audio data of each frequency band, thereby ensuring the security of the property information of the call recipient during the call.

[0108] In another alternative embodiment of the present invention, the target audio data can be frequency-divided by filters of various frequency bands to obtain target audio frequency division segments, thereby ensuring the accuracy of the equalization score.

[0109] In an embodiment of the present invention, an equalization score of the target audio data is calculated based on the target audio frequency-divided segments, the target audio parameters, and the reference audio parameters, and whether the target audio data is voice-changed is identified from the dimension of timbre, thereby ensuring the security of the property information of the receiving party during a call.

[0110] Further, as an optional embodiment of the present invention, refer to Figure 3 As shown, calculating the equalization score of the target audio data according to the target audio frequency-divided segments, the target audio parameters, and the reference audio parameters includes steps S51 - S57:

[0111] S51. Respectively construct an audio high-frequency decibel coordinate system and an audio low-frequency decibel coordinate system according to the target audio frequency-divided segments and the target audio parameters;

[0112] S52. Construct a reference audio decibel coordinate system according to the reference audio parameters;

[0113] S53. Judge whether the coordinate curve in the audio high-frequency decibel coordinate system is raised and whether the coordinate curve in the audio low-frequency decibel coordinate system is lowered according to the reference audio decibel coordinate system;

[0114] S54. When the coordinate curve in the audio high-frequency decibel coordinate system is not raised and the coordinate curve in the audio low-frequency decibel coordinate system is not lowered, it is determined that the target audio data has not been equalized;

[0115] S55. When the coordinate curve in the audio high-frequency decibel coordinate system is raised and the coordinate curve in the audio low-frequency decibel coordinate system is lowered, respectively intercept the raised part in the audio high-frequency decibel coordinate system and the lowered part in the audio low-frequency decibel coordinate system;

[0116] S56. Calculate the time length and stretching amplitude of the raised part in the audio high-frequency decibel coordinate system and the lowered part in the audio low-frequency decibel coordinate system;

[0117] S57. Calculate the equalization score of the target audio data according to the time length and the stretching amplitude.

[0118] In an alternative embodiment of the present invention, an audio equalization coordinate system is constructed according to the target audio frequency division segment and the target audio parameters, and then the audio equalization coordinate system is compared with the reference audio coordinate system, so as to find out whether there is a sign of being raised in the high-frequency region of the target audio and whether there is a sign of being lowered in the low-frequency region of the target audio, thereby determining whether the target audio data has been equalized. Furthermore, the equalization degree of the target audio data is scored according to the time length and stretching amplitude of the raised or / and lowered part in the audio high-frequency decibel coordinate system or / and the audio low-frequency decibel coordinate system.

[0119] S6. Determine whether the equalization score is greater than a second preset score.

[0120] In an embodiment of the present invention, the second preset score may be the maximum score of the equalization scores in the reference audio data.

[0121] In an alternative embodiment of the present invention, by comparing the numerical values of the equalization score and the second preset score, it is determined whether the equalization score is greater than the second preset score.

[0122] S7. When the equalization score is greater than the second preset score, it is determined that the target audio data has been equalized, and a warning prompt is given to the user.

[0123] In an alternative embodiment of the present invention, when the equalization score is greater than the second preset score, it proves that the coordinate curve in the audio high-frequency decibel coordinate system is raised or / and the coordinate curve in the audio low-frequency decibel coordinate system is lowered. Therefore, it can be determined that the target audio data has been equalized. Furthermore, a warning prompt is given to the user to ensure the security of the user's property information.

[0124] S8. When the equalization score is not greater than the second preset score, the target audio parameters are matched with the reference audio parameters to obtain a matching score, and the reverberation score, the equalization score, and the matching score are weighted to obtain an accumulated score.

[0125] In an alternative embodiment of the present invention, when the equalization score is not greater than the second preset score, it proves that the coordinate curve in the audio high-frequency decibel coordinate system is not raised and the coordinate curve in the audio low-frequency decibel coordinate system is not lowered. Therefore, it can be determined that the target audio data has not been equalized. Furthermore, the target audio parameters are matched with the reference audio parameters to obtain a matching score, and it is further determined whether the initiator of the strange call has performed voice transformation processing to ensure the security of the user's property information.

[0126] In the embodiment of the present invention, the target audio parameters are matched with the reference audio parameters to obtain a matching score, and it is judged whether the target audio data is voice-changed from the dimension of sound parameters, thereby ensuring the property information security of the receiving party during the call.

[0127] Further, as an optional embodiment of the present invention, the matching of the target audio parameters with the reference audio parameters to obtain a matching score includes:

[0128] Calculate the parameter similarity of each parameter in the target audio parameters with the corresponding reference audio parameter one by one;

[0129] According to the parameter similarity, calculate the audio similarity between the reference audio data corresponding to the reference audio parameters and the target audio data, and use the audio similarity as the matching score of the target audio data.

[0130] In the embodiment of the present invention, the parameter similarity can be obtained by the cosine similarity calculation method. The audio similarity is similar to the parameter similarity. In an optional embodiment of the present invention, first convert the target audio parameters and reference audio parameters into vector forms, and then use the cosine similarity calculation formula to calculate the parameter similarity of each parameter in the target audio parameters with the corresponding reference audio parameter. Using the same method, calculate the audio similarity between the reference audio data corresponding to the reference audio parameters and the target audio data according to the parameter similarity, and finally take the highest score from all audio similarities as the matching score of the target audio data.

[0131] Further, as another optional embodiment of the present invention, the reverberation score, the equalization score, and the matching score are respectively weighted by a preset weight value, and the weighted values are accumulated to obtain an accumulated score, so that the standard for judging whether the target audio data is voice-changed is more comprehensive and the judgment result is more accurate.

[0132] S9. Judge whether the accumulated score is greater than a third preset value.

[0133] In the embodiment of the present invention, the second preset value may be a critical value obtained by comparing a large amount of voice-changed audio data with reference audio data.

[0134] In an optional embodiment of the present invention, by comparing the numerical values of the accumulated score and the third preset value, it is realized to judge whether the accumulated score is greater than the third preset value.

[0135] S10. When the accumulated score is greater than the third preset value, it is determined that the target audio data has been voice-changed, and a warning prompt is given to the user.

[0136] In an alternative embodiment of the present invention, when the cumulative score is greater than the third preset score, it proves that the matching degree between the target audio data and the voice-transformed audio data is higher. Therefore, it can be determined that the target audio data has been voice-transformed. Furthermore, a warning prompt is given to the user to ensure the security of the user's property information.

[0137] S11. When the cumulative score is not greater than the third preset score, it is determined that the target audio data has not been voice-transformed.

[0138] In an alternative embodiment of the present invention, when the cumulative score is not greater than the third preset score, it proves that the matching degree between the target audio data and the reference audio data is higher. Therefore, it can be determined that the target audio data has not been voice-transformed. Furthermore, the security of the user's property information is ensured.

[0139] In an embodiment of the present invention, by obtaining reference audio data from a preset voice model library, extracting reference audio parameters from the reference audio data, and having a rich voice model library, the breadth of the reference audio data is improved, thereby improving the accuracy of voice transformation recognition. Further, when the user receives a strange call, the target audio data of the initiator of the strange call is collected, the target audio data is analyzed to obtain target audio parameters, according to the target audio parameters, the reverberation score of the target audio data is calculated using a voice audio data analysis algorithm, the target audio data is equalized to obtain audio equalization parameters, the target audio data is frequency-divided to obtain target audio frequency-divided segments, according to the target audio frequency-divided segments, the target audio parameters, and the reference audio parameters, the equalization score of the target audio data is calculated, the target audio parameters are matched with the reference audio parameters to obtain a matching score, the target audio data is scored from three dimensions of reverberation, timbre, and voice parameters to improve the accuracy of voice transformation recognition. Finally, the reverberation score, the equalization score, and the matching score are weighted to obtain a cumulative score, and a warning prompt is given to the user according to the cumulative score, so that the standard for judging whether the target audio data has been voice-transformed is more comprehensive and the judgment result is more accurate. Therefore, a fraud prevention warning method based on voice transformation recognition provided by the present invention can improve the accuracy of target audio voice transformation recognition.

[0140] As Figure 4 shown, it is a functional module diagram of a fraud prevention warning device based on voice transformation recognition of the present invention.

[0141] The anti-fraud warning device 100 based on voice conversion recognition according to the present invention can be installed in an electronic device. According to the functions achieved, the anti-fraud warning device 100 based on voice conversion recognition can include a reverberation scoring module 101, an equalization scoring module 102, and an accumulated scoring judgment module 103. The modules in the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by the processor of the electronic device and can complete fixed functions, and are stored in the memory of the electronic device.

[0142] In this embodiment, the functions of each module / unit are as follows:

[0143] The reverberation scoring module 101 is used to obtain reference audio data from a preset sound model library, extract reference audio parameters from the reference audio data. When the user receives an unfamiliar call, collect the target audio data of the originator of the unfamiliar call, parse the target audio data to obtain target audio parameters, and calculate the reverberation score of the target audio data using a sound audio data analysis algorithm according to the target audio parameters.

[0144] The equalization scoring module 102 is used to determine that the target audio data has been reverberation processed and give a warning prompt to the user when the reverberation score is greater than a first preset score. When the reverberation score is not greater than the first preset score, perform frequency division processing on the target audio data to obtain target audio frequency division segments, and calculate the equalization score of the target audio data according to the target audio frequency division segments, the target audio parameters, and the reference audio parameters.

[0145] The accumulated scoring judgment module 103 is used to determine that the target audio data has been equalization processed and give a warning prompt to the user when the equalization score is greater than a second preset score. When the equalization score is not greater than the second preset score, match the target audio parameters with the reference audio parameters to obtain a matching score, and weight the reverberation score, the equalization score, and the matching score to obtain an accumulated score. When the accumulated score is greater than a third preset score, determine that the target audio data has been voice-converted and give a warning prompt to the user. When the accumulated score is not greater than the third preset score, determine that the target audio data has not been voice-converted.

[0146] As Figure 5 shown, it is a schematic structural diagram of an electronic device for implementing the anti-fraud warning method based on voice conversion recognition according to the present invention.

[0147] The electronic device may include a processor 10, a memory 11, a communication bus 12, and a communication interface 13, and may also include a computer program stored in the memory 11 and executable on the processor 10, such as an anti-fraud warning program based on voice conversion recognition.

[0148] Among them, the memory 11 at least includes one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 11 can be an internal storage unit of the electronic device, such as the mobile hard disk of the electronic device. In some other embodiments, the memory 11 can also be an external storage device of the electronic device, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device. Further, the memory 11 can also include both the internal storage unit of the electronic device and the external storage device. The memory 11 can be used not only to store the application software installed on the electronic device and various types of data, such as the code of the anti-fraud warning program based on voice conversion recognition, etc., but also to temporarily store the data that has been output or will be output.

[0149] In some embodiments, the processor 10 can be composed of integrated circuits. For example, it can be composed of a single packaged integrated circuit, or can be composed of multiple integrated circuits with the same or different functions packaged, including a combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 10 is the control core (Control Unit) of the electronic device, connecting all components of the entire electronic device through various interfaces and lines, and by running or executing the programs or modules stored in the memory 11 (such as the anti-fraud warning program based on voice conversion recognition, etc.), and calling the data stored in the memory 11, to execute various functions of the electronic device and process data.

[0150] The communication bus 12 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The communication bus 12 is set to realize the connection and communication between the memory 11 and at least one processor 10, etc. For the convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0151] Figure 5Only an electronic device with components is shown. Those skilled in the art can understand that Figure 5 the shown structure does not constitute a limitation on the electronic device, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0152] For example, although not shown, the electronic device may further include a power source (such as a battery) for supplying power to each component. Preferably, the power source can be logically connected to the at least one processor 10 through a power management device, so as to implement functions such as charge management, discharge management, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or inverter, and a power status indicator. The electronic device may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.

[0153] Optionally, the communication interface 13 may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), and is generally used to establish a communication connection between this electronic device and other electronic devices.

[0154] Optionally, the communication interface 13 may further include a user interface. The user interface may be a display (Display), an input unit (such as a keyboard (Keyboard)). Optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, and is used to display the information processed in the electronic device and to display a visual user interface.

[0155] It should be understood that the above embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.

[0156] The anti-fraud warning program based on voice conversion recognition stored in the memory 11 of the electronic device is a combination of multiple computer programs. When running in the processor 10, it can implement:

[0157] Obtain reference audio data from a preset voice model library and extract reference audio parameters from the reference audio data;

[0158] When a user receives an unfamiliar call, collect the target audio data of the originator of the unfamiliar call, analyze the target audio data to obtain target audio parameters, and calculate the reverberation score of the target audio data using a sound audio data analysis algorithm according to the target audio parameters;

[0159] When the reverberation score is greater than a first preset score, determine that the target audio data has been reverberated, and give a warning prompt to the user;

[0160] When the reverberation score is not greater than the first preset score, perform frequency division processing on the target audio data to obtain target audio frequency division segments, and calculate the equalization score of the target audio data according to the target audio frequency division segments, the target audio parameters, and the reference audio parameters;

[0161] When the equalization score is greater than a second preset score, determine that the target audio data has been equalized, and give a warning prompt to the user;

[0162] When the equalization score is not greater than the second preset score, match the target audio parameters with the reference audio parameters to obtain a matching score, and weight the reverberation score, the equalization score, and the matching score to obtain an accumulated score;

[0163] When the accumulated score is greater than a third preset score, determine that the target audio data has been voice-transformed, and give a warning prompt to the user;

[0164] When the accumulated score is not greater than the third preset score, determine that the target audio data has not been voice-transformed.

[0165] Specifically, for the specific implementation method of the above computer program by the processor 10, reference can be made to Figure 1 the description of the relevant steps in the corresponding embodiment, which will not be elaborated here.

[0166] Furthermore, if the modules / units integrated in the electronic device are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable medium can be non-volatile or volatile. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory).

[0167] An embodiment of the present invention can also provide a computer-readable storage medium. The readable storage medium stores a computer program, and when the computer program is executed by a processor of an electronic device, it can implement:

[0168] Obtain reference audio data from a preset sound model library, and extract reference audio parameters from the reference audio data;

[0169] When the user receives an unfamiliar call, collect the target audio data of the originator of the unfamiliar call, parse the target audio data to obtain target audio parameters, and calculate the reverberation score of the target audio data according to the target audio parameters using a sound audio data analysis algorithm;

[0170] When the reverberation score is greater than a first preset score, determine that the target audio data has been reverberation processed, and give a warning prompt to the user;

[0171] When the reverberation score is not greater than the first preset score, perform frequency division processing on the target audio data to obtain target audio frequency division segments, and calculate the equalization score of the target audio data according to the target audio frequency division segments, the target audio parameters, and the reference audio parameters;

[0172] When the equalization score is greater than a second preset score, determine that the target audio data has been equalization processed, and give a warning prompt to the user;

[0173] When the equalization score is not greater than the second preset score, match the target audio parameters with the reference audio parameters to obtain a matching score, and weight the reverberation score, the equalization score, and the matching score to obtain an accumulated score;

[0174] When the accumulated score is greater than a third preset score, determine that the target audio data has been voice-transformed, and give a warning prompt to the user;

[0175] When the accumulated score is not greater than the third preset score, determine that the target audio data has not been voice-transformed.

[0176] Furthermore, the computer-usable storage medium may mainly include a storage program area and a storage data area. Among them, the storage program area may store an operating system, application programs required for at least one function, etc.; the storage data area may store data created according to the use of blockchain nodes, etc.

[0177] In several embodiments provided by the present invention, it should be understood that the disclosed electronic devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.

[0178] The module described as a separation component may or may not be physically separated. The component shown as a module may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0179] In addition, in each embodiment of the present invention, each functional module can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a hardware plus software functional module.

[0180] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.

[0181] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claimed claim.

[0182] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, essentially a decentralized database, is a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer, etc.

[0183] In addition, obviously, the word "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices stated in the system claims can also be implemented by one unit or device through software or hardware. Words such as second are used to denote names and do not denote any particular order.

[0184] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not restrictive. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. An anti-fraud warning method based on voice conversion recognition, characterized in that, The method includes: Obtain reference audio data from a preset sound model library, and extract reference audio parameters from the reference audio data; When the user receives an unfamiliar call, collect target audio data of the initiator of the unfamiliar call, analyze the target audio data to obtain target audio parameters, and calculate a reverberation score of the target audio data using a sound audio data analysis algorithm according to the target audio parameters; When the reverberation score is greater than a first preset score, determine that the target audio data has been reverberated, and give a warning prompt to the user; When the reverberation score is not greater than the first preset score, perform frequency division processing on the target audio data to obtain target audio frequency division segments, and calculate an equalization score of the target audio data according to the target audio frequency division segments, the target audio parameters, and the reference audio parameters; When the equalization score is greater than a second preset score, determine that the target audio data has been equalized, and give a warning prompt to the user; When the equalization score is not greater than the second preset score, match the target audio parameters with the reference audio parameters to obtain a matching score, and weight the reverberation score, the equalization score, and the matching score to obtain an accumulated score; When the accumulated score is greater than a third preset score, determine that the target audio data has been voice-transformed, and give a warning prompt to the user; When the accumulated score is not greater than the third preset score, determine that the target audio data has not been voice-transformed; Among them, calculating the equalization score of the target audio data according to the target audio frequency division segments, the target audio parameters, and the reference audio parameters includes: respectively constructing an audio high-frequency decibel coordinate system and an audio low-frequency decibel coordinate system according to the target audio frequency division segments and the target audio parameters; constructing a reference audio decibel coordinate system according to the reference audio parameters; judging whether the coordinate curve in the audio high-frequency decibel coordinate system is raised and whether the coordinate curve in the audio low-frequency decibel coordinate system is lowered according to the reference audio decibel coordinate system; when the coordinate curve in the audio high-frequency decibel coordinate system is not raised and the coordinate curve in the audio low-frequency decibel coordinate system is not lowered, determine that the target audio data has not been equalized; when the coordinate curve in the audio high-frequency decibel coordinate system is raised and the coordinate curve in the audio low-frequency decibel coordinate system is lowered, respectively intercept the raised part in the audio high-frequency decibel coordinate system and the lowered part in the audio low-frequency decibel coordinate system; calculate the time length and stretching amplitude of the raised part in the audio high-frequency decibel coordinate system and the lowered part in the audio low-frequency decibel coordinate system; calculate the equalization score of the target audio data according to the time length and the stretching amplitude.

2. The anti-fraud early warning method based on voice conversion recognition according to claim 1, wherein Calculating the reverberation score of the target audio data using a sound audio data analysis algorithm according to the target audio parameters includes: Calculate the real-time sound pressure level of the target audio data using a sound audio data analysis algorithm according to the sound pressure in the target audio parameters; Convert the real-time sound pressure level into a real-time decibel value; Generate an audio decibel graph according to the real-time decibel value; Determine whether the duration of the high decibel value in the audio decibel graph exceeds a preset threshold; When the duration of the high decibel value in the audio decibel graph does not exceed the preset threshold, determine that the target audio data does not have reverberation added; When the duration of the high decibel value in the audio decibel graph exceeds the preset threshold, determine that the target audio data has reverberation added, and obtain the duration and decibel value of the high decibel value in the audio decibel graph; According to a preset rule, calculate the difference between the duration and decibel value of the high decibel value in the audio decibel graph and the duration and decibel value of the high decibel value in the reference audio data, and obtain the reverberation score of the target audio data.

3. The anti-fraud early warning method based on voice conversion recognition according to claim 2, wherein The calculating the real-time sound pressure level of the target audio data by using a sound audio data analysis algorithm according to the sound pressure in the target audio parameters includes: Calculate the real-time sound pressure level of the target audio data using the following formula :[[]]END]] Among them, is the sound pressure among the target audio parameters, is the constant reference sound pressure value, generally Pa.

4. The anti-fraud early warning method based on voice conversion recognition according to claim 1, wherein The matching the target audio parameters with the reference audio parameters to obtain a matching score includes: Calculate the parameter similarity of each parameter in the target audio parameters with the corresponding reference audio parameter one by one; According to the parameter similarity, calculate the audio similarity between the reference audio data corresponding to the reference audio parameters and the target audio data, and use the audio similarity as the matching score of the target audio data.

5. The anti-fraud early warning method based on voice conversion recognition according to claim 1, characterized in that, The extracting the reference audio parameters in the reference audio data includes: Perform analog-to-digital conversion on the reference audio data by using a pulse code modulation method to obtain a reference audio digital signal; Analyze the reference audio digital signal to obtain analyzed audio data; Use the analyzed audio data to draw an audio waveform graph of the reference audio data; According to the audio waveform graph, use a preset acoustic algorithm to calculate various parameters of the reference audio data to obtain reference audio parameters.

6. The anti-fraud warning method based on voice conversion recognition according to claim 5, characterized in that, The performing analog-to-digital conversion on the reference audio data by using a pulse code modulation method to obtain a reference audio digital signal includes: Perform periodic scanning on the reference audio data to obtain a reference audio discrete signal; Divide the instantaneous values in the reference audio discrete signal into different levels according to the magnitude of the instantaneous values in the reference audio discrete signal; Convert the levels into binary codes to obtain a reference audio digital signal.

7. An anti-fraud early warning device based on voice conversion recognition, which is used to implement the anti-fraud early warning method based on voice conversion recognition according to any one of claims 1 to 6, and is characterized in that, The device includes: A reverberation scoring module, configured to obtain reference audio data from a preset sound model library, extract reference audio parameters in the reference audio data, collect target audio data of the initiator of the strange call when the user receives a strange call, analyze the target audio data to obtain target audio parameters, and calculate the reverberation score of the target audio data by using a sound audio data analysis algorithm according to the target audio parameters; An equalization scoring module, which is used to determine that the target audio data has been reverberated and give a warning prompt to the user when the reverberation score is greater than a first preset score. When the reverberation score is not greater than the first preset score, perform frequency division processing on the target audio data to obtain target audio frequency division segments, and calculate the equalization score of the target audio data according to the target audio frequency division segments, the target audio parameters and the reference audio parameters; An accumulated score judgment module, which is used to determine that the target audio data has been equalized and give a warning prompt to the user when the equalization score is greater than a second preset score. When the equalization score is not greater than the second preset score, match the target audio parameters with the reference audio parameters to obtain a matching score, and weight the reverberation score, the equalization score and the matching score to obtain an accumulated score. When the accumulated score is greater than a third preset score, determine that the target audio data has been voice-changed and give a warning prompt to the user. When the accumulated score is not greater than the third preset score, determine that the target audio data has not been voice-changed.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the anti-fraud warning method based on voice change recognition according to any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the anti-fraud warning method based on voice change recognition according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Sound changing method, device, equipment and medium

    CN113241082A

  • Method of recognising a sound event

    US10878840B1