An audio authentication method with acquisition signal analysis function

By constructing a comprehensive audio analysis model and combining multiple factors for audio identification, the problems of misjudgment and identification errors in the existing technology are solved, and more accurate and objective audio authenticity identification is achieved.

CN116884434BActive Publication Date: 2025-08-29CHINA ACADEMY OF INFORMATION & COMM
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310908977.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-24
Publication Date
2025-08-29
Estimated Expiration
2043-07-24

AI Technical Summary

Technical Problem

The existing audio identification methods have misjudgment and identification errors in the authenticity identification process, and fail to fully consider the impact of environmental and equipment factors.

Method used

Build a comprehensive audio analysis model, combine meteorological data, signal data, audio equipment brand and recording quality to conduct multi-dimensional analysis, generate internal and external factors influence marks, and judge the authenticity of the audio through integrated analysis.

Benefits of technology

It realizes more accurate and objective audio authenticity identification, reduces the influence of subjective factors, and provides more comprehensive identification results and solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116884434B_ABST
    Figure CN116884434B_ABST
Patent Text Reader

Abstract

The present invention discloses an audio authentication method with a signal acquisition analysis function, which relates to the technical field of audio authentication. The method comprises the following steps: S1, constructing an audio comprehensive analysis model, transmitting authentication audio, authentication audio recording device data and recording location data; S2, extracting audio data from the authentication audio; S3, analyzing and processing the audio data to generate audio data processing information, and performing data comparison on the audio recording device data to generate device matching information; S4, performing integrated analysis based on the audio data processing information and the device matching information to generate an internal influence identifier; S5, analyzing the recording location data, obtaining external influence data of the recording period through a network terminal and performing analysis and processing to generate an external influence identifier; and S6, performing combined analysis on the internal influence identifier and the external influence identifier to generate a matching target for the authentication audio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of audio authentication, and in particular to an audio authentication method with a function of collecting and analyzing signals. Background Art

[0002] After searching, it was found that a comparative document with publication number CN1 07564534A proposed an audio quality identification method and device. According to the method provided in the embodiment of the comparative document, the target audio is analyzed in the frequency domain to obtain the cutoff frequency of each audio frame in the target audio, and the quality of the target audio is identified based on the cutoff frequency; this method fundamentally analyzes the target audio and determines whether the target audio is high-quality audio based on the basic characteristics of the audio; in the case where a record company re-encodes low-bitrate audio into high-bitrate audio, or in the case where audio in a low-bitrate, lossy audio compression encoding format is played and re-recorded into audio in a high-bitrate, lossless audio compression encoding format, the authenticity of high-quality audio can be identified through this method.

[0003] With reference to the comparative documents, it is found that the existing technology still has the following deficiencies:

[0004] 1. When summarizing the process of audio authenticity identification, single data is often used for analysis and comparison. The detected audio has misjudgments and the data processing in the detection process is simplified, which leads to omissions in the authenticity identification.

[0005] 2. In the process of authenticity identification, pre-processing is basically used to remove noise. However, if the forged audio also includes the removal of excess noise and subsequent splicing and forgery, there will be certain identification errors, and there is no specific analysis and judgment based on the environment.

[0006] In order to solve the above-mentioned problems, an audio authentication method with acquisition signal analysis function is proposed. Summary of the Invention

[0007] The purpose of the present invention is to provide an audio authentication method with a collected signal analysis function to solve the shortcomings of the background technology.

[0008] The audio authentication method with the acquisition signal analysis function comprises the following steps:

[0009] Step S1: construct an audio comprehensive analysis model, transmit the identification audio, identification audio recording device data and recording location data;

[0010] Step S2, extracting audio data from the identification audio;

[0011] Step S3: Analyze and process the audio data to generate audio data processing information, and compare the audio recording device data to generate device matching information;

[0012] Step S4: Perform an integrated analysis based on the audio data processing information and the device matching information to generate an internal factor impact identifier;

[0013] Step S5: Analyze the recording location data, obtain the external factor influencing data of the recording period through the network, analyze and process it, and generate an external factor influencing identifier;

[0014] Step S6: Perform a combined analysis on the internal influence identifier and the external influence identifier, and generate a matching target for the identification audio.

[0015] In a preferred embodiment, the audio data includes a mixing duration ratio, a spectrum smoothing duration ratio, and an abnormal spectrum duration ratio. A larger mixing duration ratio indicates a greater probability of modification in the identified audio, while the spectrum smoothing duration ratio indicates the proportion of time during which the spectrum is smooth and continuous, and the abnormal spectrum duration ratio indicates the proportion of time during which abnormal peaks in the spectrum exist.

[0016] The audio recording device data includes a device brand and a device signal. The device brand is the brand of the device, and the device signal is the average signal value when the recording device is used in a specific area.

[0017] In a preferred embodiment, the steps of generating the audio data processing information are:

[0018] The audio data processing information includes low-level audio processing information, medium-level audio processing information and high-level audio processing information;

[0019] Obtain the mixing duration ratio b, spectrum smoothing duration ratio n, and abnormal spectrum duration ratio m in the identified audio, and calculate the audio impact factor γ through formula analysis;

[0020] Set audio impact reference numbers γ1 and γ2, where γ1<γ2, and substitute the audio impact factor γ into the audio impact reference numbers γ1 and γ2 for comparative analysis. When the audio impact factor γ is greater than 0 and less than the audio impact reference number γ1, high audio processing information is generated for the identification audio; when the audio impact factor γ is greater than the audio impact reference number γ1 and less than the audio impact reference number γ2, medium audio processing information is generated for the identification audio; when the audio impact factor γ is greater than the audio impact reference number γ2, low audio processing information is generated for the identification audio.

[0021] In a preferred embodiment, the logic for generating device matching information is as follows:

[0022] The device matching information includes matching information and difference information;

[0023] Set the audio reference brand and audio signal comparison range in the audio comprehensive analysis model. The audio reference brands include but are not limited to X1, X2, and X3. The audio signal comparison range corresponding to the audio reference brand X1 in the audio comprehensive analysis model is The audio signal comparison range of the audio reference brand X2 in the audio comprehensive analysis model is The audio signal comparison range of the audio reference brand X3 in the audio comprehensive analysis model is Wherein, Y1, Y2, Y3, Y4, Y5, and Y6 are audio reference signals preset in the comprehensive audio analysis model. The audio reference signals Y1, Y2, Y3, Y4, Y5, and Y6 are all greater than 0, but the audio reference signal Y1 is less than the audio reference signal Y2, the audio reference signal Y3 is less than the audio reference signal Y4, and the audio reference signal Y5 is less than the audio reference signal Y6.

[0024] The device brand X and device signal Y in the input audio recording device data are substituted into the audio signal comparison range for classification and the device brand result X is generated. Y , set the device brand result to X Y Perform matching analysis with the input device brand X. If the brand result is X Y The brand X of the input device is matched to the same brand, and the matching information is generated for the identification audio. If the brand result is X Y If the brand matched with the input device brand X is not the same brand, difference information is generated for the authentication audio.

[0025] In a preferred embodiment, the internal factor impact identifier includes a high internal factor impact identifier, a medium internal factor impact identifier, and a low internal factor impact identifier. The steps for generating the internal factor impact identifier are specifically as follows:

[0026] When the same identification audio contains both difference information and high audio processing information, difference information and medium audio processing information, or matching information and high audio processing information, the identification audio will be marked as having a high internal impact; when the same identification audio contains both difference information and low audio processing information, matching information and medium audio processing information, the identification audio will be marked as having a medium internal impact; when the same identification audio contains both matching information and low audio processing information, the identification audio will be marked as having a low internal impact.

[0027] In a preferred embodiment, the external influence data includes meteorological data and signal data, and the processing steps of the external influence data are:

[0028] The meteorological data includes temperature, wind speed and light intensity;

[0029] The metadata in the audio is retrieved and identified through the audio comprehensive analysis model. The recording date s is retrieved from the metadata in the audio. The temperature t, wind speed w and light intensity g of the three days of recording dates s+1, s, and s-1 are retrieved through the network to obtain the meteorological factor α.

[0030] The signal data includes the proportion of buildings using sound-absorbing materials in the surrounding environment and surrounding electromagnetic signal data;

[0031] The recording location input by the recorder is analyzed, and signal data within a range of one kilometer but not limited to one kilometer is collected using the network terminal at the input recording location. The signal data includes the proportion k of buildings using sound-absorbing materials in the surrounding environment and the surrounding electromagnetic signal data h. The signal factor δ is obtained through formulaic analysis.

[0032] In a preferred embodiment, the external factor impact identifier includes an external factor high impact identifier, an external factor medium impact identifier, and an external factor low impact identifier. The generation logic of the external factor impact identifier is:

[0033] In the audio comprehensive analysis model, the meteorological factor α and the signal factor δ are weighted and analyzed to obtain the external influence factor ξ;

[0034] Set external influence factor comparison thresholds ζ1 and ζ2, where the external influence factor comparison threshold ζ1 is greater than the external influence factor comparison threshold ζ2, and substitute the external influence factor ζ into the external influence factor comparison thresholds ζ1 and ζ2 for analysis. When the external influence factor ζ is greater than 0 and the external influence factor ζ is less than the external influence factor comparison threshold ζ2, generate an external factor low influence mark for the identification audio; when the external influence factor ζ is greater than the external influence factor comparison threshold ζ2 and the external influence factor ζ is less than the external influence factor comparison threshold ζ1, generate an external factor moderate influence mark for the identification audio; when the external influence factor ζ is greater than the external influence factor comparison threshold ζ1, generate an external factor high influence mark for the identification audio.

[0035] In a preferred embodiment, the matching targets include real matching targets, questionable matching targets, and false matching targets. The steps for generating matching targets for the identification audio are as follows:

[0036] The audio comprehensive analysis model is used to integrate statistics and analyze the internal and external influence indicators generated in the identification audio. The analysis logic is as follows:

[0037] When the same appraisal audio contains both the internal cause low impact flag and the external cause low impact flag, the appraisal audio is marked as a true matching target; when the same appraisal audio contains both the internal cause low impact flag and the external cause medium impact flag, the internal cause medium impact flag and the external cause low impact flag, or the internal cause medium impact flag and the external cause medium impact flag, the appraisal audio is marked as a doubtful matching target; when the same appraisal audio contains both the internal cause low impact flag and the external cause high impact flag, the internal cause high impact flag and the external cause low impact flag, the internal cause medium impact flag and the external cause high impact flag, the internal cause high impact flag and the external cause medium impact flag, or the internal cause high impact flag and the external cause high impact flag, the appraisal audio is marked as a false matching target.

[0038] In the above technical solution, the technical effects and advantages provided by the present invention are:

[0039] By constructing a comprehensive audio analysis model and combining meteorological data, signal data, audio equipment brand and recording quality for combined analysis, multi-dimensional information processing is achieved. This takes into account multiple factors, including meteorological data, signal data, equipment brand and its adapted equipment signal, as well as recording quality. These factors provide rich information for a more comprehensive assessment of the authenticity of recorded audio.

[0040] By comprehensively analyzing the combination of different factors, the authenticity of audio can be judged more accurately based on environmental influencing factors. Based on scientific principles and statistical methods, the identification process is more objective and reliable. By comprehensively considering multiple factors, the influence of subjective factors on the identification results can be reduced.

[0041] Corresponding solutions can be provided based on the final identification results. For example, if the audio is determined to be forged based on the analysis results, further investigation measures or technical means can be provided to ensure the authenticity of the audio. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction to the drawings required for use in the embodiments will be given below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0043] Figure 1 The present invention is a flowchart of an audio authentication method with a collection signal analysis function. DETAILED DESCRIPTION

[0044] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention:

[0045] See also Figure 1 As shown, the audio authentication method with the acquisition signal analysis function described in this embodiment includes the following steps:

[0046] Step S1: construct an audio comprehensive analysis model, transmit the identification audio, identification audio recording device data and recording location data;

[0047] Step S2, extracting audio data from the identification audio;

[0048] Step S3: Analyze and process the audio data to generate audio data processing information, and compare the audio recording device data to generate device matching information;

[0049] Step S4: Perform an integrated analysis based on the audio data processing information and the device matching information to generate an internal factor impact identifier;

[0050] Step S5: Analyze the recording location data, obtain the external factor influencing data of the recording period through the network, analyze and process it, and generate an external factor influencing identifier;

[0051] Step S6: Perform a combined analysis on the internal influence identifier and the external influence identifier, and generate a matching target for the identification audio.

[0052] It should be explained that the signals in the acquisition signal analysis function include data signals, environmental signals, and device signals;

[0053] The audio comprehensive analysis model includes an audio data processing model, an audio recording device data matching model, a network meteorological model, and a periodic signal model. The specific construction steps of the above four models are as follows:

[0054] Audio data processing model:

[0055] Collect multiple audio samples, including authentic and fake audio data, and preprocess the audio samples, including denoising, noise reduction, equalization and other technologies to eliminate noise and improve audio quality. Extract features from the preprocessed audio samples, such as the proportion of mixing duration, the proportion of spectrum smoothing duration, and the proportion of abnormal spectrum duration. According to the requirements of specific tasks, select appropriate features for subsequent analysis. Generate audio data processing information by analyzing and processing the extracted features.

[0056] Audio recording device data matching model:

[0057] Collect relevant data of various audio recording devices, including device type, specifications, recording characteristics, etc., clean and preprocess the collected device data to ensure data accuracy and consistency, extract features related to device brand, device signal, etc. from the device data, and generate device matching information, including device brand, device signal, etc., by comparing the features extracted from the audio data with the features in the device data.

[0058] Network meteorological model: This model collects meteorological data related to the audio recording location, such as temperature, wind speed, and light intensity. It cleans and preprocesses the collected meteorological data to ensure its accuracy and consistency. It then extracts features related to audio recording from the meteorological data. By analyzing the meteorological data at the audio recording location, it derives features related to audio quality, thereby analyzing the impact of weather conditions on audio clarity. The specific algorithm involved uses random forest, a regression and classification algorithm used to analyze the relationship between meteorological factors and audio quality.

[0059] Periodic signal model: Collects periodic signal data related to the audio recording date and time, such as a period of about two days before and after the date. The collected signal data is cleaned and preprocessed to ensure data accuracy. Signal features related to the audio recording date and time are extracted from the signal data, such as the percentage of buildings using sound-absorbing materials in the surrounding environment and surrounding electromagnetic signal data. By analyzing the percentage of buildings using sound-absorbing materials in the surrounding environment and the surrounding electromagnetic signal data, features related to audio quality are derived, such as the impact of the surrounding electromagnetic signals on recording quality.

[0060] The audio data includes the mixing duration ratio, the spectrum smoothing duration ratio, and the abnormal spectrum duration ratio. A larger mixing duration ratio indicates a greater probability of modification in the identified audio, while the spectrum smoothing duration ratio indicates the proportion of time when the spectrum is smooth and continuous, and the abnormal spectrum duration ratio indicates the proportion of time when abnormal peaks in the spectrum exist.

[0061] The audio recording device data includes the device brand and device signal. The device brand refers to the manufacturer of the device, such as Shure, Rode, Zoom, Tascam, etc., and the device signal refers to the average signal value when the recording device is used in a specific area.

[0062] It should be noted that: the greater the proportion of mixing time, the greater the probability of modification in the identified audio. The spectrum smoothness reflects the stability and continuity of the audio. The real audio has a smoother waveform in the time domain. It can be inferred that the greater the proportion of spectrum smoothing time, the higher the fluency and smoothness of the identified audio. Forged audio may have sudden changes or discontinuous parts. The more abnormal peaks in the spectrum, the longer the proportion of the duration of the abnormal peaks in the spectrum, which indicates a greater degree of audio forgery.

[0063] The steps for generating audio data processing information are:

[0064] The audio data processing information includes low-level audio processing information, medium-level audio processing information and high-level audio processing information;

[0065] Obtain the mixing duration ratio b, spectrum smoothing duration ratio n, and abnormal spectrum duration ratio m in the identified audio. Through formulaic analysis, calculate the audio influence factor γ. The specific formula is:

[0066]

[0067] Set audio impact reference numbers γ1 and γ2, where γ1<γ2, and substitute the audio impact factor γ into the audio impact reference numbers γ1 and γ2 for comparative analysis. When the audio impact factor γ is greater than 0 and less than the audio impact reference number γ1, high audio processing information is generated for the identification audio; when the audio impact factor γ is greater than the audio impact reference number γ1 and less than the audio impact reference number γ2, medium audio processing information is generated for the identification audio; when the audio impact factor γ is greater than the audio impact reference number γ2, low audio processing information is generated for the identification audio.

[0068] It should be noted that the degree of modification and forgery in the identified audio corresponding to the highly processed audio information is higher than that in the moderately processed audio information. In addition, this method can also be used to conduct a preliminary analysis of the audio quality.

[0069] The logic for generating device matching information is as follows:

[0070] The device matching information includes matching information and difference information;

[0071] Set the audio reference brand and audio signal comparison range in the audio comprehensive analysis model. The audio reference brands include but are not limited to X1, X2, and X3. The audio signal comparison range corresponding to the audio reference brand X1 in the audio comprehensive analysis model is The audio signal comparison range of the audio reference brand X2 in the audio comprehensive analysis model is The audio signal comparison range of the audio reference brand X3 in the audio comprehensive analysis model is Wherein, Y1, Y2, Y3, Y4, Y5, and Y6 are audio reference signals preset in the comprehensive audio analysis model. The audio reference signals Y1, Y2, Y3, Y4, Y5, and Y6 are all greater than 0, but the audio reference signal Y1 is less than the audio reference signal Y2, the audio reference signal Y3 is less than the audio reference signal Y4, and the audio reference signal Y5 is less than the audio reference signal Y6.

[0072] The device brand X and device signal Y in the input audio recording device data are substituted into the audio signal comparison range for classification and the device brand result X is generated. Y , set the device brand result to X Y Perform matching analysis with the input device brand X. If the brand result is X Y The brand X of the input device is matched to the same brand, and the matching information is generated for the identification audio. If the brand result is X Y If the brand matched with the input device brand X is not the same brand, difference information is generated for the authentication audio.

[0073] It should be noted that the values ​​of the audio signals Y1, Y2, Y3, Y4, Y5 and Y6 may be the same or different;

[0074] The difference information indicates that the recorded information may have been recorded through other devices and subsequently transferred to this recorded audio, or directly downloaded from the Internet to this recorded audio for processing. Compared with the matching information, the identification audio with difference information is more likely to be forged and stolen.

[0075] The internal factor impact identification includes an internal factor high impact identification, an internal factor medium impact identification, and an internal factor low impact identification. The steps for generating the internal factor impact identification are as follows:

[0076] When the same identification audio contains both difference information and high-level audio processing information, difference information and medium-level audio processing information, or matching information and high-level audio processing information, the identification audio will be marked as having a high internal impact; when the same identification audio contains both difference information and low-level audio processing information, matching information and medium-level audio processing information, the identification audio will be marked as having a medium internal impact; when the same identification audio contains both matching information and low-level audio processing information, the identification audio will be marked as having a low internal impact.

[0077] It should be noted that: compared with the identification audio with the internal cause medium impact label, the identification audio with the internal cause low impact label is more authentic and has better audio quality, and so on.

[0078] The external factor impact data includes meteorological data and signal data, and the processing steps of the external factor impact data are as follows:

[0079] The meteorological data includes temperature, wind speed and light intensity;

[0080] The metadata in the audio is retrieved through the audio comprehensive analysis model, and the recording date s is retrieved from the metadata in the audio. The temperature t, wind speed w and light intensity g of the three days of recording dates s+1, s, and s-1 are retrieved through the network to obtain the meteorological factor α. The specific formula is: s-1 *g s-1 *W s-1

[0081] α=(t s-1 *g s-1 *W s-1 +t s *g s *W s +t s+1 *g s+1 *w s+1 ) / 3 (α>0)

[0082] It should be noted that:

[0083] The greater the wind force w, the stronger the wind and the greater the impact of wind noise. The greater the light intensity g and temperature t, the greater the thermal noise in the recording equipment, thereby affecting the recording quality of the recording equipment. In addition, high temperature will also affect the normal operation of the components in the recording equipment, thereby affecting the recording quality. From this analysis, the larger the meteorological factor α, the greater the environmental impact of audio identification.

[0084] The signal data includes the proportion of buildings using sound-absorbing materials in the surrounding environment and surrounding electromagnetic signal data;

[0085] The recording location input by the recorder is analyzed, and signal data within a range of one kilometer (but not limited to one kilometer) is collected using the network terminal at the input recording location. The signal data includes the proportion of buildings using sound-absorbing materials in the surrounding environment (k) and the surrounding electromagnetic signal data (h). The signal factor δ is obtained through formula analysis. The calculation formula of the signal factor δ is:

[0086]

[0087] It should be noted that:

[0088] The larger the signal factor δ is, the worse the surrounding environmental noise processing is and the greater the impact of the surrounding electromagnetic signals is.

[0089] The external factor impact identifier includes a high external factor impact identifier, a medium external factor impact identifier, and a low external factor impact identifier. The generation logic of the external factor impact identifier is:

[0090] In the audio comprehensive analysis model, the meteorological factor α and the signal factor δ are weighted and analyzed to obtain the external influence factor ζ, which is obtained by the following formula:

[0091] ζ=α*p1+δ*p2 (ζ>0, and p1+p2=1.296, p1 and p2 are both greater than 0)

[0092] Set external influence factor comparison thresholds ζ1 and ζ2, where the external influence factor comparison threshold ζ1 is greater than the external influence factor comparison threshold ζ2, and substitute the external influence factor ζ into the external influence factor comparison thresholds ζ1 and ζ2 for analysis. When the external influence factor ζ is greater than 0 and the external influence factor ζ is less than the external influence factor comparison threshold ζ2, generate an external factor low influence mark for the identification audio; when the external influence factor ζ is greater than the external influence factor comparison threshold ζ2 and the external influence factor ζ is less than the external influence factor comparison threshold ζ1, generate an external factor moderate influence mark for the identification audio; when the external influence factor ζ is greater than the external influence factor comparison threshold ζ1, generate an external factor high influence mark for the identification audio.

[0093] It should be noted that: compared with the identification audio with the external cause medium impact mark, the identification audio with the external cause low impact mark has higher authenticity and better audio quality, and so on.

[0094] The matching targets include real matching targets, questionable matching targets, and false matching targets. The steps for generating matching targets for the identification audio are as follows:

[0095] The audio comprehensive analysis model is used to integrate statistics and analyze the internal and external influence indicators generated in the identification audio. The analysis logic is as follows:

[0096] When the same appraisal audio contains both the internal cause low impact flag and the external cause low impact flag, the appraisal audio is marked as a true matching target; when the same appraisal audio contains both the internal cause low impact flag and the external cause medium impact flag, the internal cause medium impact flag and the external cause low impact flag, or the internal cause medium impact flag and the external cause medium impact flag, the appraisal audio is marked as a doubtful matching target; when the same appraisal audio contains both the internal cause low impact flag and the external cause high impact flag, the internal cause high impact flag and the external cause low impact flag, the internal cause medium impact flag and the external cause high impact flag, the internal cause high impact flag and the external cause medium impact flag, or the internal cause high impact flag and the external cause high impact flag, the appraisal audio is marked as a false matching target.

[0097] It should be noted that: compared with the identification audio of the questionable matching target, the identification audio of the false matching target has lower authenticity and worse audio quality, and the greater the sound impact of the surrounding environment, the greater the environmental impact, the more it can cover up the traces of forgery through the impact, and so on.

[0098] By constructing a comprehensive audio analysis model and combining meteorological data, signal data, audio equipment brand and recording quality for combined analysis, multi-dimensional information processing is achieved. This takes into account multiple factors, including meteorological data, signal data, equipment brand and its adapted equipment signal, as well as recording quality. These factors provide rich information for a more comprehensive assessment of the authenticity of recorded audio.

[0099] By comprehensively analyzing the combination of different factors, the authenticity of audio can be judged more accurately based on environmental influencing factors. Based on scientific principles and statistical methods, the identification process is more objective and reliable. By comprehensively considering multiple factors, the influence of subjective factors on the identification results can be reduced.

[0100] Corresponding solutions can be provided based on the final identification results. For example, if the audio is determined to be forged based on the analysis results, further investigation measures or technical means can be provided to ensure the authenticity of the audio.

[0101] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.

[0102] It should be understood that the term "and / or" as used herein simply describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. A and B can be singular or plural. Furthermore, the character " / " as used herein generally indicates an "or" relationship between the related objects, but it may also indicate an "and / or" relationship. For specific understanding, please refer to the context.

[0103] In this application, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.

[0104] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0105] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0106] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, and may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment as needed.

[0107] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0108] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. An audio authentication method with a signal acquisition and analysis function, characterized in that: The method comprises the following steps: Step S1: construct an audio comprehensive analysis model, transmit the identification audio, identification audio recording device data and recording location data; Step S2, extracting audio data from the identification audio; Step S3: Analyze and process the audio data to generate audio data processing information, and compare the audio recording device data to generate device matching information; The audio data processing information includes low-level audio processing information, medium-level audio processing information and high-level audio processing information; The device matching information includes matching information and difference information; Step S4: Perform an integrated analysis based on the audio data processing information and the device matching information to generate an internal factor impact identifier; The specific steps for generating the internal impact identification are as follows: When the same identification audio contains both difference information and high-level audio processing information, difference information and medium-level audio processing information, or matching information and high-level audio processing information, the identification audio will be marked as having a high internal impact; when the same identification audio contains both difference information and low-level audio processing information, matching information and medium-level audio processing information, the identification audio will be marked as having a medium internal impact; when the same identification audio contains both matching information and low-level audio processing information, the identification audio will be marked as having a low internal impact. Step S5: Analyze the recording location data, obtain the external factor influencing data of the recording period through the network, analyze and process it, and generate an external factor influencing identifier; The external factor impact data includes meteorological data and signal data, and the processing steps of the external factor impact data are as follows: The meteorological data includes temperature, wind speed and light intensity; The metadata in the audio is retrieved and identified through the audio comprehensive analysis model. The recording date s is retrieved from the metadata in the audio. The temperature t, wind speed w and light intensity g of the three days of recording dates s+1, s, and s-1 are retrieved through the network to obtain the meteorological factor α. The signal data includes the proportion of buildings using sound-absorbing materials in the surrounding environment and the surrounding electromagnetic signal data; the signal factor δ is obtained through formula analysis; The generation logic of the external factor impact identifier is: In the audio comprehensive analysis model, the meteorological factor α and the signal factor δ are weighted and analyzed to obtain the external influence factor ζ, which is obtained by the following formula: ζ=αp1+δp2 (ζ>0, and p1+p2=1.296, p1 and p2 are both greater than 0) Setting external influence factor comparison thresholds ζ1 and ζ2, wherein the external influence factor comparison threshold ζ1 is greater than the external influence factor comparison threshold ζ2, substituting the external influence factor ζ into the external influence factor comparison thresholds ζ1 and ζ2 for analysis, when the external influence factor ζ is greater than 0 and the external influence factor ζ is less than the external influence factor comparison threshold ζ2, generating an external factor low influence mark for the identification audio; when the external influence factor ζ is greater than the external influence factor comparison threshold ζ2 and the external influence factor ζ is less than the external influence factor comparison threshold ζ1, generating an external factor moderate influence mark for the identification audio; when the external influence factor ζ is greater than the external influence factor comparison threshold ζ1, generating an external factor high influence mark for the identification audio; Step S6: Perform a combined analysis on the internal influence identifier and the external influence identifier, and generate a matching target for the identification audio.

2. The audio authentication method with the function of collecting signal analysis according to claim 1, characterized in that: The audio data includes the mixing duration ratio, the spectrum smoothing duration ratio, and the abnormal spectrum duration ratio. A larger mixing duration ratio indicates a greater probability of modification in the identified audio, while the spectrum smoothing duration ratio indicates the proportion of time when the spectrum is smooth and continuous, and the abnormal spectrum duration ratio indicates the proportion of time when abnormal peaks in the spectrum exist. The audio recording device data includes a device brand and a device signal. The device brand is the brand of the device, and the device signal is the average signal value when the recording device is used in a specific area.

3. The audio authentication method with the function of collecting signal analysis according to claim 2, characterized in that: The steps for generating audio data processing information are: Obtain the mixing duration ratio b, spectrum smoothing duration ratio n, and abnormal spectrum duration ratio m in the identified audio, and calculate the audio impact factor γ through formula analysis; Set audio impact reference numbers γ1 and γ2, where γ1 < γ2, and substitute the audio impact factor γ into the audio impact reference numbers γ1 and γ2 for comparative analysis. When the audio impact factor γ is greater than 0 and less than the audio impact reference number γ1, high audio processing information is generated for the identification audio; when the audio impact factor γ is greater than the audio impact reference number γ1 and less than the audio impact reference number γ2, medium audio processing information is generated for the identification audio; when the audio impact factor γ is greater than the audio impact reference number γ2, low audio processing information is generated for the identification audio.

4. The audio authentication method with the function of collecting and analyzing signals according to claim 3, characterized in that: The logic for generating device matching information is as follows: An audio reference brand and an audio signal comparison range are set in the audio comprehensive analysis model, where the audio reference brands include but are not limited to X1, X2, and X3. The audio signal comparison range corresponding to the audio reference brand X1 in the audio comprehensive analysis model is Y1∽Y2, the audio signal comparison range corresponding to the audio reference brand X2 in the audio comprehensive analysis model is Y3∽Y4, and the audio signal comparison range corresponding to the audio reference brand X3 in the audio comprehensive analysis model is Y5∽Y6, where Y1, Y2, Y3, Y4, Y5, and Y6 are audio reference signals preset in the audio comprehensive analysis model, and the audio reference signals Y1, Y2, Y3, Y4, Y5, and Y6 are all greater than 0, but the audio reference signal Y1 is less than the audio reference signal Y2, the audio reference signal Y3 is less than the audio reference signal Y4, and the audio reference signal Y5 is less than the audio reference signal Y6; The device brand X and device signal Y in the input audio recording device data are substituted into the audio signal comparison range for classification and a device brand result XY is generated. The device brand result XY is matched and analyzed with the input device brand X. If the brand result XY matches the input device brand X and is the same brand, matching information is generated for the identification audio. If the brand result XY does not match the input device brand X and is the same brand, difference information is generated for the identification audio.

5. The audio authentication method with the function of collecting signal analysis according to claim 4, characterized in that: The matching targets include real matching targets, questionable matching targets, and false matching targets. The steps for generating matching targets for the identification audio are as follows: The audio comprehensive analysis model is used to integrate statistics and analyze the internal and external influence indicators generated in the identification audio. The analysis logic is as follows: When the same appraisal audio contains both the internal cause low impact flag and the external cause low impact flag, the appraisal audio is marked as a true matching target; when the same appraisal audio contains both the internal cause low impact flag and the external cause medium impact flag, the internal cause medium impact flag and the external cause low impact flag, or the internal cause medium impact flag and the external cause medium impact flag, the appraisal audio is marked as a doubtful matching target; when the same appraisal audio contains both the internal cause low impact flag and the external cause high impact flag, the internal cause high impact flag and the external cause low impact flag, the internal cause medium impact flag and the external cause high impact flag, the internal cause high impact flag and the external cause medium impact flag, or the internal cause high impact flag and the external cause high impact flag, the appraisal audio is marked as a false matching target.

Citation Information

Patent Citations

  • Audio quality authentication method and device

    CN107564534A

  • Electronic adapter unit for selectively modifying audio or video data for use with an output device

    CN102884797A

  • Multi-scale sub-band energy set feature-based muzzle wave recognition method

    CN108269566A