A counterfeit detection analysis system with advantages in analyzing text content
Patent Information
- Application Number
- CN202310814449.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-04
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2043-07-04
AI Technical Summary
[0004]1.对音频中的内容分析过程过于单一,具体过程为对音频内容进行分类,而内容优势无法进行判定分析;
[0046] By directly analyzing the audio and then reanalyzing it after it has been converted to text, the system performs auditory judgment on the audio and text judgment on the converted audio, thus achieving a two-way function for judging the advantages of the audio content. The system also correlates the two and performs further integrated analysis, which enhances the logic of the analysis and the rigor of the judgment of the advantages of the audio content. In the process of judging the advantages of the audio content, the system also realizes the function of sorting out the advantages of the text content.
Smart Images

Figure CN116825091B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio content authentication technology, specifically to a counterfeit detection analysis system with the advantage of analyzing text content. Background Technology
[0002] The end-to-end text-to-speech conversion scheme proposed in CN110476206A is used to generate speech at the frame level. The system described in this invention can generate speech from text faster than other systems, while generating speech with comparable or even better quality. In addition, this system can reduce model size, training time and inference time, and can also significantly improve convergence speed. The system described in this invention can also generate high-quality speech without the need for manually designed language features or complex components, such as hidden Markov model aligners, reducing complexity and using less computational resources, while generating high-quality speech.
[0003] Based on the above solutions and in conjunction with existing technologies, the following shortcomings were found in the existing technologies:
[0004] 1. The content analysis process in audio is too simplistic, specifically classifying audio content without analyzing or determining its strengths.
[0005] 2. In the process of audio content analysis, there is a certain error in identifying whether the advantageous content mentioned in the audio is true information. That is, there is a difference between the content advantage judgment when the audio is heard by the human ear and the content advantage judgment when the audio is converted into text for text analysis.
[0006] To address the aforementioned issues, a counterfeit detection analysis system with the advantage of analyzing text content is proposed. Summary of the Invention
[0007] The purpose of this invention is to provide a counterfeit detection and analysis system with the advantage of sorting out text content, so as to solve the shortcomings of the background technology.
[0008] To achieve the above objectives, the present invention provides the following technical solution: the anti-counterfeiting analysis system with the advantage of sorting out text content includes the following modules:
[0009] The audio input module is used to input audio content;
[0010] The audio content analysis module is used to preprocess and process the input audio content to generate speech content coefficients that can analyze the advantages of the audio content.
[0011] The audio content transcription module is used to transcribe the input audio content into text and generate transcribed text content.
[0012] Text content feature extraction module: used to extract and compare relevant features of transcribed text content, and generate text content coefficients to process the advantages of text content.
[0013] Model Analysis Module: Used to perform model analysis on speech content coefficients and text content coefficients, and generate comparison coefficients for the correlation between audio content and transcribed text content;
[0014] The authenticity marking module is used to compare the comparison coefficients by threshold, generate authenticity markings for advantageous content, identify the advantageous content of the matched audio, and output real content objects and fake content objects.
[0015] In a preferred embodiment, the audio content processing module is specifically an audio analysis platform, and the speech content coefficients include the superior sentence fluency coefficient α and the superior sentence voice coefficient β.
[0016] In a preferred embodiment, the steps for generating the fluency coefficient α of the advantageous content segment are as follows:
[0017] The audio content is identified using a speech recognition tool, and a recognition result is generated. The recognition result includes the frequency X of the advantageous keywords, the total number of words M of the phrases containing the keywords, and the total duration S of the phrases containing the keywords.
[0018] The fluency of the audio content when the keyword appears is determined by the quotient between the total number of words M in the phrase description containing the keyword and the total duration S of the phrase description containing the keyword. The fluency coefficient α of the sentence segment is then obtained by formulaic processing of the quotient and the frequency X of the advantageous keyword.
[0019] A larger α indicates a higher fluency of narration when dominant phrases appear in the audio, and vice versa.
[0020] In a preferred embodiment, the steps for generating the voice coefficient β of the dominant content segment are as follows:
[0021] By analyzing audio content using audio processing tools and keywords, sentences in a paragraph are categorized into dominant and non-dominant sentences, and the average pitch of dominant sentences is obtained. and the average pitch of the whole sentence Among them, the mean pitch of the dominant sentence This corresponds to the average pitch of the paragraphs related to the content advantage, while the average pitch of the entire sentence... The average pitch of the entire sentence segment is used to derive the voice coefficient β of the advantageous sentence segment through formulaic analysis.
[0022] The larger the β value, the greater the pitch change when the dominant phrase appears in the audio, and vice versa. This indirectly indicates the pitch change when elaborating on the dominant phrase and the range of voice changes when elaborating on the dominant content.
[0023] In a preferred embodiment, the speech content coefficient generation step is as follows:
[0024] The speech-to-text platform is used to integrate and analyze the fluency coefficient α and the voice coefficient β of the advantageous sentence segments in a formulaic way. Specifically, let the speech content coefficient be γ, and use the formula γ=α*N1+β*N2, where γ>0, N1+N2=0.8634, and both N1 and N2 are greater than 0.
[0025] The larger the speech content coefficient γ, the greater the change in voice and the higher the clarity when the superior audio content is displayed. The integration and analysis of the two indicate that the superior audio content is more authentic.
[0026] In a preferred embodiment, the relevant features extracted from the transcribed text content include word frequency H and word vector F;
[0027] The steps for generating the text content coefficients are as follows:
[0028] Relevant features extracted from transcribed text content include dominant keywords in paragraphs, keyword frequency (H), and word vectors.
[0029] The extracted word vectors are combined with the keywords of the advantageous content to determine the performance level, and the performance level is assigned a value, specifically K;
[0030] The process of determining the representation level of the extracted word vectors and assigning values to the representation levels specifically includes:
[0031] Based on the transcribed text content data, a vocabulary list is constructed and a word vector model is trained to obtain the word vector representation of each word or phrase;
[0032] The text content is divided into n identification regions according to each sentence. The evaluation criteria and indicators are determined, and the criteria and indicators for judging the level of advantage are defined. The specific criteria and indicators adopted are the word vector performance level. The word vector performance level represents the degree of content advantage between two sentences. The word vectors are assigned to the content advantage level in each sentence and classified into three levels: high, medium and low. Let the content advantage level in the first sentence be W1 and the content advantage level in the second sentence be W2. Then the degree of content advantage between the two sentences is W1-W2, that is, the word vector performance level is W1-W2.
[0033] If the word vector performance level W1-W2 is high-medium, high-high, or medium-high, it is defined as the first dominant performance level; if W1-W2 is medium-medium, it is defined as the second dominant performance level; if W1-W2 is low-medium, low-low, or medium-low, it is defined as the third dominant performance level.
[0034] The third advantage level is lower than the second advantage level, meaning it is judged as a disadvantage. Similarly, the second advantage level is considered normal content, and the first advantage level is considered advantageous content.
[0035] According to the rating rules, word vectors are assigned to the corresponding content advantage levels. By combining the content advantage levels of two consecutive sentences, the corresponding word vector performance levels are generated.
[0036] The word vector performance levels are assigned a level value K to generate word vector performance level values. The word vector performance level values K include K1, K2 and K3. Specifically, the word vector performance level of the first dominant performance level is assigned the value K1, the word vector performance level of the second dominant performance level is assigned the value K2, and the word vector performance level of the third dominant performance level is assigned the value K3, where K1 > K2 > K3 > 0.
[0037] Let the text content coefficient be δ. The text content coefficient δ is obtained by performing a formulaic analysis on the word frequency H and the word vector performance level K.
[0038] A larger δ value indicates a higher degree of authenticity in the identification of textual superiority content, while a smaller δ value indicates a lower degree of authenticity in the identification of textual superiority content.
[0039] In a preferred embodiment, the comparison coefficient generation step is as follows:
[0040] Let the comparison coefficient be ζ. The comparison coefficient ζ is generated by integrating the speech content coefficient γ and the text content coefficient δ through a linear regression model. The specific formula for generating the comparison coefficient ζ is as follows:
[0041] ζ = u1*γ + u2*δ, u1 and u2 are weighting factors, u1 < u2, u1 + u2 = 2.463, u1 and u2 represent the degree of composition of speech content and text content. The value of the weighting factor is determined by the different levels of understanding when acquiring text content and audio content.
[0042] In a preferred embodiment, the advantage content authenticity marker includes a fake advantage content marker and a genuine advantage content marker, and the generation process of the advantage content authenticity marker is as follows:
[0043] The comparison coefficient ζ is analyzed by comparing the true and false tagging models. A comparison threshold KH is set. KH is greater than 0. The comparison coefficient ζ is substituted into the comparison threshold KH for analysis and processing. If the comparison coefficient ζ is greater than the comparison threshold KH, a true advantageous content tag is generated; if the comparison coefficient ζ is less than the comparison threshold KH, a false advantageous content tag is generated.
[0044] Audio files matching the true advantage content tag are identified as true content objects and output, while audio files matching the false advantage content tag are identified as false content objects and output.
[0045] The technical effects and advantages provided by the present invention in the above technical solution are as follows:
[0046] By directly analyzing the audio and then reanalyzing it after it has been converted to text, the system performs auditory judgment on the audio and text judgment on the converted audio, thus achieving a two-way function for judging the advantages of the audio content. The system also correlates the two and performs further integrated analysis, which enhances the logic of the analysis and the rigor of the judgment of the advantages of the audio content. In the process of judging the advantages of the audio content, the system also realizes the function of sorting out the advantages of the text content. Attached Figure Description
[0047] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0048] Figure 1 This is a flowchart of a counterfeit detection analysis system with the advantage of sorting out text content, according to the present invention. Detailed Implementation
[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0050] Please see Figure 1 As shown in this embodiment, a counterfeit detection analysis system with the advantage of analyzing text content is described. This system includes the following modules:
[0051] The audio input module is used to input audio content;
[0052] The audio content analysis module is used to preprocess and process the input audio content to generate speech content coefficients that can analyze the advantages of the audio content.
[0053] The audio content transcription module is used to transcribe the input audio content into text and generate transcribed text content.
[0054] Text content feature extraction module: used to extract and compare relevant features of transcribed text content, and generate text content coefficients to process the advantages of text content.
[0055] Model Analysis Module: Used to perform model analysis on speech content coefficients and text content coefficients, and generate comparison coefficients for the correlation between audio content and transcribed text content;
[0056] The authenticity marking module is used to compare the comparison coefficients by threshold, generate authenticity markings for advantageous content, identify the advantageous content of the matched audio, and output real content objects and fake content objects.
[0057] The audio content processing module is specifically an audio analysis platform, and the speech content coefficients include the fluency coefficient α and the voice coefficient β of the dominant content segments.
[0058] It is important to note that the audio input module is specifically a microphone device, the audio content processing module is specifically an audio analysis platform, and the audio analysis platform specifically includes speech recognition tools and audio processing tools.
[0059] The preprocessing process includes steps such as audio sampling rate conversion, noise reduction, and noise removal. Audio preprocessing can be done using audio editing software, such as Adobe Audition, Audacity, and GarageBand.
[0060] The steps for generating the fluency coefficient α of the advantageous content segment are as follows:
[0061] The audio content is identified using a speech recognition tool, and a recognition result is generated. The recognition result includes the frequency X of the advantageous keywords, the total number of words M of the phrases containing the keywords, and the total duration S of the phrases containing the keywords.
[0062] The fluency of the audio content when the keyword appears is determined by the quotient between the total number of words M in the phrase description containing the keyword and the total duration S of the phrase description containing the keyword. The fluency coefficient α of the sentence segment is then obtained by formulaic processing of the quotient and the frequency X of the advantageous keyword.
[0063] (p is the fluency error correction constant, P>0, m, s and x are all greater than 0);
[0064] A larger α indicates a higher fluency of narration when dominant phrases appear in the audio, and vice versa.
[0065] It is important to note that speech recognition tools can transcribe audio into text, such as Google Cloud Speech-to-Text, Microsoft Azure Speech to Text, and IBM Watson Speech to Text. These tools specifically employ deep neural network (DNN) models as a type of machine learning model. Through training, they can learn the mapping relationship between input data and output labels, and model the relationship between acoustic features and text labels to achieve speech-to-text conversion.
[0066] However, the ability to calculate the frequency X of advantageous keywords, the total number of words M in phrases containing keywords, and the total duration S of phrases containing keywords in speech recognition is not a function of the DNN model itself. Instead, it requires additional processing techniques on top of the DNN model. The following are the methods and techniques for obtaining the corresponding data:
[0067] Advantages: Keyword frequency X: After obtaining text from the speech recognition results, text processing techniques can be used to count the frequency of specific keywords in the audio. Common methods include using regular expressions, string matching, and other techniques to search for and count the number of times keywords appear.
[0068] Total word count M of phrases containing keywords: By segmenting the recognized speech, the total word count of phrases containing keywords can be counted. Segmentation techniques can use traditional rule-based or statistical methods, or modern deep learning models. Total duration S of phrases containing keywords: Segmenting is performed using speech processing technology, and the duration of words or syllables in phrases containing keywords is estimated based on the sampling rate and frame rate of the speech signal. Then, the duration of phrases containing keywords is obtained by summing the durations of words in the phrases or sentences.
[0069] The steps for generating the voice coefficient β of the advantageous content segment are as follows:
[0070] By analyzing audio content using audio processing tools and keywords, sentences in a paragraph are categorized into dominant and non-dominant sentences, and the average pitch of dominant sentences is obtained. and the average pitch of the whole sentence Among them, the mean pitch of the dominant sentence This corresponds to the average pitch of the paragraphs related to the content advantage, while the average pitch of the entire sentence... The average pitch of the entire sentence segment is used to derive the voice coefficient β of the advantageous sentence segment through formulaic analysis.
[0071] (L is the voice coefficient error correction constant) (and L are both greater than 0);
[0072] The larger the β value, the greater the pitch change when the dominant phrase appears in the audio, and vice versa. This indirectly indicates the pitch change when elaborating on the dominant phrase and the range of voice changes when elaborating on the dominant content.
[0073] Audio processing tools can analyze audio, including syntactic analysis and entity recognition. Specific tools include NLTK (Natural Language Toolkit), SpaCy, and Stanford CoreNLP, and the specific algorithms involved are as follows:
[0074] This involves combining algorithms or models for keyword selection, paragraph classification, and sentence pitch mean analysis. Below are some commonly used algorithms and models:
[0075] Keyword selection:
[0076] Deep learning models, specifically Convolutional Neural Networks (CNNs) or Long Short-Term Memory Networks (LSTMs), can be used to directly process audio and extract key information and features, including keywords. This method requires a large amount of labeled data for training.
[0077] Classify the sentences in the paragraph into dominant and non-dominant sentences:
[0078] Text classification models: These use machine learning or deep learning models, specifically Naive Bayes classifiers, support vector machines, or recurrent neural networks (RNNs), to classify audio segments into dominant and non-dominant sentences.
[0079] Mean pitch in dominant sentences and the average pitch of the whole sentence analyze:
[0080] Fundamental frequency extraction: Using algorithms such as autocorrelation method and HMM-based tone model, fundamental frequency or fundamental tone period information is extracted from audio.
[0081] Fundamental frequency mean calculation: The mean of the fundamental frequency extracted from each paragraph is calculated to obtain the average pitch of the paragraph.
[0082] Fundamental frequency pitch classifier: A pitch classification model is built, trained using training data, and classifies segments into different pitch types based on their mean pitch.
[0083] The steps for generating speech content coefficients are as follows:
[0084] The speech-to-text platform is used to integrate and analyze the fluency coefficient α and the voice coefficient β of the advantageous sentence segments in a formulaic way. Specifically, let the speech content coefficient be γ, and use the formula γ=α*N1+β*N2, where γ>0, N1+N2=0.8634, and both N1 and N2 are greater than 0.
[0085] The larger the speech content coefficient γ, the greater the change in voice and the higher the clarity when the superior audio content is displayed. The integration and analysis of the two indicate that the superior audio content is more authentic.
[0086] The relevant features extracted from the transcribed text content include word frequency H and word vector F;
[0087] The steps for generating the text content coefficients are as follows:
[0088] Relevant features extracted from transcribed text content include dominant keywords in paragraphs, keyword frequency (H), and word vectors.
[0089] The extracted word vectors are combined with the keywords of the advantageous content to determine the performance level, and the performance level is assigned a value, specifically K;
[0090] It should be noted that word frequency can be obtained using the bag-of-words model, specifically using word frequency statistics to calculate the frequency of each word in the document. Word vectors, on the other hand, use the pre-trained word vector model Word2Vec for feature extraction. Specifically, Word2Vec learns word vectors by training a neural network model, which includes two methods: the continuous bag-of-words model and the skip-word model.
[0091] Here's an example illustrating the process of combining advantageous content keywords with word vector analysis:
[0092] Suppose the text contains statements with the following two characteristics:
[0093] Statement A: "This plan is very interesting, rich in content, and captivating."
[0094] Statement B: "This plan is dull and empty of content; it's not worth reading."
[0095] The frequency of keywords in a sentence is calculated using a word vector model.
[0096] Suppose that the keywords used in the combined analysis are the following: ["interesting", "rich", "captivating", "boring", "empty", "not worth reading"].
[0097] Calculate the frequency of the keyword in the text. For example, if the keyword appears frequently in statement A, it indicates that statement A may have a higher advantage. Analyze other statements in the text in turn.
[0098] In text B, these keywords have a low frequency or are negative, indicating that text B may lack advantages or have negative evaluations.
[0099] Furthermore, word vectors can be used to calculate the similarity between keywords. For example, "interesting" and "captivating" may have a high similarity, while "boring" and "empty" may have a high similarity. By calculating the similarity between keywords, the strengths of the text can be further analyzed.
[0100] The analysis yielded the following conclusions regarding the strengths of the text content: Text A may be considered advantageous because it contains high-frequency keywords, has a high degree of similarity between keywords, and has a more positive sentiment polarity. Conversely, Text B may be considered lacking in advantage because it contains low-frequency or negative keywords.
[0101] The process of determining the representation level of the extracted word vectors and assigning values to the representation levels specifically includes:
[0102] Based on the transcribed text content data, a vocabulary list is constructed and a word vector model is trained to obtain the word vector representation of each word or phrase;
[0103] The text content is divided into n identification regions according to each sentence. The evaluation criteria and indicators are determined, and the criteria and indicators for judging the level of advantage are defined. The specific criteria and indicators adopted are the word vector performance level. The word vector performance level represents the degree of content advantage between two sentences. The word vectors are assigned to the content advantage level in each sentence and classified into three levels: high, medium and low. Let the content advantage level in the first sentence be W1 and the content advantage level in the second sentence be W2. Then the degree of content advantage between the two sentences is W1-W2, that is, the word vector performance level is W1-W2.
[0104] If the word vector performance level W1-W2 is high-medium, high-high, or medium-high, it is defined as the first dominant performance level; if W1-W2 is medium-medium, it is defined as the second dominant performance level; if W1-W2 is low-medium, low-low, or medium-low, it is defined as the third dominant performance level.
[0105] The third advantage level is lower than the second advantage level, meaning it is judged as a disadvantage. Similarly, the second advantage level is considered normal content, and the first advantage level is considered advantageous content.
[0106] According to the rating rules, word vectors are assigned to the corresponding content advantage levels. By combining the content advantage levels of two consecutive sentences, the corresponding word vector performance levels are generated.
[0107] The word vector performance levels are assigned a level value K to generate word vector performance level values. The word vector performance level values K include K1, K2 and K3. Specifically, the word vector performance level of the first dominant performance level is assigned the value K1, the word vector performance level of the second dominant performance level is assigned the value K2, and the word vector performance level of the third dominant performance level is assigned the value K3, where K1 > K2 > K3 > 0.
[0108] Let the text content coefficient be δ. By performing a formulaic analysis on the word frequency H and the word vector representation level K, the text content coefficient δ is obtained. Specifically;
[0109] M is the text content coefficient correction constant, M>0, n>0;
[0110] A larger δ value indicates a higher degree of authenticity in the identification of textual superiority content, while a smaller δ value indicates a lower degree of authenticity in the identification of textual superiority content.
[0111] The comparison coefficient generation step is as follows:
[0112] Let the comparison coefficient be ζ. The comparison coefficient ζ is generated by integrating the speech content coefficient γ and the text content coefficient δ through a linear regression model. The specific formula for generating the comparison coefficient ζ is as follows:
[0113] ζ=u1*γ+u2*δ, u1 and u2 are weighting factors, u1<u2, u1+u2=2.463, u1 and u2 represent the degree of judgment of speech content and text content. The value of the weighting factor is determined by the different levels of understanding when acquiring text content and audio content.
[0114] It should be noted that the larger the comparison coefficient ζ, the higher the accuracy of the comprehensive analysis of the audio's advantages, and vice versa. The linear regression model specifically uses the normal equation algorithm, which directly calculates the closed-form solution of the regression coefficient by solving the inverse of the matrix.
[0115] The authenticity markers for advantageous content include markers for false advantageous content and markers for genuine advantageous content. The generation process for these markers is as follows:
[0116] The comparison coefficient ζ is analyzed by comparing the true and false tagging models. A comparison threshold KH is set. KH is greater than 0. The comparison coefficient ζ is substituted into the comparison threshold KH for analysis and processing. If the comparison coefficient ζ is greater than the comparison threshold KH, a true advantageous content tag is generated; if the comparison coefficient ζ is less than the comparison threshold KH, a false advantageous content tag is generated.
[0117] Audio files matching the true advantage content tag are identified as true content objects and output, while audio files matching the false advantage content tag are identified as false content objects and output.
[0118] It should be noted that the true / false labeling model specifically uses the SVM model to analyze and process different coefficients. During the training process, the trained SVM model and the set threshold are used to label the audio targets in the test set. The basic algorithms involved include kernel functions, optimization algorithms, soft margins and regularization, and decision functions.
[0119] By directly analyzing the audio and then reanalyzing it after it has been converted to text, the system performs auditory judgment on the audio and text judgment on the converted audio, thus achieving a two-way function for judging the advantages of the audio content. The system also correlates the two and performs further integrated analysis, which enhances the logic of the analysis and the rigor of the judgment of the advantages of the audio content. In the process of judging the advantages of the audio content, the system also realizes the function of sorting out the advantages of the text content.
[0120] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0121] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.
[0122] In this application, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0123] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0124] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0125] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0126] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0127] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A counterfeit detection analysis system with the advantage of analyzing text content, characterized in that, The anti-counterfeiting analysis system with the advantage of analyzing text content includes the following modules: The audio input module is used to input audio content; The audio content analysis module is used to preprocess and process the input audio content to generate speech content coefficients that can analyze the advantages of the audio content. The audio content transcription module is used to transcribe the input audio content into text and generate transcribed text content. Text content feature extraction module: used to extract and compare relevant features of transcribed text content, and generate text content coefficients to optimize the text content; including: The relevant features extracted from the transcribed text content include dominant keywords in the paragraph, keyword frequency H, and word vector F; The steps for generating the text content coefficients are as follows: The extracted word vectors are combined with the keywords of the advantageous content to determine the performance level, and the performance level is assigned a value, specifically K; The process of determining the representation level of the extracted word vectors and assigning values to the representation levels specifically includes: Based on the transcribed text content data, a vocabulary list is constructed and a word vector model is trained to obtain the word vector representation of each word or phrase; The text content is divided into n identification regions according to each sentence. The standards and indicators for judging the level of advantage are defined. The specific standard and indicator adopted is the word vector performance level. The word vector performance level represents the degree of content advantage between two sentences. The word vectors are assigned to the content advantage level in each sentence and classified into three levels: high, medium and low. Let the content advantage level in the first sentence be W1 and the content advantage level in the second sentence be W2. Then the degree of content advantage between the two sentences is W1-W2, that is, the word vector performance level is W1-W2. If the word vector performance level W1-W2 is high-medium, high-high, or medium-high, it is defined as the first dominant performance level; if W1-W2 is medium-medium, it is defined as the second dominant performance level; if W1-W2 is low-medium, low-low, or medium-low, it is defined as the third dominant performance level. The third advantage level is lower than the second advantage level, meaning it is judged as a disadvantage. Similarly, the second advantage level is considered normal content, and the first advantage level is considered advantageous content. According to the rating rules, word vectors are assigned to the corresponding content advantage levels. Combining the content advantage levels of two consecutive sentences, the corresponding word vector performance levels are generated. The word vector performance levels are assigned a level value K to generate word vector performance level values. The word vector performance level values K include K1, K2 and K3. Specifically, the word vector performance level of the first dominant performance level is assigned the value K1, the word vector performance level of the second dominant performance level is assigned the value K2, and the word vector performance level of the third dominant performance level is assigned the value K3, where K1 > K2 > K3 > 0. Let the text content coefficient be δ. The word frequency H and the word vector performance level K are processed by formula analysis to obtain the text content coefficient δ. Model Analysis Module: Used to perform model analysis on speech content coefficients and text content coefficients, and generate comparison coefficients for the correlation between audio content and transcribed text content; The authenticity tagging module is used to perform threshold comparison on the comparison coefficients, generate authenticity tags for the advantageous content, generate authenticity tags for the matched audio, and output real content objects and fake content objects.
2. The anti-counterfeiting analysis system with the advantage of sorting out text content as described in claim 1, characterized in that, The audio content processing module is specifically an audio analysis platform, and the speech content coefficients include the fluency coefficient α and the voice coefficient β of the dominant content segments.
3. The anti-counterfeiting analysis system with the advantage of sorting out text content as described in claim 2, characterized in that, The steps for generating the fluency coefficient α of the advantageous content segment are as follows: The audio content is identified using a speech recognition tool, and a recognition result is generated. The recognition result includes the frequency X of the advantageous keywords, the total number of words M of the phrases containing the keywords, and the total duration S of the phrases containing the keywords. The fluency of the audio content when the keyword appears is determined by the quotient between the total number of words M in the phrase description containing the keyword and the total duration S of the phrase description containing the keyword. The fluency coefficient α of the advantageous content segment is then obtained by formulaic processing of the quotient and the frequency X of the advantageous keyword. A larger α indicates a higher fluency of narration when dominant phrases appear in the audio, and vice versa.
4. The anti-counterfeiting analysis system with the advantage of sorting out text content as described in claim 3, characterized in that, The steps for generating the voice coefficient β of the advantageous content segment are as follows: By analyzing audio content using audio processing tools and keywords, sentences in a paragraph are categorized into dominant and non-dominant sentences, and the average pitch of dominant sentences is obtained. and the average pitch of the whole sentence Among them, the average pitch of the dominant sentence This corresponds to the average pitch of the sentences that dominate the paragraph, while the average pitch of the entire sentence is... The average pitch of the entire paragraph is used to derive the voice coefficient β of the dominant content segments through formulaic analysis. The larger β is, the greater the pitch change value in the audio when the dominant phrase appears, and vice versa. This indirectly represents the range of voice changes when the dominant content is elaborated.
5. The anti-counterfeiting analysis system with the advantage of sorting out text content as described in claim 4, characterized in that, The steps for generating speech content coefficients are as follows: The speech-to-text platform uses a formulaic integration analysis to integrate the fluency coefficient α and the voice coefficient β of the advantageous content segments. Specifically, let the speech content coefficient be γ, and then use the formula γ=α. N1+β N2, where γ>0, N1+N2=0.8634, and both N1 and N2 are greater than 0; The larger the voice content coefficient γ, the greater the authenticity of the superior audio content obtained through the integrated analysis of voice and clarity when the superior audio content is displayed.
6. The anti-counterfeiting analysis system with the advantage of sorting out text content as described in claim 5, characterized in that, The comparison coefficient generation step is as follows: Let the comparison coefficient be... By integrating the speech content coefficient γ and text content coefficient δ using a linear regression model, a comparison coefficient is generated. Generate comparison coefficients The specific formula is as follows; =u1 γ+u2 δ, u1, u2 are weighting factors, u1 < u2, u1 + u2 = 2.
463. u1 and u2 represent the degree of composition of speech content and text content. The different levels of understanding when acquiring text content and audio content determine the value of the weighting factors.
7. The anti-counterfeiting analysis system with the advantage of sorting out text content as described in claim 6, characterized in that, The authenticity markers for advantageous content include markers for false advantageous content and markers for genuine advantageous content. The generation process for these markers is as follows: Comparing coefficients using a true / false labeling model The analysis is performed, and a comparison threshold KH is set. When KH is greater than 0, the comparison coefficients are... Substitute the comparison threshold KH into the data for analysis and processing. If the comparison coefficient... If the comparison coefficient is greater than the comparison threshold KH, a true advantage content tag is generated; if the comparison coefficient is greater than the comparison threshold KH, a true advantage content tag is generated. If the content is less than the comparison threshold KH, a false dominant content marker is generated. Audio files matching the true advantage content tag are identified as true content objects and output, while audio files matching the false advantage content tag are identified as false content objects and output.
Citation Information
Patent Citations
End-to-end text-to-speech conversion
CN110476206A
Popularity analysis method and system for voice data
CN107507627A