Compression speech information hiding detection method based on association rule mining
By constructing an LPC steganalysis-sensitive codeword association rule library based on association rule mining, and analyzing the intra-frame and inter-frame correlation of speech signals, this method solves the problem of low detection accuracy under low duration and low embedding rate in existing technologies, and achieves efficient speech information hiding detection.
Patent Information
- Application Number
- CN202210930504.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-04
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2042-08-04
AI Technical Summary
Existing methods for detecting hidden information in compressed speech have low accuracy under low duration and low embedding rate conditions, making real-time detection difficult.
This paper adopts a method based on association rule mining. By constructing an LPC steganalytic sensitive codeword association rule library, using a linear prediction model to predict speech parameters, analyzing the intra-frame and inter-frame correlation of speech signals, constructing a codeword transaction database, performing codeword itemset support counting and association rule mining, establishing a steganalytic sensitivity analysis model, and detecting whether secret information is embedded in speech signals.
It achieves high accuracy in detecting hidden information in compressed speech under short duration and low embedding rate, and can detect whether secret information is embedded in speech signals in real time, which is superior to existing methods.
Smart Images

Figure CN115238812B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular to a compression speech information hiding detection method based on association rule mining. BACKGROUND
[0002] The essence of information hiding technology, also known as steganography, is that the hider selects the relatively insensitive area in the human auditory and visual system in the carrier data by using the redundancy of information, and embeds secret information in it through a certain method to achieve unknowable covert information transmission. Information hiding technology includes steganography method and steganalysis method, and the two are closely related. At present, the two research directions show a spiral upward trend. In recent years, steganography methods and steganalysis methods based on various mainstream multimedia such as text, speech, audio, image and video have been proposed. Among them, voice over IP (VoIP) is a widely used information hiding carrier due to its advantages in application range, transmission efficiency and use scenario.
[0003] At present, network compression speech coding information hiding mainly includes two methods. The first method directly modifies the code stream after encoding to achieve information hiding. This method often selects a carrier with less impact on the overall quality of the speech by speech perceptual quality evaluation, signal-to-noise ratio and other methods. In the process of direct modulation of code elements, various algorithms are used to improve the embedding rate and concealment of the code stream; the second method changes the encoding result by modifying the encoding coefficient to achieve information hiding. The main methods include pitch modulation information hiding method, fixed codebook coefficient steganography method, LPC coefficient-based quantization index modulation steganography method, etc. Among them, the QIM compression speech steganography algorithm has the advantages of low distortion, high embedding rate and high anti-interference, and is very suitable for information hiding with VoIP speech as the carrier; a compression speech information hiding detection method based on association rule mining is proposed. By taking G.723.1 (6.3kbits / s) speech coding as an example, a transaction database is constructed based on the LPC quantization index value of adjacent frames and the steganographic sensitive association rules are mined to realize real-time detection of QIM information hiding. The detection accuracy is higher at low time length and low embedding rate. SUMMARY
[0004] In view of the above defects, the technical problem to be solved by the present application is to provide a compression speech information hiding detection method based on association rule mining, which includes the following detection process:
[0005] S1: LPC steganographic sensitive code word association rule mining;
[0006] S2: steganalysis;
[0007] In S1, there are also code word transaction database construction, intra-frame and inter-frame code word association combination analysis, and code word item set support count.
[0008] In the code word item set support count, there are the following processes: hash bucket construction, support count, and code word association analysis.
[0009] Secondly, S2 also includes steganographic sensitive association rule library construction and steganographic detection process.
[0010] In the technical scheme of the compression speech information hiding detection method based on association rule mining, preferably, the LPC steganographic sensitive code word association rule mining in S1 can use a linear prediction model to provide very accurate speech parameter prediction for a speech signal, and an LPC synthesis filter is as shown in formula 2-1:
[0011] (2-1)
[0012] wherein is the order LPC prediction coefficient of the speech signal, in the G.723.1 network compression speech coding mode, = 10, the filter coefficient is obtained based on the minimum mean square error criterion, and an LPC synthesis filter is constructed to obtain a residual signal, then the LPC coefficient is converted into a line spectrum frequency coefficient (LSP), and finally, the LSP is subjected to prediction splitting vector quantization to obtain three quantized code words , , Each code word is represented by 8 bits.
[0013] In the technical scheme of the compression speech information hiding detection method based on association rule mining, preferably, the code word transaction database construction includes the following processes:
[0014] The speech signal is composed of phonemes, wherein the voiced sound carrying effective information has intra-frame short-time stationarity after speech coding, which means that the LPC index VQ code word has certain intra-frame association, and the voiced part of the speech signal expresses specific information and has continuity characteristics in the time domain, that is, the LPC index VQ code words of adjacent two frames have certain inter-frame association. The intra-frame and inter-frame VQ code words are used to construct a transaction database, and on this basis, association rule mining analysis is performed to realize detection of information hiding.
[0015] Suppose that the speech code stream contains frames, ={ =1, 2, …, }, ={ , , }, wherein , ∈ {1, 2, 3} represents the first frame of the VQ code word, which ranges from [0, 255], and is denoted by The code word transaction database constructed by = { , … }, wherein represents the code word set of the first frame and the first +1 frame of the speech stream, which is as formula (2-2):
[0016] (2-2)
[0017] A single transaction in the transaction database is composed of the LPC code words of two adjacent frames, The size of is 6 × (n-1), which is converted to .
[0018] In the technical scheme of the compression speech information hiding detection method based on the association rule mining, preferably, the intra-frame and inter-frame code word association combination analysis comprises the following process:
[0019] The set of items is called an item set, which contains The item set containing items is called an item set, which is denoted as:
[0020] . The speech frame contains two kinds of association relationships, intra-frame and inter-frame, which are represented by patterns represents an intra-frame code word association rule, represents an inter-frame code word association rule, and the intra-frame and inter-frame code word association rules satisfy formula 2-3:
[0021] (2-3)
[0022] There are kinds of intra-frame association rule code word combination cases, and each code word has 256 values, so there are at most kinds of intra-frame association rules. There are kinds of inter-frame association rule code word combination cases, so there are at most kinds of inter-frame association rules.
[0023] In the technical scheme of the compression speech information hiding detection method based on the association rule mining, preferably, the code word item set support degree counting comprises the following process:
[0024] The association rule mining comprises two steps of item set generation and rule generation, the mode and The corresponding 3-item set is consistent, and therefore, only one of them needs to be considered when the item set is generated. In order to count the support degree of the code word item set, the frequency of the occurrence of the corresponding code word item set in the is recorded based on the hash counting, the code word item set mined in the is hashed into different hash buckets by using a hash function to realize the support degree counting of the association rule;
[0025] The hash bucket is constructed as follows:
[0026] For the association rule mode , the code word combination of each transaction is recorded, and the code word value combination has kinds, and the and have the sizes as 2-4:
[0027] (2-4)
[0028] In order to realize the statistical technology of the association rule, the length of the address pool of the hash table required by the association rule is also , and the theoretical size of the address pool of the hash table is the sum of the combination conditions of all the association rules as formula 2-5:
[0029] (2-5)
[0030] When the length of the code word association rule is , , the memory opening is difficult, and therefore, in order to effectively mine the association rules sensitive to steganography, the case of is analyzed, which contains 3 kinds of intra-frame 2-item association rules, 1 kind of intra-frame 3-item association rule, 9 kinds of inter-frame 2-item association rules, and 18 kinds of inter-frame 3-item association rules;
[0031] The support degree counting process is as follows:
[0032] Speech signals are composed of phonemes. To express the information they contain, speech signals exhibit a certain continuity between phonemes in the time and frequency domains of the speech stream. This correlation is reflected in the correlation characteristics of codeword values within and between frames. These characteristics can be demonstrated through the support of codeword intra- and inter-frame association rules in a transaction database. Strong correlations between codeword values lead to uneven distribution of association rule support; the stronger the correlation between codeword values, the higher the support of the corresponding association rule. As shown in equation 2-6:
[0033] (2-6)
[0034] After the hash table is built, the transaction database is scanned. Every transaction in The values of the acquired transaction elements are arranged and combined according to the codeword frame intra-frame and inter-frame combination table. The unique hash value of the association rule is calculated by the hash function. The corresponding hash bucket is found through the hash value and accumulated to realize the count. Finally, the support of a single association rule in the overall transaction database is obtained.
[0035] Association rules The confidence level is calculated according to Equation 2-7, and its meaning is the transaction database. Includes The number of transactions and containing The ratio of the number of transactions:
[0036] (2-7)
[0037] The codeword value correlation analysis process is as follows:
[0038] The correlation between intra-frame and inter-frame codeword values will lead to a non-uniform distribution of confidence scores for different codeword association rules in the transaction database. For example, observe Distribution in the transaction database, and by Taking an example, we observe the confidence distribution of each codeword value in the transaction database for inter-frame association rules; based on the support obtained from the codeword value statistics, the confidence of the corresponding association rule can be calculated. To intuitively illustrate the differences in the association between different codeword association rules in speech information, we take intra-frame association rules as an example. For example, among which The axis corresponds to the value of the first codeword. The axis corresponds to the value of the second codeword. The axis is the confidence of the association rule corresponding to the two codewords, wherein the codeword takes a value in a range of 0-255, the confidence takes a value in a range of 0-1, and the higher the confidence of an association rule is, the stronger the association of the corresponding intraframe codeword value is, and the stronger the association of the speech frame signal in the intraframe is.
[0039] In the technical scheme of the compression speech information hiding detection method based on association rule mining, preferably, the construction of the steganographic sensitive association rule library in the steganalysis includes the following construction process:
[0040] The codeword transaction database The steganographic transaction database is randomly generated by using the CNV-QIM steganographic algorithm The confidence of the association rule The confidence of the association rule changes after steganography, so the steganographic sensitive association rule library can be constructed by using the rules with obvious confidence changes, and then the steganographic analysis is realized.
[0041] The association rule The confidence of the association rule in and is and respectively. The steganographic sensitivity is quantified by the change of the confidence of the association rule before and after steganography , which is defined as formula 3-1.
[0042] (3-1)
[0043] It can be known that when >0 and =0, is positive infinity, at this time, it is indicated that the association rule disappears after information hiding, and this kind of association rule is a missing association rule. Similarly, when =0 and >0, is also positive infinity, at this time, it is indicated that a new association rule appears after information hiding, and this kind of association rule is a derived association rule. When is positive infinity, the steganographic sensitive association rule library is constructed for the corresponding association rule, the missing association rule library generated by is denoted as , and the derived association rule library generated by is denoted as
[0044] Obviously, the confidence level of the same association rule will change before and after steganography. A two-dimensional scatter plot can be used to represent the distribution of association rules before and after steganography. By comprehensively comparing the association rules with non-zero confidence levels before and after steganography, we can intuitively discover the existence of steganography-sensitive association rules. Missing rules existed before steganography but have zero confidence levels after steganography. Derived association rules are new association rules added after steganography.
[0045] Similarly, through comparison The confidence levels of each association rule before and after steganography yield the corresponding steganalysis-sensitive association rule library as follows, where the axes represent the corresponding symbol values in the rule. If an association rule has a non-zero confidence level before steganography but a zero confidence level after steganography, it is a missing inter-frame association rule. Conversely, if an association rule has a zero confidence level before steganography but appears in the transaction database after steganography, it is a derived inter-frame association rule. Based on the distribution of missing and derived association rule sets, the final inter-frame steganalysis-sensitive association rule library is obtained.
[0046] In the above-mentioned technical solution of the compressed speech information hiding detection method based on association rule mining, preferably, the steganalysis detection process is as follows:
[0047] Based on speech samples with different embedding rates and speech durations, different steganalysis-sensitive association rule bases are constructed. and Then, the association rules for codeword values in the speech samples to be detected can be used to... and The distribution of association rules within the speech sample is used to predict whether secret information is embedded. Association rules can be categorized into intra-frame and inter-frame association rules. 'a' and 'b' can be used to cumulatively count an association rule among missing and derived association rules. By enumerating the number of association rules, their distribution within the steganalysis-sensitive association rule base is determined, initially completing the segmentation of speech frames. By combining the association rules extracted from all speech frames in the speech sample to be detected, steganalysis of the speech sample can be achieved.
[0048] As can be seen from the above technical solution, the compressed speech information hiding detection method and apparatus based on association rule mining provided by the present invention has the following beneficial effects compared with the prior art:
[0049] The application analyzes the relevance of the vector quantization code word value in the time domain and the frequency domain, establishes a steganography sensitive analysis model, and uses adjacent frame VQ code words to construct a transaction database and carry out code word association rule analysis, constructs a steganography sensitive code word association rule library based on the confidence level change before and after steganography based on the code word association rule, realizes steganography analysis by analyzing the distribution of the corresponding association rule in the steganography sensitive association rule library of the to-be-detected speech sample, and the experimental results show that CARM is superior to the comparison method in the short time and low embedding rate indicators, and can achieve real-time detection effect. BRIEF DESCRIPTION OF DRAWINGS
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the description of the embodiments of the present application or the prior art will be briefly introduced and described below. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creating any creative labor.
[0051] Figure 1 Layout diagram of the compressed speech information hiding detection method based on association rule mining in the present application;
[0052] Figure 2 Support degree counting algorithm flowchart in the compressed speech information hiding detection method based on association rule mining in the present application;
[0053] Figure 3 Transaction database construction diagram in the compressed speech information hiding detection method based on association rule mining in the present application;
[0054] Figure 4 VQ1 and VQ2 association rule confidence level difference in the compressed speech information hiding detection method based on association rule mining in the present application;
[0055] Figure 5 VQ1 and VQ3 association rule confidence level difference in the compressed speech information hiding detection method based on association rule mining in the present application;
[0056] Figure 6 VQ2 and VQ3 association rule confidence level difference in the compressed speech information hiding detection method based on association rule mining in the present application;
[0057] Figure 7 Inter-frame association rule confidence level distribution in the compressed speech information hiding detection method based on association rule mining in the present application;
[0058] Figure 8The confidence distribution diagram of the hidden frame in the steganography detection method of compressed speech information hiding based on the association rule mining in the application;
[0059] Figure 9 The distribution diagram of the missing association rule in the steganography detection method of compressed speech information hiding based on the association rule mining in the application;
[0060] Figure 10 The distribution diagram of the derived association rule in the steganography detection method of compressed speech information hiding based on the association rule mining in the application;
[0061] Figure 11 The distribution diagram of the missing rule in the steganography detection method of compressed speech information hiding based on the association rule mining in the application;
[0062] Figure 12 The distribution diagram of the derived rule in the steganography detection method of compressed speech information hiding based on the association rule mining in the application;
[0063] Figure 13 The steganography detection process diagram in the steganography detection method of compressed speech information hiding based on the association rule mining in the application;
[0064] Figure 14 The detection accuracy diagram under different association rule modes in the steganography detection method of compressed speech information hiding based on the association rule mining in the application. DETAILED DESCRIPTION
[0065] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the following described embodiments are only some of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative work are within the protection scope of the application.
[0066] In order to make the technical solutions and implementation modes of the application clearer and more understandable, the following introduces several preferred specific embodiments for implementing the technical solutions of the application. EMBODIMENT
[0067] Data preparation:
[0068] To ensure the diversity of speech samples to expand the accuracy of steganalysis algorithm, 10000 speech samples in the literature are selected as experimental data set, the data set is composed of 10 kinds of mixed speech, the ten kinds of speech include Chinese and English male and female voices, French voice, German voice, Japanese voice, string music, piano music, symphony and other different timbres, a single type of speech contains 1000 speech samples, each sample has a duration of 10s, and a frame length of 333 frames, in this paper, each kind of speech is divided into training set and test set in the ratio of 3:2 for experiment.
[0069] Performance analysis of different association rule modes:
[0070] Because the steganography-sensitive transaction databases generated under different association rule modes are different, the steganalysis accuracy is also different. In this section, the CNV-QIM steganography detection algorithm is taken as an example, the steganography-sensitive rule base is constructed under 5 different embedding rates with a speech sample duration of 2s, and the performance of the association rule steganography under 5 modes is analyzed, and the results are shown in Figure 14
[0071] Under the condition of the same number of item sets, the detection accuracy of the steganography association rule base constructed by the inter-frame association rule is higher than that of the intra-frame association rule, because the inter-frame association rule is more than the intra-frame association rule, and the rules used to distinguish whether steganography are also more. Similarly, the intra-frame or inter-frame mode is higher than the detection accuracy of the steganography association rule base constructed by the mode , because generates more association rules than . Among them, The detection accuracy of the mode is slightly higher than , and the detection accuracy is close to 90% under the condition of speech duration of 2s and embedding rate of 0.4, which shows that the method proposed in the application achieves satisfactory steganalysis performance under low embedding and short duration.
[0072] Performance analysis of comparative experiments:
[0073] In order to further prove the detection performance of the code word association rule mining method proposed in this paper, three typical steganalysis methods are compared, which are IDC steganalysis method, RNN-SM steganalysis method and CBN steganalysis method, and the best performance of the five association rule modes is selected to compare the performance in terms of speech duration and embedding rate.
[0074] ①Performance analysis of different embedding rates:
[0075] The impact of the hiding algorithm on the quality of the speech will decrease with the decrease of the steganographic embedding rate, thereby improving the steganographic security, which also causes the accuracy of the steganographic detection method to decrease significantly when analyzing speech samples with low embedding rate. In order to measure the accuracy of the steganographic analysis of speech samples with different embedding rates by the steganographic detection method in this paper, speech samples with 5 embedding rates and a length of 10s were used for the experiment. The detection accuracy of the CNV-QIM and NPP-QIM steganographic algorithms by the IDC, RNN-SM, CBN and CARM methods under different embedding rates is shown in Table 1:
[0076]
[0077] Table 1
[0078] As can be seen from Table 1, for the CNV-QIM steganographic method, the accuracy of the CARM detection method proposed in this paper is higher than 86% under 5 embedding rates, and is better than other detection methods. When the embedding rate is greater than or equal to 0.4, the detection accuracy of the CARM method is greater than 95%; when the embedding rate is 0.2, the detection accuracy of the CARM method is 86.06%, and the detection accuracy of other methods is less than 80%, which shows that the CARM method proposed in this paper has a high detection accuracy for the CNV-QIM steganographic method under low embedding rate. Since NPP-QIM uses the nearest projection point to embed the steganographic information, the steganographic security is higher and the concealment is stronger, and the detection accuracy of the 4 detection methods in the comparative experiment decreases to a certain extent, mainly because the NPP-QIM steganographic method changes fewer code values. Under 5 embedding rates, the detection performance of the CARM is obviously better than that of the IDC and RNN-SM methods, and is slightly better than that of the CBN method.
[0079] ② Performance analysis of different speech lengths
[0080] Whether the short-length speech samples can be effectively detected is an important indicator for measuring the information hiding detection method. In order to test the influence of the speech length on the steganographic detection result, the comparative experiment of the multiple-length speech samples was conducted, and the accuracy was used for measurement. The CNV-QIM and NPP-QIM steganography were used to generate speech samples with 5 embedding lengths and an embedding rate of 1.0 for steganographic detection, and the detection results of the 4 steganographic analysis methods for the 2 steganographic methods are shown in Table 2:
[0081]
[0082] Table 2
[0083] From Table 2, it can be seen that the CARM method has stable detection performance for the CNV-QIM steganography method, and the detection accuracy is more than 99% under five speech durations, which is better than other comparative methods. When detecting the speech samples of the NPP-QIM steganography method, the detection accuracy of the CARM method for speech samples longer than 4 seconds is more than 90%. In summary, CARM can effectively detect both CNV-QIM and NPP-QIM speech steganography methods. Under different embedding rates and different speech durations, the CARM method is better than other steganalysis methods, especially when it can still achieve a high detection accuracy under short duration and low embedding rate. However, the method is limited by the size of the hash bucket, and cannot construct association rules with more than 3 items.
[0084] Finally, it should be noted that the structures, proportions, sizes, etc. shown in the drawings of the present specification are only used to assist in understanding and reading the content disclosed by the present specification for those skilled in the art, and do not define the limiting conditions for the implementation of the present application, and therefore do not have technical substantive significance. Any modification of the structure, change of the proportion relationship or adjustment of the size, without affecting the effects and purposes that can be achieved by the present application, should still fall within the scope of the technical content disclosed by the present application.
[0085] The terms "comprise", "contain", or any other variant thereof, as used in this document, are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or device. Without more limitations, the element defined by the statement "comprises a" does not exclude the presence of other identical elements in the process, method, article or device that includes the element.
[0086] The present application is not limited to the above-mentioned best mode, and anyone should know that any structural changes made under the inspiration of the present application fall within the protection scope of the present application, and any technical solutions with the same or similar technical solutions as the present application also fall within the protection scope of the present application.
Claims
1. A method for detecting steganographic compression speech information based on association rule mining, characterized in that: The detection process comprises the following steps: S1: mining of LPC steganography-sensitive code word association rules; S2: steganalysis; The S1 further comprises the following steps: construction of a code word transaction database, intra-frame and inter-frame code word association combination analysis, and code word item set support degree counting. The code word item set support degree counting comprises the following steps: hash bucket construction, support degree counting, and code word association analysis. Secondly, the S2 further comprises the following steps: construction of a steganography-sensitive association rule library and steganography detection. The code word transaction database construction comprises the following steps: Speech signals are composed of phonemes, wherein the voiced sound carrying effective information has intra-frame short-time stationarity after speech coding, which means that the LPC index VQ code word has certain intra-frame association, and the voiced part of the speech signal expresses specific information and has continuity in the time domain, i.e., the LPC index VQ code words of adjacent two frames have certain inter-frame association. A transaction database is constructed by using intra-frame and inter-frame VQ code words, and on this basis, association rule mining analysis is performed to realize detection of information hiding. Assuming the voice bitstream Include frame, ={ =1, 2, …, }, ={ , , }, where W i,j , ∈ {1, 2, 3} represents the th... The first frame The values of the VQ codewords range from [0, 255], denoted by […]. The constructed codeword transaction database is = { ,… …, },in Indicates the first voice stream Frame and the The codewords of +1 frame constitute the set as shown in equation (2-2): (2-2) A single transaction in the transaction database consists of the LPC code words of two adjacent frames, of size 6 x (n - 1), from to .
2. The compression speech information hiding detection method based on association rule mining according to claim 1, wherein, The LPC steganography-sensitive code word association rule mining in the S1 can use a linear prediction model to provide very accurate speech parameter prediction for speech signals. An LPC synthesis filter is as shown in formula (2-1): (2-1) wherein is a speech signal LPC prediction coefficients, in G.723.1 network compression speech coding mode, = 10, the filter coefficients are calculated based on the minimum mean square error criterion and the residual signal is obtained by constructing the LPC synthesis filter, then the LPC coefficients are converted into line spectrum frequency coefficients LSP, finally the LSP is quantized by the prediction splitting vector to obtain three quantized code words 、 、 Each code word is represented by 8 bits.
3. The compression speech information hiding detection method based on association rule mining according to claim 1, characterized in that, The intra-frame and inter-frame code word association combination analysis comprises the following steps: A set of items is called an itemset, containing An itemset of k codewords is called a k-itemset An itemset is denoted by The speech frame contains two kinds of association relations, intra-frame and inter-frame, and adopts a mode represents an intra-frame code word association rule, represents an inter-frame code word association rule, then the intra-frame and inter-frame code word association rules satisfy formula (2-3): (2-3) For the intra-frame association rule code word combination, there are totally For the intra-frame association rule code word combination, there are totally For the inter-frame association rule code word combination, there are totally For the inter-frame association rule code word combination, there are totally For the inter-frame association rule code word combination, there are totally 4. The compression speech information hiding detection method based on association rule mining according to claim 1, characterized in that, The code word item set support degree counting comprises the following steps: The association rule mining includes two steps of item set generation and rule generation and The corresponding 3-item set is consistent, only one of which is considered when the item set generation is performed, in order to count the support degree of the code word item set, the frequency of the corresponding code word item set in is recorded based on the hash count, the code word item set mined in is hashed into different hash barrels by using a hash function to realize the support degree counting of the association rule. The hash bucket construction is as follows: For association rule pattern , record each transaction The corresponding code word combination has , the code word value combination has , then And The size of formula (2-4): (2-4) To implement the statistical technique of association rules, the length of the address pool required by the hash table is also The theoretical size of the address pool of the hash table is The total sum of all combinations of association rules is as formula (2-5): (2-5) When the codeword association rule length hour, Memory allocation is difficult. In order to effectively discover association rules that are sensitive to steganography, [the following is necessary]. The analysis includes 3 types of intra-frame 2-item association rules, 1 type of intra-frame 3-item association rule, 9 types of inter-frame 2-item association rules, and 18 types of inter-frame 3-item association rules. The support degree counting calculation process is as follows: The speech signal is composed of phonemes, and the speech signal has certain continuity in the time domain and frequency domain of the speech flow in order to express the information contained therein. The correlation will be reflected in the correlation characteristics of the code word values in the frame and between the frames. These characteristics can be reflected by the support degree of the code word frame and inter-frame association rules in the transaction database. Strong correlation between code word values will lead to uneven distribution of rule support degrees. The stronger the correlation of the code word values, the higher the support degree of the corresponding association rules. The rule support degree As formula (2-6): (2-6) After the hash table is built, the transaction database is scanned. Every transaction in The values of the acquired transaction elements are arranged and combined according to the codeword frame intra-frame and inter-frame combination table. The unique hash value of the association rule is calculated by the hash function. The corresponding hash bucket is found through the hash value and accumulated to realize the count. Finally, the support of a single association rule in the overall transaction database is obtained. Confidence of association rules is calculated according to equation (2-7), which means the ratio of the number of transactions containing to the number of transactions containing . (2-7) The code word value association analysis process is as follows: The correlation between intra-frame and inter-frame codeword values will lead to a non-uniform distribution of confidence scores for different codeword association rules in the transaction database. Whether or not this is used... For example, observe Distribution in the transaction database, and by Taking an example, we observe the confidence distribution of each codeword value in the transaction database for inter-frame association rules; based on the support obtained from the codeword value statistics, the confidence of the corresponding association rule can be calculated. To intuitively illustrate the differences in the association between different codeword association rules in speech information, we take intra-frame association rules as an example. For example, among which The axis corresponds to the value of the first codeword. The axis corresponds to the value of the second codeword. The axis represents the confidence level of the association rule corresponding to two codewords. The codeword value ranges from 0 to 255, and the confidence level ranges from 0 to 1. The higher the confidence level of an association rule, the stronger the correlation between the codeword values in the corresponding frame, and the stronger the correlation between the speech frame signal in the frame.
5. The compression speech information hiding detection method based on association rule mining according to claim 1, characterized in that, The construction of the steganography-sensitive association rule library in the steganalysis comprises the following construction process: Transaction database for codewords Steganographic transaction database generated randomly using CNV-QIM steganographic algorithm , the The confidence of the association rules changes before and after steganography, and the confidence of the association rules changes after steganography. Therefore, the steganographic sensitive association rule library can be constructed by means of the rules with obvious confidence changes, and then steganographic analysis is realized. Association rules In and The confidence is and respectively, the confidence of association rules The steganographic sensitivity is quantified by the change of the confidence of association rules , is defined as equation (3-1): (3-1) It can be seen that when > 0 and = 0, is positive infinity, which means that the association rule disappears after information hiding. Such association rules are missing association rules. Similarly, when = 0 and > 0, is also positive infinity, which means that new association rules appear after information hiding. Such association rules are derived association rules. When is positive infinity, the corresponding association rule constructs a steganographic sensitive association rule base. The missing association rule base generated by is denoted as , and the derived association rule base generated by is denoted as Obviously, the confidence of the same association rule will change before and after steganography. The distribution of the association rules before and after steganography is represented by a two-dimensional plane scatter plot. By comprehensively comparing the association rules with a confidence of 0 before and after steganography, the existence of steganography-sensitive association rules can be intuitively found. Missing rules exist before steganography but have a confidence of 0 after steganography, and derived association rules are newly added association rules after steganography. Similarly, by comparing The confidence of each association rule before and after steganography can obtain the corresponding steganography-sensitive association rule library as follows, where the coordinate axes represent the corresponding symbol values in the rules. If the confidence of an association rule before steganography is not 0 but the confidence after steganography is 0, it is an inter-frame missing association rule. Conversely, if the confidence of an association rule before steganography is 0 but it appears in the transaction database after steganography, it is an inter-frame derived association rule. According to the distribution of the missing association rules and the derived association rules, the inter-frame steganography-sensitive association rule library is finally obtained.
6. The compression speech information hiding detection method based on association rule mining according to claim 1, wherein, The steganography detection process is as follows: According to different embedding rates and different speech lengths of speech samples, different steganalysis sensitive association rule libraries are constructed and After that, whether the speech sample is embedded with secret information can be predicted by the distribution of codeword value association rules in and The types of association rules can be divided into intra-frame association rules and inter-frame association rules, a and b can be used for the cumulative count of an association rule in the missing association rule and the derived association rule, the number of items of the association rule is enumerated to determine its distribution in the steganalysis sensitive association rule library, the division of the speech frame is preliminarily completed, and the steganalysis of the speech sample to be detected can be realized by comprehensively extracting the association rules of all speech frames in the speech sample to be detected.
Citation Information
Patent Citations
Voice information hiding algorithm for dynamic codebook based on complete binary tree grouping
CN107527621A
SILK security steganography method based on LSF coefficient statistical distribution characteristics
CN110097887A