A method for detecting encryption data vulnerability based on discrete neural network
By using frequency domain transformation, bit correlation detection, and time-frequency joint analysis based on discrete neural networks, a structured distributed representation is generated, which solves the problem of difficulty in identifying subtle abnormal patterns in encrypted data in existing technologies and achieves efficient detection of defects in encrypted data streams.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 北京建恒信安科技有限公司
- Filing Date
- 2026-03-27
- Publication Date
- 2026-07-03
Smart Images

Figure CN122333495A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information technology, and in particular to a method for detecting the vulnerability of encrypted data based on discrete neural networks. Background Technology
[0002] In the field of information security, the robustness of encryption technology is directly related to the core defense line of data protection, and its importance is self-evident.
[0003] As cyberattack techniques continue to evolve, detecting potential vulnerabilities in encrypted data has become a crucial step in ensuring communication security and protecting privacy.
[0004] Whether it's financial transactions or personal privacy data, the reliability of encryption algorithms must withstand various complex threats. Discovering hidden flaws without decryption has become an urgent research need.
[0005] However, existing methods often struggle to capture subtle anomaly patterns when dealing with highly complex encrypted data.
[0006] These methods often start from the overall statistical characteristics, ignoring the local biases of the data at the micro level, resulting in a lack of accurate identification of the subtle defects hidden in the encryption algorithm.
[0007] Especially when the data is highly chaotic, traditional analysis methods often fail to delve into the fine structure at the bit level, missing the opportunity to discover potential risks.
[0008] A deeper technical challenge lies in the fact that the highly disordered nature of encrypted data makes it difficult to effectively represent its characteristics in a continuous space.
[0009] This characteristic makes it difficult to detect tiny non-random deviations within the data, which are often weaknesses exposed by encryption algorithms under specific conditions.
[0010] Furthermore, the existence of these minute deviations is closely related to the abnormal distribution of data in different mapping spaces. If the data cannot be transformed from a high-dimensional chaotic state into a more easily analyzable structured form, it will be difficult to reveal the patterns behind these deviations.
[0011] For example, in certain encryption processes, specific key combinations may lead to subtle uneven distributions in the data permutation process. Although this unevenness is not statistically significant, it may become a vulnerability that attackers can exploit in specific scenarios.
[0012] Therefore, how to transform highly chaotic encrypted data into an analyzable structured form without relying on decryption, and accurately capture the subtle deviations hidden within it, has become a key problem that this research urgently needs to solve. Summary of the Invention
[0013] This invention provides a method for detecting the vulnerability of encrypted data based on discrete neural networks, mainly comprising: Obtain the bit sequence of the encrypted data stream, use frequency domain transformation technology to extract the frequency domain distribution characteristics of the bit sequence, determine the preliminary vulnerable regions based on the distribution non-uniformity in the frequency domain distribution characteristics, and form a preliminary vulnerable region set. For the initial set of vulnerable regions, bit correlation detection technology is used to group the bits in the region. Based on the persistence of defects in the grouped bit combinations, potential risk regions are determined and a risk region list is formed. Local bit structures are extracted from the risk area list. Combining the processing depth and operational complexity of the encryption algorithm, the local bit structures are mapped to a simplified feature space through a spatial dimensionality reduction transformation method. Based on the deviation of the mapped distribution from the expected distribution pattern, the abnormal distribution structure is determined and an abnormal distribution set is formed. For the set of abnormal distributions, the temporal stability and transformation sensitivity features of the structure are extracted by time-frequency joint analysis. Based on the case where the noise interference in the features exceeds a preset threshold, high-risk abnormal patterns are determined and a list of high-risk patterns is formed. According to the high-risk pattern list, obtain the bit subsequences with persistent defects, use sequence similarity comparison technology to group the subsequences, determine the encryption defect patterns based on the pattern predictability of the grouped subsequences, and form defect pattern groups. The associated bit sequences are extracted from the defect pattern group, and the time-domain stability and frequency-domain concealment features are fused together. A structured distribution representation is generated through bit recombination technology. The hidden encryption defects are determined based on the distribution non-uniformity and local dependency in the structured distribution representation, and the final defect identifier set is obtained.
[0014] Furthermore, the step of obtaining bit subsequences with persistent defects based on the high-risk pattern list, grouping the subsequences using sequence similarity comparison technology, and determining the encryption defect pattern and forming defect pattern groups based on the pattern predictability of the grouped subsequences includes: Obtain the bit subsequence with defect persistence based on the list of high-risk patterns; Sequence alignment techniques are used to calculate the similarity distance between subsequences; If the similarity distance is less than a preset threshold, the corresponding subsequences will be grouped into the same group. The encryption defect pattern is determined based on the pattern predictability score of the subsequences within each group; The group with the highest predictability score was identified as the defect pattern group; Obtain representative bit sequences from the defect pattern group as feature patterns.
[0015] Furthermore, the step of extracting associated bit sequences from the defect pattern group, fusing time-domain stability and frequency-domain concealment features, generating a structured distribution representation through bit recombination technology, determining hidden encryption defects based on the distribution non-uniformity and local dependency in the structured distribution representation, and obtaining the final defect identifier set includes: By obtaining initial data from the defect pattern group, a bit extraction tool is used to separate the associated bit sequences, resulting in a preliminary sequence set. Based on the preliminary set of sequences and combined with the temporal stability characteristics, a preset threshold is used for screening. If the temporal fluctuation of a sequence exceeds the threshold, the unstable part is removed and the stable bit segment is determined. For stable bit segments, frequency domain concealment features are fused, and frequency domain transformation method is used to analyze concealment strength. Segments whose concealment meets preset conditions are identified, and the filtered feature sequence is obtained. By using the filtered feature sequences, a structured distribution representation is constructed using bit recombination technology. If the recombined distribution is uneven, the abnormal distribution area is marked and the abnormal distribution identifier is obtained. Based on the abnormal distribution identifier, analyze the local dependency characteristics. If the dependency strength of a local area is higher than the preset standard, it is determined to be a hidden encryption defect, and a preliminary defect set is obtained. For the initial defect set, a second verification is performed by integrating defect identification methods. A comparison tool is then used to match the features of the final identifier set to determine the final defect identifier set.
[0016] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects: This invention discloses a defect detection method for encrypted data streams based on frequency domain transformation and multidimensional feature analysis. Addressing the challenge of identifying hidden defects in encrypted data streams in business scenarios, it proposes a systematic solution encompassing the frequency domain distribution characteristics of bit sequences, time-frequency joint analysis, and sequence similarity comparison. The invention initially locates vulnerable regions by extracting frequency domain distribution non-uniformity, and then accurately identifies potential risk areas and abnormal distribution structures by combining bit correlation detection and spatial dimensionality reduction transformation. Further, through time-frequency joint analysis and high-risk pattern screening, it identifies bit subsequences with persistent defects. Finally, it fuses time-domain stability and frequency-domain concealment features to generate a structured distribution representation, thus identifying hidden encryption defects. This invention effectively solves the defect detection problem caused by non-uniform distribution and local dependencies in encrypted data streams, significantly improving the accuracy and comprehensiveness of defect identification and providing strong technical protection for data security. Attached Figure Description
[0017] Figure 1 This is a flowchart of the encrypted data vulnerability detection method of the present invention.
[0018] Figure 2 This is a schematic diagram of the encrypted data vulnerability detection method of the present invention.
[0019] Figure 3 This is another schematic diagram of the encrypted data vulnerability detection method of the present invention.
[0020] Figure 4 This is a schematic diagram illustrating the principle of frequency domain transformation and vulnerable region localization in this invention.
[0021] Figure 5 This is a visualization diagram of the spatial dimensionality reduction transformation mapping of the present invention.
[0022] Figure 6 This is a diagram of the discrete neural network feature mapping mechanism of the present invention.
[0023] Figure 7 This is a diagram showing the physical deployment of the encrypted data vulnerability detection system of the present invention in a real network environment.
[0024] Figure 8 This is a cross-sectional view of the bit sequence recombination into a two-dimensional matrix structure according to the present invention.
[0025] Figure 9 This is a comparative analysis chart showing the accuracy of vulnerable area detection in this invention.
[0026] Figure 10 This is a heatmap showing the relationship between the defect persistence threshold and the detection rate in this invention.
[0027] Figure 11 This is a comprehensive analysis chart of data processing throughput and latency in this invention. Detailed Implementation
[0028] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0029] Terminology definition: Unless otherwise specified, the custom technical terms used in this application are uniformly defined as follows: Frequency domain concealment: The degree to which the power spectral density distribution of an encrypted bit sequence closely approximates an ideal random white noise distribution after frequency domain transformation is the core indicator characterizing the concealment of encryption features in the frequency domain. Local dependency: The degree of non-random statistical correlation between bits in a local region after a bit sequence is recombined into a structured distribution representation is a key feature for judging local permutation defects in encryption algorithms; Pattern predictability: The accuracy of predicting the values of subsequent bits based on the historical distribution characteristics of bit subsequences, quantified by the normalized value of conditional entropy; the higher the value, the stronger the predictability. Noise interference: The degree of interference of random noise components on effective features in the features extracted by joint time-frequency analysis is quantified and characterized by signal-to-noise ratio and spectral flatness. Structured distribution representation: The bit distribution form formed by reorganizing a one-dimensional encrypted bit sequence into a multi-dimensional matrix structure according to preset rules, which is used to reveal the spatial correlation characteristics between bits.
[0030] like Figure 1 As shown, the encrypted data vulnerability detection method provided by the present invention includes steps S101 to S106. Step S101: Obtain the bit sequence of the encrypted data stream, extract frequency domain distribution features using frequency domain transformation technology, determine preliminary vulnerable regions based on distribution non-uniformity, and form a preliminary vulnerable region set. Step S102: For the preliminary vulnerable region set, combine and group the bits using bit correlation detection technology, determine potential risk regions based on defect persistence, and form a risk region list. Step S103: Extract local bit structures from the risk region list, combine the encryption algorithm processing depth and operational complexity, map to a simplified feature space through spatial dimensionality reduction transformation, determine abnormal distribution structures, and form an abnormal distribution set. Step S104: For the abnormal distribution set, extract time domain stability and transformation sensitivity features using time-frequency joint analysis technology, determine high-risk abnormal patterns based on noise interference exceeding a preset threshold, and form a high-risk pattern list. Step S105: Obtain bit subsequences with defect persistence based on the high-risk pattern list, group them using sequence similarity comparison technology, determine encryption defect patterns based on pattern predictability, and form defect pattern groups. Step S106 extracts associated bit sequences from the defect pattern group, integrates time-domain stability and frequency-domain concealment features, generates a structured distributed representation through bit recombination technology, determines hidden encryption defects based on distribution non-uniformity and local dependency, and obtains the final defect identifier set. Steps S101 to S106 constitute a top-down linear processing flow, with the output of each step serving as the input for the next step.
[0031] like Figure 2As shown, the data processing pipeline of the encrypted data vulnerability detection method of the present invention includes six functional modules. The encrypted data stream, as input in the form of a bit sequence, is first processed by a frequency domain transformation module, outputting a preliminary set of vulnerable regions; then it enters a bit-inter-correlation detection module, outputting a list of risk regions; next, it is processed by a spatial dimensionality reduction transformation module, outputting a set of abnormal distributions; then it is processed by a time-frequency joint analysis module, outputting a list of high-risk patterns; next, it is processed by a sequence similarity comparison module, outputting a group of defective patterns; finally, it is comprehensively processed by a bit recombination and discrete neural network module, outputting a final set of defect identifiers. The modules are sequentially connected via data streams, with the output of one module serving as the input of the next, forming a complete detection and processing pipeline.
[0032] like Figure 7 As shown, the deployment of the encrypted data vulnerability detection system based on discrete neural networks described in this invention in a real network environment includes the following physical devices and their connections. Data source device 701 includes a group of terminal devices such as IoT devices, industrial equipment, and servers that generate encrypted data streams. Each data source device sends the encrypted data stream to network switching device 702 via a network link. Network switching device 702 is responsible for forwarding and aggregating the encrypted data streams and copies the data traffic to data acquisition probe 703 through a mirror port. Data acquisition probe 703 is deployed at the mirror port of network switching device 702 and is responsible for non-interference acquisition of encrypted data flowing through the network, transmitting the acquired data to detection server 704. Detection server 704 has a built-in discrete neural network computing unit responsible for performing core detection tasks on the acquired encrypted data, including frequency domain transformation, bit correlation detection, spatial dimensionality reduction, time-frequency joint analysis, sequence alignment, and discrete neural network defect identification. After detection, detection server 704 outputs the detection results and defect report to management terminal 705 for visualization and analysis by security management personnel. In addition, a management channel is provided between the network switching device 702 and the detection server 704 for device management and configuration communication.
[0033] Example 1 uses the application scenario of "testing the security of a custom stream encryption algorithm for a certain type of industrial protocol": In the specific implementation process, the system first acquires the original bit sequence of the encrypted communication stream under test. It then extracts the frequency domain distribution characteristics, such as the power spectral density, using the Fast Fourier Transform (FFT) technique. If an abnormal concentration of energy in a specific high-frequency component is observed in the spectrum (i.e., exhibiting uneven distribution), the system inversely maps this frequency point back to the bit offset position, identifying preliminary vulnerable regions and aggregating them into a preliminary vulnerable region set. Next, for this set, the system uses mutual information to calculate the inter-bit correlation detection technique, grouping the bits into sliding window combinations. If a specific bit combination is found to occur at a fixed offset in multiple consecutive encrypted packets with a probability exceeding 70% (i.e., satisfying the defect persistence requirement), it is identified as a potential risk region, and a risk region list containing offset addresses and lengths is generated. Subsequently, local bit structures are extracted from the list. Combining the complexity of the S-box permutation operation with the 12 rounds of iterations used by the algorithm and the high nonlinearity, the high-dimensional bit vector is projected to a low-dimensional simplified feature space using Principal Component Analysis (PCA), a spatial dimensionality reduction transformation method. By calculating the deviation of the projection point distribution from the ideal random white noise (expected distribution pattern), the abnormal distribution structure is determined and an abnormal distribution set is formed.
[0034] Next, the set is analyzed using a joint time-frequency analysis technique called Short-Time Fourier Transform (STFT). Temporal stability is extracted by observing the energy envelope of the signal on the time axis, and transform sensitivity features are extracted by observing the drastic changes in the output through input bit flipping. If the calculated signal-to-noise ratio (SNR) and other noise interference indicators exceed a preset 6dB threshold, the feature pattern is identified as a high-risk anomalous pattern and added to the high-risk pattern list. Based on this, bit subsequences with persistent defects are extracted, and sequence similarity comparison techniques such as Hamming distance are applied to calculate the clustering distance between subsequences. If the distance is less than a preset threshold, they are grouped into the same group, and their pattern predictability score is evaluated based on information entropy. If the score is below 3.0 bits, it is determined to be an encryption defect pattern with statistical bias, thus extracting representative feature patterns. Finally, the associated sequences are stripped using bit extraction tools. After evaluating time-domain stability using the sliding variance method to remove fluctuating random bits and combining frequency-domain hidden features (such as evaluating spectral flatness) for screening, the screened bits are rearranged into a 16×16 matrix structure using bit recombination technology. The uneven distribution of 1-bits in the recombination matrix (such as local clustering) and the strength of local dependencies between adjacent rows / columns (such as Pearson correlation coefficient) are analyzed. If the strength is greater than 0.85, an algorithm defect is confirmed. Finally, a set of defect identifiers containing vulnerability type identifiers is generated by matching a preset feature library.
[0035] Example 2 uses "testing the security of a certain type of blockchain transaction data encryption protocol" as an application scenario: This embodiment provides a security detection method for a custom stream encryption protocol for blockchain transaction data. First, 10,000 consecutive transaction data packets (totaling 2,048,000 bits of encrypted payload) are collected from the test environment. After being converted into a ±1 signal sequence, a Fast Fourier Transform (FFT) with a Hanning window is used for frequency domain transformation. The mean energy at each frequency point is calculated, and the skewness (1.85) and kurtosis (4.73) features are extracted. When the skewness exceeds 1.2 and the kurtosis exceeds 3.5, it is determined that there is uneven distribution, and two preliminary vulnerable regions are located (Region A: 312,450–315,200; Region B: 820,100–824,600). The frequency domain distribution characteristics here are specifically manifested as abnormal energy clustering in specific frequency bands (k=64–80 and 192–224), which can be accurately traced back to the original bit position through frequency-time mapping.
[0036] For the aforementioned regions, a bit correlation detection technique is implemented: sampling is performed using a 4-bit sliding window, and bit pair dependencies are quantified by combining mutual information and linear correlation coefficients. Bit combination groups are formed through hierarchical clustering. Furthermore, the concept of defect persistence (i.e., the probability that a specific bit combination maintains the same pattern in consecutive samples) is introduced, and the pattern consistency score of each combination in tens of thousands of samples is calculated (e.g., the persistence of combination A7 reaches 0.91). When the score exceeds the 0.65 threshold and the region density is higher than 0.4%, two potential risk sub-regions are identified. Subsequently, combining the processing depth of the 16-round Feistel structure of the encryption algorithm with the operational complexity defined by the ratio of S-box / P-permutation operations (3:1), a weighting coefficient (W=9.9) was constructed, and spatial dimensionality reduction transformation (PCA projection to 8-dimensional feature space, cumulative variance contribution rate >95%) was performed on the 64-bit blocks of the risk region. The dimensionality reduction result was compared with the expected distribution pattern generated by Monte Carlo simulation (8-dimensional Gaussian mixture model) using KL divergence (threshold 0.35). Combined with DBSCAN clustering, three abnormal distribution structures were identified, whose spatial clustering and distribution deviation were significantly different from the ideal random state.
[0037] A joint time-frequency analysis technique was applied to anomalous structures: Short-time Fourier Transform (STFT, 128-bit window, 75% overlap) was used to simultaneously extract temporal stability features (stability score calculated based on inter-frame variance correlation coefficient, ideal value close to 0.5) and transform sensitivity features (evaluation of Euclidean distance change of feature vectors through a 5% bit perturbation experiment); simultaneously, signal-to-noise ratio (SNR) and spectral flatness (thresholds of 8dB and 0.65) were calculated. Structures with all three exceeding the thresholds were identified as high-risk anomalous patterns (e.g., structure_B1: SNR=14.7dB, flatness=0.49). Based on the high-risk pattern, 2324-bit defect persistence bit subsequences were extracted. The sequence similarity was accurately calculated by MinHash initial screening and Dynamic Time Warping (DTW), and the sequence was divided into 4 groups by spectral clustering. Markov models were constructed for each group to evaluate the predictability of the pattern (using conditional entropy normalization score, threshold 0.3). Among them, 3 groups had predictability exceeding the threshold (up to 0.82), and were identified as the cryptographic defect pattern group. Their representative bit pattern (such as "11110000") has a logical relationship with the transaction field.
[0038] Finally, the associated sequences of the defect pattern group were double-verified: first, segments with insufficient temporal stability were removed using the sliding window variance method (variance threshold 0.06); then, frequency domain concealment features (Welch power spectrum flatness factor and wavelet high-frequency coefficient distribution, comprehensive score threshold 0.88) were fused, retaining 1874 stable features; after bit recombination into a 19×19 matrix, the distribution non-uniformity was quantified by the chi-square test (threshold 15.5), and the conditional mutual information of adjacent bits was calculated to evaluate local dependencies (threshold 0.75). It was found that the dependency strength of the lower right sub-block reached 0.89, pointing to the intermediate state defect before the S-box replacement. After secondary verification with the cryptographic vulnerability feature library (weighted Jaccard similarity > 85%), a final defect identifier set (e.g., "BC-LEAK-2023-S01") containing defect ID, location, severity level, and remediation suggestions was generated, fully revealing the linear approximation vulnerability in the encryption protocol caused by insufficient S-box nonlinearity. This embodiment enables those skilled in the art to directly reproduce the technical solution of this invention through a complete process operation characterized by adjustable parameters, clearly defined thresholds, and tightly integrated algorithms.
[0039] More specifically, the method for detecting the vulnerability of encrypted data based on discrete neural networks according to the present invention may include: Step S101: Obtain the bit sequence of the encrypted data stream, extract the frequency domain distribution characteristics of the bit sequence using frequency domain transformation technology, determine the preliminary vulnerable regions based on the distribution non-uniformity in the frequency domain distribution characteristics, and form a preliminary vulnerable region set.
[0040] The encrypted data stream is acquired and the bit sequence is parsed. A Fourier transform is applied to the bit sequence to obtain the frequency domain amplitude spectrum. The statistical distribution of amplitude values is extracted from the frequency domain amplitude spectrum as the frequency domain distribution feature. The skewness coefficient and kurtosis coefficient of the frequency domain distribution feature are calculated. If the skewness coefficient is greater than a preset threshold or the kurtosis coefficient is greater than another preset threshold, a distribution non-uniformity is determined. The frequency points corresponding to the distribution non-uniformity are mapped back to the bit sequence index to determine the initial vulnerable regions. Adjacent indices are merged to form continuous regions, constituting the initial vulnerable region set.
[0041] like Figure 4 As shown, this invention achieves precise location of vulnerable areas in encrypted data through frequency domain transformation. Figure 4 (a) shows a time-domain waveform of a 128-bit encrypted data stream, with a stepped waveform representing the 0 and 1 values in the bit sequence, where gray shaded areas mark bit segments with periodic patterns. Figure 4 (b) shows the frequency domain amplitude spectrum obtained after performing a Fast Fourier Transform on the bit sequence. The horizontal axis represents the normalized frequency, and the vertical axis represents the amplitude value. The DC component has been removed from the figure to highlight the effective frequency characteristics. The dashed line represents the anomaly detection threshold, which is calculated based on the mean amplitude and standard deviation. Frequency components exceeding the threshold are marked with circles as abnormally high amplitude frequency points. The significant peaks at frequencies such as f=0.250 and f=0.281 indicate the presence of periodic structures in the bit sequence. These periodic structures deviate from the uniform random distribution characteristics that ideal encrypted data should possess. Figure 4 (c) shows the results of vulnerable region localization after inversely mapping the frequency domain anomaly information back to the bit sequence position. Dark blocks mark the bit intervals A and B identified as vulnerable, corresponding to bits 48 to 63 and 96 to 111 in the original sequence, respectively. The three sub-graphs are connected by dashed arrows, representing the complete detection process from time domain to frequency domain transformation and from frequency domain anomaly inverse mapping back to time domain position.
[0042] In one possible implementation, acquiring the encrypted data stream and parsing the bit sequence is the foundation of the entire process.
[0043] For example, in the scenario of IoT device security testing, the monitoring device's communication link captures a TCP packet payload that is encrypted with AES.
[0044] Specifically, the analysis engine first strips the protocol header from the data packet, treating the payload as continuous binary data. Then, following the typical block size of encryption algorithms, such as 128 bits per unit, the entire payload is sequentially segmented and concatenated to form a long original bit sequence. This process ensures that subsequent analysis focuses on the bit stream carrying the information itself. After obtaining the bit sequence, a Fourier transform is used to convert it to the frequency domain.
[0045] Understandably, the Fourier transform can reveal the intensity of periodic or quasi-periodic patterns in a sequence.
[0046] For example, a bit sequence of length 1024, where each bit takes the value 0 or 1, can be considered as a discrete signal. Performing a Fast Fourier Transform on this signal yields 512 complex results; taking their magnitudes gives the frequency domain amplitude spectrum. Each point in the amplitude spectrum corresponds to a frequency component, and its amplitude value reflects the energy of the original bit sequence at that frequency fluctuation. Next, statistical distribution characteristics are extracted from the frequency domain amplitude spectrum.
[0047] Specifically, the 512 amplitude values obtained above are used as a sample set, and their histogram distribution is calculated.
[0048] In one embodiment, the analysis system counts the frequency of these amplitude values falling within different numerical ranges, thus forming a distribution profile. This distribution characteristic describes the concentration and dispersion of the encrypted data's energy in the frequency domain. Based on this distribution, its skewness and kurtosis coefficients are calculated. The skewness coefficient measures the asymmetry of the distribution, while the kurtosis coefficient reflects how steep or flat the distribution curve is compared to a normal distribution.
[0049] For example, when detecting a bit sequence of an encrypted video stream, the system calculates that its frequency domain amplitude distribution has a skewness coefficient of 1.8 and a kurtosis coefficient of 5.2. If the preset skewness and kurtosis thresholds are 1.5 and 4.0 respectively, both conditions are met, and the system determines that there is significant distribution non-uniformity. This non-uniformity is often associated with specific frequency points. The system will locate those frequency points with abnormally high or low amplitudes and map their indices back to the original bit sequence time indices.
[0050] For example, if the amplitude of high-frequency bands (corresponding to indices 300 to 350) is found to be generally abnormally high, the mapping relationship can be used to locate the region near bits 6000 to 7000 in the original bit sequence. These regions are marked as preliminary vulnerable areas. Finally, the system merges adjacent or close indices to form a continuous set of vulnerable areas. For example, the above region can be merged with another region near index 7200 because they are very close together, thus forming a preliminary set of vulnerable areas from bit 6000 to bit 7300, providing a clear target range for subsequent in-depth security analysis.
[0051] Step S102: For the preliminary set of vulnerable regions, bit correlation detection technology is used to group the bits in the region, and potential risk regions are determined and a risk region list is formed based on the defect persistence of the grouped bit combinations.
[0052] For the initially identified vulnerable areas, bit correlation detection technology is used to analyze the bit data within the areas. Multiple bit combinations are formed through grouping, resulting in preliminary combination grouping results. Based on these preliminary grouping results, the defect persistence characteristics of each bit combination are obtained and compared using a preset threshold standard. If the defect persistence of a combination exceeds the threshold, it is marked as a high-risk combination, and a high-risk combination list is determined. Using the high-risk combination list, the corresponding vulnerable area location information is obtained, and a mapping method is used to associate high-risk combinations with specific areas to determine the distribution of potential risk areas. Based on the distribution of potential risk areas, the proportion of high-risk combinations in each area is obtained. If the proportion exceeds a preset range, the area is added to the risk area candidate set, resulting in a candidate area set. For the candidate area set, statistical analysis tools are used to perform a secondary verification of the bit combination characteristics of each area. By comparing the defect persistence with historical data, the final risk area list is determined. Based on the final risk area list, detailed bit combination information for each risk area is obtained, and a structured area list is formed through classification and organization, determining the priority ranking results.
[0053] In the context of this application, "defect persistence" refers to the frequency with which a specific bit combination or statistical deviation pattern recurs at the same position in multiple consecutive encrypted data packets; "transformation sensitivity" refers to the degree to which the extracted time-frequency features fluctuate in the frequency domain when the original input bits undergo a very small change (such as a single bit flip). Quantitative evaluation of these terms allows for an effective distinction between normal statistical randomness and genuine algorithmic vulnerability.
[0054] like Figure 8 As shown, the process of reconstructing the bit sequence of encrypted data into a two-dimensional matrix structure is visually presented in the form of paper tape folding. The bit sequence paper tape 801 is a one-dimensional bit stream of the original encrypted data, where black blocks represent bit values of 1 and white blocks represent bit values of 0. In the folding process 802, the bit sequence paper tape is folded row by row according to the rule of 16 bits per row, with each row arranged sequentially to form a multi-layered structure. After folding, a structured matrix 803 is formed, which is a two-dimensional structure of 16 columns and 8 rows, corresponding to 128 basic units. The column direction represents 16 bits per row, and the row direction represents 8 consecutive 16-bit segments. In the structured matrix 803, through detection and analysis by a discrete neural network, abnormal clustering regions 804 can be identified, that is, bit values exhibiting non-random clustering characteristics at specific spatial locations in the matrix. Abnormal clustering regions 804 are marked with dashed ellipses, where dark gray blocks represent the detected defect clustering locations. This two-dimensional matrix structure allows spatial correlation defects that were originally hidden in the one-dimensional bitstream to be visually displayed on a two-dimensional plane, providing a structured data foundation for subsequent spatial dimensionality reduction analysis and defect localization.
[0055] In one possible implementation, for a preliminary set of vulnerable regions, the inter-bit correlation detection technique can be specifically implemented as mutual information analysis.
[0056] For example, in a business scenario involving encrypted video stream data, the initial vulnerable area is identified as a continuous segment of bits after a specific offset in the packet header. To analyze the inherent relationships between these bits, the system divides the bit sequence within this area into groups of four bits each, forming multiple bit combinations.
[0057] Specifically, the system calculates the mutual information value between any two bits within each combination to quantify their statistical dependence. If most bit pairs within a combination exhibit high mutual information values, it indicates that the bits within that combination are not completely independent and may exhibit a fixed pattern introduced by the characteristics of the encryption algorithm or the data format. This combination is then recorded as the preliminary combination grouping result.
[0058] It should be noted that obtaining the persistence characteristics of defects is a crucial step. It does not refer to the error of a single bit, but rather describes the degree to which a non-random, predictable pattern exhibited by a combination of bits repeats itself in multiple consecutive data units.
[0059] For example, when analyzing network protocol encryption payloads, the system tracks whether a preliminarily identified bit combination appears in the same location region in the subsequent 100 data packets. If a specific value sequence of this combination (such as "1010") remains unchanged in more than 70% of the data packets, then its persistence eigenvalue is high. A preset threshold might be set at a persistence rate of 60%. When a combination's persistence eigenvalue exceeds this threshold, it is marked as a high-risk combination, and all such combinations constitute a high-risk combination list. Using this high-risk combination list, the system can perform reverse mapping. Each high-risk combination corresponds to a bit location index from which it originates; these indices, when aggregated, reveal which consecutive segments in the bit sequence generated a large number of high-risk combinations.
[0060] For example, after mapping, high-risk combinations were found to be concentrated in bits 200 to 215 and bits 400 to 410 of the bit sequence. These two regions were identified as potential risk areas, meaning that the bit stream at these positions may not have been sufficiently randomized by the encryption algorithm. Based on the distribution of potential risk areas, the system further calculated the proportion of high-risk combinations in each region to the total number of combinations in that region.
[0061] For example, the first potential region (bits 200-215) has four bit combinations, three of which are marked as high-risk, a proportion as high as 75%. If the preset risk proportion exceeds 50%, this region will undoubtedly be included in the candidate set of high-risk regions. The second region (bits 400-410) may have two combinations, one of which is high-risk, with a proportion of 50%. This proportion just reaches or does not exceed the threshold, and may not be included, depending on the specific threshold setting. For the candidate region set, a second check will introduce more detailed statistical analysis.
[0062] For example, a chi-square test is used to compare the observed persistence pattern of defects with the random distribution at the corresponding positions in historical normal encrypted data streams to see if there is a significant difference. The system will call the historical database to obtain the feature distribution of a large number of encrypted data samples in this area under the same business conditions (such as the same video encoding format and the same key length) as a baseline. If the bit combination characteristics of the current candidate area are statistically significantly inconsistent with the historical baseline, the area is finally identified as a risk area and added to the final list. Finally, the system will perform structured organization based on the final list of risk areas.
[0063] For example, for identified risk areas, the start and end bit indices, the specific value patterns of the high-risk combinations they contain, their persistence strength, and the data packet type (e.g., I-frame data packets) to which the area belongs are recorded. Subsequently, these risk areas are prioritized based on a weighted score of the proportion and persistence strength of the high-risk combinations within the area. Areas with high proportions and strong persistence are given higher priority for priority processing in subsequent encryption algorithm optimization or security audits.
[0064] Step S103: Extract local bit structures from the risk area list, and map the local bit structures to a simplified feature space by combining the processing depth and operational complexity of the encryption algorithm through a spatial dimensionality reduction transformation method. Determine the abnormal distribution structure and form an abnormal distribution set based on the deviation of the mapped distribution from the expected distribution pattern.
[0065] Initial data is obtained from a list of risk areas. Preliminary extraction of local bit structures is performed, and key data units are separated using block processing techniques to obtain preliminary extracted bit fragments. Based on these preliminary bit fragments, and considering the processing depth and operational complexity of the encryption algorithm, layered encryption is implemented to transform the bit fragments into encrypted data units, determining their structural characteristics. For these encrypted data units, a spatial dimensionality reduction transformation method is used to map high-dimensional data to a low-dimensional simplified feature space, obtaining the mapped distribution pattern. This mapped distribution pattern is compared with a pre-established expected distribution model. If the distribution deviation exceeds a preset threshold, it is identified as an anomalous distribution unit and collected into an anomalous distribution set. Further classification processing is performed on the anomalous distribution set, using a support vector machine algorithm to categorize the anomalous distribution units, resulting in classified anomalous subsets. For each of these classified anomalous subsets, data tracing is performed, and the specific source location of the anomalous distribution is determined by combining it with the original records of the risk areas. Based on the specific source location of the anomalous distribution, a corresponding identifier is generated and stored in a pre-set database, completing the location and recording of the anomalous distribution.
[0066] The processing depth of the encryption algorithm refers to the number of logical iteration rounds of the encryption scheme under test; the operational complexity refers to the amount of nonlinear computation (such as the scale of S-box substitution) involved in each round of transformation. In known algorithm detection scenarios, these parameters are directly obtained from the algorithm specifications; in black-box detection scenarios, these parameters are estimated by assessing the local entropy change rate and linear skewness of the bit sequence. By introducing these two parameters as constraints for spatial dimensionality reduction transformation, the scale mapped to the simplified feature space can be adjusted, ensuring that the detection process matches the encryption strength of the data under test.
[0067] Before the spatial dimensionality reduction transformation is performed, the extracted local bit structure is first converted into a high-dimensional numerical vector of fixed dimensions, where each element of the vector is the normalized value of the corresponding bit. The weighting coefficients obtained by combining the processing depth and operation complexity of the encryption algorithm are then multiplied element-wise with the high-dimensional numerical vector to obtain a weighted high-dimensional vector. Principal component analysis is then performed on the weighted high-dimensional vector, and principal components whose cumulative variance contribution rate reaches a preset threshold are selected to form a low-dimensional feature space, thus completing the mapping of the local bit structure.
[0068] like Figure 5 As shown, this invention achieves spatial dimensionality reduction transformation and anomaly distribution detection of high-dimensional bit feature vectors through principal component analysis. Figure 5 (a) A three-dimensional scatter plot shows the distribution of encrypted data bit features in a high-dimensional space, where circular markers represent feature points of normal encrypted data and triangle markers represent feature points of abnormal data. Normal data points show a compact clustered distribution, while abnormal data points are significantly deviated from the normal cluster center. Figure 5The middle section illustrates the PCA dimensionality reduction transformation process, which uses the processing depth and operational complexity of the encryption algorithm as weighting parameters to map the high-dimensional feature space to the low-dimensional representation space. Figure 5 (b) illustrates the data distribution in the two-dimensional feature space after PCA dimensionality reduction. The horizontal and vertical axes correspond to the directions of the first principal component PC1 and the second principal component PC2, respectively. The solid ellipse encloses the expected distribution range of normal data, while the dashed ellipse encloses the deviation area of abnormal data. The spatial separation between the two clearly reflects the significant deviation between the abnormal distribution structure and the expected distribution. This dimensionality reduction visualization method can effectively identify abnormal distribution structures caused by vulnerabilities in encrypted data, providing intuitive low-dimensional feature inputs for subsequent discrete neural network pattern recognition.
[0069] For example, when processing a list of risk areas, initial data is first obtained from the list, such as for a cybersecurity dataset containing multiple bit sequences, which might originate from binary streams in server logs. Preliminary extraction of local bit structures is then performed.
[0070] Specifically, block processing technology can divide a bit sequence into blocks of fixed size, such as groups of 8 bits each. By using a scanning algorithm, key data units can be separated out. These units may be combinations of bits representing abnormal patterns, thus obtaining preliminary extracted bit fragments. For example, in actual business, if the dataset involves encrypted communication traffic, block processing can isolate potentially vulnerable bit fragments, ensuring the targeted nature of subsequent analysis.
[0071] In one embodiment, layered encryption is performed based on these initially extracted bit fragments, taking into account the processing depth and operational complexity of the encryption algorithm.
[0072] Specifically, layered encryption can be understood as a multi-level encryption process. First, a simple XOR operation is used in the shallow layer to process the basic structure of bit segments. Then, the complex rounds of transformation of the AES algorithm are introduced in the deep layer to transform the bit segments into encrypted data units.
[0073] For example, in data protection operations, if a bit fragment represents user privacy data, the encrypted structural characteristics are determined by assessing operational complexity, such as the number of computation rounds. This includes checking the bit entropy value of the encrypted unit to verify the encryption strength, thereby providing a secure intermediate data form for subsequent steps. The goal of this approach is to improve data confidentiality and effectively prevent risks from unauthorized access in business operations.
[0074] For example, spatial dimensionality reduction transformation is used to process encrypted data units.
[0075] Specifically, spatial dimensionality reduction transformation is a technique that maps high-dimensional data to a low-dimensional space. Its principle is to preserve the main variance through principal component analysis (PCA) or similar methods. For example, projecting a 128-dimensional encrypted unit vector onto a 2-dimensional plane yields the mapped distribution. In network security monitoring, this simplifies the visualization of complex data and helps identify distribution patterns.
[0076] In one embodiment, the mapped distribution pattern is compared with a pre-established expected distribution model. The expected model may be a statistical model based on historical normal data, such as a Gaussian mixture model. If the distribution deviation, such as KL divergence, exceeds a preset threshold of 0.5, it is judged as an abnormal distribution unit and collected into the abnormal distribution set.
[0077] For example, in practical applications, if the deviation indicates a potential intrusion signal, this step can quickly screen for anomalies.
[0078] Specifically, further classification processing is carried out based on the abnormal distribution set. The support vector machine algorithm is used to classify the abnormal distribution units. The support vector machine is a supervised learning model. Its principle is to separate different categories by finding the maximum margin hyperplane. For example, in the training phase, the hyperplane parameters are optimized using labeled abnormal samples to obtain the classified abnormal subsets, such as high-risk and low-risk subsets.
[0079] For example, after obtaining the classified abnormal subsets, data tracing is performed for each subset, and the specific source location of the abnormal distribution is determined by combining the original records of the risk area.
[0080] Specifically, the tracing process involves reverse mapping, tracing back from subset features to the original bit position, such as matching timestamps and region IDs in a database query.
[0081] In one embodiment, a corresponding identifier is generated based on the specific source location of the abnormal distribution and stored in a preset database to complete the location and recording. For example, the identifier could be “abnormal source - region A - bit offset 32”, which facilitates subsequent auditing and repair in business operations.
[0082] Step S104: For the set of abnormal distributions, extract the temporal stability and transformation sensitivity features of the structure using time-frequency joint analysis technology, determine high-risk abnormal patterns and form a list of high-risk patterns based on the case where the noise interference in the features exceeds a preset threshold.
[0083] For anomalous distribution data, a time-frequency joint analysis technique is used to decompose the data and obtain structural time-domain stability and transformation sensitivity related feature data. Based on the obtained time-domain stability and transformation sensitivity feature data, the specific value of noise interference is calculated and compared with a preset threshold. If the noise interference value exceeds the preset threshold, it is identified as a potential high-risk anomalous feature. Through further classification of the potential high-risk anomalous features, feature combinations that conform to high-risk patterns are extracted to obtain a preliminary high-risk pattern set. For the preliminary high-risk pattern set, structural analysis methods are used to perform time-domain and frequency-domain cross-validation to determine the core high-risk anomalous patterns in the pattern set. Based on the determined core high-risk anomalous patterns and the interference judgment results, the patterns are prioritized to obtain a ranked list of high-risk patterns. By integrating the ranked high-risk pattern list, a structured anomaly identification result is generated for subsequent business process calls.
[0084] The specific calculation steps for noise interference are as follows: First, calculate the signal-to-noise ratio (SNR) of the feature data after time-frequency joint analysis and normalize it to the 0-1 interval to obtain the normalized SNR value; Second, calculate the spectral flatness of the feature data and normalize it to the 0-1 interval to obtain the normalized spectral flatness value; Third, sum the two normalized values with equal weights to obtain the specific value of noise interference. If the value is greater than 0.5, it is determined that the noise interference exceeds the preset threshold.
[0085] For example, in network security monitoring systems, abnormally distributed data is first processed using time-frequency joint analysis. This technique combines time-domain and frequency-domain analysis, enabling the decomposition of data signals into different frequency components, thereby extracting characteristic data related to the time-domain stability and transform sensitivity of the structure.
[0086] Specifically, the basic principle of time-frequency joint analysis technology is to represent the signal in both time and frequency dimensions simultaneously through methods such as short-time Fourier transform or wavelet transform, thereby avoiding the limitations of single time-domain or frequency-domain analysis.
[0087] In one possible implementation, for an abnormal data stream from a financial trading platform, the system inputs the data sequence into the time-frequency analysis module. First, it performs time-domain decomposition to observe the stability of the data on the time axis, such as whether the fluctuation amplitude is uniform. Then, it extracts transformation-sensitive features, such as rapidly changing frequency peaks, through frequency domain transformation, thereby obtaining a set of feature vectors for subsequent anomaly detection.
[0088] In one possible implementation, a specific value for noise interference is calculated based on these temporal stability and transform sensitivity characteristics. Here, noise interference can be understood as a quantitative indicator of unexpected interference in the data, typically assessed by calculating the signal-to-noise ratio or entropy value.
[0089] For example, in monitoring industrial IoT devices, the system takes the standard deviation of the feature vector as the noise value and compares it with a preset threshold, such as 0.5. If the noise exceeds the threshold, it is identified as a potentially high-risk anomaly. This comparison process helps to identify interference sources that may cause system crashes early on.
[0090] For example, by further classifying these potentially high-risk abnormal features, feature combinations that conform to high-risk patterns can be extracted.
[0091] In one possible implementation, clustering algorithms such as K-means are used to group features to obtain a preliminary set of high-risk patterns. For example, in power grid risk detection, high-frequency change sensitive features are combined into a "sudden load anomaly" pattern.
[0092] In one possible implementation, structural analysis is used to perform cross-validation in the time and frequency domains for a preliminary set of high-risk modes. This approach involves checking the continuity of modes in the time domain and verifying harmonic consistency in the frequency domain to identify core high-risk anomalous modes.
[0093] For example, in medical data analysis, cross-validation can filter out false positives and ensure that core patterns such as "signs of data tampering" are accurately identified.
[0094] For example, based on the identified core high-risk anomaly patterns and the interference assessment results, the patterns are prioritized. The process of obtaining the ranked list of high-risk patterns can be achieved through a weighted scoring mechanism.
[0095] In one possible implementation, an interference weight is assigned to each pattern, such as giving higher priority to high-noise patterns. In a supply chain management system, this helps to prioritize the “supply chain disruption risk” pattern.
[0096] In one possible implementation, structured anomaly identification results are generated by integrating the sorted list of high-risk patterns.
[0097] For example, in intelligent transportation systems, integrating the inventory into a JSON-formatted report can be used by subsequent business processes such as alarm triggers. This integration can improve response efficiency and ensure business continuity.
[0098] Step S105: Obtain bit subsequences with persistent defects according to the high-risk pattern list, group the subsequences using sequence similarity comparison technology, determine the encryption defect pattern based on the pattern predictability of the grouped subsequences, and form a defect pattern group.
[0099] Identify persistent bit subsequences based on a high-risk pattern list. Calculate the similarity distance between subsequences using sequence alignment techniques. If the similarity distance is less than a preset threshold, group the corresponding subsequences into the same group. Determine the encryption defect pattern based on the pattern predictability score of the subsequences within each group. Identify groups with high predictability scores as defect pattern groups. Obtain representative bit sequences from the defect pattern groups as feature patterns.
[0100] The sequence similarity comparison uses Hamming distance as the similarity distance quantification index. Before calculation, all bit subsequences are unified to the same length. Subsequences that are not long enough are padded with 0s at the end, and subsequences that are too long are truncated by the preceding bits according to a fixed window. When calculating the Hamming distance, the bit values of the two subsequences are compared bit by bit, and the number of bits with different values is counted. This number is the similarity distance between the two subsequences. If the distance is less than a preset threshold, they are judged to be similar subsequences and classified into the same group.
[0101] For example, in cryptographic analysis within the cybersecurity field, the first step is to extract bit subsequences that exhibit persistent vulnerabilities from an existing list of high-risk patterns. These subsequences typically refer to bit patterns that recur in the data stream, potentially stemming from weaknesses in the cryptographic algorithm, causing the vulnerability to persist across multiple data samples.
[0102] Specifically, assuming we are dealing with an encrypted data stream from a wireless communication system, the list of high-risk patterns may include identified anomalous bit combinations, such as repeated weak key patterns.
[0103] In one possible implementation, patterns in the list can be scanned to identify subsequences that appear consecutively for more than five periods in the time series. For example, a bit sequence "101010" repeatedly appearing in multiple data packets indicates flaw persistence, as normal encryption should ensure randomness, and such repetition suggests a potential algorithmic vulnerability. This extraction allows focus on persistent flaws that could be exploited by attackers, laying the foundation for subsequent analysis.
[0104] In one possible implementation, sequence alignment techniques are then used to calculate the similarity distance between these subsequences. This technique essentially treats bit sequences as strings and uses algorithms such as edit distance or Hamming distance to quantify the differences between them.
[0105] For example, in a cryptographic flaw detection system, two subsequences, such as "11001100" and "11001101", are selected. Using Hamming distance calculation, they differ only in the last character, with a distance of 1. If a preset threshold of 2 is set, then when the distance is less than the threshold, these subsequences will be grouped into the same group.
[0106] Specifically, sequence alignment involves comparing bit values bit by bit and counting the number of mismatches, without involving complex numerical calculations. This method helps identify clusters of similar patterns; for example, in encrypted banking transaction data, multiple similar subsequences may correspond to the same weak encryption key generation logic, thus revealing systemic flaws.
[0107] For example, continuing this logic, if the similarity distance is less than a preset threshold, the corresponding subsequences are grouped into the same group.
[0108] In one possible implementation, for the example above, the grouping process aggregates all subsequences with a Hamming distance less than 2 into a cluster, such as a cluster containing "11001100", "11001101", and "11001110". This ensures that similar defects are handled centrally, avoiding scattered analysis. The encryption defect pattern is determined based on the pattern predictability score of the subsequences within each group.
[0109] Specifically, predictability scores can be assessed by statistically analyzing the repetition frequency and pattern regularity of subsequences. For example, the entropy of subsequences appearing in a group can be calculated; low entropy indicates high predictability. In banking encryption, if the average score of a group is higher than 0.8 (out of 1), it is considered a flawed pattern because high predictability means that attackers can more easily guess subsequent bits, thus posing a security risk.
[0110] In one possible implementation, groups with high predictability scores are identified as defect pattern groups.
[0111] For example, in the above groupings, if a score reaches 0.85, it is marked as a defective pattern group, which helps to prioritize high-risk items. Finally, representative bit sequences from the defective pattern groups are obtained as feature patterns.
[0112] Specifically, the most frequently occurring sequence within the group, such as "11001100", is selected as a representative for subsequent encryption strengthening.
[0113] For example, in actual business operations, this characteristic pattern can be input into a security audit system to generate alerts or optimize algorithms, thereby improving overall encryption robustness. This method can effectively reduce false alarms and focus on real threats when processing large-scale data.
[0114] Step S106: Extract associated bit sequences from the defect pattern group, fuse time-domain stability and frequency-domain concealment features, generate a structured distribution representation through bit recombination technology, determine the hidden encryption defects based on the distribution non-uniformity and local dependency in the structured distribution representation, and obtain the final defect identifier set.
[0115] Initial data is obtained from the defect pattern group, and associated bit sequences are separated using a bit extraction tool to obtain a preliminary sequence set. Based on this preliminary sequence set, and combined with temporal stability characteristics, a preset threshold is used for filtering. If the temporal fluctuation of a sequence exceeds the threshold, unstable parts are removed, and stable bit segments are identified. For stable bit segments, frequency domain concealment features are integrated, and frequency domain transformation methods are used to analyze the concealment strength. Segments whose concealment meets preset conditions are identified, resulting in filtered feature sequences. Using the filtered feature sequences, bit recombination techniques are applied to construct a structured distribution representation. If the recombined distribution exhibits a non-uniform state, abnormal distribution regions are marked, and abnormal distribution identifiers are obtained. Based on the abnormal distribution identifiers, local dependency characteristics are analyzed. If the dependency strength of a local region is higher than a preset standard, it is determined to be a hidden encryption defect, resulting in a preliminary defect set. For the preliminary defect set, a secondary verification is performed using defect identification methods. A comparison tool is used to match the features of the final identifier set to determine the final defect identifier set.
[0116] Specifically, the bit rearrangement technique refers to rearranging a one-dimensional bit sequence into a multi-dimensional matrix structure according to a preset width and depth. For example, associated bit sequences can be folded into rows of 16 or 32 bits each to construct a two-dimensional bitmap. The purpose of this rearrangement is to transform logical dependencies that originally spanned large distances in the sequence (such as permutation relationships between encryption rounds) into spatial adjacency relationships, thereby generating a structured distributed representation. This allows hidden algorithmic defects to exhibit clustering characteristics that can be recognized by neural networks within a local space.
[0117] When performing bit recombination, if the encrypted data to be tested is streaming encrypted data, it is recombined in units of 128 bits each, with each unit arranged in a 16×8 matrix structure. If it is block encrypted data, the number of rows and columns of the matrix is determined directly according to the block length of the encryption algorithm. A block length of 128 bits is a 16×16 matrix, and 256 bits is a 16×16 or 32×8 matrix. The original order of the bits is maintained during the recombination process, and no additional permutation or scrambling operations are performed.
[0118] In one possible implementation, the initial data obtained from the defect pattern group might be a set of binary data blocks marked as suspicious. The core function of the bit extraction tool here is to extract the associated bit stream from these data blocks based on predefined association rules, such as specific offsets or specific bit operation patterns.
[0119] For example, when analyzing the handshake data packets of an encrypted communication protocol, the tool identifies and extracts all bit fields related to key exchange parameters, and concatenates these bits scattered in different locations in sequence to form a preliminary set of sequences that represent the original bit stream that may contain defects.
[0120] Specifically, for this initial set of sequences, evaluating the temporal stability characteristics is crucial. A sliding window analysis process can be introduced here.
[0121] For example, the sequence set is divided into multiple consecutive segments by timestamp, and the distribution statistics of bits 0 and 1 within each segment are calculated, such as Hamming weight. A preset threshold can be the upper limit of the standard deviation of this statistic. If the statistic of a segment fluctuates drastically and exceeds the threshold range, it indicates that the segment may be affected by random noise interference or normal fluctuations without defects, and is therefore judged as an unstable part and removed. What remains are stable bit segments whose statistical characteristics remain relatively constant across multiple time windows. Next, frequency domain concealment analysis needs to be performed on these stable segments. One embodiment uses the Discrete Fourier Transform (DFT) method. The bit sequence is treated as a binary signal, and it is transformed to the frequency domain to observe its power spectral density distribution. Defect patterns with strong concealment often tend to have frequency domain characteristics similar to white noise, i.e., the power is uniformly distributed across all frequency components. A preset condition can be a spectral flatness threshold. By calculating the spectral flatness of a segment, if its value is higher than the threshold, the segment is considered to have good concealment in the frequency domain and meets the screening criteria, thus obtaining a batch of feature sequences that are both stable and difficult to detect in the frequency domain. Subsequently, these feature sequences are processed using bit recombination techniques. This technique is not a simple splicing, but rather reorganizes the bits in the sequence according to a structured template, such as a matrix or a specific tree structure. After construction, the uniformity of the new structure's distribution needs to be analyzed.
[0122] For example, in a reconstructed two-dimensional bit matrix, the density of 1s in each row or column is counted. If the density of certain rows or columns is found to be significantly higher or lower than the average level, exhibiting a clear "cluster" or "strip"-like uneven distribution, these areas are marked as anomalous distribution indicators, suggesting the possible existence of non-random structural features. Based on these anomalous distribution indicators, their local dependencies are further analyzed.
[0123] For example, within marked abnormal rows, the correlation coefficient between adjacent bits or bits at specific intervals can be calculated. If bits in a local region exhibit strong correlation, exceeding a preset standard expected based on random sequences, then this region is likely to contain traces left by human design or algorithmic flaws, rather than truly random numbers. This step identifies hidden encryption flaws, forming a preliminary flaw set. Finally, a secondary verification is performed on the preliminary flaw set. The flaw identification method here can be a matching verification process based on a known flaw feature library. A comparison tool is used to match the feature patterns of each flaw in the preliminary set with the features in the final identifier set (i.e., an authoritatively verified flaw feature database). Only those flaws that meet strict matching standards are confirmed as genuine, reproducible encryption flaws, thus outputting the final flaw identifier set to guide subsequent algorithm hardening or security auditing.
[0124] In one possible implementation, determining the hidden encryption defects and obtaining the final defect identifier set based on the distribution non-uniformity and local dependency in the structured distribution representation can be: inputting the structured distribution representation into a preset discrete neural network for pattern recognition, using the discrete weight matrix in the discrete neural network to perform nonlinear feature mapping on the distribution representation, and determining the hidden encryption defects based on the distribution non-uniformity and local dependency reflected in the mapping result, thereby obtaining the final defect identifier set.
[0125] like Figure 3 As shown, the discrete neural network used in this invention includes an input layer, two discrete hidden layers, and an output layer. The input layer receives a structured distribution representation, expanding the bit matrix into a one-dimensional feature vector containing multiple input neuron nodes x1 to xn. Discrete hidden layer 1 uses a discrete weight matrix W1 with weights ranging from {-1, 0, 1}, and performs feature transformation using a step activation function. Discrete hidden layer 2 uses a discrete weight matrix W2, also with weights ranging from {-1, 0, 1}, and similarly performs further feature extraction using a step activation function. The output layer uses a sigmoid discrete approximation activation function to output a defect identifier, where 1 indicates a defect and 0 indicates normal operation. Neurons in each layer are interconnected via fully connected layers, with inter-layer connections representing the corresponding discrete weights. The output results are used for distribution non-uniformity detection and local dependency detection, thereby identifying hidden defects in encrypted data.
[0126] In one possible implementation, the discrete neural network employed in this application serves as the core discriminative tool for detecting the target. This network consists of an input layer, multiple discrete hidden layers, and an output layer. Its key features are: first, the discretization of weights; the connection weights in the network are not continuous floating-point numbers, but are constrained to sets such as {-1, 0, 1} or specific fixed-point numbers. This allows the network to more sensitively capture extreme and discrete statistical deviations in the encrypted bitstream; second, the discretization of the activation function; a step function is used as the activation method for neurons, mapping complex time-frequency features into discrete risk level signals. Through this structure, the discrete neural network can effectively filter statistical noise caused by pseudo-randomness in encrypted data, focusing on the nonlinear vulnerability patterns caused by algorithm design flaws, thereby providing accurate defect identification.
[0127] The nonlinear feature mapping execution process of the discrete neural network described in this application is as follows: the structured distribution representation is expanded into a one-dimensional feature vector by rows, and input into the input layer of the discrete neural network. After multiplication of the discrete weight matrix of the input layer and the first discrete hidden layer, it is input into the step-type discrete activation function for nonlinear transformation. The activation result is passed to the next discrete hidden layer to repeat the above operation. Finally, it is processed by the sigmoid discrete approximation activation function of the output layer. When the output value of the output layer is 1, it is determined that the corresponding region has uneven distribution or local dependence. When the output value is 0, it is determined to be a normal region. In this case, if the weight matrix is updated to a value other than {-1,0,1} during network training, the discrete value closest to the value is directly taken as the final weight to ensure the discrete characteristics of the network.
[0128] like Figure 6 As shown, the discrete neural network feature mapping mechanism of the present invention includes a complete processing chain from input to output. Figure 6 The upper left section shows a 16x8 bit matrix as the input form for a structured distributed representation. Black cells represent bit values of 1, and white cells represent bit values of 0. This matrix is formed by folding and rearranging a one-dimensional bit sequence. After being expanded, the bit matrix is transformed into a 128-dimensional one-dimensional vector and then input into a discrete neural network. Figure 6 The middle section displays a visualization of the weight distribution of two discrete weight matrices W1 and W2, where the weights are strictly constrained to three discrete values: negative 1, 0, and positive 1, represented by black, gray, and white color codes, reflecting the core feature of discretization of network weights. Figure 6 The upper right part shows the discretization effect of the step activation function. The function outputs 1 when the weighted sum of the inputs is greater than or equal to zero, indicating that a defect has been detected, and outputs 0 when the sum is less than zero, indicating that there is no defect. This achieves complete discretization of the activation output. Figure 6The lower section displays the defect recognition results of the network output layer, presented as a binary heatmap. Gray areas represent the locations of detected defects, and dashed boxes mark the spatial locations of defect regions A and B, respectively. The overall mechanism diagram highlights two core features of discrete neural networks: weight discretization and activation discretization, which enable the network to perform efficient pattern recognition on structured distribution representations.
[0129] like Figure 9 As shown, this invention conducted a comparative experimental analysis of the vulnerability region detection accuracy of six mainstream encryption algorithms (AES-128, AES-256, DES, 3DES, Blowfish, and ChaCha20). The left vertical axis of the figure displays the detection accuracy of traditional detection methods and the method of this invention on each encryption algorithm in the form of a bar chart, while the right vertical axis displays the percentage improvement in accuracy in the form of a line graph. Experimental results show that the detection accuracy of traditional methods on different encryption algorithms ranges from 65% to 78%, while the detection accuracy of the method of this invention reaches over 90% in all algorithms, with the highest accuracy of 96% for the DES algorithm. In terms of the improvement in accuracy, the method of this invention improves upon traditional methods by 23.1% to 40.0% across all algorithms, with the most significant improvement of 40.0% for the AES-256 algorithm. These results fully verify that the technical solution of this invention, which uses frequency domain transformation to initially locate vulnerable regions and combines bit-to-bit correlation detection with spatial dimensionality reduction transformation, can significantly improve the accuracy of vulnerability detection in encrypted data and has good algorithmic universality.
[0130] like Figure 10 As shown, this invention uses a heatmap to illustrate the combined influence of the defect persistence threshold and the data sample size on the defect pattern detection rate. The horizontal axis represents the defect persistence threshold, ranging from 0.3 to 0.9; the vertical axis represents the data sample size, covering different scales from 1K to 500K. The grayscale value and labeled value of each cell in the heatmap represent the defect pattern detection rate under the corresponding conditions. The experimental results show a clear gradient distribution: when the defect persistence threshold is low and the sample size is large, the detection rate can reach 98%; while when the threshold is high and the sample size is small, the detection rate drops to about 45%. Within the optimal working area (threshold 0.5 to 0.7, sample size 10K to 100K) marked by the dashed box in the figure, the detection rate is stable between 75% and 93%, balancing detection accuracy and computational efficiency. This result shows that the time-frequency joint analysis and high-risk pattern screening mechanism of this invention can effectively lock bit subsequences with defect persistence under reasonable parameter configuration, providing a clear parameter selection basis for practical engineering deployment.
[0131] like Figure 11As shown, this invention comprehensively analyzes the processing throughput and end-to-end detection latency under different input data stream rates. The horizontal axis represents the input data stream rate, covering a range from 10 Mbps to 1000 Mbps. The left vertical axis displays the throughput distribution of the three processing stages—frequency domain analysis, feature extraction, and neural network recognition—in the form of a stacked area plot, while the right vertical axis displays the end-to-end detection latency in the form of a broken line. Experimental results show that as the input data stream rate increases, the processing throughput of all three stages exhibits a near-linear growth trend, with the neural network recognition stage exhibiting the highest throughput, reaching 900 Mbps at an input rate of 1000 Mbps. Regarding latency, when the input rate is 100 Mbps, the end-to-end detection latency is only 12 ms; even under the high-rate condition of 1000 Mbps, the latency is controlled within 95 ms. These results verify that the technical solution of this invention, which integrates multi-dimensional features and achieves efficient recognition through discrete neural networks, has excellent processing efficiency and good scalability, and can meet the application requirements of real-time detection of large-scale encrypted data streams.
[0132] It should be noted that the above examples are merely some specific embodiments of the present invention. Obviously, the present invention is not limited to the above embodiments and many variations are possible. All variations that can be directly derived or conceived by those skilled in the art from the content disclosed in this invention should be considered within the scope of protection of this invention.
Claims
1. A method for detecting encryption data vulnerability based on a discrete neural network, characterized in that, The method includes: Obtain the bit sequence of the encrypted data stream, use frequency domain transformation technology to extract the frequency domain distribution characteristics of the bit sequence, determine the preliminary vulnerable regions based on the distribution non-uniformity in the frequency domain distribution characteristics, and form a preliminary vulnerable region set; For the initial set of vulnerable regions, bit correlation detection technology is used to group the bits in the region. Based on the persistence of defects in the grouped bit combinations, potential risk regions are determined and a risk region list is formed. Local bit structures are extracted from the risk area list. Combining the processing depth and operational complexity of the encryption algorithm, the local bit structures are mapped to a simplified feature space through a spatial dimensionality reduction transformation method. Based on the deviation of the mapped distribution from the expected distribution pattern, the abnormal distribution structure is determined and an abnormal distribution set is formed. For the set of abnormal distributions, the temporal stability and transformation sensitivity features of the structure are extracted by time-frequency joint analysis. Based on the case where the noise interference in the features exceeds a preset threshold, high-risk abnormal patterns are determined and a list of high-risk patterns is formed.
2. The method of claim 1, wherein, Also includes: According to the high-risk pattern list, obtain the bit subsequences with persistent defects, use sequence similarity comparison technology to group the subsequences, and determine the encryption defect patterns and form defect pattern groups based on the pattern predictability of the grouped subsequences. The associated bit sequences are extracted from the defect pattern group, and the time-domain stability and frequency-domain concealment features are fused together. A structured distribution representation is generated through bit recombination technology. The hidden encryption defects are determined based on the distribution non-uniformity and local dependency in the structured distribution representation, and the final defect identifier set is obtained.
3. The method of claim 2, wherein, The step of determining hidden encryption defects and obtaining the final defect identifier set based on the distribution non-uniformity and local dependency in the structured distribution representation includes: The structured distribution representation is input into a preset discrete neural network for pattern recognition. The discrete weight matrix in the discrete neural network is used to perform nonlinear feature mapping on the distribution representation. The hidden encryption defects are determined based on the distribution non-uniformity and local dependence reflected in the mapping result, thereby obtaining the final defect identifier set.
4. The method of claim 2, wherein, The process of obtaining the bit sequence of the encrypted data stream involves using frequency domain transformation technology to extract the frequency domain distribution characteristics of the bit sequence, determining preliminary vulnerable regions based on the distribution non-uniformity in the frequency domain distribution characteristics, and forming a preliminary vulnerable region set, including: Acquire the encrypted data stream and parse out the bit sequence; The frequency domain amplitude spectrum is obtained by processing the bit sequence using Fourier transform; Extract the statistical distribution of amplitude values from the frequency domain amplitude spectrum as the frequency domain distribution characteristics; Calculate the skewness and kurtosis coefficients of the frequency domain distribution characteristics; If the skewness coefficient is greater than a preset threshold or the kurtosis coefficient is greater than another preset threshold, it is determined that there is uneven distribution. By mapping the frequency points corresponding to the uneven distribution back to the bit sequence index, the initial vulnerable areas can be determined; Merging adjacent indexes forms a contiguous region, constituting a preliminary set of vulnerable regions.
5. The method of claim 2, wherein, For the initial set of vulnerable regions, bit correlation detection technology is used to group the bits within the region. Based on the persistence of defects in the grouped bit combinations, potential risk regions are determined and a risk region list is formed, including: For the initially identified vulnerable areas, bit correlation detection technology is used to analyze the bit data within the areas, and multiple bit combinations are formed by grouping to obtain preliminary combined grouping results; Based on the preliminary combination grouping results, the defect persistence characteristics of each bit combination are obtained, and a preset threshold standard is used for comparison. If the defect persistence of a combination exceeds the threshold, it is marked as a high-risk combination, and a list of high-risk combinations is determined. By using a list of high-risk combinations, the location information of the corresponding vulnerable areas is obtained. A mapping method is then used to associate the high-risk combinations with specific areas to determine the distribution of potential risk areas. Based on the distribution of potential risk areas, the proportion of high-risk combinations in each area is obtained. If the proportion exceeds the preset range, the area is included in the risk area candidate set, thus obtaining the candidate area set. For the candidate region set, statistical analysis tools are used to perform a secondary check on the bit combination characteristics of each region. By comparing the persistence of defects with the consistency of historical data, the final list of risk regions is determined. Based on the final list of risk areas, detailed bit combination information for each risk area is obtained, and a structured list of areas is formed by classifying and organizing them to determine the priority ranking result.
6. The method of claim 2, wherein, The process involves extracting local bit structures from the risk region list, combining the processing depth and operational complexity of the encryption algorithm, mapping the local bit structures to a simplified feature space using a spatial dimensionality reduction transformation method, and determining the abnormal distribution structure and forming an abnormal distribution set based on the deviation of the mapped distribution from the expected distribution pattern. Initial data is obtained from the risk area list, and preliminary extraction is performed on the local bit structure. Key data units are separated using block processing technology to obtain the preliminary extracted bit fragments. Based on the initially extracted bit fragments, and considering the processing depth and operational complexity of the encryption algorithm, a layered encryption operation is performed to transform the bit fragments into encrypted data units and determine the structural characteristics of the encrypted data. For the encrypted data units, a spatial dimension reduction transformation method is used to map the high-dimensional data to a low-dimensional simplified feature space to obtain the distribution pattern after mapping. The mapped distribution pattern is compared with the pre-established expected distribution model. If the distribution deviation exceeds the preset threshold, it is judged as an abnormal distribution unit and collected into the abnormal distribution set. Based on the set of abnormal distributions, further classification processing is carried out. The support vector machine algorithm is used to classify the abnormal distribution units to obtain a classified subset of abnormalities. Obtain the categorized subsets of anomalies, trace the data for each subset, and combine the original records of the risk areas to determine the specific source of the anomaly distribution. By identifying the specific source location of the abnormal distribution, a corresponding identifier is generated and stored in a pre-defined database, thus completing the location and recording of the abnormal distribution.
7. The method according to claim 2, characterized in that, For the set of abnormal distributions, the temporal stability and transformation sensitivity features of the structure are extracted using time-frequency joint analysis. Based on the cases where noise interference in these features exceeds a preset threshold, high-risk anomaly patterns are identified and a list of high-risk patterns is generated, including: For anomalous distribution data, time-frequency joint analysis technology is used to decompose the data and obtain the temporal stability and transformation sensitivity characteristics of the structure. Based on the acquired temporal stability and transformation sensitivity characteristic data, the specific value of noise interference is calculated and compared with a preset threshold. If the noise interference value exceeds the preset threshold, it is judged as a potential high-risk anomaly. By further classifying and processing the potential high-risk abnormal features, feature combinations that match the high-risk patterns are extracted to obtain a preliminary set of high-risk patterns. For the initial set of high-risk patterns, structural analysis was used to perform cross-validation in the time and frequency domains to identify the core high-risk anomaly patterns in the set. Based on the identified core high-risk anomaly patterns and the interference judgment results, the patterns are prioritized and a list of high-risk patterns is obtained. By integrating the sorted list of high-risk patterns, a structured anomaly identification result is generated for subsequent business processes.