Redundant information deduplication processing method and system based on data feature recognition
By analyzing the correlation contribution and redundancy confidence of Wi-Fi router terminal feature data, the problem of redundancy identification of high-dimensional sparse feature data in existing technologies is solved, and a more accurate redundancy deduplication effect is achieved.
Patent Information
- Application Number
- CN202610773145.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-01
- Publication Date
- 2026-06-30
AI Technical Summary
Existing technologies cannot effectively identify nonlinear correlations and redundant feature combinations when processing high-dimensional sparse feature data from smart home Wi-Fi routers, leading to missed and false judgments and affecting the deduplication effect of redundant data.
By reconstructing the behavioral trajectory of features, calculating the feature activity index, hysteresis energy value, and response misalignment offset, generating correlation contribution, analyzing the redundancy strength of feature combinations, and combining local clustering density and unidirectional locking value to eliminate environmental interference and accurately identify effective redundant feature combinations.
It effectively preserves the nonlinear correlation and local structural information of high-dimensional sparse data, reduces missed and false judgments, and improves the redundancy deduplication accuracy of Wi-Fi router terminal feature data.
Smart Images

Figure CN122310065A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data deduplication technology, specifically to a method and system for deduplicating redundant information based on data feature recognition. Background Technology
[0002] Currently, when identifying the data characteristics of Wi-Fi routers, a combination of feature dimensionality reduction and similarity comparison is often used. That is, high-dimensional features are compressed using dimensionality reduction algorithms such as PCA and LDA, and then the similarity of the feature vectors after dimensionality reduction is calculated. The data redundancy is then determined based on the similarity threshold. However, the above deduplication methods still have the following drawbacks: In smart homes, the terminal connection feature data of Wi-Fi routers usually exhibits high-dimensional sparse characteristics. Each data point contains features of multiple dimensions, and there are non-linear correlations and redundancies between some features. When existing technologies use linear dimensionality reduction methods such as PCA to process such data, they will lose the local structure and high-order correlation information between features, resulting in the inability to effectively identify redundant feature combinations at the correlation level, which in turn leads to missed and misjudged features, affecting the deduplication effect of redundant data. Summary of the Invention
[0003] To address the shortcomings of existing technologies, this invention provides a method and system for deduplicating redundant information based on data feature recognition, thus solving the aforementioned problems.
[0004] The above-mentioned technical objective of the present invention is achieved through the following technical solution: Redundancy removal methods based on data feature recognition include: Step S1: Obtain terminal feature data of the target object; based on the preprocessed terminal feature data, analyze the correlation between each feature and generate a representation of the correlation contribution between features. Step S2: Based on the correlation contribution, analyze the distribution and density of feature combinations to generate a feature correlation redundancy set representing the redundancy strength of each feature combination. Step S3: Based on the feature association redundancy set, analyze the redundancy authenticity of each feature combination to obtain the redundancy confidence level representing the redundancy probability of each feature combination. Step S4: Analyze the redundancy confidence and determine whether the features are effective redundant feature combinations to obtain the deduplicated feature set.
[0005] Furthermore, based on the preprocessed terminal feature data, the correlation between various features is analyzed, and the contribution of the correlation between features is generated, including: For the preprocessed terminal feature data, the trajectory of each feature is reconstructed, and the oscillation amplitude and duration of the rate of change in the trajectory are analyzed to obtain the feature activity index representing the chaotic amplitude of feature changes.
[0006] Furthermore, based on the preprocessed terminal feature data, the correlation between various features is analyzed, and a contribution degree representing the correlation between features is generated, which also includes: Based on the feature activity index, the impact of the feature change on the other features is analyzed, and a hysteresis energy value is generated to represent the hysteresis effect of the feature change on the other features. Based on the terminal feature data, the response of the features to interference is analyzed, and a response misalignment offset representing the degree of response of the two features to external interference is generated.
[0007] Furthermore, based on the preprocessed terminal feature data, the correlation between various features is analyzed, and a contribution degree representing the correlation between features is generated, which also includes: The feature activity index, hysteresis energy value, and response misalignment offset are fused to generate a representation of the correlation contribution between features.
[0008] Furthermore, based on the correlation contribution, the distribution and density of feature combinations are analyzed to generate a feature correlation redundancy set representing the redundancy strength of each feature combination, including: Analyze the distribution pattern of the correlation contribution of each feature combination, determine the inflection point interval of the correlation contribution transitioning from sparse to dense regions, and obtain the redundancy perturbation range representing the change in the correlation strength between feature combinations. Within each redundancy perturbation range, the local clustering density of each feature combination is calculated and divided to obtain a redundancy stacking hierarchy representing the density of feature combinations under different redundancy intensities.
[0009] Furthermore, based on the correlation contribution, the distribution and density of feature combinations are analyzed to generate a feature correlation redundancy set representing the redundancy strength of each feature combination, which also includes: By fusing the range of redundancy perturbation with the redundancy stacking level, a feature-associated redundancy set representing the degree of redundancy strength of each feature combination is generated.
[0010] Furthermore, based on the feature association redundancy set, the redundancy authenticity of each feature combination is analyzed to obtain the redundancy confidence level representing the redundancy probability of each feature combination, including: For each feature combination in the feature-associated redundant set, analyze the same-direction and opposite-direction changes of the two to obtain the same-direction locking value representing the degree of feature synchronization; Based on the signal strength of the terminal feature data, the fluctuations between feature combinations are analyzed to obtain the stripping binding coefficient, which represents the strength of the feature combination association binding after interference removal.
[0011] Furthermore, based on the feature association redundancy set, the redundancy authenticity of each feature combination is analyzed to obtain the redundancy confidence level representing the redundancy probability of each feature combination, which also includes: By fusing the same-direction locking value with the stripping binding coefficient, the true redundancy of each feature combination is analyzed, and the redundancy confidence of the redundancy probability of each feature combination is obtained.
[0012] Furthermore, the redundancy confidence is analyzed, and it is determined whether the features are effective redundant feature combinations, resulting in a deduplicated feature set, including: The redundancy confidence of each feature combination is sorted and analyzed, and then removed and retained to obtain the deduplicated feature set.
[0013] Furthermore, a redundancy deduplication system based on data feature recognition, applied to the above processing method, includes: The data analysis unit is used to acquire terminal feature data of the target object, analyze the correlation between features based on the preprocessed terminal feature data, and generate a representation of the correlation contribution between features. The redundancy analysis unit is used to analyze the distribution and density of feature combinations based on their correlation contribution, and generate a feature correlation redundancy set representing the degree of redundancy of each feature combination. The probability analysis unit is used to analyze the redundancy authenticity of each feature combination based on the feature association redundancy set, and obtain the redundancy confidence level representing the redundancy probability of each feature combination. The deduplication unit is used to analyze the redundancy confidence and determine whether the features are effective redundant feature combinations, thus obtaining the deduplicated feature set.
[0014] In summary, the present invention has the following main beneficial effects: By reconstructing the behavioral trajectory of each feature and calculating the feature activity index, hysteresis energy value, and response misalignment offset, the nonlinear correlation and local structural information in high-dimensional sparse data can be fully preserved, avoiding the loss of higher-order correlations caused by linear projection. This accurately reflects the correlation contribution between features. At the same time, the inflection point interval of the correlation contribution transitioning from sparse to dense is adaptively located through the redundancy perturbation range, and the redundancy stacking level is divided by combining local clustering density, effectively overcoming the misjudgment problem caused by fixed thresholds. This allows the feature correlation redundancy set to truly reflect different redundancy strengths. Moreover, the same-direction locking value can accurately capture the consistency of the sign change of features within the local time window. Combined with the stripping binding coefficient to remove public environmental interference, the true fluctuation characteristics of the terminal itself are preserved, thereby reducing false redundancy judgments caused by environmental noise. Finally, the natural segmentation point and reverse feature pair removal method are used to obtain the deduplicated feature set. This solution improves the redundancy deduplication accuracy of smart home Wi-Fi router terminal feature data, reduces missed judgments and false judgments, and ensures the deduplication effect of redundant data. Attached Figure Description
[0015] Figure 1This is a flowchart illustrating the steps of the redundant information deduplication method based on data feature recognition of the present invention. Figure 2 This is a schematic diagram of the redundant information deduplication system based on data feature recognition of the present invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] refer to Figure 1 and Figure 2 Redundancy removal methods based on data feature recognition include: Step S1: Obtain terminal feature data of the target object; based on the preprocessed terminal feature data, analyze the correlation between each feature and generate a representation of the correlation contribution between features. The target is Wi-Fi routers in the home; Terminal characteristic data includes: connection duration, connection frequency, signal strength, and data transmission volume between the Wi-Fi router and each terminal. The terminal can be a mobile phone, computer, or other terminal device. Step S2: Based on the correlation contribution, analyze the distribution and density of feature combinations to generate a feature correlation redundancy set representing the redundancy strength of each feature combination. Step S3: Based on the feature association redundancy set, analyze the redundancy authenticity of each feature combination to obtain the redundancy confidence level representing the redundancy probability of each feature combination. Step S4: Analyze the redundancy confidence and determine whether the features are effective redundant feature combinations to obtain the deduplicated feature set.
[0018] In one embodiment, based on the preprocessed terminal feature data, the correlation between various features is analyzed to generate a contribution degree representing the correlation between features, including: For the preprocessed terminal feature data, the trajectory of each feature is reconstructed, and the oscillation amplitude and duration of the rate of change in the trajectory are analyzed to obtain the feature activity index representing the chaotic amplitude of feature changes. Specifically, this includes: for the preprocessed terminal feature data, the behavior trajectory sequence of each feature is independently reconstructed according to the sampling point order, with the horizontal axis representing the sampling point number and the vertical axis representing the feature value; the rate of change between two adjacent sampling points in the trajectory sequence is calculated; after traversing all adjacent sampling points, a set of rate of change sequences is obtained, and the median of the rate of change sequence is found. Whenever the rate of change is greater than the median, it is recorded as an effective fluctuation. The sampling point number of each effective fluctuation is recorded, and the interval between two adjacent effective fluctuations is calculated. The interval is represented by the number of sampling points. The average of the reciprocals of all intervals is taken as the oscillation amplitude. Re-traverse the rate of change sequence, starting from the first sampling point. Whenever a sampling point with a rate of change ≤ the median is encountered, it is marked as the starting point of a continuous time period. Then continue traversing until a sampling point with a rate of change greater than the median is encountered. At this point, all sampling points from the starting point to the previous sampling point constitute a continuous time period. The number of sampling points in this continuous time period is the length of the time period. Find all such continuous time periods according to this rule, sum up the lengths of each time period, divide by the total number of sampling points, and get the duration percentage. Subtracting the duration percentage from 1 and multiplying by the oscillation amplitude, then normalizing the result to the 0-1 range, yields the feature activity index, which represents the degree of disorder in feature changes. The larger the feature activity index, the more disordered the feature changes.
[0019] In one embodiment, based on the preprocessed terminal feature data, the correlation between various features is analyzed to generate a contribution degree representing the correlation between features, which further includes: Based on the feature activity index, the impact of the feature change on other features is analyzed, and a hysteresis energy value is generated to represent the hysteresis effect of the feature change on other features. Specifically, this includes: extracting the sampling point numbers of all effective fluctuations from the feature's behavior trajectory, numbering them in chronological order, and adding up the number of effective fluctuations to obtain the total number. Set a fixed hysteresis search window length, which is an integer greater than 0. For each valid fluctuation, sequentially obtain the rate of change of the feature at the 1st, 2nd, and up to the window length sampling point after the fluctuation occurs. Under a fixed number of shift steps, add the absolute values of the feature rate of change corresponding to the total number of times, and then divide by the total number of times to obtain the average hysteresis response value under that number of shift steps. Calculate the average hysteresis response value for each number of shift steps in this way. Find the maximum value of the average hysteresis response value under all shift steps, and multiply the maximum value by the feature activity index of the feature to obtain the hysteresis energy value. The larger the hysteresis energy value, the stronger the hysteresis effect of the feature on other features due to its own chaotic changes.
[0020] Based on the terminal feature data, analyze the response of the features to interference, and generate a response misalignment offset that represents the degree of response of the two features to external interference. Specifically, this includes: selecting two features to be analyzed, denoted as feature A and feature B respectively, finding the sampling point number of all effective fluctuations of each feature, and sorting them according to their occurrence time. Set the matching time window length to an integer greater than 0. For each valid fluctuation moment of feature A, in the valid fluctuation moment sequence of feature B, search for a fluctuation moment whose absolute value of the difference between its sampling point index and the fluctuation moment of feature A does not exceed the window length. If it exists, mark this pair of fluctuations as a matching pair and record the difference between the two fluctuation moments. The difference is in units of sampling points and is also positive or negative. A positive sign indicates that feature B lags behind feature A, and a negative sign indicates that feature B is ahead. After traversing all valid fluctuation moments of feature A, a set of difference sequences is obtained. The absolute value of the difference with the highest frequency in this set of difference sequences is taken as the response misalignment offset. If there are multiple differences with the highest frequency, the arithmetic mean of the absolute values of these parallel differences is taken as the response misalignment offset. The response misalignment offset represents the average time misalignment of the responses of two features to the same external disturbance, and is expressed in units of sampling points.
[0021] In one embodiment, based on the preprocessed terminal feature data, the correlation between various features is analyzed to generate a contribution degree representing the correlation between features, which further includes: The feature activity index, hysteresis energy value, and response misalignment offset are fused to generate a value representing the correlation contribution between features. Specifically, for all feature pairs to be analyzed, such as feature A and feature B, feature A is taken as the source feature and feature B is the feature of the receiving target. The feature activity index of the source feature is multiplied by the hysteresis energy value of the source feature to obtain the first product value. Then, 1 is divided by the sum of 1 and the response misalignment offset of the feature pair to obtain the second product value. The first product value and the second product value are multiplied to obtain the temporary fusion value of the feature pair. Normalize the temporary fusion values of all feature pairs to the 0-1 range. This value represents the correlation contribution between features. The larger the correlation contribution value, the stronger the correlation between the feature pairs.
[0022] By reconstructing the behavioral trajectory of each feature, a feature activity index is calculated to reflect the degree of disorder in feature changes. The hysteresis energy value is combined to determine the hysteresis effect of feature changes on other features. At the same time, the response misalignment offset is used to capture the response time misalignment of features to external disturbances. Finally, the correlation contribution between features is obtained by fusion. This can effectively preserve the nonlinear correlation and local structural information in high-dimensional sparse data, avoid the loss of high-order correlation information between features due to dimensionality reduction, and thus more accurately identify redundant feature combinations, reduce missed judgments and false judgments, and improve the redundancy deduplication effect of Wi-Fi router terminal feature data in smart home scenarios.
[0023] In one embodiment, based on the correlation contribution, the distribution and density of feature combinations are analyzed to generate a feature correlation redundancy set representing the redundancy strength of each feature combination, including: Analyze the distribution pattern of the correlation contribution of each feature combination, determine the inflection point interval of the correlation contribution from the sparse region to the dense region, and obtain the redundancy perturbation range representing the change in the correlation strength between feature combinations. Specifically, for the correlation contribution between all feature pairs, arrange these correlation contributions in ascending order to obtain an ordered sequence. The difference between two adjacent values in an ordered sequence is calculated to obtain a difference sequence. These differences reflect the density variation of the correlation contribution on the numerical axis. The larger the difference, the wider the interval between the values, i.e., the sparse region. The smaller the difference, the denser the values, i.e., the dense region. Calculate the arithmetic mean and standard deviation of the difference sequence. Use the mean plus the standard deviation as a judgment threshold to filter out all differences greater than the judgment threshold. The positions corresponding to these differences are significant transition points from sparse to dense regions. Construct a numerical interval for the correlation contribution corresponding to the minimum and maximum values in these positions to obtain the inflection point interval. The correlation contribution values within this inflection point interval are in the transition zone between sparse and dense regions. The inflection point interval is the redundant perturbation range representing the change in the correlation strength between feature combinations. The correlation contribution within the redundant perturbation range indicates that the correlation strength of the corresponding feature combination is in an unstable state of change, and its redundancy is easily affected by data fluctuations.
[0024] Within each redundancy perturbation range, the local clustering density of each feature combination is calculated and divided to obtain a redundancy stacking level representing the density of feature combinations under different redundancy intensities. Specifically, for all correlation contributions within the redundancy perturbation range, each correlation contribution corresponds to a feature combination. For the number of these correlation contributions, if the number is less than 3, all feature combinations within the redundancy perturbation range are classified into a single level and are not further subdivided. For each feature combination within the redundancy perturbation range, with its correlation contribution as the center, calculate the absolute distance from the correlation contribution to all other correlation contributions within the redundancy perturbation range. Find the minimum value among these distances. If the minimum value is greater than 0, then half of the minimum value is taken as the local clustering radius of the point. If the minimum value is 0, that is, there is at least one other point with the same value as the point, then find the minimum value among all absolute distances greater than 0 for the point, and half of the minimum value is taken as the local clustering radius. If the point has no non-zero distances, that is, it means that the values of all other points are exactly the same as the point, and all distances are 0, then the local clustering radius is 0.01. The number of other related contributions within the redundancy perturbation range whose distance from the center is less than or equal to the local cluster radius is the local cluster density of the feature combination. By traversing all feature combinations within the redundancy perturbation range, multiple local cluster density values are obtained. Arrange these local cluster density values in ascending order, calculate the difference between two adjacent local cluster density values, find the position with the largest difference, and use the local cluster density value corresponding to that position as the division threshold. Feature combinations with local cluster density values ≤ the division threshold are classified as low-density layers, and feature combinations with local cluster density values > the division threshold are classified as high-density layers. Low-density levels and high-density levels are denoted as the first and second levels in the redundant stacking hierarchy, respectively. The second level indicates a higher degree of density. Each level contains several feature combinations, thus obtaining the redundant stacking hierarchy that represents the density of feature combinations under different redundancy intensities.
[0025] In one embodiment, based on the correlation contribution, the distribution and density of feature combinations are analyzed to generate a feature correlation redundancy set representing the redundancy strength of each feature combination, which further includes: The redundancy perturbation range is fused with the redundancy stacking level to generate a feature-related redundancy set representing the redundancy strength of each feature combination. Specifically, the lower limit and upper limit of the redundancy perturbation range are denoted as the low boundary value and the high boundary value, respectively. For feature combinations whose correlation contribution is within the range of redundancy perturbation, the redundancy value is calculated according to their redundancy stacking level: if it belongs to the first level, its local cluster density is divided by the maximum value of all local cluster densities within the range of redundancy perturbation to obtain the ratio, and then the ratio is multiplied by the low boundary value to obtain the redundancy value of the feature combination. If it belongs to the second level, divide its local cluster density by the maximum value of all local cluster densities within the redundancy perturbation range to obtain the ratio. Then multiply the ratio by the high boundary value to obtain the redundancy value. The redundancy value represents the strength. If all local cluster densities within the redundancy perturbation range are zero, the average of the low boundary value and the high boundary value is used as the redundancy value. Normalize the redundant values of all feature combinations to the 0-1 range. The set of redundant values is the feature association redundancy set that represents the degree of redundancy of each feature combination.
[0026] By analyzing the distribution pattern of the contribution of feature combination associations, the range of redundancy perturbation is determined, and the transition zone from sparse to dense association strength is accurately captured, avoiding misjudgment caused by fixed thresholds. At the same time, by combining local clustering density and redundancy stacking level, redundancy values are calculated for feature combinations with different densities, which can adaptively distinguish feature association relationships under high and low redundancy strength. This scheme does not require compression of high-dimensional sparse data, and fully preserves the nonlinear association and local structural information between features. It can effectively identify redundant feature combinations at the association level and reduce the rate of missed and false judgments.
[0027] In one embodiment, based on the feature association redundancy set, the redundancy authenticity of each feature combination is analyzed to obtain a redundancy confidence level representing the redundancy probability of each feature combination, including: For each feature combination in the feature-associative redundant set, the co-directional and inverse changes of the two are analyzed to obtain a co-directional locking value representing the degree of feature synchronization. Specifically, for feature A and feature B to be analyzed, the signed rate of change sequence of each feature is calculated. That is, according to the sampling point order, the feature value of the next sampling point is subtracted from the feature value of the previous sampling point to obtain the difference. Then, the difference is divided by the feature value of the previous sampling point to obtain the signed rate of change. The signed rate of change can be positive or negative. Arrange the signed rates of change according to the sampling point order to form a signed rate of change sequence. Map the signed rates of change greater than 0 in the signed rate of change sequences of feature A and feature B to +1, the signed rates of change less than zero to -1, and the signed rates of change equal to zero to 0, and you can get two sign sequences. Set the length of the sliding window to the square root of the total number of sampling points, round the square root to the nearest integer, add 1 to adjust it to an odd number if the integer is even, and keep it unchanged if it is odd. This length adapts to the number of sampling points. For each window center position, extract the symbol values within the current window coverage area from the two symbol sequences to obtain two symbol segments. Calculate the symbol consistency between these two segments by dividing the number of identical symbols at corresponding positions within the window by the window length to obtain the local consistency rate of the window. Simultaneously, calculate the cumulative sum of symbols in the two segments within the window: sum the +1 and -1 values within the window to obtain cumulative sum A and cumulative sum B. If the product of the two cumulative sums is greater than zero, it indicates that the two sequences within the window are biased in the same direction; if the product of the two cumulative sums is less than zero, it indicates that the two sequences within the window are biased in opposite directions; if the product of the two cumulative sums is zero, ignore the contribution of this window. For each window that is not ignored, calculate the standard deviation of its two sign segments, then multiply the two standard deviations to get the product value, and multiply the product value by the local consistency rate of the window to get the weighted local consistency rate of the window. The smaller the product value, the more stable the sign change within the window, that is, the weak the fluctuation, and the smaller the weighted local consistency rate, thereby suppressing the influence of the stationary segment on the final result. The weighted local consistency rates of all windows biased towards the same direction are summed to obtain the cumulative consistency rate of the same direction. The weighted local consistency rates of all windows biased towards the opposite direction are summed to obtain the cumulative consistency rate of the opposite direction. Divide the cumulative consistency rate in the same direction by the sum of the cumulative consistency rates in the same direction and the cumulative consistency rates in the opposite direction, and normalize the result to the 0-1 range. This is the same-direction locking value, which represents the degree of feature synchronization. If the sum of the cumulative consistency rates in the same direction and the cumulative consistency rates in the opposite direction is 0, then the same-direction locking value is directly 0.5. The closer the same-direction locking value is to 1, the stronger the consistency of the sign change trend of the two features within the local time window and the higher the overall proportion of the same direction.
[0028] Based on the signal strength of the terminal feature data, the fluctuations between feature combinations are analyzed to obtain the stripping binding coefficient, which represents the strength of the association binding of the feature combination after removing interference. Specifically, for the feature combination to be analyzed, when both features in the feature combination to be analyzed are signal strength features, the following calculation is performed; otherwise, the stripping binding coefficient is directly set to 0.5, because the fluctuation binding strength after removing interference cannot be evaluated. A neutral value of 0.5 can indicate that the feature combination does not have a special association binding tendency, neither biased towards strong binding nor weak binding, so that when calculating with the same direction locking value, it will not affect the neutral judgment of the redundancy confidence. The calculation process is as follows: extract the signal strength of the two features at all sampling points, arrange them according to the sampling point number, and at the same time take the median of the signal strength of all terminals at the same sampling point to form a new sequence, which is the environmental interference median sequence. For the signal strength sequence of terminal X, the signal strength sequence of terminal Y, and the median sequence of environmental interference, the following extreme value detection is performed sequentially: starting from the second sampling point and ending at the second to last sampling point, the value is judged point by point. If the signal strength of the current sampling point is greater than the signal strength of the previous sampling point and greater than the signal strength of the next sampling point, then the sampling point is marked as a local peak point. If the signal strength of the current sampling point is less than the signal strength of the previous sampling point and less than the signal strength of the next sampling point, then the sampling point is marked as a local valley point. Each marked point records its sampling point number, signal strength, and type, which is either peak or valley. The detected local peak points and local valley points are collectively referred to as extreme points. The extreme points of each sequence are arranged in the order of the sampling points to form an extreme point list. After traversal, each of the three sequences obtains an extreme point list. For each extreme point of terminal X, check whether there is an extreme point in the extreme point list of the median sequence of environmental interference that meets two conditions: first, the sampling point number of the environmental extreme point and the extreme point of terminal X differs by 0 or ±1, that is, no more than one sampling point before and after; second, the two are of the same type, that is, both are peaks or both are valleys. If such an environmental extreme point exists, the extreme point of terminal X is determined to be caused by environmental interference and is deleted from the list. The same removal operation is performed on the extreme point list of terminal Y. The remaining extreme points after removal are regarded as the true fluctuation characteristics of each terminal itself. Calculate the amplitude of the remaining extreme points of terminal X, i.e., the arithmetic mean of the signal strength values. Keep the extreme points whose amplitude is greater than the arithmetic mean to obtain the main fluctuation points of terminal X. If terminal X has no remaining extreme points, the main fluctuation points are empty. The same operation can be performed to obtain the main fluctuation points of terminal Y. Calculate the absolute value of the difference between the sampling point numbers of all adjacent points in the main fluctuation points of terminal X, take the median of these differences as the median of the interval, and then divide the median of the interval by 2 to obtain the allowable deviation of the position matching. If the number of main fluctuation points of terminal X is less than 2, the allowable deviation is directly set to 1 sampling point. For each major fluctuation point of terminal X, search the set of major fluctuation points of terminal Y to find a point that simultaneously satisfies the following three conditions: First, the absolute value of the difference between the sampling point index of this point and the current major fluctuation point of terminal X does not exceed the allowable deviation; second, the two are of the same type; third, the amplitudes of both are positive. Each time a matching pair that meets the conditions is found, it is considered a successful match. After traversing all major fluctuation points of terminal X, the total number of successful matches is obtained. The number of successful matches is divided by the total number of major fluctuation points of terminal X to obtain the original matching rate. If the total number of major fluctuation points of terminal X is zero, the original matching rate is 0. Taking the cube root of the original matching rate and normalizing the result to the 0-1 range yields the stripping binding coefficient, which represents the strength of the feature combination association after removing interference. The larger the stripping binding coefficient, the higher the similarity between the two terminals in terms of the time, type, and magnitude of the local extrema after removing common environmental interference, i.e., the stronger the association binding strength of the feature combination.
[0029] In one embodiment, based on the feature association redundancy set, the redundancy authenticity of each feature combination is analyzed to obtain a redundancy confidence level representing the redundancy probability of each feature combination, and the method further includes: The co-directional locking value and the stripping binding coefficient are fused to analyze the true redundancy of each feature combination and obtain the redundancy confidence of the redundancy probability of each feature combination. Specifically, for each feature combination, the co-directional locking value and the stripping binding coefficient are multiplied to obtain the product result. The square root of the product result is taken to obtain the geometric mean. The geometric mean is used as the redundancy confidence of the feature combination. The larger the redundancy confidence value, the higher the true redundancy probability of the feature combination.
[0030] In one embodiment, the redundancy confidence is analyzed, and it is determined whether the features are a valid combination of redundant features, resulting in a deduplicated feature set, including: For each feature combination, the redundancy confidence scores are sorted and analyzed, and then removed or retained to obtain a deduplicated feature set. Specifically, this involves: arranging all feature combinations in descending order of redundancy confidence scores to obtain an ordered list; calculating the difference between two adjacent redundancy confidence scores in the ordered list to obtain a difference sequence; finding the maximum value in the difference sequence and using the position of this maximum value as the natural split point. If the difference sequence is empty, the split point is set at the middle position of the ordered list. All feature combinations before the split point are marked as candidate retention combinations, and all feature combinations after the split point are directly eliminated. Among the candidate retention combinations, it is checked whether there are mutually inverse feature pairs, that is, feature A to B and feature B to A exist at the same time. For each pair of mutually inverse feature pairs, the one with higher redundancy confidence is retained and the one with lower confidence is eliminated. If the two are equal, either one is retained. By merging all the features in the retained feature combinations, removing duplicate features, and arranging them according to their original indices, we can obtain the deduplicated feature set.
[0031] By accurately reflecting the same-direction and opposite-direction change trends of features within a local time window through the same-direction locking value, and removing interference from the public environment by stripping the binding coefficient, the true fluctuation characteristics of the terminal itself are retained, and finally the redundancy confidence is obtained. This allows for an effective assessment of the redundancy probability of each feature combination. Compared with existing linear dimensionality reduction methods such as PCA, this scheme can avoid the loss of higher-order associations caused by linear projection. At the same time, by using natural segmentation points and opposite feature pairs to remove features, it accurately identifies and retains effective redundant feature combinations, reduces missed and false judgments, and improves the deduplication accuracy of feature data of smart home Wi-Fi router terminals.
[0032] In one embodiment, a redundancy deduplication system based on data feature recognition is applied to the above-described processing method, including: The data analysis unit is used to acquire terminal feature data of the target object, analyze the correlation between features based on the preprocessed terminal feature data, and generate a representation of the correlation contribution between features. The redundancy analysis unit is used to analyze the distribution and density of feature combinations based on their correlation contribution, and generate a feature correlation redundancy set representing the degree of redundancy of each feature combination. The probability analysis unit is used to analyze the redundancy authenticity of each feature combination based on the feature association redundancy set, and obtain the redundancy confidence level representing the redundancy probability of each feature combination. The deduplication unit is used to analyze the redundancy confidence and determine whether the features are effective redundant feature combinations, thus obtaining the deduplicated feature set.
[0033] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for processing redundant information deduplication based on data feature recognition, characterized in that, include: Step S1: Obtain terminal feature data of the target object; based on the preprocessed terminal feature data, analyze the correlation between features and generate a contribution degree representing the correlation between features, including: For the preprocessed terminal feature data, the trajectory of each feature is reconstructed, and the oscillation amplitude and duration of the rate of change in the trajectory are analyzed to obtain the feature activity index representing the chaotic amplitude of feature changes. Based on the feature activity index, the impact of the feature change on the other features is analyzed, and a hysteresis energy value is generated to represent the hysteresis effect of the feature change on the other features. Based on the terminal feature data, the response of the features to interference is analyzed, and a response misalignment offset representing the degree of response of the two features to external interference is generated. The feature activity index, hysteresis energy value and response misalignment offset are fused to generate a representation of the correlation contribution between features; Step S2: Based on the correlation contribution, analyze the distribution and density of feature combinations to generate a feature correlation redundancy set representing the redundancy strength of each feature combination. In this step, analyze the density distribution of the correlation contribution of all feature combinations to find the inflection point interval from sparse to dense regions, i.e., the redundancy perturbation range. Feature combinations falling within the redundancy perturbation range are more susceptible to data fluctuations in their correlation strength. Within the redundancy perturbation range, calculate the local cluster density of each feature combination. The higher the local cluster density, the stronger the redundancy of the feature combination; the lower the local cluster density, the weaker the redundancy. Step S3: Based on the feature association redundancy set, analyze the redundancy authenticity of each feature combination to obtain the redundancy confidence level representing the redundancy probability of each feature combination. Step S4: Analyze the redundancy confidence and determine whether the features are effective redundant feature combinations to obtain the deduplicated feature set.
2. The method for deduplicating redundant information based on data feature recognition according to claim 1, characterized in that, Based on the correlation contribution, the distribution and density of feature combinations are analyzed to generate a feature correlation redundancy set representing the redundancy strength of each feature combination, including: Analyze the distribution pattern of the correlation contribution of each feature combination, determine the inflection point interval of the correlation contribution transitioning from sparse to dense regions, and obtain the redundancy perturbation range representing the change in the correlation strength between feature combinations. Within each redundancy perturbation range, the local clustering density of each feature combination is calculated and divided to obtain a redundancy stacking hierarchy representing the density of feature combinations under different redundancy intensities.
3. The method for deduplicating redundant information based on data feature recognition according to claim 2, characterized in that, Based on the correlation contribution, the distribution and density of feature combinations are analyzed to generate a feature correlation redundancy set representing the redundancy strength of each feature combination, which also includes: By fusing the range of redundancy perturbation with the redundancy stacking level, a feature-associated redundancy set representing the degree of redundancy strength of each feature combination is generated.
4. The method for deduplicating redundant information based on data feature recognition according to claim 3, characterized in that, Based on the feature association redundancy set, the redundancy authenticity of each feature combination is analyzed to obtain the redundancy confidence level representing the redundancy probability of each feature combination, including: For each feature combination in the feature-associated redundant set, analyze the same-direction and opposite-direction changes of the two to obtain the same-direction locking value representing the degree of feature synchronization; Based on the signal strength of the terminal feature data, the fluctuations between feature combinations are analyzed to obtain the stripping binding coefficient, which represents the strength of the feature combination association binding after interference removal.
5. The method for deduplicating redundant information based on data feature recognition according to claim 4, characterized in that, Based on the feature association redundancy set, the redundancy authenticity of each feature combination is analyzed to obtain the redundancy confidence level representing the redundancy probability of each feature combination, which also includes: By fusing the same-direction locking value with the stripping binding coefficient, the true redundancy of each feature combination is analyzed, and the redundancy confidence of the redundancy probability of each feature combination is obtained.
6. The method for deduplicating redundant information based on data feature recognition according to claim 5, characterized in that, Redundancy confidence is analyzed, and it is determined whether the features are effective redundant feature combinations, resulting in a deduplicated feature set, including: The redundancy confidence of each feature combination is sorted and analyzed, and then removed and retained to obtain the deduplicated feature set.
7. A redundant information deduplication system based on data feature recognition, applied in the processing method described in any one of claims 1-6, characterized in that, include: The data analysis unit is used to acquire terminal feature data of the target object, analyze the correlation between features based on the preprocessed terminal feature data, and generate a representation of the correlation contribution between features. The redundancy analysis unit is used to analyze the distribution and density of feature combinations based on their correlation contribution, and generate a feature correlation redundancy set representing the degree of redundancy of each feature combination. The probability analysis unit is used to analyze the redundancy authenticity of each feature combination based on the feature association redundancy set, and obtain the redundancy confidence level representing the redundancy probability of each feature combination. The deduplication unit is used to analyze the redundancy confidence and determine whether the features are effective redundant feature combinations, thus obtaining the deduplicated feature set.