Self-adaptive watermarking method and system for power system data
By building a watermark database and using a random forest model for adaptive analysis, selecting a suitable watermark algorithm, the problem of difficulty in taking into account both robustness and obscurity in power system data is solved, and adaptive watermark operations with high reliability and high applicability are achieved.
Patent Information
- Application Number
- CN202510462454.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-25
AI Technical Summary
The existing power system data watermarking methods are difficult to balance both robustness and obscurity, and the existing solutions fail to effectively layer the watermark rules and embedding locations, resulting in mismatch between selection and data features.
The watermark database is constructed, and a variety of watermark schemes based on the rule layer and the position layer are adopted, combined with the random forest model to perform adaptive analysis of power system data, and the most suitable watermark algorithm is selected for embedding, including zero watermark, differential expansion, histogram translation, least significant bit, invisible characters, multi-encoding combinations and tuple permutations.
It realizes adaptive watermark operation of power system data, improves reliability, accuracy and applicability, reduces manual configuration costs, and improves the matching degree between watermark algorithms and business scenarios.
Smart Images

Figure CN120372585A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of electrical automation, and particularly relates to an adaptive watermarking method and system for power system data. Background Art
[0002] With the development of economic technology and the improvement of people's living standards, electric energy has become an essential secondary energy source in people's production and life, bringing endless convenience to people's production and life. Therefore, ensuring the stable and reliable supply of electric energy has become one of the most important tasks of the power system.
[0003] The data security of the power system is crucial for the power system. Therefore, it is of great significance to perform watermarking processing on power system data. At present, the data types in the data center of the power system are diverse. Taking power data as an example, it includes numerical measurement data, character-based device codes, spatio-temporal sequence data, etc. Traditional watermarking schemes are difficult to balance robustness and invisibility due to their simplicity. For example, the histogram shift watermarking technique is suitable for high-precision floating-point data, but this scheme will destroy the statistical characteristics of discrete coding; the least significant bit (LSB) embedding watermarking technique is effective for integers but may cause loss of floating-point data accuracy. In addition, existing watermarking methods do not hierarchically process watermarking rules (such as feature extraction, difference expansion) and embedding positions (such as LSB, tuple permutation), resulting in a mismatch between the selection of existing schemes and data characteristics. Summary of the Invention
[0004] One of the purposes of the present invention is to provide an adaptive watermarking method for power system data with high reliability, good accuracy and good applicability.
[0005] Another purpose of the present invention is to provide a system for implementing the adaptive watermarking method for power system data.
[0006] The adaptive watermarking method for power system data provided by the present invention includes the following steps:
[0007] S1. Based on the watermarking scheme, construct a watermark database;
[0008] S2. Obtain the data information of the target power system data;
[0009] S3. Based on the random forest model, analyze the application scenario, processing preference, watermark content and data characteristics of the target power system data;
[0010] S4. According to the analysis results obtained in step S3, perform watermark embedding operation on the target power system data to complete the watermark operation of the target power system data.
[0011] Based on the watermarking scheme described in step S1, construct a watermark database, which specifically includes the following steps:
[0012] The constructed database includes a rule layer and a location layer;
[0013] The rule layer is used to determine how to generate watermarks or how to adjust the data structure; the location layer is used to determine the physical embedding location or logical embedding location of the watermarks;
[0014] The addition rule algorithms of the rule layer include the zero-watermark-based scheme, the differential expansion-based watermarking scheme, and the histogram translation-based watermarking scheme; the addition location algorithms of the location layer include the least significant bit-based watermarking scheme, the invisible character-based watermarking scheme, the pseudo-row and pseudo-column-based watermarking scheme, the multi-coding combination-based watermarking scheme, and the tuple permutation-based watermarking scheme.
[0015] The zero-watermark-based scheme specifically includes the following steps:
[0016] Calculate the Pearson correlation coefficient of each attribute column in the target data with other attribute columns;
[0017] Calculate the average value of the Pearson correlation coefficients of each attribute column with other attribute columns, and select the attribute column with the smallest average value as the candidate feature column;
[0018] Divide the data of the candidate feature column into two parts: the first part is used for feature extraction, and the second part is used for watermark embedding: set the number of blocks b as is the ceiling function, m is the number of tuple values; arrange the data of the candidate feature column in ascending order of tuple values and divide it into k consecutive blocks, is the floor function; select the blocks with odd serial numbers for feature extraction and the blocks with even serial numbers for watermark embedding; if m mod b≠0, evenly distribute the remaining data to the feature extraction part and the watermark embedding part;
[0019] Convert the m tuple values of the candidate feature column into an h×h square matrix, where is the floor function;
[0020] Compare the value of each data point in the obtained square matrix with the values of its 8 neighborhood data points: if the values of the neighborhood data points are all greater than or equal to the value of this data point, mark the neighborhood data value as 1 in the LBP mode, otherwise mark the neighborhood data value as 0 in the LBP mode; for each point in the square matrix, obtain the corresponding 8-bit binary number;
[0021] Count the number of occurrences of each possible LBP model in the matrix to obtain a corresponding histogram vector consisting of 256 elements, where each element represents the frequency of a specific LBP pattern;
[0022] Generate a feature vector and compare each element of the feature vector with the corresponding average value: generate a new binary vector and use this new binary vector as the perceptual hash;
[0023] Perform an exclusive OR operation on the watermark information represented in binary and the perceptual hash to obtain the zero-watermark feature code.
[0024] The watermarking scheme based on differential expansion specifically includes the following steps:
[0025] Calculate the difference d between adjacent data pairs as d = x i - x j where x i is the i-th data and x j is the j-th data;
[0026] Expand the difference d into a watermark bit sequence d' as d' = 2d + b, where b is the watermark information to be embedded;
[0027] For the obtained watermark bit sequence d', expand the original data pair x i and x j to ensure that d' = x i ' - x' j where is the ceiling function.
[0028] The watermarking scheme based on histogram shifting specifically includes the following steps:
[0029] Construct a histogram: count the frequency of occurrence of the values in all target columns, construct a histogram, and identify the peak points and the distribution of adjacent values;
[0030] Select the value with the highest frequency as the peak point P, shift the values from P + 1 to P + k one bit to the right to create k empty positions for embedding watermark information;
[0031] Allocate the watermark information to be embedded to the empty positions in the set order.
[0032] The watermarking scheme based on the least significant bit specifically includes the following steps:
[0033] For integer data, replace the least significant bit of the quaternary representation of the data with the watermark bit; for floating-point data, replace the mantissa bit of the quaternary representation of the data with the watermark bit.
[0034] The described watermarking scheme based on invisible characters specifically includes the following steps:
[0035] Select invisible characters including 0, 9, 10, and 127 in ASCII code, as well as zero-width space, zero-width non-joiner, zero-width joiner, and zero-width no-break space in zero-width characters;
[0036] The watermark sequence is set as a string of quaternary codes, expressed as W = (w1, w2,..., w m ), where w i = 0 corresponds to the invisible character C0, w i = 1 corresponds to the invisible character C1, w i = 2 corresponds to the invisible character C2, w i = 3 corresponds to the invisible character C3; through the mapping relationship, the watermark sequence W is converted into a string composed of invisible characters; the string is embedded at the end of the target attribute column to achieve the addition of the watermark.
[0037] The described watermarking scheme based on multi-coding combination specifically includes the following steps:
[0038] Select several attribute columns, classify different character coding formats for each attribute column; use the assigned coding format to embed the watermark information into the data;
[0039] The described character coding formats include ASCII, UTF-8, UTF-16, and Unicode.
[0040] The described watermarking scheme based on tuple permutation specifically includes the following steps:
[0041] Set the database table T to include n tuples, n > 3; the initial order is expressed as T = [t1, t2,..., t n ; t i is the i-th row tuple;
[0042] The watermark sequence W is a quaternary code, expressed as W = (w1, w2,..., w m );
[0043] Define the permutation rule as follows:
[0044] When w k = 0, then swap the first tuple t1 and the last tuple t n ;
[0045] When w k = 1, then swap the first tuple t1 and the second-to-last tuple t n-1 ;
[0046] When w kIf w = 2, then swap the first tuple t1 and the second tuple t2;
[0047] When w k = 3, then swap the first tuple t1 and the third tuple t3.
[0048] Based on the random forest model described in step S3, perform application scenario, processing preference, watermark content, and data feature analysis on the target power system data, specifically including the following steps:
[0049] For the target power system data, perform application scenario analysis, processing preference analysis, watermark content analysis, and data feature analysis; among them, application scenario analysis includes mid-platform to library application and export to file application; processing preference analysis includes processing function preference, processing object preference, and user-specified preference; watermark content analysis includes full traceability of the data link and traceability of the data exporter; data feature analysis includes character-type data and numerical-type data;
[0050] For character-type data, directly convert it to numerical-type data;
[0051] For numerical-type data, perform standardization processing;
[0052] Based on the random forest model, construct and train a processing preference prediction model; among them, the inputs of the processing preference prediction model include:
[0053] Data features: including data types, such as character type, numerical type, etc.;
[0054] Processing preferences: including the historical operation function set (such as SELECT, GROUP BY, CONCAT, etc.) and the function call frequency (such as the number of times SUBSTR is used);
[0055] Application scenario tags: including mid-platform to library (tag 0) or export to file (tag 1);
[0056] User-specified preferences: including the keyword field list marked by the user and the watermark strength requirements defined by the user (the watermark strength includes low, medium, and high);
[0057] The outputs of the processing preference prediction model include four types of prediction labels:
[0058] Application scenario prediction: including mid-platform to library or export to file;
[0059] Processing preference prediction: including function sensitive path, field sensitive path, and user-specified path;
[0060] Watermark content prediction: including full traceability of the data link or traceability of the data exporter;
[0061] Data feature prediction: including numerical type dominance or character type dominance;
[0062] The sources of training data for the processing preference prediction model include:
[0063] Historical data processing records: Extract historical data tables and processing operations (including cleaning, aggregation, etc.) and watermark embedding records from the logs of the target power system;
[0064] Annotated samples: The optimal watermark algorithm combinations in different scenarios are annotated by experts (for example, invisible characters are preferred in the scenario of exporting files);
[0065] User feedback data: Record the protection requirements of users for specific fields and the tolerance for watermark perturbation;
[0066] Input the target power system data into the pre-trained processing preference prediction model to obtain the corresponding processing preference prediction results; the processing preference prediction results include application scenario prediction results, processing preference prediction results, watermark content prediction results, and data feature prediction results. Dynamically map data features and application requirements through machine learning to guide the adaptive selection of watermark algorithms and avoid sub-optimal decisions caused by relying on fixed rules. This model can improve the matching degree between watermark algorithms and business scenarios, reduce the manual configuration cost, and achieve automated decision-making.
[0067] According to the analysis result obtained in step S3, perform watermark embedding operation on the target power system data, which specifically includes the following steps:
[0068] (1) Data features:
[0069] Set the data table T to include n attribute columns A1 to A n , and calculate the correlation coefficient matrix using the following formula:
[0070]
[0071] In the formula, C ij is the element in the i-th row and j-th column of the correlation coefficient matrix; cov(A i , A j ) is the covariance of column A i and column A j ; is the standard deviation of column A i ;
[0072] For the numerical attribute column A k , calculate the corresponding kurtosis K k using the following formula:
[0073]
[0074] In the formula, E() is the variance;
[0075] The following rules are used for judgment:
[0076] If is greater than the set correlation threshold, a watermarking scheme based on differential expansion is adopted; represents the maximum value of the average of the correlation coefficients in the i-th column of the correlation coefficient matrix;
[0077] If K k is greater than the set peak threshold, a watermarking scheme based on histogram shifting is adopted for the k-th column;
[0078] If it is character-type data, a zero-watermarking-based scheme or an invisible-character-based watermarking scheme is adopted;
[0079] (2) Application scenario adaptation:
[0080] For the scenario of the middle platform to the library application, a watermarking scheme based on tuple permutation or a watermarking scheme based on the least significant bit is adopted;
[0081] For the scenario of exporting to a file application, an invisible-character-based watermarking scheme or a watermarking scheme based on multi-encoding combination is adopted;
[0082] (3) Processing preference priority:
[0083] Set the user-specified field set as F, and the importance vector is calculated using the following formula:
[0084]
[0085] In the formula is the importance weight of the user-specified field f i , with a value range of [0, 1]; H(f i ) is the information entropy of the field f i , and the lower the entropy, the higher the importance, and H(f i ) = -∑p(x)log(p(x)), where p(x) is the value probability; |F| is the size of the user-specified field set F;
[0086] The penalty value is calculated using the following formula:
[0087]
[0088] In the formula, penalty is the penalty value for the sensitive operation path, used to avoid destroying the data processing logic; F sens is the sensitive operation function set (such as SUBSTR, CONCAT); is the sensitivity coefficient of the function f, fitted through historical attack data (such as The value of () is 0.9); Senstivity(f,algo) is the sensitivity score of algorithm algo to function f. If the algorithm perturbation may affect the output of f, the score is higher;
[0089] The preference score is calculated using the following formula:
[0090] Pr = (ω f ) T ·u algo -penalty
[0091] In the formula, Pr is the preference score. The higher the score, the more suitable for the current scenario; ω f is the importance weight vector of the user-specified field; u algo is the coverage vector of algorithm algo for the field. If the algorithm acts on field f i then u algo = 1, otherwise u algo = 0;
[0092] For the user-specified field, a watermarking scheme based on histogram translation or a watermarking scheme based on tuple permutation is adopted;
[0093] For the function sensitive path, a zero-watermarking-based scheme is adopted;
[0094] (4) Watermark content weight:
[0095] The combined score of the algorithms is calculated using the following formula:
[0096] TS(algo1,algo2) = γ1·Resilience(algo1) + γ2·Stealth(algo2)
[0097] In the formula, TS(algo1,algo2) is the combined score of rule layer algorithm algo1 and location layer algorithm algo2; γ1 is the first weight coefficient, γ2 is the second weight coefficient, and γ1 + γ2 = 1; Resilience(algo1) is the anti-attack value of algorithm algo1, and Resilience(algo1) = 1 - BER, where BER is the bit error rate; Stealth(algo2) is the stealth value of algorithm algo2, and Stealth(algo2) = 1 - Detectbility, where Detectbility is the detection probability;
[0098] For full traceability of the data link, a zero-watermarking-based scheme is adopted in the rule layer, and a watermarking scheme based on tuple permutation is adopted in the location layer;
[0099] For the traceability of data exporters, a watermarking scheme based on differential expansion is adopted at the rule layer, and a watermarking scheme based on invisible characters is adopted at the location layer;
[0100] (5) Scheme selection:
[0101] Use the analytic hierarchy process to calculate the scores of each combined algorithm; the combined algorithm is the paired combination of the rule layer algorithm and the location layer algorithm; for the above data characteristics, application scenarios, processing preferences, and watermark content, conduct pairwise importance comparisons to generate an importance matrix P; by decomposing the eigenvalues of the importance matrix P, obtain the weight Λ = [λ1, λ2, λ3, λ4] T , λ i is the sub - weight, i takes values from 1 to 4, and For each combined algorithm, use the following formula to calculate the final score:
[0102]
[0103] In the formula, TotalScore is the final score of each combined algorithm; S1 is the data feature matching degree, which is obtained by scoring according to the correlation coefficient matrix and kurtosis; S2 is the scene adaptation degree score; S3 is the processing preference matching degree, and the value is the preference score Pr; S4 is the watermark content weight, and the value is the algorithm combination score;
[0104] Finally, select the combined algorithm with the highest total score for watermark embedding.
[0105] The present invention also provides a system for implementing the adaptive watermarking method for power system data, including a database construction module, a data acquisition module, a data analysis module, and an adaptive watermarking module; the database construction module, the data acquisition module, the data analysis module, and the adaptive watermarking module are connected in series in sequence; the database construction module is used to construct a watermark database based on the watermarking scheme and upload the data information to the data acquisition module; the data acquisition module is used to obtain the data information of the target power system data according to the received data information and upload the data information to the data analysis module; the data analysis module is used to analyze the application scenario, processing preference, watermark content, and data characteristics of the target power system data based on the random forest model according to the received data information and upload the data information to the adaptive watermarking module; the adaptive watermarking module is used to perform watermark embedding operations on the target power system data according to the received data information and the obtained analysis results to complete the watermark operation of the target power system data.
[0106] The adaptive watermarking method and system for power system data provided by the present invention construct a watermark database, comprehensively analyze the target power system data, and adaptively select a watermarking scheme, which not only realizes the adaptive watermarking operation for power system data, but also has higher reliability, better accuracy and better applicability. BRIEF DESCRIPTION OF THE DRAWINGS
[0107] Figure 1 It is a schematic flow chart of the method of the present invention.
[0108] Figure 2 It is a schematic diagram of the functional modules of the system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0109] As Figure 1 shown, it is a schematic flow chart of the method of the present invention: The adaptive watermarking method for power system data disclosed by the present invention includes the following steps:
[0110] S1. Based on the watermarking scheme, construct a watermark database; specifically, it includes the following steps:
[0111] The constructed database includes a rule layer and a position layer;
[0112] The rule layer is used to determine how to generate the watermark or how to adjust the data structure; the position layer is used to determine the physical embedding position or logical embedding position of the watermark;
[0113] The addition rule algorithms of the rule layer include a zero-watermark-based scheme, a differential expansion-based watermarking scheme, and a histogram shifting-based watermarking scheme; the addition position algorithms of the position layer include a least significant bit-based watermarking scheme, an invisible character-based watermarking scheme, a pseudo-row and pseudo-column-based watermarking scheme, a multi-coding combination-based watermarking scheme, and a tuple permutation-based watermarking scheme;
[0114] Specifically in implementation:
[0115] The zero-watermark-based scheme specifically includes the following steps:
[0116] Calculate the Pearson correlation coefficient between each attribute column in the target data and other attribute columns;
[0117] Calculate the average value of the Pearson correlation coefficients between each attribute column and other attribute columns, and select the attribute column with the smallest average value as the candidate feature column to ensure that the change of the feature column has the least impact on the overall data;
[0118] Divide the data of the candidate feature column into two parts: the first part is used for feature extraction, and the second part is used for watermark embedding: Set the number of blocks b to be is the ceiling function, and m is the number of tuple values; the data of the candidate feature column is sorted according to the size of the tuple values and divided into k consecutive blocks. is the floor function; the blocks with odd serial numbers are selected for feature extraction, and the blocks with even serial numbers are selected for watermark embedding; if m mod b ≠ 0, the remaining data is evenly distributed to the feature extraction part and the watermark embedding part.
[0119] Convert the m tuple values of the candidate feature column into an h×h square matrix, where is the floor function;
[0120] Compare the value of each data point in the obtained square matrix with the values of its 8 neighboring data points: if the values of the neighboring data points are all greater than or equal to the value of this data point, mark the neighboring data value as 1 in the LBP mode, otherwise mark the neighboring data value as 0 in the LBP mode; for each point in the square matrix, obtain the corresponding 8-bit binary number.
[0121] Count the number of times each possible LBP model appears in the matrix to obtain the corresponding histogram vector including 256 elements, where each element represents the frequency of a specific LBP pattern.
[0122] Generate a feature vector and compare each element of the feature vector with the corresponding average value: generate a new binary vector and use this new binary vector as the perceptual hash.
[0123] Perform an exclusive OR operation on the watermark information represented in binary and the perceptual hash to obtain the zero-watermark feature code.
[0124] The watermarking scheme based on differential expansion specifically includes the following steps:
[0125] Calculate the difference d between adjacent data pairs as d = x i -x j , where x i is the i-th data and x j is the j-th data;
[0126] Expand the difference d into a watermark bit sequence d' as d' = 2d + b, where b is the watermark information to be embedded;
[0127] For the obtained watermark bit sequence d', expand the original data pair x i and x j to ensure that d' = x i '-x' j , where is the ceiling function; if the extended value exceeds the data range, the data is directly marked as unavailable and the watermark embedding is skipped;
[0128] The watermarking scheme based on histogram shifting specifically includes the following steps:
[0129] Construct a histogram: Count the occurrence frequencies of the values in all target columns, construct a histogram, and identify the distribution of peak points and adjacent values;
[0130] Select the value with the highest frequency as the peak point P, shift the values from P + 1 to P + k one bit to the right, creating k empty positions for embedding watermark information;
[0131] Allocate the watermark information to be embedded to the empty positions in the set order;
[0132] The watermarking scheme based on the least significant bit specifically includes the following steps:
[0133] For integer data, replace the least significant bit of the quaternary representation of the data with the watermark bit; for floating-point data, replace the mantissa bit of the quaternary representation of the data with the watermark bit;
[0134] The watermarking scheme based on invisible characters specifically includes the following steps:
[0135] Select the invisible characters including 0, 9, 10, and 127 in ASCII code, as well as zero-width space, zero-width non-joiner, zero-width joiner, and zero-width non-breaking space in zero-width characters;
[0136] The watermark sequence is set as a string of quaternary codes, expressed as W = (w1, w2,..., w m )), where w i = 0 corresponds to the invisible character C0, w i = 1 corresponds to the invisible character C1, w i = 2 corresponds to the invisible character C2, w i = 3 corresponds to the invisible character C3; through the mapping relationship, the watermark sequence W is converted into a string composed of invisible characters; the string is embedded at the end of the target attribute column to achieve the addition of the watermark;
[0137] The watermarking scheme based on multi-coding combination specifically includes the following steps:
[0138] Select several attribute columns, classify different character coding formats for each attribute column; use the allocated coding formats to embed the watermark information into the data;
[0139] The character coding formats include ASCII, UTF-8, UTF-16, and Unicode;
[0140] The watermarking scheme based on tuple permutation specifically includes the following steps:
[0141] Set the database table T to include n tuples, where n > 3; the initial order is expressed as T = [t1, t2,..., t n ; t i is the i-th row tuple;
[0142] The watermark sequence W is in quaternary encoding, expressed as W = (w1, w2,..., w m );
[0143] Define the permutation rule as follows:
[0144] When w k = 0, then swap the first tuple t1 and the last tuple t n ;
[0145] When w k = 1, then swap the first tuple t1 and the second-to-last tuple t n-1 ;
[0146] When w k = 2, then swap the first tuple t1 and the second tuple t2;
[0147] When w k = 3, then swap the first tuple t1 and the third tuple t3;
[0148] S2. Obtain the data information of the target power system data;
[0149] S3. Based on the random forest model, analyze the application scenario, processing preference, watermark content, and data characteristics of the target power system data; specifically include the following steps:
[0150] For the target power system data, conduct application scenario analysis, processing preference analysis, watermark content analysis, and data characteristics analysis; among them, the application scenario analysis includes the application from the middle platform to the library and the application of exporting to a file; the processing preference analysis includes processing function preference, processing object preference, and user-specified preference; the watermark content analysis includes full traceability of the data link and traceability of the data exporter; the data characteristics analysis includes character-type data and numerical-type data;
[0151] For character-type data, directly convert it to numerical-type data;
[0152] For numerical-type data, perform standardization processing (including Z-score standardization and min-max standardization);
[0153] Based on the random forest model, construct and train a processing preference prediction model; among them, the input of the processing preference prediction model includes:
[0154] Data characteristics: including data types, such as character type, numerical type, etc.;
[0155] Processing preferences: including a set of historical operation functions (such as SELECT, GROUP BY, CONCAT, etc.) and function call frequencies (such as the number of times SUBSTR is used);
[0156] Application scenario tags: including from the middle platform to the library (tag 0) or export to a file (tag 1);
[0157] User-specified preferences: including a list of keyword fields marked by the user and the user-defined watermark strength requirements (watermark strength includes low, medium, and high);
[0158] The output of the processing preference prediction model includes four types of prediction tags:
[0159] Application scenario prediction: including from the middle platform to the library or export to a file;
[0160] Processing preference prediction: including function sensitive paths, field sensitive paths, and user-specified paths;
[0161] Watermark content prediction: including full traceability of the data link or traceability of the data exporter;
[0162] Data characteristic prediction: including numerical type dominance or character type dominance;
[0163] The training data sources of the processing preference prediction model include:
[0164] Historical data processing records: Extract historical data tables and processing operations (including cleaning, aggregation, etc.), and watermark embedding records from the logs of the target power system;
[0165] Annotated samples: The optimal watermark algorithm combinations in different scenarios are annotated by experts (such as preferentially using invisible characters in the scenario of exporting files);
[0166] User feedback data: Record the protection requirements of users for specific fields and the tolerance for watermark perturbation;
[0167] Input the data of the target power system into the pre-trained processing preference prediction model to obtain the corresponding processing preference prediction results; the processing preference prediction results include application scenario prediction results, processing preference prediction results, watermark content prediction results, and data characteristic prediction results; Dynamically map data characteristics and application requirements through machine learning to guide the adaptive selection of watermark algorithms, avoiding suboptimal decisions caused by relying on fixed rules. This model can improve the matching degree between the watermark algorithm and the business scenario, reduce the manual configuration cost, and achieve automated decision-making;
[0168] S4. According to the analysis results obtained in step S3, perform watermark embedding operation on the target power system data to complete the watermark operation of the target power system data; specifically, it includes the following steps:
[0169] (1) Set the data table T to include n attribute columns A1 to A n , and calculate the correlation coefficient matrix using the following formula:
[0170]
[0171] where C ij is the element in the i-th row and j-th column of the correlation coefficient matrix; cov(A i , A j ) is the covariance of column A i and column A j ; is the standard deviation of column A i ;
[0172] For the numerical attribute column A k , calculate the corresponding kurtosis K k using the following formula:
[0173]
[0174] where E() is the variance;
[0175] Use the following rules for judgment:
[0176] If is greater than the set correlation threshold, then adopt a watermarking scheme based on differential expansion; represents the maximum value of the average of the correlation coefficients in the i-th column of the correlation coefficient matrix;
[0177] If K k is greater than the set peak threshold, then adopt a watermarking scheme based on histogram shifting for the k-th column;
[0178] If it is character-type data, then adopt a zero-watermarking scheme or a watermarking scheme based on invisible characters;
[0179] (2) Application scenario adaptation:
[0180] For the scenario of middle platform to library application, adopt a watermarking scheme based on tuple permutation or a watermarking scheme based on the least significant bit;
[0181] For the scenario of exporting to file application, adopt a watermarking scheme based on invisible characters or a watermarking scheme based on multi-encoding combination;
[0182] (3) Processing preference priority:
[0183] Set the user-specified field set as F, then the importance vector is calculated using the following formula:
[0184]
[0185] In the formula is the importance weight of the user-specified field f, and the value range is [0,1]; H(f i i ) is the information entropy of the field f i , the lower the entropy, the higher the importance, and H(f i ) = -∑p(x)log(p(x)), where p(x) is the value probability; |F| is the size of the user-specified field set F;
[0186] The penalty value is calculated using the following formula:
[0187]
[0188] In the formula, penalty is the penalty value for the sensitive operation path, which is used to avoid destroying the data processing logic; F sens is the set of sensitive operation functions (such as SUBSTR, CONCAT); is the sensitivity coefficient of the function f, which is fitted through historical attack data (such as the value of SUBSTR is 0.9); Senstivity(f,algo) is the sensitivity score of the algorithm algo for the function f. If the algorithm perturbation may affect the output of f, the score is higher;
[0189] The preference score is calculated using the following formula:
[0190] Pr = (ω f ) T ·u algo -penalty
[0191] In the formula, Pr is the preference score, and the higher the score, the more suitable for the current scenario; ω f is the importance weight vector of the user-specified field; u algo is the coverage vector of the algorithm algo for the field. If the algorithm acts on the field f i then u algo = 1, otherwise u algo = 0;
[0192] For the user-specified field, a watermarking scheme based on histogram translation or a watermarking scheme based on tuple permutation is adopted;
[0193] For the function sensitive path, a zero-watermarking-based scheme is adopted;
[0194] (4) Watermark content weight:
[0195] The combined score of the algorithm is calculated using the following formula:
[0196] TS(algo1,algo2) = γ1·Resilience(algo1)+γ2·Stealth(algo2)
[0197] Where TS(algo1,algo2) is the combined score of the rule-layer algorithm algo1 and the location-layer algorithm algo2; γ1 is the first weight coefficient, γ2 is the second weight coefficient, and γ1 + γ2 = 1; Resilience(algo1) is the anti-attack value of the algorithm algo1, and Resilience(algo1) = 1 - BER, where BER is the bit error rate; Stealth(algo2) is the invisibility value of the algorithm algo2, and Stealth(algo2) = 1 - Detectbility, where Detectbility is the detection probability;
[0198] For full traceability of the data link, a zero-watermark-based scheme is adopted in the rule layer, and a tuple permutation-based watermarking scheme is adopted in the location layer;
[0199] For tracing the data exporter, a differential expansion-based watermarking scheme is adopted in the rule layer, and an invisible character-based watermarking scheme is adopted in the location layer;
[0200] (5) Scheme selection:
[0201] The analytic hierarchy process is used to calculate the scores of each combined algorithm; the combined algorithms are the paired combinations of the rule-layer algorithms and the location-layer algorithms; for the above data characteristics, application scenarios, processing preferences, and watermark content, pairwise importance comparisons are made to generate an importance matrix P; by decomposing the eigenvalues of the importance matrix P, the weights Λ = [λ1, λ2, λ3, λ4] T , λ i is the sub-weight, i takes values from 1 to 4, and For each combined algorithm, the following formula is used to calculate the final score:
[0202]
[0203] Where TotalScore is the final score of each combined algorithm; S1 is the data feature matching degree, which is scored according to the correlation coefficient matrix and kurtosis; S2 is the scenario adaptation degree score; S3 is the processing preference matching degree, and the value is the preference score Pr; S4 is the watermark content weight, and the value is the algorithm combined score;
[0204] Finally, the combined algorithm with the highest total score is selected for watermark embedding.
[0205] After watermark embedding, in subsequent steps, at the watermark extraction stage, reverse operations are performed on the watermark addition steps of each watermark algorithm in the watermark algorithm library to extract watermark information. Finally, the information extracted according to each extraction method is decoded, and the successfully decoded information is the watermark information.
[0206] As Figure 2 The following is a schematic diagram of the functional modules of the system of the present invention: The system for implementing the adaptive watermark method for power system data disclosed in the present invention includes a database construction module, a data acquisition module, a data analysis module, and an adaptive watermark module; the database construction module, the data acquisition module, the data analysis module, and the adaptive watermark module are connected in series in sequence; the database construction module is used to construct a watermark database based on the watermark scheme and upload the data information to the data acquisition module; the data acquisition module is used to obtain the data information of the target power system data according to the received data information and upload the data information to the data analysis module; the data analysis module is used to analyze the application scenario, processing preference, watermark content, and data characteristics of the target power system data based on the random forest model according to the received data information and upload the data information to the adaptive watermark module; the adaptive watermark module is used to perform watermark embedding operations on the target power system data according to the received data information and the obtained analysis results to complete the watermark operation of the target power system data.
Claims
1. An adaptive watermarking method for power system data, comprising the following steps: S1. Based on the watermarking scheme, construct a watermark database; S2. Obtain the data information of the target power system data; S3. Based on the random forest model, analyze the application scenario, processing preference, watermark content, and data characteristics of the target power system data; S4. According to the analysis results obtained in step S3, perform a watermark embedding operation on the target power system data to complete the watermark operation of the target power system data.
2. The adaptive watermarking method for power system data according to claim 1, wherein The step of constructing a watermark database based on the watermarking scheme in step S1 specifically includes the following steps: The constructed database includes a rule layer and a location layer; The rule layer is used to determine how to generate watermarks or how to adjust the data structure; the location layer is used to determine the physical embedding location or logical embedding location of the watermark; The addition rule algorithms of the rule layer include a zero-watermark-based scheme, a differential expansion-based watermarking scheme, and a histogram shift-based watermarking scheme; The addition location algorithms of the location layer include a least significant bit-based watermarking scheme, an invisible character-based watermarking scheme, a pseudo-row and pseudo-column-based watermarking scheme, a multi-coding combination-based watermarking scheme, and a tuple permutation-based watermarking scheme.
3. The adaptive watermarking method for power system data according to claim 2, characterized in that The zero-watermark-based scheme specifically includes the following steps: Calculate the Pearson correlation coefficient between each attribute column in the target data and other attribute columns; Calculate the average value of the Pearson correlation coefficients between each attribute column and other attribute columns, and select the attribute column with the smallest average value as the candidate feature column; Divide the data of the candidate feature column into two parts: the first part is used for feature extraction, and the second part is used for watermark embedding. Set the number of blocks b as is the ceiling function, m is the number of tuple values; arrange the data of the candidate feature column according to the size of the tuple values, and divide it into k consecutive blocks, is the floor function; select the blocks with odd serial numbers for feature extraction, and select the blocks with even serial numbers for watermark embedding; if m mod b ≠ 0, then evenly distribute the remaining data to the feature extraction part and the watermark embedding part; Convert the m tuple values of the candidate feature column into a square matrix of h×h, where is the floor function; Compare the value of each data point in the obtained square matrix with the values of its 8 neighboring data points: if the values of the neighboring data points are all greater than or equal to the value of this data point, mark the neighboring data value as 1 in the LBP mode, otherwise mark the neighboring data value as 0 in the LBP mode; for each point in the square matrix, obtain the corresponding 8-bit binary number; Count the number of occurrences of each possible LBP model in the matrix to obtain the corresponding histogram vector including 256 elements, where each element represents the frequency of a specific LBP mode; Generate a feature vector, and compare each element of the feature vector with the corresponding average value: generate a new binary vector, and use this new binary vector as the perceptual hash; Perform an exclusive OR operation on the watermark information represented in binary and the perceptual hash to obtain the zero-watermark feature code.
4. The adaptive watermarking method for power system data according to claim 3, wherein The differential expansion-based watermarking scheme specifically includes the following steps: Calculate the difference d between adjacent data pairs as d = x i - x j , x i is the i-th data, x j is the j-th data; Expand the difference d into a watermark bit sequence d' as d' = 2d + b, where b is the watermark information to be embedded; For the obtained watermark bit sequence d', for the original data pair x i and x j are extended to ensure that d' = x i '- x' j , where is the ceiling function.
5. The adaptive watermarking method for power system data according to claim 4, characterized in that The histogram shift-based watermarking scheme specifically includes the following steps: Construct a histogram: count the occurrence frequencies of the numerical values of all target columns, construct a histogram, and identify the peak points and the distribution of adjacent values; Select the value with the highest frequency as the peak point P, shift the values from P + 1 to P + k one bit to the right to create k empty positions for embedding watermark information; Allocate the watermark information to be embedded to the vacated empty positions in the set order; The least significant bit-based watermarking scheme specifically includes the following steps: For integer data, replace the least significant bit of the quaternary representation of the data with the watermark bit; for floating-point data, replace the mantissa bit of the quaternary representation of the data with the watermark bit.
6. The adaptive watermarking method for power system data according to claim 5, characterized in that The watermarking scheme based on invisible characters specifically includes the following steps: Select invisible characters including 0, 9, 10, and 127 in ASCII code, and zero-width space, zero-width non-joiner, zero-width joiner, and zero-width non-breaking space in zero-width characters; The watermark sequence is set as a string of quaternary codes, denoted as W=(w1, w2,..., w m ), where w i =0 corresponds to the invisible character C0, w i =1 corresponds to the invisible character C1, w i =2 corresponds to the invisible character C2, w i =3 corresponds to the invisible character C3; through the mapping relationship, the watermark sequence W is converted into a string composed of invisible characters; the string is embedded at the end of the target attribute column to achieve the addition of the watermark; The watermarking scheme based on multi-encoding combination specifically includes the following steps: Select several attribute columns and classify different character encoding formats for each attribute column; use the assigned encoding format to embed the watermark information into the data; The character encoding formats include ASCII, UTF-8, UTF-16, and Unicode.
7. The adaptive watermarking method for power system data according to claim 6, characterized in that The watermarking scheme based on tuple permutation specifically includes the following steps: It is assumed that the database table T includes n tuples, where n > 3; the initial order is expressed as T = [t1, t2,..., t n ; t i is the tuple in the i-th row; The watermark sequence W is a quaternary code, expressed as W = (w1, w2,..., w m ); Define the permutation rule as follows: When w k = 0, the first tuple t1 and the last tuple t n will be swapped; When w k = 1, swap the first tuple t1 and the second last tuple t n-1 ; When w k = 2, swap the first tuple t1 and the second tuple t2; When w k = 3, swap the first tuple t1 and the third tuple t3.
8. The adaptive watermarking method for power system data according to claim 7, characterized in that The application scenario, processing preference, watermark content, and data feature analysis of the target power system data based on the random forest model described in step S3 specifically include the following steps: For the target power system data, conduct application scenario analysis, processing preference analysis, watermark content analysis, and data feature analysis; among them, the application scenario analysis includes the application from the middle platform to the library and the application of exporting to a file; the processing preference analysis includes the preference for processing functions, the preference for processing objects, and the user-specified preference; the watermark content analysis includes the full traceability of the data link and the traceability of the data exporter; the data feature analysis includes character-type data and numerical-type data; For character-type data, directly convert it into numerical-type data; For numerical-type data, perform normalization processing; Based on the random forest model, construct and train a processing preference prediction model; among them, the inputs of the processing preference prediction model include: Data features: including data types, such as character type, numerical type, etc.; Processing preferences: including the historical operation function set and the function call frequency; Application scenario markers: including the application from the middle platform to the library or exporting to a file; User-specified preferences: including the list of keyword fields marked by the user and the user-defined watermark strength requirement; The outputs of the processing preference prediction model include four types of prediction labels: Application scenario prediction: including the application from the middle platform to the library or exporting to a file; Processing preference prediction: including the function-sensitive path, the field-sensitive path, and the user-specified path; Watermark content prediction: including the full traceability of the data link or the traceability of the data exporter; Data feature prediction: including numerical-type dominance or character-type dominance; The training data sources of the processing preference prediction model include: Historical data processing records: Extract historical data tables, processing operations, and watermark embedding records from the logs of the target power system; Annotated samples: The optimal watermark algorithm combinations in different scenarios are annotated by experts; User feedback data: Record the protection requirements of users for specific fields and the tolerance of users to watermark perturbation; Input the target power system data into the pre-trained processing preference prediction model to obtain the corresponding processing preference prediction results; the processing preference prediction results include the application scenario prediction results, the processing preference prediction results, the watermark content prediction results, and the data feature prediction results.
9. The adaptive watermarking method for power system data according to claim 8, characterized in that Performing a watermark embedding operation on the target power system data according to the analysis result obtained in step S4, specifically including the following steps: (1) The set data table T includes n attribute columns A1 to A n , and the correlation coefficient matrix is calculated using the following formula: where C ij is the element in the i-th row and j-th column of the correlation coefficient matrix; cov(A i , A j ) is the covariance of the i -th column and the j -th column of A; is the standard deviation of the i -th column of A; For the numerical attribute column A k , the corresponding kurtosis K is calculated using the following formula k : Where E() is the variance; The following rules are used for judgment: If is greater than the set correlation threshold, then a watermarking scheme based on differential expansion is adopted; represents the maximum value of the average of the correlation coefficients in the i-th column of the correlation coefficient matrix; If K k is greater than the set peak threshold, then a watermarking scheme based on histogram translation is adopted for the k-th column; If it is character-type data, a zero-watermark-based scheme or an invisible-character-based watermark scheme is adopted; (2) Application scenario adaptation: For the scenario of middle platform to library application, a tuple permutation-based watermark scheme or a least significant bit-based watermark scheme is adopted; For the scenario of exporting to a file application, an invisible-character-based watermark scheme or a multi-coding combination-based watermark scheme is adopted; (3) Processing preference priority: Set the user-specified field set as F, then the importance vector is calculated using the following formula: where is the importance weight of the user-specified field f i ; H(f i ) is the information entropy of the field f i , and H(f i ) = -∑p(x)log(p(x)), where p(x) is the value-taking probability; |F| is the size of the user-specified field set F The penalty value is calculated using the following formula: where penalty is the penalty value for sensitive operation paths; F sens is the set of sensitive operation functions; is the sensitivity coefficient of function f; Senstivity(f,algo) is the sensitivity score of algorithm algo for function f; The preference score is calculated using the following formula: Pr = (ω f ) T ·u algo -penalty where Pr is the preference score; ω f is the importance weight vector of the user-specified field; u algo is the coverage vector of the field by the algorithm algo. If the algorithm acts on the field f i then u algo = 1, otherwise u algo = 0; For the user-specified field, a histogram translation-based watermark scheme or a tuple permutation-based watermark scheme is adopted; For the function sensitive path, a zero-watermark-based scheme is adopted; (4) Watermark content weight: The algorithm combination score is calculated using the following formula: TS(algo1,algo2) = γ1·Resilience(algo1) + γ2·Stealth(algo2) Where TS(algo1,algo2) is the combined score of the rule layer algorithm algo1 and the position layer algorithm algo2; γ1 is the first weight coefficient, γ2 is the second weight coefficient, and γ1 + γ2 = 1; Resilience(algo1) is the anti-attack value of the algorithm algo1, and Resilience(algo1) = 1 - BER, where BER is the bit error rate; Stealth(algo2) is the invisibility value of the algorithm algo2, and Stealth(algo2) = 1 - Detectbility, where Detectbility is the detection probability; For full traceability of the data link, a zero-watermark-based scheme is adopted in the rule layer, and a tuple permutation-based watermark scheme is adopted in the position layer; For tracing the data exporter, a differential expansion-based watermark scheme is adopted in the rule layer, and an invisible-character-based watermark scheme is adopted in the position layer; (5) Scheme selection: The analytic hierarchy process is used to calculate the scores of each combined algorithm; the combined algorithm is the paired combination of the rule layer algorithm and the position layer algorithm; for the above data characteristics, application scenarios, processing preferences, and watermark content, pairwise importance comparisons are made to generate the importance matrix P; by decomposing the eigenvalues of the importance matrix P, the weights Λ = [λ1, λ2, λ3, λ4] are obtained T , λ i is the sub - weight, i takes values from 1 to 4, and For each combined algorithm, the following formula is used to calculate the final score: Where TotalScore is the final score of each combined algorithm; S1 is the data feature matching degree, which is scored according to the correlation coefficient matrix and kurtosis; S2 is the scenario adaptation score; S3 is the processing preference matching degree, and the value is the preference score Pr; S4 is the watermark content weight, and the value is the algorithm combination score; Finally, the combined algorithm with the highest total score is selected for watermark embedding.
10. A system for implementing the adaptive watermarking method for power system data according to any one of claims 1 to 9, characterized in that Including a database construction module, a data acquisition module, a data analysis module, and an adaptive watermark module; the database construction module, the data acquisition module, the data analysis module, and the adaptive watermark module are connected in series in sequence; the database construction module is used to construct a watermark database based on the watermark scheme and upload the data information to the data acquisition module; the data acquisition module is used to obtain the data information of the target power system data according to the received data information and upload the data information to the data analysis module; The data analysis module is used to analyze the application scenario, processing preference, watermark content, and data characteristics of the target power system data based on the received data information and the random forest model, and upload the data information to the adaptive watermark module; The adaptive watermark module is used to perform a watermark embedding operation on the target power system data according to the received data information and the obtained analysis results to complete the watermark operation of the target power system data.