Intelligent recognition method and system for website tampering and storage medium

By collecting full and real-time website page source code data, and after privacy protection and anonymization processing, a support vector machine model is used to identify website tampering. This solves the problem of insufficient self-optimization and generalization ability in existing technologies, and achieves efficient and accurate website tampering detection and self-learning.

CN120930187APending Publication Date: 2025-11-11BEIJING AN XIN TIAN XING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510997128.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing website tampering detection technologies cannot self-optimize or self-learn, require regular updates to the original content, lack generalization ability, and may generate false positives for legitimate updates.

Method used

By collecting full, real-time website page source code data, and after privacy protection and anonymization processing, a support vector machine model is used to construct a decision function by maximizing the classification margin to identify website tampering. The model parameters are then optimized through cross-validation and grid search to achieve self-learning and dynamic adjustment.

Benefits of technology

It improved detection accuracy, reduced false alarm rate, enhanced adaptability to dynamic changes on websites, protected user privacy, and achieved self-optimization and self-learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120930187A_ABST
    Figure CN120930187A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent recognition method for website tampering, and the method comprises the following steps: S1, website data collection: carrying out the full-amount real-time collection of the source code data of all levels of pages of a target website through a computer system, and enabling the collection depth to be configurable; s2, website data preprocessing: carrying out privacy protection and desensitization processing on the collected data, and sending the processed data to a message queue for decoupling caching; s3, website data feature extraction: extracting key features used for identifying abnormal changes of website contents from the preprocessed data; and S4, intelligent identification of website tampering: the extracted features are classified by using the trained support vector machine model, whether the website is tampered is identified, and the support vector machine model constructs a decision function by maximizing a classification interval. Detection is carried out depending on a support vector machine algorithm, and the website tampering problem is quickly responded; the detection process is self-optimized and self-learned, and the adaptive capacity of dynamic changes of the website is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of network security technology, and in particular relates to an intelligent identification method, system and storage medium for website tampering. Background Technology

[0002] Existing website tampering detection technologies, such as traditional detection software, mostly determine whether a website has been tampered with by comparing its page content with known original content or templates in real time. This typically involves text matching and hash value verification for content comparison; once a discrepancy is found, the website is considered tampered with, and corresponding measures are immediately taken. This method can accurately identify subtle changes in page content, including text, images, and links; it is highly flexible, allowing for customized detection rules and response strategies based on actual needs; and it monitors page content changes in real time, enabling rapid response to tampering events. However, it generally suffers from the following problems:

[0003] It cannot perform self-optimization or self-learning and needs to regularly update the original content or template to adapt to the dynamic changes in website content.

[0004] The generalization ability is insufficient, which may generate false alarms for legitimate updates (such as normal operations of the content management system) and cannot provide evidence of the changes. Summary of the Invention

[0005] The purpose of this invention is to provide an intelligent identification method, system, and storage medium for website tampering, so as to solve the problems existing in the prior art.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] A method for intelligently identifying website tampering, the method comprising the following steps:

[0008] S1: Website data collection, which uses a computer system to collect all levels of page source code data of the target website in real time, with configurable collection depth;

[0009] S2: Website data preprocessing, which involves privacy protection and anonymization of the collected data, and sending the processed data to a message queue for decoupling and caching;

[0010] S3: Website data feature extraction, extracting key features from preprocessed data to identify abnormal changes in website content;

[0011] S4: Website tampering intelligent identification. The extracted features are classified using a trained support vector machine model to identify whether the website has been tampered with. The support vector machine model constructs a decision function by maximizing the classification margin.

[0012] Furthermore, the website data collection described in S1 involves the full and real-time collection of source code data from the first-level, second-level, third-level, and so on pages of the website.

[0013] The data collection depth can be set manually during the collection process.

[0014] Furthermore, the website data preprocessing described in S2 is a privacy protection process for data, conducted in accordance with the Cybersecurity Law, the Data Security Law, and the Personal Information Protection Law, all centered around data security.

[0015] Furthermore, the support vector machine model is constructed in the following manner:

[0016] The website content tampering detection problem is transformed into a binary classification problem, with positive class representing normal content and negative class representing tampered content.

[0017] The two classes of data are separated by maximizing the margin hyperplane.

[0018] The Sequence Minimum Optimization (SMO) algorithm was used for model training.

[0019] Furthermore, the decision function of the support vector machine model described in S4 is:

[0020]

[0021] w T It is the transpose of the weight vector w, b is the bias term, and λ is the weight vector w. i It is a Lagrange multiplier, x i y is the feature vector of the i-th sample. i w is the class label of the i-th sample. T x i They are vectors w and x i The dot product of the samples, where m represents the number of samples.

[0022] Furthermore, the intelligent recognition method also includes a model optimization step: automatically finding the optimal parameter combination through cross-validation and grid search, and continuously optimizing the model parameters through a feedback mechanism.

[0023] Furthermore, the intelligent recognition method also includes a real-time detection and feedback optimization mechanism to dynamically adjust model parameters based on the detection results.

[0024] The present invention also provides an intelligent identification system for website tampering, the intelligent identification system comprising:

[0025] The website data collection module is configured to collect all page source code data of the target website in real time through a computer system, with configurable collection depth;

[0026] The website data preprocessing module is configured to perform privacy protection and anonymization processing on the collected data, and send the processed data to a message queue for decoupling and caching.

[0027] The website data feature extraction module is configured to extract key features from preprocessed data to identify abnormal changes in website content.

[0028] The website tampering intelligent identification module is configured to use a trained support vector machine model to classify the extracted features and identify whether the website has been tampered with. The support vector machine model constructs a decision function by maximizing the classification margin.

[0029] The present invention also provides a computer-readable storage medium storing computer program instructions, which, when executed by a computer, enable the computer to perform the steps of any of the above-described intelligent identification methods for website tampering.

[0030] By adopting the above technical solution, the present invention has the following beneficial effects:

[0031] 1. Multiple nodes and multiple systems are deployed with separate data collection programs. The collected data is sent to a message queue in a unified manner to achieve data decoupling, establish an intermediate cache, and achieve the purpose of peak smoothing and valley filling. Subsequent programs can directly pull data from the message queue.

[0032] 2. In accordance with laws and regulations, the collected data is anonymized to ensure that no information is leaked during data processing and to protect users' personal privacy.

[0033] 3. It replaces the original tampering detection method, relies on the support vector machine algorithm for detection, and responds quickly to website tampering issues; it improves generalization ability, improves detection accuracy, and reduces false positive rate; it enables the detection process to self-optimize and self-learn, eliminating excessive manual maintenance and improving the adaptability to dynamic changes in websites. Attached Figure Description

[0034] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0035] Figure 1 This is an overall flowchart of the method of the present invention;

[0036] Figure 2 This is a schematic diagram illustrating the principle of support vector machine classification.

[0037] Figure 3 This is a schematic diagram of the maximum margin hyperplane;

[0038] Figure 4 This is a block diagram of the system of the present invention. Detailed Implementation

[0039] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0040] Combination Figure 1 As shown, this invention provides an intelligent identification method for website tampering.

[0041] The intelligent recognition method includes the following steps:

[0042] S1: Website data collection, which uses a computer system to collect all levels of page source code data of the target website in real time, with configurable collection depth;

[0043] S2: Website data preprocessing, which involves privacy protection and anonymization of the collected data, and sending the processed data to a message queue for decoupling and caching;

[0044] S3: Website data feature extraction, extracting key features from preprocessed data to identify abnormal changes in website content;

[0045] The key features include text features, structural features, and link features;

[0046] S4: Website tampering intelligent identification. The extracted features are classified using a trained support vector machine model to identify whether the website has been tampered with. The support vector machine model constructs a decision function by maximizing the classification margin.

[0047] Specifically, the website data collection mentioned in S1 is a full, real-time collection of the source code data of the first-level, second-level, third-level, etc. pages of the website;

[0048] The data collection depth can be set manually during the collection process.

[0049] Specifically, according to the intelligent identification method according to claim 1, the website data preprocessing in S2 is to perform privacy protection processing on the data in accordance with the Cybersecurity Law, the Data Security Law and the Personal Information Protection Law, focusing on data security.

[0050] The key features described in S3 can reduce data dimensionality, remove noise, and improve the model's generalization ability; the extracted features can be used in tampering detection, security assessment, and content recommendation scenarios.

[0051] Specifically, the support vector machine model is constructed in the following manner:

[0052] The website content tampering detection problem is transformed into a binary classification problem, with positive class representing normal content and negative class representing tampered content.

[0053] The two classes of data are separated by maximizing the margin hyperplane.

[0054] The Sequence Minimum Optimization (SMO) algorithm was used for model training.

[0055] Specifically, the decision function of the support vector machine model described in S4 is:

[0056]

[0057] w T It is the transpose of the weight vector w, b is the bias term, and λ is the weight vector w. i It is a Lagrange multiplier, x i y is the feature vector of the i-th sample. i w is the class label of the i-th sample. T x i They are vectors w and x i The dot product of the samples, where m represents the number of samples.

[0058] Support Vector Machines (SVM): A binary classification algorithm model. Its purpose is to find a hyperplane to segment the samples. The principle of segmentation is to maximize the margin. Ultimately, it is transformed into a convex quadratic programming problem to solve.

[0059] Support Vector Machines (SVMs) are used to solve data classification problems in the field of pattern recognition and belong to supervised learning algorithms. They can be divided into two main categories: linear and nonlinear.

[0060] The Support Vector Machine (SVM) is divided into two main categories: linear and nonlinear.

[0061] Linear support vector machines are mainly used to process linearly separable or approximately linearly separable datasets. The goal is to find an optimal hyperplane that maximizes the separation of data points from different classes. This hyperplane is determined by support vectors, which are the data points closest to the hyperplane.

[0062] Nonlinear support vector machines are mainly used to process linearly inseparable datasets. They use kernel functions to map the original data to a high-dimensional feature space, making the data linearly separable in the high-dimensional space. The principle of linear support vector machines is then applied in the high-dimensional space to find the optimal hyperplane.

[0063] See Figure 2 As shown, Figure 2 The image shows a two-dimensional plane with two different types of data, represented by circles and squares. It's easy to find a straight line that perfectly separates the two types of data. However, there are more than one straight line that can completely separate the data points. So, which line should we choose from so many? To select the "optimal" line, we need to consider a key metric: the margin. The margin refers to the perpendicular distance from the point closest to the decision boundary (the line) to that line. Our goal is to find a line that maximizes this margin. Such a line is called the maximum margin hyperplane.

[0064] See Figure 3 As shown, Figure 3 Therefore, the requirement is that the interval between the two outer lines is maximized.

[0065] In the sample space Find a hyperplane w in T x i +b=0 is used to distinguish different samples. The dataset D consists of m samples, each sample being a pair (xi) consisting of a feature vector xi and a label yi. i ,y i The sample index i ranges from 1 to m, indicating that there are m samples in the dataset; each feature vector x i It is an n-dimensional vector belonging to the real number space R. n This means that each sample has n features; each label y i It is a scalar, taking the value -1 or 1. This type of binary label is typically used in binary classification problems, where 1 represents the positive class and -1 represents the negative class; w T : Represents the transpose of the weight vector w, which is usually a column vector. T Then it is its corresponding row vector, x i : represents the feature vector of the i-th sample, b: is the bias term, which is a scalar, w T x i This is vector w. T sum vector x i The inner product (dot product) operation. If w and x i They are all n-dimensional vectors.

[0066] Binary classification uses two parallel support vector planes. The hyperplane lies between these two planes.

[0067]

[0068] In this part, the numerator “∣f(x)∣” represents the absolute value of the function f(x), and the absolute value symbol “∣∣” is in the form of a vertical bar; the denominator “||w||” represents the norm of the vector w (the mapping from the vector space or matrix space to the real number field, usually referring to the L2 norm), and the norm symbol “||||” is in the form of a double vertical bar.

[0069] The margin can be expressed as:

[0070]

[0071] Where ||x1-x2|| represents the norm (length) of the difference vector between vectors x1 and x2, and cosθ represents the cosine of angle θ, where θ is the angle between the two vectors x1 and x2.

[0072] We need to maximize the margin:

[0073]

[0074] Right now

[0075]

[0076] Formula (1) is marked as “ab:”: where w T The expression represents the transpose of vector w, x1 and x2 represent two vectors respectively, the dot product symbol "." represents the vector dot product operation, and the right side of the equation is a constant 2.

[0077] Formula (2) is an equivalent transformation of Formula (1): ||w|| represents the norm (length) of vector w, ||x1-x2|| represents the norm of vector x1-x2, cosθ represents the cosine of the angle θ between the two vectors w and x1-x2, and the right side of the equation is also 2.

[0078] Formula (3) defines "margin": that is, the margin is equal to 2 divided by the norm of vector w. The problem of finding the maximum margin is transformed into finding the minimum value of w:

[0079] Mark the area above the hyperplane as 1 and the area below as -1. Based on the established planes a and b, the constraints are as follows:

[0080]

[0081] y i (wT x i +b)≥1

[0082] Where i is the index of the sample, used to identify the i-th sample in the dataset. y i Sample x i Category labels, usually taking y i =1 or y i =-1, indicating whether the sample belongs to the positive or negative class, respectively. T : The transpose of the weight vector w. x i : The feature vector of the i-th sample. b: The bias term. w T x i +b: The value of the decision function, used to determine the value of sample x. i The category.

[0083] By introducing the Lagrange multiplier / lambda:

[0084] F(x,y,λ)=f(x,y)+λh(x,y)

[0085] The original maximum margin problem with inequality constraints is transformed into a dual problem. The dual problem is an optimization problem with Lagrange multipliers, which is generally easier to solve.

[0086] F(x,y,λ) is a function of variables x, y, and parameter λ. f(x,y) is a bivariate function of x and y. λ is a parameter, typically appearing as a Lagrange multiplier in optimization problems and Lagrange multiplier methods. h(x,y) is also a bivariate function of x and y.

[0087] Lagrange:

[0088] ①

[0089]

[0090] Where L(w,b,λ) is a Lagrangian function, where w is the weight vector, b is the bias term, λ is the Lagrange multiplier vector, and λi is the i-th component of the vector λ. 2 x represents the square of the norm of the weight vector w. m is a positive integer representing the number of samples. i y is the feature vector of the i-th sample. i is the label of the i-th sample (usually 1 or -1). T x i It is the transpose of vector w and vector x i The dot product.

[0091] ②

[0092]

[0093] y represents the partial derivative of the Lagrange function L with respect to the variable w; w is a variable; λi are Lagrange multipliers, i ranging from 1 to m; i This is the label of the sample (usually 1 or -1); x i is the feature vector of the sample. This formula represents taking the partial derivative of the Lagrange function L with respect to w and setting it equal to 0.

[0094] λ represents the partial derivative of the Lagrange function L with respect to the variable b; b is a variable; i y i The meaning is the same as in the first set of formulas. This formula represents taking the partial derivative of the Lagrange function L with respect to b and setting it equal to 0.

[0095] ③ Substituting the terms with partial derivatives of zero back into the Lagrange function yields:

[0096]

[0097] Then its dual problem is:

[0098]

[0099] λ i ≥0, i=1,…,m

[0100] The final decision function is constructed as follows:

[0101]

[0102] Here, f(x) is a function, typically represented as a decision function in a Support Vector Machine (SVM), used to classify and predict the input sample x. T It is the transpose of the weight vector w, x i w is the feature vector of the i-th sample. T x i They are vectors w and x i The dot product of . b is the bias term. m represents the number of samples. λi is the Lagrange multiplier.

[0103] y i x is the class label of the i-th sample, typically taking the value 1 or -1. i T It is the sample feature vector x i transpose of x i T x i It is a vector x i The dot product with itself, i.e., x i The square of the L2 norm.

[0104] The following conditions must be met for KKT to be met:

[0105] 1.λ i >=0

[0106] 2. yi f(x i )>=1

[0107] 3.λ i (y i f(x i )-1)=0

[0108] 1. λ i >=0 indicates that the Lagrange multiplier λi is greater than or equal to zero.

[0109] 2, y i f(x i )>=1, where y i It is sample x i The category label (usually with a value of 1 or -1), f(x) i ) is the decision function in sample x i The output value at this point, the condition represents y i With f(x) i) The product of is greater than or equal to 1.

[0110] 3. λ i (y i f(x i )-1)=0, which is a complementary relaxation condition, meaning that for each sample i, either λi=0 or y i f(x i )-1=0.

[0111] Therefore, we can conclude that:

[0112] When y i f(x i When )=1, λ i ≠0, sample x i These are support vectors.

[0113] When y i f(x i When λ > 1, i =0, sample x i It is not a support vector (located outside the margin).

[0114] Specifically, the intelligent recognition method also includes a model optimization step: automatically finding the optimal parameter combination through cross-validation and grid search, and continuously optimizing the model parameters through a feedback mechanism.

[0115] Specifically, the intelligent recognition method also includes a real-time detection and feedback optimization mechanism to dynamically adjust model parameters based on the detection results.

[0116] In yet another embodiment, see Figure 4 As shown, an intelligent identification system for website tampering includes:

[0117] The website data collection module is configured to collect all page source code data of the target website in real time through a computer system, with configurable collection depth;

[0118] The website data preprocessing module is configured to perform privacy protection and anonymization processing on the collected data, and send the processed data to a message queue for decoupling and caching.

[0119] The website data feature extraction module is configured to extract key features from preprocessed data to identify abnormal changes in website content.

[0120] The website tampering intelligent identification module is configured to use a trained support vector machine model to classify the extracted features and identify whether the website has been tampered with. The support vector machine model constructs a decision function by maximizing the classification margin.

[0121] In yet another embodiment, a computer-readable storage medium stores computer program instructions that, when executed by a computer, cause the computer to perform the steps of the intelligent identification method for website tampering as described above.

[0122] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for intelligently identifying website tampering, characterized in that, The intelligent recognition method includes the following steps: S1: Website data collection, which uses a computer system to collect all levels of page source code data of the target website in real time, with configurable collection depth; S2: Website data preprocessing, which involves privacy protection and anonymization of the collected data, and sending the processed data to a message queue for decoupling and caching; S3: Website data feature extraction, extracting key features from preprocessed data to identify abnormal changes in website content; S4: Website tampering intelligent identification. The extracted features are classified using a trained support vector machine model to identify whether the website has been tampered with. The support vector machine model constructs a decision function by maximizing the classification margin.

2. The intelligent recognition method according to claim 1, characterized in that, The website data collection described in S1 is a full, real-time collection of the source code data of the first-level, second-level, third-level, etc. pages of the website that has been entered; The data collection depth can be set manually during the collection process.

3. The intelligent recognition method according to claim 1, characterized in that, The website data preprocessing described in S2 is a privacy protection process for data, conducted in accordance with the Cybersecurity Law, the Data Security Law, and the Personal Information Protection Law.

4. The intelligent recognition method according to claim 1, characterized in that, The support vector machine model is constructed in the following way: The website content tampering detection problem is transformed into a binary classification problem, with positive class representing normal content and negative class representing tampered content. The two classes of data are separated by maximizing the margin hyperplane. The Sequence Minimum Optimization (SMO) algorithm was used for model training.

5. The intelligent recognition method according to claim 1, characterized in that, The decision function of the support vector machine model described in S4 is: w T It is the transpose of the weight vector w, b is the bias term, and λ is the weight vector w. i It is a Lagrange multiplier, x i y is the feature vector of the i-th sample. i w is the class label of the i-th sample. T x i They are vectors w and x i The dot product of the samples, where m represents the number of samples.

6. The intelligent recognition method according to claim 1, characterized in that, It also includes a model optimization step: automatically finding the optimal parameter combination through cross-validation and grid search, and continuously optimizing the model parameters through a feedback mechanism.

7. The intelligent recognition method according to claim 1, characterized in that, It also includes a real-time detection and feedback optimization mechanism to dynamically adjust model parameters based on the detection results.

8. An intelligent identification system for website tampering, characterized in that, The intelligent recognition system includes: The website data collection module is configured to collect all page source code data of the target website in real time through a computer system, with configurable collection depth; The website data preprocessing module is configured to perform privacy protection and anonymization processing on the collected data, and send the processed data to a message queue for decoupling and caching. The website data feature extraction module is configured to extract key features from preprocessed data to identify abnormal changes in website content. The website tampering intelligent identification module is configured to use a trained support vector machine model to classify the extracted features and identify whether the website has been tampered with. The support vector machine model constructs a decision function by maximizing the classification margin.

9. A computer-readable storage medium storing computer program instructions, which, when executed by a computer, enable the computer to perform the steps of the intelligent identification method for website tampering as described in any one of claims 1-7.