A website content identification method and system based on federated learning
Patent Information
- Application Number
- CN202511925312.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-19
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2045-12-19
AI Technical Summary
[0004]为克服现有技术中网站内容识别训练数据因严格安全和地域管辖限制而形成数据孤岛、跨域数据分布的高度异构性导致协作模型准确性低、以及缺乏安全合规机制保障模型信息跨域共享的合法性的不足,从而提供一种基于联邦学习的网站内容识别方法及系统
本发明通过采用联邦学习架构,解决了网站内容识别的跨域间训练的数据隐私和数据孤岛问题。通过结合偏差机制,在保证客户端数据隐私的前提下,使得全局模型能够自适应地补偿和修正高度异构的本地数据分布带来的偏差,从而确保跨域协作后的模型在各网站内容识别上的高准确性、高稳定性和强鲁棒性。最终,本方法实现了在不暴露任何域间敏感数据的前提下,安全、高效地聚合跨域模型经验,全面增强了特定网站内容识别等预测方面的数据可用性能力及其整体效能。
Smart Images

Figure CN121750304B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data security technology, and in particular relates to a method and system for website content recognition based on federated learning. Background Technology
[0002] With the development of the internet, mobile payment, and anonymity technologies, website content has become increasingly complex and concealed. Cross-border and cross-platform content sources pose a serious threat to the economy, public opinion, and cybersecurity. Although traditional security systems, such as threat intelligence sharing platforms and multi-agency cooperation mechanisms, have achieved information sharing to some extent, their core model training and decision-making capabilities remain fundamentally limited. This is because the underlying, sensitive transaction data and user behavior data possessed by different institutions, operators, and service providers remain within their respective regions or institutions, forming insurmountable data silos. This data isolation hinders the recognition of global threat patterns in specific website content and severely weakens the accuracy and generalization ability of models in cross-domain availability assessments.
[0003] Therefore, there is an urgent need for a usability framework that effectively combines cross-domain data content for trusted synchronous collaboration to identify website content. Summary of the Invention
[0004] To overcome the shortcomings of existing technologies, such as data silos formed by strict security and regional jurisdiction restrictions on website content recognition training data, low accuracy of collaborative models due to the high heterogeneity of cross-domain data distribution, and lack of security and compliance mechanisms to ensure the legality of cross-domain sharing of model information, this paper provides a website content recognition method and system based on federated learning.
[0005] This invention proposes a website content recognition method based on federated learning, which is executed collaboratively by a federated server and multiple clients, and includes the following steps: The federated server distributes the global model for the current round to multiple selected clients; Each client performs local training based on the global model and locally collected website content data. During the training process, a local personalized bias term is calculated to compensate for differences in data distribution, and a corresponding local model update is generated. Each client performs compliance processing on local model updates according to preset local security policies before uploading them to the federated server; The federated server aggregates local model updates from multiple clients to generate a new round of global models; Repeat the process of distribution, local training, and uploading aggregation until the training stopping condition is met to obtain the final global model for website content recognition.
[0006] Optionally, the process of each client performing local training based on the global model and locally collected website content data includes: The client collects local raw data covering user behavior, transaction characteristics, risk tags, and system environment characteristics; Based on the local raw data, the deviation of the local data itself and the difference deviation between clients are calculated, and then fused according to the preset weight to obtain the local deviation update amount, which is used to update the local personalized deviation item. The local model is updated by iteratively training the global model, the updated local personalized bias term, and the original local data.
[0007] Optionally, the process of calculating the local data's own bias and the client-side discrepancy bias includes: The internal deviation components of the local raw data are calculated using the orthogonal projection method. Calculate the difference between the local training gradient and the global model gradient, as the client-to-client difference bias component; The local deviation update amount is obtained by weighting and summing the internal deviation component and the differential deviation component according to a preset weight.
[0008] Optionally, the process of performing compliance processing on the local model update according to a preset local security policy includes: Based on local data protection regulations, identify the feature directions that are strongly correlated with sensitive features in local model updates; Construct a sensitive feature subspace based on the identified feature directions; Before uploading the local model update, remove the components that fall into the sensitive feature subspace.
[0009] Optionally, the process by which the federated server aggregates the local model updates from multiple clients includes: The federated server receives local model updates uploaded by each client after compliance processing; The global model update amount is obtained by calculating a weighted average of all received local model updates. The global model update amount is added to the global model parameters of the current round to generate the new round of global model.
[0010] Optionally, in each round of training, the method further includes: Each client uploads the obtained status response data to the federated server during local training. The federated server dynamically adjusts the global learning rate or the client selection strategy for the next round of training based on the received status response data.
[0011] On the other hand, the present invention also proposes a website content recognition system based on federated learning for implementing the method, including a federated server and multiple clients; The federated servers include: The service delivery module is used to deliver the global model to the client; The aggregation update module is used to aggregate local model updates uploaded by clients to generate a new global model; The service decision module is used to make decisions on training scheduling based on training objectives and client feedback. Each of the aforementioned clients includes: The data acquisition module is used to collect content data from local websites. The bias training module is used to calculate and update local personalized bias items during local training. The security policy module is used to process local model updates based on local compliance requirements; The upload module is used to update and upload the processed local model to the federated server for website content recognition.
[0012] Optionally, the deviation training module includes: The data preprocessing unit is used to perform feature standardization processing on the collected local raw data; The deviation weight calculation unit is used to calculate the weighted sum of the deviation of the local data itself and the difference deviation between the clients; The deviation update unit is used to store and cumulatively update local personalized deviation items.
[0013] Optionally, the security policy module is used for: Identify the directions of sensitive features protected by regulations in the model update and construct a sensitive feature subspace; remove components falling into the sensitive feature subspace from the local model update to be uploaded.
[0014] Optionally, the system further includes a service status response module, located on the federated server, for monitoring and reporting the performance metrics and resource status of local training to the service decision module.
[0015] Compared with the prior art, the present invention has the following advantages and technical effects: This invention addresses the data privacy and data silo issues in cross-domain training for website content recognition by employing a federated learning architecture. By incorporating a bias mechanism, it enables the global model to adaptively compensate for and correct biases introduced by highly heterogeneous local data distributions, while ensuring client-side data privacy. This ensures high accuracy, stability, and robustness of the model after cross-domain collaboration in recognizing content from various websites. Ultimately, this method achieves secure and efficient aggregation of cross-domain model experience without exposing any sensitive inter-domain data, comprehensively enhancing data availability and overall performance in prediction tasks such as website content recognition. Attached Figure Description
[0016] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a schematic diagram of the method flow according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the system structure according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the deviation training model structure according to an embodiment of the present invention. Detailed Implementation
[0017] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0018] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0019] Example 1 This embodiment provides a website content recognition method based on federated learning, which is executed collaboratively by a federated server and multiple clients, including the following steps: The federated server distributes the global model for the current round to multiple selected clients; Each client performs local training based on the global model and locally collected website content data. During the training process, a local personalized bias term is calculated to compensate for differences in data distribution, and a corresponding local model update is generated. Each client performs compliance processing on local model updates according to preset local security policies before uploading them to the federated server; The federated server aggregates local model updates from multiple clients to generate a new round of global models; Repeat the process of distribution, local training, and uploading aggregation until the training stopping condition is met to obtain the final global model for website content recognition.
[0020] The feasible process for each client to perform local training based on the global model and locally collected website content data includes: The client collects local raw data covering user behavior, transaction characteristics, risk tags, and system environment characteristics; based on the local raw data, it calculates the local data's own deviation and the difference deviation between clients, and fuses them according to preset weights to obtain the local deviation update amount, which is used to update the local personalized deviation item; iteratively training is performed by combining the global model, the updated local personalized deviation item, and the local raw data to obtain the local model update.
[0021] An feasible process for calculating the deviation of local data itself and the difference between client data includes: The internal bias component of the local original data is calculated using the orthogonal projection method; the difference between the local training gradient and the global model gradient is calculated as the client-to-client difference bias component; the internal bias component and the difference bias component are weighted and summed according to preset weights to obtain the local bias update amount.
[0022] The feasible process of performing compliance processing on the local model update according to a preset local security policy includes: According to local data protection regulations, feature directions that are strongly correlated with sensitive features in local model updates are identified; a sensitive feature subspace is constructed based on the identified feature directions; and before the local model update is uploaded, components that fall into the sensitive feature subspace are removed.
[0023] Implementable, the process by which the federated server aggregates the local model updates from multiple clients includes: The federated server receives local model updates uploaded by each client after compliance processing; it calculates a weighted average of all received local model updates to obtain the global model update amount; and it adds the global model update amount to the global model parameters of the current round to generate the global model for the new round.
[0024] Implementable, in each round of training, the method further includes: Each client uploads the obtained status response data to the federated server during local training; the federated server dynamically adjusts the global learning rate or the client selection strategy for the next round of training based on the received status response data.
[0025] As a specific implementation method, such as Figure 1As shown, this method can be executed collaboratively by the server and the client, and mainly includes the following steps: S201: Initialization and Model Configuration. The federated server loads the initial global model parameters. And set global training parameters, such as the total number of rounds. Learning rate Client sampling rate The client loads its own captured local dataset. and initialize local personalized deviation items. The security policy is initialized to reflect the legal and regulatory requirements of each local client.
[0026] S202: Client Selection and Model Deployment. The server selects a subset of clients to participate in this round of training based on the clients' status responses and preset strategies. And the current global model Distribute to selected clients .
[0027] S203: Client-side local training and gradient generation. Each client... Combination and local deviation item Forming a local complete model Based on local data The client performs local bias training and generates a local global model update. Update with local deviation items .
[0028] S204: State Response and Model Upload. Monitor local training performance and generate state feedback data. The client uploads model gradients and state response data to the federated server.
[0029] S205: Federated Aggregation and Model Updates. The federated server receives all updates from local clients. Perform weighted average aggregation to generate a global model update. The server updates the global model: .
[0030] S206: Model Decision-Making and Dynamic Adjustment. The server-side... The system receives performance feedback to assess whether the usability of the global model has improved and whether the convergence condition has been met. If convergence has not occurred, the hyperparameters (such as the learning rate and client sampling rate) are dynamically adjusted, and the system returns to step S202 to begin the next iteration, thus achieving closed-loop dynamic optimization.
[0031] S207: Model Deployment and Service Decision. After training convergence, the final model will be deployed to identify website content and related user information.
[0032] Through the above implementation methods, this embodiment can effectively solve the problem of cross-domain data trust collaboration in website content recognition and ensure data availability.
[0033] On the other hand, this embodiment also proposes a website content recognition system based on federated learning for implementing the method, including a federated server and multiple clients; The federated servers include: The service delivery module is used to deliver the global model to the client; The aggregation update module is used to aggregate local model updates uploaded by clients to generate a new global model; The service decision module is used to make decisions on training scheduling based on training objectives and client feedback. Each of the aforementioned clients includes: The data acquisition module is used to collect content data from local websites. The bias training module is used to calculate and update local personalized bias items during local training. The security policy module is used to process local model updates based on local compliance requirements; The upload module is used to update and upload the processed local model to the federated server for website content recognition.
[0034] like Figure 2 As shown, the system is deployed in a distributed network environment consisting of a federated server and multiple clients (i.e., multi-domain data acquisition), aiming to achieve secure and efficient cross-domain model training.
[0035] The system architecture consists of core modules that collaborate, including a server-side service decision-making module, service delivery module, and service status response module, as well as a client-side data acquisition module, deviation training module, security policy module, and client upload module. When the framework starts, the service decision-making module guides the service delivery module to upload the global model. Data is sent to the client. The client's data acquisition module collects raw local data. The bias training module then uses this data to personalize the local bias items. Updates are performed to compensate for model bias caused by non-independent and identically distributed models. Simultaneously, the client trains the model locally and generates global model updates. Subsequently, the client upload module updates it along with the local deviation item. The data is uploaded to the server, where it is received and federated to generate a new global model. Based on feedback from the service status response module, the system will conduct the next round of decision-making and scheduling. Ultimately, the execution decision-making process includes service decisions such as website content prediction and recognition.
[0036] Furthermore, the service decision module, acting as the central control module of the federated server, is responsible for the top-level scheduling and optimization of the entire training process. Specifically, based on preset global training objectives and convergence criteria (such as the accuracy threshold of the global model on the public validation set or the rate of decline of the loss function), this module dynamically determines the start and stop of training, the selection of clients to participate (such as based on the historical contribution quality and data volume of the client), and the aggregation frequency. After receiving performance feedback from the service status response module, this module executes adaptive adjustment strategies. For example, if it detects that the client's local loss function is oscillating drastically, it dynamically reduces the global learning rate that the service delivery module is about to deliver, or excludes the client from participating in the next round of aggregation, to ensure the availability and stability of the global model.
[0037] The service delivery module serves as the channel for information synchronization between the server and the client. Specifically, this module is responsible for delivering the latest global model weights determined by the service decision module. The current training learning rate Number of iterations And the necessary model data is distributed to the selected client groups.
[0038] The data acquisition module is responsible for comprehensively and in real-time collecting raw training data from specific content websites. These data cover user behavior, transaction characteristics, risk tags, and system environment characteristics. The collected data is cleaned and filtered locally.
[0039] Specifically, in terms of user behavior, the data collection module continuously monitors and records indicators such as user login frequency, average session duration, click depth, stability of time intervals for access and interaction, and periodic changes in user behavior patterns under the target website content.
[0040] Regarding transactions and fund flows, the data collection module gathers information on users' average transaction size, the frequency and amount fluctuations of deposits and withdrawals, changes in historical withdrawal channels, the correlation of payment methods, and potential credit card records.
[0041] Regarding risk control and security records, the data collection module aggregates the identity verification status of the website account, associated account information, number of historical abnormal login attempts, geographical location stability of the IP address and proxy usage detection records, high-risk behaviors automatically marked by the risk control system (such as multi-account association and batch registration), and records of being included in internal or external watchlists.
[0042] In terms of system environment characteristics, the data acquisition module collects specific resource information (static resources, dynamic resources), DNS resolution, IP address, registration information, etc. of the target website.
[0043] Data collection can be conducted through proactive probing or passive content, i.e., feedback information actively submitted by users.
[0044] Furthermore, the deviation training module includes: The data preprocessing unit is used to perform feature standardization processing on the collected local raw data; the deviation weight calculation unit is used to calculate the weighted sum of the deviation of the local data itself and the difference deviation between the client; the deviation update unit is used to store and cumulatively update the local personalized deviation items.
[0045] Specifically, the offset training module is responsible for addressing model bias caused by the non-independent and identically distributed nature of local data. Bias calculation is divided into two parts: the bias of the local data itself and the bias between different clients. The weight of the bias of the local data itself is... The weights for the differences and deviations between each client are: Self-bias is determined through orthogonal projection. calculate The difference in gradients between clients is the difference between the local gradient and the global gradient. The total deviation is .
[0046] Reference Figure 3 This module can be further subdivided into a data preprocessing unit, a deviation weight calculation unit, and a deviation update unit.
[0047] The data preprocessing unit is responsible for meticulously processing the raw data output by the data acquisition module. This includes statistical feature distribution differentiation, feature normalization (converting features with large differences in the units of different categories of data into a unified and comparable scale), and outlier robustness handling to ensure the quality of the data input to the deviation calculation module.
[0048] The bias weight calculation unit is the core of bias training. The implementation involves calculating the bias weight of its own data, specifically by calculating the negative correlation between its various discrepancies through orthogonal projection, reducing internal overhead. This component represents the unique feature patterns of the local data that are inconsistent with the current learning direction of the global model. Simultaneously, it calculates the bias weight of discrepancies between clients. This module aims to quantify the unique patterns of the local data and the degree of conflict with the global model, comprehensively considering the weight relationship between internal and external discrepancies to calculate the bias. .
[0049] The deviation update unit is responsible for storing and updating the client's local personalized deviation items. The update process is dynamic; this module will use the output of the deviation weight calculation module b. With the current Perform cumulative averaging to ensure It can continuously adapt to subtle changes in the local data distribution.
[0050] Implementable, the security policy module is used to identify the direction of sensitive features protected by regulations in model updates, construct a sensitive feature subspace, and remove components falling into the sensitive feature subspace from the local model update to be uploaded.
[0051] Specifically, the security policy module is a compliance assurance module for achieving cross-domain data availability. It allows clients to define preferences and constraints for model information sharing based on local laws, regulations, and industry regulatory requirements. Specifically, the core functions of this module include sensitive data compliance and data domain jurisdiction. Regarding sensitive data compliance, this module, in accordance with local data protection regulations (the Personal Information Protection Law and specific industry data security regulations), uses methods such as sensitivity analysis to identify feature vectors that must be strictly protected (such as user biometric information and precise transaction records). In terms of subspace construction, this module defines the identified feature directions as those used for sensitive feature subspaces. The core purpose of this operation is to ensure that, before uploading, model updates have removed components strongly related to sensitive information that is prohibited from being exported or shared by local laws and regulations. This ensures that model contributions meet the compliance requirements of cross-domain data collaboration, thereby activating and maintaining the availability of cross-domain data.
[0052] If feasible, the system also includes a service status response module, located on the federated server, for monitoring and reporting the performance metrics and resource status of local training to the service decision module.
[0053] Specifically, the service status response module is responsible for monitoring key status and performance metrics during local training and providing this response data to the server-side service decision module for global adjustments. Specifically, this module monitors and outputs the following key response data in real time: local model performance metrics (accuracy on the local validation set, loss function value, convergence speed), local system status (local computing resource utilization, network bandwidth / latency, training time overhead), and data status feedback (changes in local data volume, and the degree of difference in local data distribution in the latest iteration). This module sends the above status response data to the server after training is completed on the client or periodically, enabling the service decision module to dynamically adjust the global learning rate, client sampling rate, and the stopping conditions for the next training round based on multi-dimensional, real-time feedback information, thereby optimizing the overall training efficiency and availability of the framework.
[0054] The client upload module transmits the processed model update to the federated server for service decision-making and then distributes the service (model weights, gradients) to the next stage.
[0055] Through the above implementation methods, this embodiment can effectively solve the problem of cross-domain data trust collaboration in website content identification and ensure data availability.
[0056] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A website content recognition method based on federated learning, characterized in that, This process is performed collaboratively by a federated server and multiple clients, and includes the following steps: The federated server distributes the global model for the current round to multiple selected clients; Each client performs local training based on the global model and locally collected website content data. During the training process, a local personalized bias term is calculated to compensate for differences in data distribution, and a corresponding local model update is generated. Each client performs compliance processing on local model updates according to preset local security policies before uploading them to the federated server; The federated server aggregates local model updates from multiple clients to generate a new round of global models; Repeat the process of distribution, local training, and uploading aggregation until the training stopping condition is met to obtain the final global model for website content recognition; The process by which each client performs local training based on the global model and locally collected website content data includes: The client collects local raw data covering user behavior, transaction characteristics, risk tags, and system environment characteristics; Based on the local raw data, the deviation of the local data itself and the difference deviation between clients are calculated, and then fused according to the preset weight to obtain the local deviation update amount, which is used to update the local personalized deviation item. The local model update is obtained by iteratively training the global model, the updated local personalized bias term, and the local original data. The process of calculating the deviation of local data itself and the difference between clients includes: The internal deviation components of the local raw data are calculated using the orthogonal projection method. Calculate the difference between the local training gradient and the global model gradient, as the client-to-client difference bias component; The local deviation update amount is obtained by weighting and summing the internal deviation component and the differential deviation component according to a preset weight.
2. The method according to claim 1, characterized in that, The process of performing compliance processing on the local model update according to the preset local security policy includes: Based on local data protection regulations, identify the feature directions that are strongly correlated with sensitive features in local model updates; Construct a sensitive feature subspace based on the identified feature directions; Before uploading the local model update, remove the components that fall into the sensitive feature subspace.
3. The method according to claim 1, characterized in that, The process by which the federated server aggregates local model updates from multiple clients includes: The federated server receives local model updates uploaded by each client after compliance processing; The global model update amount is obtained by calculating a weighted average of all received local model updates. The global model update amount is added to the global model parameters of the current round to generate the new round of global model.
4. The method according to claim 1, characterized in that, In each round of training, the method further includes: Each client uploads the obtained status response data to the federated server during local training. The federated server dynamically adjusts the global learning rate or the client selection strategy for the next round of training based on the received status response data.
5. A website content recognition system based on federated learning according to any one of claims 1-4, characterized in that, Includes a federated server and multiple clients; The federated servers include: The service delivery module is used to deliver the global model to the client; The aggregation update module is used to aggregate local model updates uploaded by clients to generate a new global model; The service decision module is used to make decisions on training scheduling based on training objectives and client feedback. Each of the aforementioned clients includes: The data acquisition module is used to collect content data from local websites. The bias training module is used to calculate and update local personalized bias items during local training. The security policy module is used to process local model updates based on local compliance requirements; The upload module is used to update and upload the processed local model to the federated server for website content recognition.
6. The system according to claim 5, characterized in that, The deviation training module includes: The data preprocessing unit is used to perform feature standardization processing on the collected local raw data; The deviation weight calculation unit is used to calculate the weighted sum of the deviation of the local data itself and the difference deviation between the clients; The deviation update unit is used to store and cumulatively update local personalized deviation items.
7. The system according to claim 5, characterized in that, The security policy module is used for: Identify the directions of sensitive features protected by regulations in the model update and construct a sensitive feature subspace; remove components falling into the sensitive feature subspace from the local model update to be uploaded.
8. The system according to claim 5, characterized in that, The system also includes a service status response module, which is located on the federated server and is used to monitor and report the performance metrics and resource status of local training to the service decision module.
Citation Information
Patent Citations
Federal learning method based on adaptive differential privacy
CN117874829A
Personalized federal learning method, device and system based on gradient similarity dynamic model fusion
CN120106246A