Cross-domain privacy protection method, system and device for advertisement recommendation and medium

By encrypting and storing user behavior data on local devices and generating cross-domain joint feature vectors using privacy intersection and differential privacy techniques, the advertising recommendation model is split and deployed to edge nodes and local devices. This solves the problems of privacy leakage, security risks, and network burden in advertising recommendation systems, and improves the accuracy of recommendations and user experience.

CN120874122APending Publication Date: 2025-10-31ANHUI SANQI JIYU NETWORK TECH CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202511072200.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing advertising recommendation systems rely on third-party cookies, leading to privacy leaks, security risks, network burden, and user experience issues, which affect user trust and recommendation effectiveness.

Method used

By encrypting and storing user behavior data on local devices, privacy intersection and differential privacy techniques are used to generate cross-domain joint feature vectors. The advertising recommendation model is split and deployed to edge nodes and local devices for encrypted computation and federated learning, ensuring that the data does not leave the domain while generating effective feature vectors.

Benefits of technology

It enables cross-domain data utilization and user privacy protection, improves the accuracy and real-time performance of advertising recommendations, enhances user trust and security, and reduces network burden and the impact of frequent data exchange.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120874122A_ABST
    Figure CN120874122A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-domain privacy protection method and system for advertisement recommendation, equipment and a medium, and the method specifically comprises the steps: carrying out the encryption matching of user behavior data and anonymized equipment data, and generating an initial cross-domain joint feature vector; fusing the initial cross-domain joint feature vector and the disturbance feature vector to form a target cross-domain joint feature vector; splitting a pre-trained advertisement recommendation model into a feature coding sub-module and a reasoning sub-module, deploying the feature coding sub-module to a user adjacent edge node, and retaining the reasoning sub-module in a user local device; and based on the target cross-domain joint feature vector, performing calculation of the feature coding sub-module and calculation of the reasoning sub-module, and uploading the encrypted hidden layer feature vector to a federated learning aggregation server for global model updating. According to the method, cross-domain data utilization and user privacy protection in an advertisement recommendation process are realized, and effective feature vectors are generated for personalized advertisement recommendation while data are guaranteed not to be out of a domain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of advertising recommendation technology, and in particular to a cross-domain privacy protection method, system, device, and medium for advertising recommendation. Background Technology

[0002] In today's era of rapid digital information development, advertising recommendation systems have become an indispensable part of the internet business ecosystem. They aim to accurately push personalized advertising content to users based on their interests and behaviors, thereby improving the effectiveness and conversion rate of advertising and bringing higher revenue to advertisers and platforms.

[0003] Currently, most advertising recommendation systems rely on third-party cookies to track user behavior, which serves as a key basis for personalized advertising recommendations. Third-party cookies can record a user's browsing history, click behavior, and other information as they browse different websites. By collecting and analyzing this data, advertising recommendation systems can build user profiles and then push advertisements that match their interests and preferences.

[0004] However, while this method of advertising recommendation relying on third-party cookies has improved the accuracy and effectiveness of ad targeting to some extent, it has also caused a series of serious problems, as follows: 1. Privacy Issues: Third-party cookies comprehensively and continuously track users' browsing behavior. Every webpage visit and every click a user makes may be recorded and analyzed. This puts users' personal privacy in a state of extreme transparency, making it difficult for users to control how their information is collected and used. This raises concerns about privacy leaks and leads to user distrust of advertising recommendation systems.

[0005] 2. Security Issues: Third-party cookies are stored on the user's local device, and their data transmission and processing may contain security vulnerabilities. Malicious attackers may exploit these vulnerabilities to steal user data stored in third-party cookies, thereby obtaining sensitive user information such as login credentials and personal preferences, causing user data leaks and a series of security problems, resulting in potential economic losses and security risks to users.

[0006] 3. Efficiency Issues: Systems relying on third-party cookies need to frequently exchange data with the server to update and synchronize user behavior data. This frequent data interaction not only increases the network burden, leading to network congestion and bandwidth waste, but also significantly prolongs the system's response time, affecting the real-time performance and accuracy of advertising recommendations, and reducing user experience.

[0007] 4. User Experience Issues: Due to the aforementioned privacy and security concerns, as well as the performance impact caused by frequent data exchange, users may encounter problems such as slow ad loading, inaccurate content recommendations, and even receive a large number of irrelevant or offensive ads while browsing the web. These issues severely affect the user's online experience and further reduce users' trust and acceptance of the advertising recommendation system. Summary of the Invention

[0008] The purpose of this invention is to provide a cross-domain privacy protection method, system, device, and medium for advertising recommendation, which realizes cross-domain data utilization and user privacy protection in the advertising recommendation process. While ensuring that the data does not leave the domain, it generates effective feature vectors for personalized advertising recommendation, thus balancing recommendation effect and privacy security, and solving at least one of the above-mentioned problems of the prior art.

[0009] In a first aspect, the present invention provides a cross-domain privacy protection method for advertising recommendations, the method specifically comprising: When a user visits a first-party website, user behavior data is collected and encrypted and stored on the local device. By using privacy intersection technology, user behavior data is encrypted and matched with anonymized device data provided by security hardware manufacturers, and an initial cross-domain joint feature vector is generated without leaving the domain of the data from each party. The intensity of the Laplace noise is dynamically adjusted according to the user's preset sensitivity level to generate a perturbation feature vector that satisfies differential privacy. The initial cross-domain joint feature vector and the perturbation feature vector are then fused to form the target cross-domain joint feature vector. The pre-trained ad recommendation model is split into a feature encoding sub-module and an inference sub-module. The feature encoding sub-module is deployed to the edge nodes near the user through a CDN-level edge computing network, while the inference sub-module is kept on the user's local device. Based on the target cross-domain joint feature vector, the feature encoding submodule and the inference submodule are calculated, and the encrypted hidden layer feature vector is uploaded to the federated learning aggregation server for global model update. Real-time ad matching is performed based on the encoded features returned by the user's neighboring edge nodes, and the local privacy budget counter is updated in conjunction with the user's interaction with the ads.

[0010] Secondly, the present invention provides a cross-domain privacy protection system for advertising recommendations, the system specifically comprising: The first privacy protection module is used to collect user behavior data and encrypt and store it to the local device when the user visits a first-party website; The second privacy protection module is used to collect user behavior data and encrypt and store it on the local device when the user visits the first-party website; through privacy intersection technology, the user behavior data is encrypted and matched with the anonymized device data provided by the security hardware manufacturer, and an initial cross-domain joint feature vector is generated under the premise that the data of each party does not leave the domain. The third privacy protection module is used to collect user behavior data and encrypt and store it to the local device when the user visits a first-party website; dynamically adjust the Laplace noise intensity according to the user's preset sensitivity level, generate a perturbation feature vector that satisfies differential privacy, and fuse the initial cross-domain joint feature vector and the perturbation feature vector to form the target cross-domain joint feature vector. The fourth privacy protection module is used to collect user behavior data and encrypt and store it on the local device when the user visits a first-party website; the pre-trained advertising recommendation model is split into a feature encoding sub-module and an inference sub-module. The feature encoding sub-module is deployed to the edge nodes near the user through a CDN-level edge computing network, while the inference sub-module is kept on the user's local device. The fifth privacy protection module is used to collect user behavior data and encrypt and store it on the local device when the user visits a first-party website; based on the target cross-domain joint feature vector, it performs feature encoding submodule calculation and inference submodule calculation, and uploads the encrypted hidden layer feature vector to the federated learning aggregation server for global model update; The sixth privacy protection module is used to collect user behavior data and encrypt and store it on the local device when the user visits a first-party website; perform real-time ad matching based on the encoded features returned by the user's neighboring edge nodes, and update the local privacy budget counter based on the user's interaction with the ads.

[0011] Thirdly, the present invention provides a computer device comprising: a memory and a processor, and a computer program stored in the memory, wherein when the computer program is executed on the processor, it implements a cross-domain privacy protection method for advertising recommendations as described in any of the above methods.

[0012] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements a cross-domain privacy protection method for advertising recommendations as described in any of the above methods.

[0013] Compared with the prior art, the present invention has at least one of the following technical effects: 1. This invention realizes cross-domain data utilization and user privacy protection in the process of advertising recommendation. While ensuring that the data does not leave the domain, it generates effective feature vectors for personalized advertising recommendation, thus balancing recommendation effect and privacy security.

[0014] 2. This invention collects user behavior data and encrypts and stores it on a local device. It also uses privacy intersection technology and differential privacy technology to process the data, ensuring that the joint feature vector is generated and fused without leaving the domain. This avoids the leakage of user privacy information and enhances users' trust in the advertising recommendation system.

[0015] 3. This invention employs encryption technology and a secure multi-party computation protocol to encrypt data transmission and processing, preventing user data from being maliciously stolen and tampered with during transmission and processing, effectively ensuring the security of user data and reducing security risks.

[0016] 4. This invention splits the pre-trained advertising recommendation model into a feature encoding sub-module and an inference sub-module, and deploys the feature encoding sub-module to edge nodes near the user through a CDN-level edge computing network. This reduces the amount of data interaction between the user's local device and the server, reduces the network burden, and improves the system's response speed and the real-time performance of advertising recommendations.

[0017] 5. This invention provides users with more accurate and personalized ad recommendation services by real-time ad matching and updating the local privacy budget counter based on user interaction behavior. At the same time, it avoids the impact of frequent data exchange on user experience and improves user satisfaction and acceptance of the ad recommendation system.

[0018] 6. This invention uses multi-round encryption and homomorphic operation matching to accurately generate an initial cross-domain joint feature vector without leaving the domain of the data from each party, effectively protecting the security of user behavior data and anonymized device data.

[0019] 7. This invention dynamically adjusts the noise intensity to generate and fuse perturbation feature vectors based on the user's preset sensitivity level, thereby meeting differential privacy requirements and enhancing the protection of user privacy.

[0020] 8. This invention splits the advertising recommendation model and deploys it separately to the user's nearby edge nodes and local devices, reducing data transmission volume, lowering network latency, and improving the response speed of advertising recommendations.

[0021] 9. This invention calculates and encrypts gradient data based on target feature vectors, and achieves global model updates through secure multi-party computation, thereby improving the accuracy and generalization ability of the model while protecting user privacy.

[0022] 10. This invention matches advertisements in real time based on encoding features and updates the privacy budget counter in conjunction with user interaction behavior, effectively controlling privacy consumption and preventing excessive privacy leakage while providing personalized advertisements.

[0023] 11. This invention creates a privacy consumption mapping table by comprehensively considering multiple factors, and defines reasonable consumption weight coefficients for different interactive behaviors, thereby achieving accurate quantification and management of user privacy consumption. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a flowchart illustrating a cross-domain privacy protection method for advertising recommendations provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a cross-domain privacy protection system for advertising recommendation provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0026] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0027] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0028] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0029] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0030] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0031] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0032] In this application embodiment, the entity executing the process includes a terminal device. This terminal device includes, but is not limited to, devices capable of executing the methods disclosed in this application, such as servers, computers, smartphones, and tablets. Figure 1 A flowchart illustrating a cross-domain privacy protection method for advertising recommendations according to an embodiment of the present invention is shown below in detail: S101 collects user behavior data and encrypts and stores it on a local device when a user visits a first-party website.

[0033] In this embodiment, when a user accesses a first-party website through a browser, a lightweight data collection module embedded in the website's front end is automatically loaded. This module is pre-deployed by the first-party website developer and is restricted to accessing only resources under the same origin domain through the browser's Content Security Policy (CSP) to prevent cross-domain script injection attacks. Upon the user's first visit, the browser will clearly inform the user of the purpose, scope, and encrypted storage method of data collection through a pop-up window or privacy settings page, and initiate the data collection process after obtaining explicit authorization from the user. The collection module only collects non-sensitive data related to user behavior, including the URL of the currently accessed page, page title, access timestamp, mouse click position, page scroll depth, and dwell time. It also obtains device environment information (such as browser type, operating system version, and screen resolution) through the browser's API, but all information that directly identifies the user (such as IP address and device MAC address) is excluded from the collection scope.

[0034] To prevent user behavior data from being associated with real identities, the data collection module de-identifies the raw data. For sensitive fields in device environment information (such as screen resolution), a hash algorithm (such as SHA-256) is used to generate fixed-length de-identified values, and a random salt is added to prevent rainbow table attacks. Simultaneously, the collection module uses a browser-provided encrypted random number generator (such as the `crypto.getRandomValues()` method of the Web Crypto API) to generate a unique temporary identifier (Session ID) for each user session. This identifier is only valid until the current browser tab or window is closed and is not bound to the user account or device hardware information. All collected data is associated using this temporary identifier as an index, ensuring that it is impossible to trace back to a specific user through a single data point.

[0035] The data collection module employs asymmetric encryption to encrypt and store user behavior data. Upon a user's first visit, the first-party website server sends a public key to the browser via HTTPS. This public key is generated by pairing the server's private key and is used solely for encryption. The collection module locally converts the de-identified user behavior data into JSON format and calls the browser's built-in encryption interface (such as the `subtle.encrypt()` method of the Web Crypto API) to encrypt the data using the server's public key, generating ciphertext data blocks. These encrypted data blocks are stored in binary form in the browser's IndexedDB or LocalStorage. Simultaneously, the collection module maintains an encrypted data index table in memory, recording the mapping between temporary identifiers and the storage locations of ciphertext data blocks; however, the index table itself does not store any plaintext information.

[0036] To prevent unauthorized access or tampering of encrypted data, the browser sets the IndexedDB or LocalStorage area storing user behavior data to allow access only to scripts from the same origin, and prohibits third-party scripts from reading or modifying the content of this area through CSP policies. Simultaneously, the collection module adds a timestamp field to the encrypted data index table to record the time of each data update and sets a data retention period (e.g., 30 days). Data exceeding this period is automatically deleted from local storage. Furthermore, the browser periodically performs integrity checks on the stored encrypted data by calculating the hash value of the encrypted data blocks and comparing it with the hash value at the time of initial storage to ensure that the data has not been tampered with. If data anomalies are detected, the collection module immediately clears local storage and reinitializes the data collection process.

[0037] When a user subsequently visits a first-party website, the data collection module first checks its local storage for any unuploaded encrypted data blocks. If found, the module uploads the encrypted data block and its corresponding temporary identifier to the first-party website's server via HTTPS. The server decrypts the data using its private key and extracts only anonymized behavioral characteristics (such as page visit frequency and click hotspot distribution) to update the user profile, without storing the original encrypted data or temporary identifier. To further protect user privacy, the data collection module allows users to manually clear encrypted data from local storage via browser settings or set automatic cleanup rules (such as deleting data immediately upon closing the browser). Simultaneously, first-party websites must clearly explain the data collection, encrypted storage, and usage methods in their privacy policies and provide user complaint channels to ensure the entire process complies with data protection regulations (such as GDPR or CCPA).

[0038] S102 uses privacy intersection technology to encrypt and match user behavior data with anonymized device data provided by security hardware manufacturers, generating an initial cross-domain joint feature vector without leaving the domain of any party's data.

[0039] In this embodiment, to achieve cross-domain data matching, the first-party website and the security hardware vendor negotiate and determine a privacy intersection protocol (such as an elliptic curve-based PSI protocol). Both parties agree to use the same elliptic curve parameters (such as secp256r1) and hash function (such as Keccak-256), and generate a joint session key for subsequent communication encryption. The first-party website reads the encrypted user behavior data ciphertext packet from its local device, and the security hardware vendor extracts the encrypted anonymized device data ciphertext set from its database. Both parties exchange public key certificates through a secure channel (such as TLS 1.3), verify each other's identities, and establish an encrypted communication link.

[0040] Under the privacy intersection protocol framework, the first-party website divides its encrypted user behavior data ciphertext into blocks, adding a random blinding factor (such as a random number) to each block to generate a set of blinded data blocks. The security hardware vendor performs similar processing on its encrypted anonymized device data ciphertext, generating a set of blinded device data blocks. Both parties use partial homomorphic encryption (such as Paillier encryption) to encrypt the blinded data blocks, generating homomorphic ciphertext. The first-party website sends its homomorphic ciphertext of user behavior data to the security hardware vendor, and the security hardware vendor sends its homomorphic ciphertext of device data to the first-party website. Both parties perform intersection operations locally on the received homomorphic ciphertexts (e.g., through polynomial interpolation or Bloom filter matching) to identify the intersection items (i.e., successfully matched device identifiers and behavioral feature combinations) that exist simultaneously in both user behavior data and device data. Due to the characteristics of homomorphic encryption, the intersection calculation can be completed in ciphertext form without decrypting the original data. Both parties verify the correctness of the intersection calculation results using zero-knowledge proofs to ensure there are no false matches or data tampering. After successful verification, the first-party website and the security hardware manufacturer each use their private keys to decrypt the ciphertext of the intersection item and obtain the matching result in plaintext form.

[0041] Based on the matching results obtained through privacy-preserving intersection, the first-party website and the security hardware vendor each extract feature data related to the matching items from their local databases. Both parties convert their respective feature data into fixed-dimensional numerical vectors (e.g., through one-hot encoding or word embedding), and then use the Secure Multi-Party Computation (SMC) protocol to concatenate and fuse the vectors. Specifically, the two parties use the SMC protocol to calculate the weighted sum or concatenation operation of the vectors in encrypted form to generate an initial cross-domain joint feature vector.

[0042] This embodiment achieves secure matching of user behavior data and anonymized device data in an encrypted state, generates an initial cross-domain joint feature vector, and strictly adheres to the principles of data minimization and privacy protection to ensure that the original data of all parties does not leave the domain and is not leaked.

[0043] S103 dynamically adjusts the Laplace noise intensity according to the user-preset sensitivity level, generates a perturbation feature vector that satisfies differential privacy, and fuses the initial cross-domain joint feature vector and the perturbation feature vector to form the target cross-domain joint feature vector.

[0044] In this embodiment, when a user first uses the advertising recommendation system, the system guides the user to set a privacy sensitivity level through an interactive interface, providing options such as "High," "Medium," and "Low," each corresponding to different levels of privacy protection. After the user selects a level, the system stores the sensitivity level parameter in the encrypted storage area of ​​the local device and binds it to the user's unique identifier. When the user subsequently accesses the advertising recommendation service, the system reads this parameter from the local storage as a basis for dynamically adjusting the noise intensity. For example, if the user selects a "High" sensitivity level, the system will assign a stronger privacy protection strategy to ensure greater anonymity of their data in cross-domain joint analysis.

[0045] The system retrieves the corresponding parameter from a pre-defined noise intensity mapping table based on the user's preset sensitivity level. For example, a "high" sensitivity level corresponds to a noise intensity parameter α=0.5, a "medium" level corresponds to α=0.3, and a "low" level corresponds to α=0.1. This parameter α is related to the dimension of the initial cross-domain joint feature vector. The system dynamically adjusts the noise generation range based on the vector dimension n to ensure that the noise intensity is inversely proportional to the data sensitivity. For example, for a 100-dimensional initial vector, if the user selects "high" sensitivity, the system will generate 100 independent Laplace random numbers. The scale parameter of each random number is determined by α and the feature value sensitivity weight, which is pre-set according to the feature type (such as geographic location or browsing duration).

[0046] The system adds generated Laplace noise to the corresponding dimensions of the initial cross-domain joint feature vector, forming a perturbed feature vector. For example, if the feature value of "time spent browsing e-commerce websites" in the initial cross-domain joint feature vector is 10 minutes, and the noise added at "high" sensitivity is -3 minutes, then after perturbation, this feature value becomes 7 minutes. The system performs differential privacy verification on the perturbed vector to ensure that it satisfies the (ε, δ)-differential privacy definition, where ε (privacy budget) is determined by both the sensitivity level and the noise intensity, and δ (failure probability) is set to a minimum value (e.g., 10^-5). If the verification passes, the perturbed vector is marked as "privacy protection effective"; otherwise, noise is regenerated until the condition is met.

[0047] The system performs a dimensional weighted fusion of the initial cross-domain joint feature vector and the perturbation feature vector. For example, for users with "high" sensitivity, the initial vector weight is set to 0.3, and the perturbation vector weight is set to 0.7. The value of each dimension of the fused target vector is the weighted sum of the initial value and the perturbation value. The fusion weights are dynamically adjusted according to the sensitivity level; the higher the sensitivity, the greater the weight of the perturbation vector, to enhance privacy protection. The fused target vector is encrypted and stored locally on the device and used as input for subsequent advertising recommendation models.

[0048] During ad recommendation, the system continuously monitors user interactions with the recommendations (such as click-through rate and dwell time). If users frequently ignore recommended ads or mark them as "not interested," the system infers that the current privacy protection level may be too high, leading to decreased recommendation accuracy. In this case, it automatically lowers the sensitivity level (e.g., from "high" to "medium") and re-executes the above steps to generate a new target cross-domain joint feature vector. Conversely, if user interaction is active, the system maintains or increases the sensitivity level to balance privacy protection and recommendation effectiveness. This dynamic feedback mechanism ensures that the system can adaptively adjust its privacy protection strategy based on actual user needs.

[0049] S104 splits the pre-trained advertising recommendation model into a feature encoding sub-module and an inference sub-module. The feature encoding sub-module is deployed to the user's nearest edge node through a CDN-level edge computing network, while the inference sub-module is kept on the user's local device.

[0050] In this embodiment, during the model training phase, a public dataset containing user historical behavior data, device features, and ad interaction tags is used to pre-train a complete ad recommendation model in a secure computing environment. This model employs a multi-layer neural network structure, including an input layer, a feature encoding layer, a hidden feature extraction layer, and an output inference layer. After training, a model structure analysis tool is used to identify the functional dependencies between each layer, determining the logical boundary between the feature encoding layer and the inference layer. For example, the feature encoding layer is responsible for converting raw user behavior data (such as browsing duration and click categories) into a high-dimensional feature vector, while the inference layer calculates ad relevance scores based on this vector. Both can be deployed independently through data stream splitting.

[0051] The pre-trained model is split into two independent sub-modules. The feature encoding sub-module contains an input layer and a feature encoding layer, responsible for receiving raw user behavior data (such as encrypted browsing history stored locally) and generating a fixed-dimensional hidden feature vector. The output of this module is a standardized feature vector, which does not contain any directly identifiable user information. The inference sub-module contains a hidden feature extraction layer and an output inference layer. Based on the feature vector output by the feature encoding sub-module, it calculates the user's matching score for candidate advertisements and generates a recommendation ranking result. This module is deployed on the user's local device and only receives encrypted feature vectors as input, avoiding the transmission of raw data. During the splitting process, an intermediate verification layer is added to ensure the compatibility of the sub-module interfaces. For example, a hash checksum is embedded at the output of the feature encoding sub-module, and the inference sub-module verifies data integrity upon receiving the data to prevent functional abnormalities due to inconsistent module versions after deployment.

[0052] Based on user geographic location and network topology information, the CDN edge node closest to the user's local device is dynamically selected as the deployment location for the feature encoding submodule. Specifically, the approximate geographic location is determined by the user's local device IP address or GPS information (after generalization processing), and 3-5 candidate edge nodes are selected by combining the CDN node coverage database. Probe packets are sent to the candidate nodes to measure round-trip time (RTT) and packet loss rate, prioritizing nodes with an RTT of less than 50ms and a packet loss rate of less than 1%. The real-time computing resource utilization (such as CPU and memory utilization) of the candidate nodes is queried, and nodes with resource utilization below 70% are selected as the final deployment target to avoid feature encoding delays due to node overload. After determining the target node, the containerized image (such as a Docker image) of the feature encoding submodule is pushed to that node, and the service is started through the node management platform, configured to only receive encrypted requests from specific user local devices to ensure controllable data access permissions.

[0053] The inference submodule is packaged as a lightweight local application (such as a mobile SDK or browser extension) and distributed to users' local devices through app stores or official channels. During deployment, the binary code of the inference submodule is obfuscated, removing debug symbols and sensitive function names to prevent reverse engineering attacks. An independent sandbox environment is allocated to the inference submodule on the user's local device, restricting its access to system resources (such as prohibiting access to the local file system, camera, and other non-essential hardware). Communication between the inference submodule and the feature encoding submodule uses an end-to-end encryption protocol (such as TLS 1.3), with the key generated locally on the user's local device and rotated periodically to ensure that the feature vector is not stolen or tampered with during transmission. After deployment, the running status of the inference submodule is periodically verified through a local health check service. If an anomaly is detected (such as process crash or illegal permission request), the service is immediately terminated and an alarm is triggered.

[0054] When a user initiates an ad recommendation request, the user's local device reads the user behavior data of the current session (such as the product categories viewed in the last 10 times) from the encrypted storage area. After local lightweight cleaning (such as removing duplicate records), a data packet to be encoded is generated. The user's local device encrypts the data packet and sends it to a nearby edge node. The feature encoding submodule on the edge node decrypts the data, extracts effective features, generates a hidden layer feature vector, and encrypts it again before returning it to the user's local device. The inference submodule on the user's local device receives the encrypted feature vector, decrypts it, inputs it into the model to calculate the ad matching score, sorts the data according to the score, generates a recommendation list, and displays it on the user interface. The inference submodule encrypts and uploads the key metrics of this calculation (such as the ID and timestamp of the ad clicked by the user) to the federated learning aggregation server. The server aggregates multi-user data, updates the global model parameters, and periodically pushes the updated inference submodule to the user's local device to achieve iterative optimization of the model.

[0055] S105, based on the target cross-domain joint feature vector, performs feature encoding submodule calculation and inference submodule calculation, and uploads the encrypted hidden layer feature vector to the federated learning aggregation server for global model update.

[0056] In this embodiment, the system generates a unique identifier (such as a hash value based on a timestamp and device ID) for the target cross-domain joint feature vector, and encapsulates the vector and identifier together into an encrypted data packet. A layered encryption strategy is employed during encapsulation: the outer layer uses the public key of the federated learning aggregation server for encryption, ensuring that only the server can decrypt it; the inner layer uses a symmetric key generated locally on the user's device for encryption, preventing interception by a man-in-the-middle during transmission. After encryption, the data packet is stored in the secure storage area of ​​the user's local device (such as a TEE trusted execution environment), awaiting subsequent computation task scheduling.

[0057] After locating a nearby edge node via a CDN selection protocol (such as DNS-based intelligent routing), the user's local device sends a computation request to that node. The request contains metadata about the encrypted data packets (such as packet size and encryption type), but not the actual data content. Upon receiving the request, the edge node first verifies the legitimacy of the user's local device (e.g., checking if the device certificate is on a whitelist). If the verification is successful, it returns a confirmation message indicating that computational resources are ready.

[0058] The user's local device fragments the encrypted data packet and transmits it to the edge node (the fragment size is dynamically adjusted according to network bandwidth, typically 10-100KB / fragment). The edge node receives and temporarily stores the fragments in its node buffer. Subsequently, the feature encoding submodule on the edge node (pre-loaded into the node's memory) uses the public key of the federated learning aggregation server to decrypt the outer encapsulation and obtain the feature vector encrypted with the inner symmetric key. At this point, the edge node initiates a key negotiation request to the user's local device. The user's local device sends the symmetric key to the edge node in two parts via a secure channel (such as TLS 1.3): one part is the plaintext key (used for current calculations), and the other part is the key hash value (used for subsequent verification).

[0059] Edge nodes use plaintext keys to decrypt inner-layer data, obtaining the target cross-domain joint feature vector, and then perform feature encoding calculations. This calculation includes features normalization (e.g., mapping browsing time to the [0,1] interval) and feature cross-combination (e.g., combining "browsed product category" and "device model" into a new feature), ultimately generating a fixed-dimensional hidden-layer feature vector. During the calculation, edge nodes monitor resource utilization (e.g., CPU utilization) in real time. If the utilization exceeds 80%, the current calculation task is paused, and high-priority requests (e.g., real-time ad bidding) are prioritized. Calculation resumes only after resources are released.

[0060] After the edge nodes complete feature encoding, they re-encrypt the hidden layer feature vectors (using the user's local device public key) and return them to the user's local device. Upon receiving the data, the user's local device decrypts it using its local private key and inputs the decrypted vectors into the inference submodule. The inference submodule performs inference calculations based on a pre-trained ad recommendation model (already deployed locally), including steps such as ad relevance scoring (e.g., calculating the probability of a user clicking on candidate ads) and ranking optimization (e.g., generating a recommendation list based on the scores). During inference, the inference submodule dynamically adjusts the results based on the real-time context information of the user's local device (e.g., current time, geographical location). For example, if the user is currently near a shopping mall, the score weight of ads from nearby merchants is increased. After calculation, the inference submodule generates the final ad recommendation list and records key interaction metrics (e.g., the ID of the ad clicked by the user, display duration). These metrics will be used for subsequent privacy budget updates and model optimization.

[0061] The user's local device encrypts the hidden layer feature vectors (including intermediate calculation results) generated during the inference submodule's computation process and uploads them to the federated learning aggregation server. Before uploading, the system performs differential privacy enhancement processing on the vectors: based on the user's preset sensitivity level (e.g., adding more noise for highly sensitive users), Laplacian noise is added again to ensure that the server cannot deduce the original user data from the vectors.

[0062] After receiving encrypted hidden feature vectors uploaded by multiple users, the federated learning aggregation server performs global model aggregation. During the aggregation process, the server uses a secure multi-party computation protocol (such as a weighted average based on homomorphic encryption) to calculate the average of all user vectors without decrypting individual user data, and updates the parameters of the global model (such as neural network weights). After the update is complete, the server generates a new model version number and pushes the updated inference submodule (containing only the parameter changes) to the user's local device via the CDN network.

[0063] After receiving the model update on the user's local device, the system verifies the integrity of the update package locally (e.g., verifying the digital signature). If the verification passes, the parameters of the current inference submodule are hot-replaced, achieving seamless model upgrades. Simultaneously, the system dynamically adjusts subsequent privacy budget allocation based on user feedback on the updated model (e.g., changes in click-through rate), balancing privacy protection and recommendation effectiveness.

[0064] S106, perform real-time ad matching based on the encoded features returned by the user's neighboring edge nodes, and update the local privacy budget counter based on the user's interaction with the ads.

[0065] In this embodiment, after completing edge computing in the feature encoding submodule, the user equipment receives the encrypted hidden feature vector returned by the neighboring edge node through a secure transmission channel (such as an encrypted connection based on TLS 1.3). Upon reception, the system first verifies the integrity of the data packet by comparing the hash value of the data packet with the signature of the edge node to ensure that the data has not been tampered with. If the verification fails, the system automatically triggers a retransmission mechanism, sends a retransmission request to the edge node, and records the transmission error log for subsequent analysis. After successful verification, the user equipment decrypts the data packet using a locally pre-stored symmetric key to obtain the hidden feature vector. The decrypted vector contains multi-dimensional feature representations of user behavior (such as browsing product categories, dwell time, click frequency, etc.), which the system stores in a local secure cache (such as isolated storage based on a TEE trusted execution environment) and sets a cache validity period (usually 24 hours). After the expiration period, the data is automatically cleared to prevent sensitive information from being retained for a long time.

[0066] User devices synchronize the currently available ad candidate pool from the Ad Content Management System (ACM). During the synchronization process, the system adopts an incremental update strategy, downloading only the ad metadata (such as ad ID, category tags, and target audience characteristics) that has been added or modified since the last synchronization, reducing network transmission volume. The ad candidate pool is stored in a local database and indexed by dimensions such as ad category and delivery time to improve subsequent matching efficiency. Based on the hidden layer feature vector returned by the edge nodes, the system performs ad matching calculations. The matching process is divided into two stages: (1) Coarse-grained screening: Based on the main dimensions in the hidden layer feature vector (such as user interest categories), irrelevant ads in the ad candidate pool are quickly filtered. For example, if a user has been frequently browsing sports equipment recently, sports ads are retained first, and irrelevant categories such as beauty and home furnishings are removed. (2) Fine-grained sorting: For the ads after coarse screening, the system calculates their similarity score with the hidden layer feature vector. The similarity calculation adopts the weighted cosine similarity algorithm, and the weight is dynamically adjusted according to the user's historical behavior (such as increasing the weight of the "price range" feature if the user is price-sensitive). Finally, the ads are sorted from high to low to generate a list of recommended ads.

[0067] When a user's device displays a list of recommended ads, an interaction behavior monitoring module is simultaneously activated. This module captures user interaction events with ads through browser extensions or built-in app hooks, including ad clicks, display duration exceeding a threshold (e.g., 3 seconds), and active ad closing. For each interaction event, the system records key information such as event type, timestamp, and ad ID, and generates a unique event identifier (Event ID). Simultaneously, the system initializes a local privacy budget counter based on the user's preset privacy sensitivity level (e.g., high, medium, and low). The privacy budget counter limits the number of times user data is used for model training, preventing excessive tracking. During initialization, high-sensitivity users are allocated a lower privacy budget (e.g., 0.1 budget per interaction, total budget 5), while low-sensitivity users are allocated a higher budget (e.g., 0.3 budget per interaction, total budget 15). The budget allocation rules can be customized by the user in the device settings.

[0068] Whenever a user interacts with an ad, a corresponding value is deducted from the privacy budget counter based on the type of interaction and the user's sensitivity level. For example, a highly sensitive user clicking an ad deducts 0.3 budget, while displaying an ad without clicking deducts 0.1 budget; a low-sensitivity user clicking an ad deducts 0.5 budget, while displaying an ad without clicking deducts 0.2 budget. If the remaining value of the budget counter falls below a threshold (e.g., 20% of the total budget), the system triggers an alert mechanism, pushing a notification to the user (e.g., "Your privacy protection mode is enabled, and ad recommendations will be less personalized"), and suspends the uploading of this interaction data to the federated learning aggregation server. The privacy budget counter automatically recovers at regular intervals (e.g., resetting to its initial value every day at midnight), or recovers through user intervention (e.g., increasing the budget by 10% after the user completes a privacy protection task, such as reading and agreeing to the privacy policy). Simultaneously with updating the privacy budget, the system encrypts and uploads the key information of this interaction event (Event ID, Ad ID, Interaction Type) to the federated learning aggregation server. Differential privacy technology is used during upload to perturb the interaction data (e.g., adding Laplacian noise to click events to prevent the server from accurately inferring the user's true behavior), ensuring user privacy is not compromised. After receiving the data, the server uses it for fine-tuning the global model to optimize the ad recommendation strategy.

[0069] In some embodiments, in step S102 above, the step of encrypting and matching user behavior data with anonymized device data provided by security hardware manufacturers using privacy intersection technology, and generating an initial cross-domain joint feature vector without leaving the domain of any party's data, specifically includes: On the first-party website client, based on user behavior data, encryption is performed using a first public key encryption algorithm to generate a first set of encrypted identifiers; On the security hardware vendor's server, the data from the anonymization device is encrypted using a second public key encryption algorithm to generate a second set of encrypted identifiers; The first encrypted identifier set is sent to the security hardware vendor's server, where a blinding factor is applied and the first encrypted identifier set is encrypted a second time to generate a blinded encrypted identifier set. The blinded encrypted identifier set and the second encrypted identifier set are matched using a homomorphic operation to generate the intersection identifier index in the encrypted state; Based on the intersection identifier index, the first-party website client and the security hardware manufacturer's server exchange homomorphic encrypted device feature data through a secure channel. The encrypted device feature data and user behavior data are then concatenated to generate an initial cross-domain joint feature vector.

[0070] In this embodiment, during a user's visit to a first-party website, the client collects user behavior data in real time. This data includes, but is not limited to, the URLs of pages viewed, links clicked, dwell time, and search keywords. The collected data is stored and organized using user identifiers (such as user IDs) as indexes to form a user behavior dataset. Simultaneously, the user behavior data is cleaned to remove invalid or erroneous data records, ensuring the accuracy and completeness of the data.

[0071] Security hardware vendors collect anonymized data related to user devices from their servers, such as device hardware model, operating system version, device usage duration, and device geographic location information (obfuscated). This data is stored and managed using device identifiers (such as device serial numbers) as indexes, forming an anonymized device dataset. Similarly, the anonymized device data undergoes preprocessing to remove duplicates and anomalies.

[0072] Choose a suitable public-key encryption algorithm as the first public-key encryption algorithm, such as RSA. Use this algorithm to encrypt user identifiers in the user behavior dataset. Specifically, the client uses the pre-obtained public key from the security hardware vendor to encrypt each user identifier, generating a corresponding first encrypted identifier. Combine all the first encrypted identifiers into a first encrypted identifier set.

[0073] Choose another public-key encryption algorithm as the second public-key encryption algorithm, such as the ElGamal algorithm. Use this algorithm to encrypt the device identifiers in the anonymized device dataset. The server uses its own public key to encrypt each device identifier, generating a corresponding second encrypted identifier. Combine all the second encrypted identifiers into a second encrypted identifier set.

[0074] The first-party website client sends the generated first set of encrypted identifiers to the security hardware vendor's server via a secure network channel (such as an encrypted connection based on the TLS protocol). Upon receiving the first set of encrypted identifiers, the security hardware vendor's server generates a set of random numbers as blinding factors. These blinding factors are used to blind each identifier in the first set of encrypted identifiers; that is, a mathematical operation (such as multiplication) is performed on the encrypted identifier to maintain certain mathematical relationships while preventing the original encrypted identifier from being directly deduced from the blinded result. Then, the server uses its own private key to perform a second encryption on the blinded identifiers, generating a blinded encrypted identifier set.

[0075] The security hardware vendor's server performs homomorphic matching on the generated set of blinded encrypted identifiers and the second set of encrypted identifiers. Homomorphism is a special type of encryption operation that allows computation on encrypted data without first decrypting it. For example, if an encryption algorithm with additive homomorphic properties is used, the server can perform some operation (such as a variation of a comparison operation) on the blinded and second encrypted identifiers without decrypting them, finding their intersection. Through homomorphic matching, the server generates an intersection identifier index in the encrypted state, indicating which blinded encrypted identifiers match the second encrypted identifier.

[0076] Based on the generated intersection identifier index, the server extracts corresponding device feature data (such as device hardware configuration information, usage habit data, etc.) from the anonymized device dataset, and encrypts this device feature data using a homomorphic encryption algorithm to generate homomorphically encrypted device feature data. Then, the server sends the homomorphically encrypted device feature data to the first-party website client through a secure channel.

[0077] Upon receiving homomorphically encrypted device feature data, the client extracts corresponding user behavior data from locally stored user behavior data based on the intersection identifier index. The extracted user behavior data is then concatenated with the received homomorphically encrypted device feature data. Because the device feature data is homomorphically encrypted, the concatenation process is performed in an encrypted state, preventing the leakage of original data information. After concatenation, an initial cross-domain joint feature vector is generated. This vector integrates user behavior data and device feature data, providing a rich information foundation for subsequent data analysis and applications.

[0078] In some embodiments, step S103 above, which involves dynamically adjusting the Laplacian noise intensity according to a user-preset sensitivity level, generating a perturbation feature vector that satisfies differential privacy, and fusing the initial cross-domain joint feature vector and the perturbation feature vector to form a target cross-domain joint feature vector, specifically includes: Based on the data sensitivity level preset by the user in the visual configuration interface, the corresponding noise intensity coefficient is generated through the local privacy management module. Based on the noise intensity coefficient, a Laplace distribution noise sequence is generated on the user's local device, and the Laplace distribution noise sequence is superimposed element-wise with the initial cross-domain joint feature vector to generate a perturbation feature vector. The initial cross-domain joint feature vector and the perturbation feature vector are weighted and fused using a preset privacy-preserving weight allocator to generate the target cross-domain joint feature vector.

[0079] In this embodiment, users preset data sensitivity levels by accessing a visual configuration interface. The interface presents different sensitivity level options in an intuitive graphical format, such as low, medium, and high, each corresponding to a different level of privacy protection. Users select the appropriate sensitivity level based on their needs and level of concern for data privacy.

[0080] The local privacy management module interacts with the visual configuration interface to obtain the user's preset data sensitivity level information. Based on the preset mapping relationship between sensitivity levels and noise intensity coefficients, the local privacy management module generates corresponding noise intensity coefficients. This mapping relationship is pre-defined during the system design phase according to differential privacy theory and practical application requirements. For example, a low sensitivity level corresponds to a smaller noise intensity coefficient, and a high sensitivity level corresponds to a larger noise intensity coefficient. In this way, different sensitivity levels can correspond to different intensities of noise addition to meet diverse user privacy protection needs.

[0081] The local privacy management module generates a Laplace-distributed noise sequence on the user's local device based on the generated noise intensity coefficient. The local device uses a random number generation algorithm, combined with the noise intensity coefficient, to generate a series of random noise values ​​following the Laplace distribution. These noise values ​​possess specific statistical properties that meet the requirements of differential privacy. For example, the larger the noise intensity coefficient, the greater the fluctuation range of the generated noise values, thus providing stronger privacy protection. The length of the generated Laplace-distributed noise sequence is the same as the dimension of the initial cross-domain joint feature vector to facilitate subsequent element-wise superposition operations.

[0082] The generated Laplace-distributed noise sequence is element-wise superimposed on the initial cross-domain joint feature vector. The initial cross-domain joint feature vector was generated during previous data processing by fusing user information from different data sources using techniques such as privacy-preserving intersection. Element-wise superposition involves adding each noise value in the noise sequence to the corresponding element in the initial cross-domain joint feature vector to obtain the perturbed element value. This operation generates the perturbed feature vector. While retaining some information from the initial cross-domain joint feature vector, the perturbed feature vector, due to the addition of Laplace noise, effectively protects individual data information, meeting the requirements of differential privacy.

[0083] A pre-defined privacy-preserving weight allocator performs a weighted fusion of the initial cross-domain joint feature vector and the perturbation feature vector. The privacy-preserving weight allocator pre-sets the weight values ​​of the initial cross-domain joint feature vector and the perturbation feature vector based on the system's privacy protection strategy and actual application requirements. For example, if data usability is prioritized, the weight of the initial cross-domain joint feature vector can be appropriately increased; if privacy protection is emphasized, the weight of the perturbation feature vector is increased. The weighted fusion process involves multiplying each element of the initial cross-domain joint feature vector by its corresponding weight, multiplying each element of the perturbation feature vector by its corresponding weight, and then adding the corresponding elements of the two result vectors to finally generate the target cross-domain joint feature vector. The target cross-domain joint feature vector contains some original data information and, through noise addition and weighted fusion, meets the requirements of differential privacy protection, allowing for safe use in subsequent data analysis and applications.

[0084] In some embodiments, step S104 above, which involves splitting the pre-trained advertising recommendation model into a feature encoding submodule and an inference submodule, deploying the feature encoding submodule to edge nodes near the user via a CDN-level edge computing network, and keeping the inference submodule on the user's local device, specifically includes: The pre-trained ad recommendation model is divided into a feature encoding submodule and an inference submodule according to the calculation function. The feature encoding submodule is responsible for the transformation of the original features into latent vectors, and the inference submodule is responsible for the mapping of latent vectors to ad matching results. By using the node selection algorithm of the CDN-level edge computing network, the feature encoding submodule is deployed to the edge node closest to the user's local device, while the inference submodule is retained on the user's local device. When a user’s local device initiates a recommendation request, the feature encoding submodule is dynamically downloaded from the edge node. The feature encoding submodule is then executed on the user’s local device to calculate the user’s feature data and obtain the first hidden layer feature vector. The first hidden layer feature vector is lightly encrypted and transmitted to the local inference submodule. The local inference submodule generates personalized ad matching results and displays recommended ads on the user's local device interface.

[0085] In this embodiment, the pre-trained advertising recommendation model is analyzed in depth and broken down into a feature encoding submodule and an inference submodule based on its computational functions. The feature encoding submodule is responsible for converting raw features into latent vectors. Raw features include multi-dimensional information such as the user's age, gender, browsing history, and purchase history. After processing by the feature encoding submodule, this information is transformed into latent vectors with fixed dimensions. These latent vectors can more effectively represent the user's feature information, providing a foundation for subsequent inference. The inference submodule is responsible for mapping the latent vectors to advertising matching results. Based on the user feature information contained in the latent vectors, it selects the advertisements that best match the user's interests and needs from the advertising database, achieving personalized advertising recommendations. For example, in an e-commerce platform's advertising recommendation system, the feature encoding submodule converts the user's original features, such as the product categories browsed and the price range purchased, into latent vectors. The inference submodule then uses these latent vectors to select advertisements that the user may be interested in from a large number of product advertisements for display.

[0086] Leveraging the node selection algorithm of CDN-level edge computing networks, this algorithm comprehensively considers multiple factors, such as the network distance between the edge node and the user's local device, the node's computing resources, and network bandwidth, to select the edge node closest to the user's local device from a large pool of edge nodes. The resulting feature encoding submodule is then deployed to this selected edge node, ensuring that the user's feature data travels through the shortest network path during transmission and reducing transmission latency. Simultaneously, the inference submodule is retained on the user's local device. Since the inference submodule is relatively small and has high real-time requirements, retaining it locally allows for rapid response to user recommendation requests, further improving recommendation efficiency. For example, on a social media platform, when a user is located in a certain city, the node selection algorithm finds an edge node near that city, deploys the feature encoding submodule on that node, and installs the inference submodule on the user's local device, such as a mobile phone.

[0087] When a user's local device initiates a recommendation request, the system first dynamically downloads the feature encoding submodule from selected edge nodes. The dynamic download method can be flexibly adjusted according to network conditions and user needs; for example, downloading quickly when network conditions are good, and using chunked downloading when the network is congested to ensure smooth downloading. After downloading, the feature encoding submodule is executed on the user's local device to calculate the user's feature data and obtain the first hidden layer feature vector. This process utilizes the computing resources of the edge nodes, reducing the computational burden on the user's local device. Furthermore, because the feature encoding submodule performs a certain degree of preprocessing at the edge nodes, the generated first hidden layer feature vector is more concise and effective. For example, a user's browsing history data may be very large; after processing by the feature encoding submodule at the edge nodes, the generated first hidden layer feature vector can more accurately reflect the user's interests and preferences.

[0088] The obtained first hidden layer feature vector is lightweight encrypted. The encryption algorithm is chosen to ensure data security while considering computational efficiency, avoiding excessive computational burden on the user's local device. The encrypted first hidden layer feature vector is transmitted to the local inference submodule, which decrypts it and generates personalized ad matching results according to preset inference rules. Finally, recommended ads are displayed on the user's local device interface. The display method can be designed according to different application scenarios, such as displaying product ads in a combination of images and text on e-commerce platforms, and displaying various ads in the form of a news feed on social media platforms. For example, on a video platform, the personalized ad matching results generated based on the user's first hidden layer feature vector may be ads related to the user's preferred video types, which will be displayed on the video playback interface in the form of interstitial ads or overlays.

[0089] In some embodiments, step S105 above, which involves performing feature encoding submodule calculation and inference submodule calculation based on the target cross-domain joint feature vector, and uploading the encrypted hidden layer feature vector to the federated learning aggregation server for global model update, specifically includes: Based on the target cross-domain joint feature vector, the feature encoding submodule is executed on the user's local device to generate the initial hidden layer feature vector; The initial hidden layer feature vector is input into the local inference submodule for forward computation. The local model gradient is generated through the backpropagation algorithm, and gradient clipping and normalization are performed on the local model gradient to obtain the processed gradient data. The processed gradient data is encrypted using an additive homomorphic encryption algorithm to obtain encrypted gradient data, which is then uploaded to the federated learning aggregation server. The federated learning aggregation server collects encrypted gradient data from multiple participants, performs weighted average aggregation operations through a secure multi-party computation protocol, and generates global model update parameters. The global model update parameters are distributed to each participant, and the parameters of the inference submodule and feature encoding submodule are updated on the user's local device and the user's neighboring edge nodes, respectively.

[0090] In this embodiment, the target cross-domain joint feature vector is input into a pre-deployed feature encoding submodule on the user's local device. The feature encoding submodule performs a series of transformations and processes on the input feature vector, such as feature mapping and dimensionality reduction, converting it into an initial hidden layer feature vector. This process aims to extract key information from the feature vector, reduce data dimensionality and redundancy, and provide more efficient input for subsequent inference computation. For example, the feature encoding submodule can employ a deep neural network structure, using multiple layers of nonlinear transformations to convert the original feature vector into an initial hidden layer feature vector with a higher level of abstraction.

[0091] The generated initial hidden layer feature vector is input into the local inference submodule for forward computation. Based on the preset model structure and parameters, the inference submodule further calculates and processes the initial hidden layer feature vector, outputting the ad recommendation result. Simultaneously, during the computation process, the inference submodule records the output and error information of each layer, providing a foundation for the subsequent backpropagation algorithm. For example, in a deep learning-based ad recommendation model, the inference submodule may contain multiple fully connected layers and activation function layers, progressively obtaining the final ad recommendation score through forward computation.

[0092] Based on the ad recommendation results output by the inference submodule and a pre-defined loss function, the backpropagation algorithm is used to calculate the local model gradient. Starting from the output layer, the backpropagation algorithm calculates the gradient of each layer forward, using a chain rule to pass error information back to the parameters of each layer, thus obtaining the gradient value of each parameter. These gradient values ​​reflect the direction and magnitude of the model parameter adjustments under the current input and are crucial for model updates. For example, if the mean squared error loss function is used, the backpropagation algorithm will calculate the gradient of each layer's parameters based on the difference between the predicted result and the true label.

[0093] Gradient clipping and normalization are performed on the generated local model gradients. Gradient clipping prevents gradient explosion by clipping gradients to a set threshold range when the absolute value exceeds it. Gradient normalization scales the gradients to a suitable numerical range, improving model training stability and convergence speed. For example, a gradient clipping threshold of 1.0 can be set; if the gradient value of a parameter is greater than 1.0, it is set to 1.0; if it is less than -1.0, it is set to -1.0. The clipped gradients are then normalized to have a mean of 0 and a variance of 1.

[0094] The process involves encrypting the processed gradient data using an additive homomorphic encryption algorithm and uploading the encrypted gradient data to the federated learning aggregation server. Specifically, a suitable additive homomorphic encryption algorithm, such as the Paillier encryption algorithm, is selected on the user's local device. Additive homomorphic encryption algorithms have the characteristic of performing addition operations in ciphertext, which allows the aggregation server to aggregate the encrypted gradient data from multiple participants without decryption during the federated learning process, effectively protecting user data privacy. For example, the Paillier encryption algorithm, based on the difficult problem of modular arithmetic of composite numbers, has high security and computational efficiency. Based on the selected additive homomorphic encryption algorithm, an encryption key pair, including a public key and a private key, is generated on the user's local device. The public key is used to encrypt the gradient data and can be publicly shared with other participants and the aggregation server; the private key is used to decrypt the data and must be strictly kept confidential, stored only on the user's local device. For example, when using the Paillier encryption algorithm, the public and private keys are generated by selecting suitable large prime numbers and generators. The generated public key is then used to encrypt the gradient data after gradient clipping and normalization. The encryption process follows the rules of additive homomorphic encryption, converting gradient data into ciphertext. The encrypted gradient data has a different representation than the original data, but still retains the property of being capable of addition. For example, for each gradient value, it is encrypted using a public key to obtain the corresponding ciphertext. The encrypted gradient data is then uploaded to the federated learning aggregation server over the network. During the upload process, secure communication protocols, such as HTTPS, are used to ensure the security and integrity of the data during transmission. Simultaneously, the uploaded data can be compressed to reduce network traffic and transmission time. For example, the user's local device can package the encrypted gradient data and send it to a designated interface of the aggregation server via HTTPS.

[0095] The federated learning aggregation server monitors and collects encrypted gradient data from multiple participants in real time. The server assigns a unique identifier to each participant to distinguish data uploaded by different participants. Simultaneously, the server validates the collected data to ensure its integrity and legitimacy. For example, the server checks whether the uploaded data conforms to predetermined format and size requirements and whether it originates from a legitimate participant.

[0096] To perform encrypted gradient data aggregation while protecting the data privacy of all participants, the aggregation server selects a suitable secure multi-party computation protocol. Secure multi-party computation protocols allow multiple participants to collaboratively complete a computational task without disclosing their individual private data. For example, a secure multi-party computation protocol based on homomorphic encryption can be chosen. This protocol combines additive homomorphic encryption and secure multi-party computation techniques, enabling weighted average aggregation of gradient data in encrypted form.

[0097] Based on the selected secure multi-party computation protocol, the aggregation server performs a weighted average aggregation operation on the collected cryptographic gradient data from multiple participants. During the aggregation process, the server assigns appropriate weights to the cryptographic gradient data of each participant based on factors such as the amount and importance of their data. Then, the cryptographic gradient data is calculated according to the weighted average rules to generate global cryptographic gradient data.

[0098] The generated global encrypted gradient data is decrypted using a private key pre-stored on the aggregation server (in some secure multi-party computation protocols, decryption may be achieved through multi-party collaboration, not just the aggregation server storing the private key). Then, based on the global model gradient and preset parameters such as the learning rate, global model update parameters are generated. These global model update parameters are used to update the parameters of the advertising recommendation model to improve its performance and accuracy.

[0099] The federated learning aggregation server distributes the generated global model update parameters to all participants over the network. During distribution, secure communication protocols are used to ensure the secure transmission of parameter data. Simultaneously, the server can record the distribution time and the participants' reception status for subsequent tracking and management. For example, the aggregation server packages the global model update parameters and sends them to the designated interfaces of each participant via HTTPS.

[0100] After receiving the global model update parameters, the user's local device updates the parameters of the local inference submodule according to the parameter type and structure. The update process follows the model parameter update rules, merging the global model update parameters with the current parameters of the local inference submodule to obtain the updated parameters. For example, if the global model update parameters are for the weight parameters of a fully connected layer in the inference submodule, the user's local device adds the updated parameters to the current weight parameters to obtain the new weight parameters.

[0101] In addition to updating the inference submodule parameters on the user's local device, the system also synchronizes the global model update parameters to the user's nearest edge nodes, updating the parameters of the feature encoding submodule deployed on these edge nodes. Upon receiving the parameters, the edge nodes update their parameters in a similar manner to the user's local device. This ensures synchronized parameter updates between the feature encoding and inference submodules, improving the overall performance of the advertising recommendation model. For example, in an advertising recommendation system based on a CDN-level edge computing network, when the user's local device updates the inference submodule parameters, the system sends the global model update parameters to the edge node closest to the user, and the edge node updates the parameters of the feature encoding submodule.

[0102] After the parameter updates are completed, the user's local device and edge nodes respectively perform performance verification on the updated ad recommendation model. By using a portion of reserved test data, metrics such as accuracy and recall can be calculated to evaluate whether the model's performance has improved. If the model performance does not meet expectations, the reasons can be further analyzed, and the model structure or parameter update strategy can be adjusted for the next round of model updates and optimizations. For example, the user's local device uses the test dataset to test the updated inference submodule, calculating the accuracy and recall of ad recommendations. If the accuracy is lower than a preset threshold, it may be necessary to adjust the learning rate or increase the amount of training data, and update the model again.

[0103] In some embodiments, step S106 above, which involves performing real-time ad matching based on the encoded features returned by the user's neighboring edge nodes and updating the local privacy budget counter in conjunction with the user's interaction with the ads, specifically includes: Obtain the encrypted hidden layer feature vector processed by the feature encoding submodule from the user's neighboring edge nodes; The inference submodule is invoked on the user's local device to perform an ad library retrieval operation based on the encrypted hidden layer feature vector, generating a candidate ad set. Real-time capture of user interaction events with displayed ads, and conversion of interaction events into corresponding privacy consumption values ​​according to a preset privacy consumption mapping table; Privacy consumption is deducted from the local privacy budget counter. When the remaining budget value is lower than the preset budget threshold, the privacy protection mode is automatically triggered to terminate data collection.

[0104] In this embodiment, when a user device accesses the network, the system first identifies the user's current geographical location and determines the nearest edge node. This edge node is equipped with a feature encoding submodule, which pre-processes the advertising content in the advertising library by extracting and encoding features. Specifically, after receiving advertising materials, the edge node extracts multi-dimensional features of the advertisement, including visual and semantic features, using a deep learning model, and converts these features into fixed-dimensional hidden feature vectors. To ensure data transmission security, the edge node uses homomorphic encryption to encrypt the hidden feature vectors, generating encrypted hidden feature vectors. The user device sends a feature request to the nearest edge node via a secure communication protocol (such as TLS / SSL), and the edge node responds to the request and transmits the encrypted hidden feature vector to the user's local device. Throughout this process, the original advertising feature data remains encrypted, and neither the edge node nor the user device can directly obtain the plaintext feature information.

[0105] The user's local device deploys an inference submodule, which includes an ad retrieval engine and a lightweight decryption component. Upon receiving the encrypted hidden feature vector returned by the edge node, the inference submodule first utilizes the operational characteristics of homomorphic encryption to calculate the similarity of the encrypted feature vector without decryption. Specifically, the inference submodule matches the encrypted feature vector with pre-encrypted ad features in the local ad index library, and obtains a similarity score through vector dot product operations supported by homomorphic encryption. Based on the score results, the inference submodule selects the top N ads with the highest similarity, generating a candidate ad set. Throughout this process, all calculations are performed within the encrypted domain, ensuring the privacy of ad features and user queries. After the candidate ad set is generated, the inference submodule further optimizes the ranking of the ads in the set, adjusting the ad display order based on user historical preference data (such as anonymized interest tags stored on the device), ultimately determining the list of ads to be displayed.

[0106] User devices monitor user interactions in real time after ad display. These interactions include, but are not limited to, clicking ads, prolonged viewing of ad content, swiping to skip ads, or ignoring ads. The system assigns different privacy weights to each interaction type. For example, clicking, due to its involvement of expressing user interest, is given a higher privacy weight, while ignoring ads is given a lower weight. Simultaneously, the system maintains a privacy cost mapping table, which establishes a correspondence between interaction types and privacy cost values ​​based on experimental data or expert experience. When a user interacts, the device's event capture module records the behavior type and queries the privacy cost mapping table to obtain the corresponding privacy cost value. For example, if a user clicks an ad, the system determines the corresponding privacy cost value to be 2 based on the mapping table; if the user ignores an ad, the corresponding cost value is 0.5. This design dynamically correlates privacy cost with the intensity of user behavior, avoiding wasted privacy budgets or insufficient protection due to fixed thresholds.

[0107] When a user device is initialized, a total privacy budget is set. This value is determined based on the user's device type, usage scenario, or user-defined preferences, for example, set to 100. Each time a user interaction is captured, the system deducts the corresponding privacy consumption value from the local privacy budget counter. For example, if a user clicks on an ad three times consecutively, each click corresponds to a privacy consumption value of 2, resulting in a cumulative deduction of 6 privacy consumption values, leaving the privacy budget counter with 94 remaining. The system continuously monitors the remaining budget value. When the remaining budget value falls below a preset budget threshold (e.g., 20% of the total budget), it automatically triggers a privacy protection mode. In privacy protection mode, the system immediately stops sending any feature requests to edge nodes, terminating the collection and transmission of ad data. Simultaneously, the local inference submodule only performs a limited number of recommendations based on anonymized ad data cached on the device, without updating the ad index library or user interest model. Furthermore, the system pushes a privacy protection prompt to the user, informing them that the current privacy budget has been exhausted and providing options to restore the budget (e.g., waiting 24 hours for automatic reset or manually adjusting budget settings). Through this mechanism, the system ensures the continuity of ad recommendations while strictly limiting the exposure of user privacy data.

[0108] This embodiment achieves encrypted processing of advertising features and dynamic adjustment of privacy awareness through the division of labor and cooperation between edge nodes and local devices, effectively balancing personalized recommendation needs and user privacy protection goals.

[0109] Furthermore, the steps for creating the privacy consumption mapping table specifically include: By analyzing historical user interaction behavior datasets, a quantitative correlation model between different advertising operation types and privacy leakage risks is established; The advertising content is divided into multiple sensitivity levels, and each sensitivity level corresponds to a preset basic privacy consumption coefficient to obtain the sensitivity level classification results; By combining the user's local device type and current geographical location information, a dynamic adjustment factor is superimposed on the basic privacy consumption coefficient to obtain a dynamic weight configuration; Based on the quantitative correlation model, sensitivity level classification results, and dynamic weight configuration, a privacy consumption mapping table is generated, which is used to define the consumption weight coefficients for different interactive behaviors.

[0110] In this embodiment, a large-scale historical user interaction behavior dataset is collected. This dataset contains records of user actions such as clicking, browsing, skipping, and sharing advertisements, as well as metadata such as device identifier, timestamp, geographic location, and advertisement category for each record. Simultaneously, the privacy leakage risk level of each record is labeled using third-party data sources or differential privacy technology. For example, advertisement interactions involving sensitive fields such as healthcare and finance are marked as high-risk, while ordinary product advertisement interactions are marked as low-risk. Based on the labeled dataset, the system uses machine learning algorithms (such as random forests or gradient boosting trees) to train a quantitative association model. During training, the model uses interaction behavior type (such as click, skip), advertisement category, and operation duration as input features, and privacy leakage risk level as the output label. Iterative optimization learns the mapping relationship between different feature combinations and risk levels. For example, the model might find a strong correlation between "medical advertisement click behavior" and "high privacy leakage risk," while the correlation between "clothing advertisement browsing behavior" and "low privacy leakage risk" is weaker. The final generated quantitative association model can output a privacy risk score for any combination of interaction behaviors, with the score range set, for example, from 0 to 10, where a higher score indicates a greater privacy leakage risk.

[0111] Based on the potential privacy sensitivity of advertising content, advertisements are categorized into multiple levels. Criteria for this categorization include, but are not limited to, the industry sector involved (e.g., healthcare, finance, education, entertainment), whether it contains personally identifiable information (e.g., name, contact information), and whether it involves sensitive topics (e.g., politics, religion). For example, medical advertisements are categorized into the highest sensitivity level (Level 3) due to the potential disclosure of users' health status; financial advertisements are categorized into the second highest level (Level 2) due to their involvement with financial information; and ordinary product advertisements are categorized into the lowest level (Level 1). For each sensitivity level, the system presets a basic privacy cost coefficient, which reflects the baseline privacy cost when a user interacts with an advertisement of that level. For example, the basic coefficient for Level 3 advertisements is set to 3.0, for Level 2 it is 2.0, and for Level 1 it is 1.0. The sensitivity level categorization results and the basic coefficient configuration table are stored in the system backend as a benchmark reference for subsequent dynamic adjustments.

[0112] Building upon the basic privacy consumption coefficient, the system introduces user's local device type and geolocation information as dynamic adjustment factors. Regarding device type, the system identifies the user's device as a smartphone, tablet, or smart wearable device, and assesses its potential for privacy leakage based on attributes such as storage capacity, computing power, and sensor permissions. For example, smartphones, due to their greater number of sensors (such as GPS and cameras) and persistent storage, are assigned a higher device risk coefficient (e.g., 1.2), while smartwatches, due to their limited functionality, are assigned a lower coefficient (e.g., 0.8). Regarding geolocation, the system obtains the user's current location through IP address or GPS positioning and associates this with the strength of privacy protection regulations in that area (e.g., EU GDPR areas, China's Personal Information Protection Law areas) or cultural privacy sensitivity (e.g., areas surrounding religious sites or government agencies). For example, user interactions in GDPR-covered areas are assigned a regulatory compliance coefficient of 0.9, while user interactions in privacy-sensitive areas are assigned an environmental sensitivity coefficient of 1.1. When calculating the dynamic adjustment factor, the system multiplies the device risk coefficient by the geolocation coefficient to obtain a comprehensive dynamic adjustment value. For example, when a smartphone user operates in a GDPR-controlled area, the dynamic adjustment value is 1.2 × 0.9 = 1.08; when a smartwatch user operates in a privacy-sensitive area, the dynamic adjustment value is 0.8 × 1.1 = 0.88.

[0113] The system integrates the privacy risk score, the base coefficient of sensitivity level, and the dynamic adjustment factor output by the quantitative correlation model to generate the final privacy consumption mapping table. The specific calculation logic is as follows: For each combination of interaction behavior type (e.g., click, skip) and ad sensitivity level, the system first queries the quantitative correlation model to obtain the privacy risk score for that combination. Then, the score is multiplied by the product of the base coefficient and the dynamic adjustment factor to obtain the final privacy consumption value. For example, if a user clicks on a level 3 medical ad, and the quantitative model outputs a risk score of 8, the device type is a smartphone (coefficient 1.2), and the geographical location is a normal area (coefficient 1.0), then the privacy consumption value = 8 × 3.0 × 1.2 × 1.0 = 28.8; if the user skips a level 1 clothing ad, and the model outputs a risk score of 2, the device is a smartwatch (coefficient 0.8), and the geographical location is a GDPR area (coefficient 0.9), then the privacy consumption value = 2 × 1.0 × 0.8 × 0.9 = 1.44. The generated privacy consumption map is indexed by interaction type, ad sensitivity level, device type, and geographic location, storing the consumption values ​​corresponding to each combination. It is periodically updated based on newly collected behavioral data and regulatory changes. In practical applications, user devices query the map in real time to obtain privacy consumption values ​​based on current interaction behavior, ad category, device information, and geographic location. These values ​​are then used to update the local privacy budget counter, enabling dynamic adjustment of privacy protection.

[0114] This embodiment uses multi-dimensional data fusion and dynamic weight configuration to enable the privacy cost mapping table to accurately reflect the privacy costs in different scenarios, providing a quantitative basis for privacy budget control in personalized advertising recommendations.

[0115] Reference Figure 2 An embodiment of the present invention provides a cross-domain privacy protection system 2 for advertising recommendations, wherein system 2 specifically includes: The first privacy protection module 201 is used to collect user behavior data and encrypt and store it to the local device when the user visits a first-party website; The second privacy protection module 202 is used to collect user behavior data and encrypt and store it to the local device when the user visits the first-party website; and to encrypt and match the user behavior data with the anonymized device data provided by the security hardware manufacturer through privacy intersection technology, and generate an initial cross-domain joint feature vector under the premise that the data of each party does not leave the domain. The third privacy protection module 203 is used to collect user behavior data and encrypt and store it to the local device when the user visits a first-party website; dynamically adjust the Laplace noise intensity according to the user's preset sensitivity level, generate a perturbation feature vector that satisfies differential privacy, and fuse the initial cross-domain joint feature vector and the perturbation feature vector to form a target cross-domain joint feature vector. The fourth privacy protection module 204 is used to collect user behavior data and encrypt and store it on the local device when the user visits a first-party website; the pre-trained advertising recommendation model is split into a feature encoding sub-module and an inference sub-module, the feature encoding sub-module is deployed to the edge nodes near the user through a CDN-level edge computing network, and the inference sub-module is kept on the user's local device; The fifth privacy protection module 205 is used to collect user behavior data and encrypt and store it on the local device when the user visits a first-party website; based on the target cross-domain joint feature vector, it performs feature encoding submodule calculation and inference submodule calculation, and uploads the encrypted hidden layer feature vector to the federated learning aggregation server for global model update; The sixth privacy protection module 206 is used to collect user behavior data and encrypt and store it to the local device when the user visits a first-party website; perform real-time ad matching based on the encoded features returned by the user's neighboring edge nodes, and update the local privacy budget counter based on the user's interaction with the ads.

[0116] It is understandable that, such as Figure 1 The content of the cross-domain privacy protection method embodiments for advertising recommendations shown herein is applicable to the cross-domain privacy protection system embodiments for advertising recommendations. The specific functions implemented by the cross-domain privacy protection system embodiments for advertising recommendations are as follows: Figure 1 The cross-domain privacy protection method for advertising recommendations shown is the same as that implemented in this example, and the beneficial effects achieved are the same as those described above. Figure 1 The beneficial effects achieved by the cross-domain privacy protection method embodiment shown in the advertisement recommendation are also the same.

[0117] It should be noted that the information interaction and execution process between the above systems are based on the same concept as the method embodiments of the present invention. For details on their specific functions and technical effects, please refer to the method embodiments section, which will not be repeated here.

[0118] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0119] Reference Figure 3 The present invention also provides a computer device 3, including: a memory 302 and a processor 301, and a computer program 303 stored on the memory 302. When the computer program 303 is executed on the processor 301, it implements the cross-domain privacy protection method for advertising recommendations as described in any of the above methods.

[0120] The computer device 3 may be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device 3 may include, but is not limited to, a processor 301 and a memory 302. Those skilled in the art will understand that... Figure 3 The computer device 3 is merely an example and does not constitute a limitation on the computer device 3. It may include more or fewer components than shown in the figure, or combine certain components, or different components, such as input / output devices, network access devices, etc.

[0121] The processor 301 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0122] In some embodiments, the memory 302 may be an internal storage unit of the computer device 3, such as a hard disk or memory of the computer device 3. In other embodiments, the memory 302 may be an external storage device of the computer device 3, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 3. Furthermore, the memory 302 may include both internal and external storage units of the computer device 3. The memory 302 is used to store the operating system, applications, boot loader, data, and other programs, such as the program code of the computer program. The memory 302 can also be used to temporarily store data that has been output or will be output.

[0123] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a cross-domain privacy protection method for advertising recommendations as described in any of the above methods.

[0124] In this embodiment, if the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0125] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0126] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0127] In the embodiments disclosed in this application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0128] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

Claims

1. A method for cross-domain privacy protection in advertising recommendations, characterized in that, The method specifically includes: When a user visits a first-party website, user behavior data is collected and encrypted and stored on the local device. By using privacy intersection technology, user behavior data is encrypted and matched with anonymized device data provided by security hardware manufacturers, and an initial cross-domain joint feature vector is generated without leaving the domain of the data from each party. The intensity of the Laplace noise is dynamically adjusted according to the user's preset sensitivity level to generate a perturbation feature vector that satisfies differential privacy. The initial cross-domain joint feature vector and the perturbation feature vector are then fused to form the target cross-domain joint feature vector. The pre-trained ad recommendation model is split into a feature encoding sub-module and an inference sub-module. The feature encoding sub-module is deployed to the edge nodes near the user through a CDN-level edge computing network, while the inference sub-module is kept on the user's local device. Based on the target cross-domain joint feature vector, the feature encoding submodule and the inference submodule are calculated, and the encrypted hidden layer feature vector is uploaded to the federated learning aggregation server for global model update. Real-time ad matching is performed based on the encoded features returned by the user's neighboring edge nodes, and the local privacy budget counter is updated in conjunction with the user's interaction with the ads.

2. The method according to claim 1, characterized in that, The method involves using privacy-preserving intersection techniques to encrypt and match user behavior data with anonymized device data provided by security hardware manufacturers. This generates an initial cross-domain joint feature vector while ensuring that the data from each party remains within its own domain. Specifically, this includes: On the first-party website client, based on user behavior data, encryption is performed using a first public key encryption algorithm to generate a first set of encrypted identifiers; On the security hardware vendor's server, the data from the anonymization device is encrypted using a second public key encryption algorithm to generate a second set of encrypted identifiers; The first encrypted identifier set is sent to the security hardware vendor's server, where a blinding factor is applied and the first encrypted identifier set is encrypted a second time to generate a blinded encrypted identifier set. The blinded encrypted identifier set and the second encrypted identifier set are matched using a homomorphic operation to generate the intersection identifier index in the encrypted state; Based on the intersection identifier index, the first-party website client and the security hardware manufacturer's server exchange homomorphic encrypted device feature data through a secure channel. The encrypted device feature data and user behavior data are then concatenated to generate an initial cross-domain joint feature vector.

3. The method according to claim 1, characterized in that, The process of dynamically adjusting the Laplace noise intensity based on a user-preset sensitivity level to generate a perturbation feature vector that satisfies differential privacy, and fusing the initial cross-domain joint feature vector and the perturbation feature vector to form a target cross-domain joint feature vector, specifically includes: Based on the data sensitivity level preset by the user in the visual configuration interface, the corresponding noise intensity coefficient is generated through the local privacy management module. Based on the noise intensity coefficient, a Laplace distribution noise sequence is generated on the user's local device, and the Laplace distribution noise sequence is superimposed element-wise with the initial cross-domain joint feature vector to generate a perturbation feature vector. The initial cross-domain joint feature vector and the perturbation feature vector are weighted and fused using a preset privacy-preserving weight allocator to generate the target cross-domain joint feature vector.

4. The method according to claim 1, characterized in that, The process of splitting the pre-trained ad recommendation model into a feature encoding submodule and an inference submodule, deploying the feature encoding submodule to edge nodes near the user via a CDN-level edge computing network, and keeping the inference submodule on the user's local device, specifically includes: The pre-trained ad recommendation model is divided into a feature encoding submodule and an inference submodule according to the calculation function. The feature encoding submodule is responsible for the transformation of the original features into latent vectors, and the inference submodule is responsible for the mapping of latent vectors to ad matching results. By using the node selection algorithm of the CDN-level edge computing network, the feature encoding submodule is deployed to the edge node closest to the user's local device, while the inference submodule is retained on the user's local device. When a user’s local device initiates a recommendation request, the feature encoding submodule is dynamically downloaded from the edge node. The feature encoding submodule is then executed on the user’s local device to calculate the user’s feature data and obtain the first hidden layer feature vector. The first hidden layer feature vector is lightly encrypted and transmitted to the local inference submodule. The local inference submodule generates personalized ad matching results and displays recommended ads on the user's local device interface.

5. The method according to claim 1, characterized in that, The process involves calculating the feature encoding submodule and the inference submodule based on the target cross-domain joint feature vector, and then uploading the encrypted hidden layer feature vector to the federated learning aggregation server for global model updates. Specifically, this includes: Based on the target cross-domain joint feature vector, the feature encoding submodule is executed on the user's local device to generate the initial hidden layer feature vector; The initial hidden layer feature vector is input into the local inference submodule for forward computation. The local model gradient is generated through the backpropagation algorithm, and gradient clipping and normalization are performed on the local model gradient to obtain the processed gradient data. The processed gradient data is encrypted using an additive homomorphic encryption algorithm to obtain encrypted gradient data, which is then uploaded to the federated learning aggregation server. The federated learning aggregation server collects encrypted gradient data from multiple participants, performs weighted average aggregation operations through a secure multi-party computation protocol, and generates global model update parameters. The global model update parameters are distributed to each participant, and the parameters of the inference submodule and feature encoding submodule are updated on the user's local device and the user's neighboring edge nodes, respectively.

6. The method according to claim 1, characterized in that, The step of performing real-time ad matching based on the encoded features returned by the user's neighboring edge nodes, and updating the local privacy budget counter in conjunction with the user's interaction with the ads, specifically includes: Obtain the encrypted hidden layer feature vector processed by the feature encoding submodule from the user's neighboring edge nodes; The inference submodule is invoked on the user's local device to perform an ad library retrieval operation based on the encrypted hidden layer feature vector, generating a candidate ad set. Real-time capture of user interaction events with displayed ads, and conversion of interaction events into corresponding privacy consumption values ​​according to a preset privacy consumption mapping table; Privacy consumption is deducted from the local privacy budget counter. When the remaining budget value is lower than the preset budget threshold, the privacy protection mode is automatically triggered to terminate data collection.

7. The method according to claim 6, characterized in that, The steps for creating the privacy consumption mapping table specifically include: By analyzing historical user interaction behavior datasets, a quantitative correlation model between different advertising operation types and privacy leakage risks was established. The advertising content is divided into multiple sensitivity levels, and each sensitivity level corresponds to a preset basic privacy consumption coefficient to obtain the sensitivity level classification results; By combining the user's local device type and current geographical location information, a dynamic adjustment factor is superimposed on the basic privacy consumption coefficient to obtain a dynamic weight configuration; Based on the quantitative correlation model, sensitivity level classification results, and dynamic weight configuration, a privacy consumption mapping table is generated, which is used to define the consumption weight coefficients for different interactive behaviors.

8. A cross-domain privacy protection system for advertising recommendations, characterized in that, The system specifically includes: The first privacy protection module is used to collect user behavior data and encrypt and store it to the local device when the user visits a first-party website; The second privacy protection module is used to collect user behavior data and encrypt and store it on the local device when the user visits the first-party website; through privacy intersection technology, the user behavior data is encrypted and matched with the anonymized device data provided by the security hardware manufacturer, and an initial cross-domain joint feature vector is generated under the premise that the data of each party does not leave the domain. The third privacy protection module is used to collect user behavior data and encrypt and store it to the local device when the user visits a first-party website; dynamically adjust the Laplace noise intensity according to the user's preset sensitivity level, generate a perturbation feature vector that satisfies differential privacy, and fuse the initial cross-domain joint feature vector and the perturbation feature vector to form the target cross-domain joint feature vector. The fourth privacy protection module is used to collect user behavior data and encrypt and store it on the local device when the user visits a first-party website; the pre-trained advertising recommendation model is split into a feature encoding sub-module and an inference sub-module. The feature encoding sub-module is deployed to the edge nodes near the user through a CDN-level edge computing network, while the inference sub-module is kept on the user's local device. The fifth privacy protection module is used to collect user behavior data and encrypt and store it on the local device when the user visits a first-party website; based on the target cross-domain joint feature vector, it performs feature encoding submodule calculation and inference submodule calculation, and uploads the encrypted hidden layer feature vector to the federated learning aggregation server for global model update; The sixth privacy protection module is used to collect user behavior data and encrypt and store it on the local device when the user visits a first-party website; perform real-time ad matching based on the encoded features returned by the user's neighboring edge nodes, and update the local privacy budget counter based on the user's interaction with the ads.

9. A computer device, characterized in that, include: The memory and processor, and the computer program stored in the memory, which, when executed on the processor, implement the cross-domain privacy protection method for advertising recommendations as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, It stores a computer program, which, when executed by a processor, implements the cross-domain privacy protection method for advertising recommendations as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Cross-domain collaborative privacy protection method and system, medium and product

    CN121864464A

  • Sports event sponsorship effect evaluation method, device and equipment and storage medium

    CN121883095A

  • Multivariable time point process analysis method based on two-dimensional local differential privacy

    CN121919913A

  • A federated learning marketing user privacy analysis method and system

    CN122548768A