Cross-platform advertisement joint modeling method driven by privacy calculation
Through the cross-platform advertising joint modeling method driven by privacy computing, the balance problem between privacy protection and data quality in advertising data generation is solved, and the efficient generation and secure application of multimodal data are achieved, which is suitable for advertising analysis in multiple scenarios.
Patent Information
- Application Number
- CN202510885641.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-14
AI Technical Summary
Existing technologies make it difficult to protect user privacy while maintaining the statistical characteristics of data and the ability to generate multimodal data in advertising data generation, resulting in poor effectiveness of advertising data in multi-scenario applications.
A privacy-focused cross-platform advertising joint modeling method is adopted. The original advertising data is desensitized through a pre-established privacy protection mechanism. Generative adversarial networks are used to extract features and learn data distribution patterns. Text and image data are encrypted in layers, a multimodal data generation sub-module is constructed, and cross-modal fusion processing and secondary desensitization are performed to ensure data security and applicability in multiple scenarios.
It achieves the generation of high-quality multimodal advertising data while protecting user privacy, which is suitable for advertising analysis and application in multiple scenarios and improves the availability and security of data.
Smart Images

Figure CN120782467A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of Internet advertising, specifically to a privacy computing driven cross-platform advertising joint modeling method. BACKGROUND
[0002] In the data-driven advertising industry, the balance between protecting user privacy and utilizing data for precise marketing has become a crucial research area. The generation and application of advertising data directly relate to the commercial value of enterprises and the trust of users. How to generate effective data without touching real personal information is a major issue that needs to be addressed. The importance of this field lies in the fact that it not only concerns technological innovation, but also involves social ethics and legal compliance.
[0003] However, existing methods have significant limitations in addressing this issue. Many solutions tend to rely too much on simple anonymization or data desensitization techniques. These methods often fail to preserve the statistical properties of data when faced with complex data analysis needs, resulting in poor performance of the generated advertising data in practical applications. The deeper problem is that existing technologies lack systematic consideration of the dynamic balance between privacy protection and data quality in the data generation process, making it difficult to adapt to diverse application scenarios. First, the implementation of privacy protection requires embedding strong mechanisms in the data generation process, and the design of such mechanisms often weakens the practicality of data due to excessive simplification of real data. Further, this lack of practicality directly affects whether the generated data can play a role in multiple scenarios and forms of advertising needs, such as the collaborative generation of multi-modal data such as text and images. SUMMARY
[0004] The purpose of the present application is to provide a privacy computing driven cross-platform advertising joint modeling method, which can generate high-quality advertising data framework on the basis of protecting privacy and ensure its support for multi-modal data generation and application.
[0005] To achieve the above-mentioned objectives, the present invention provides the following technical solutions: a cross-platform advertising joint modeling method driven by privacy computing, the method comprising performing preliminary desensitization processing on the original advertising data through a pre-established privacy protection mechanism, adopting layered encryption means for the sensitive fields contained therein, obtaining the desensitized data set after the first stage of processing, and obtaining a data set that preliminarily shields sensitive information; based on the desensitized data set after the first stage of processing, using a generative adversarial network model to extract features from the data, performing deep learning on the statistical characteristics of the data, obtaining implicit distribution laws, and determining the generated preliminary feature mapping results; through the preliminary feature mapping results, determining a preliminary synthetic data set that is adapted to the application scenario, and based on the preliminary synthetic data set, In order to meet the needs of data practicality evaluation, the preset evaluation indicators are used to detect the retention of statistical characteristics of the generated data. If the detection result is lower than the preset threshold, the generative adversarial network model is trained a second time to obtain an optimized synthetic data set; through the optimized synthetic data set, the generated text and image data are cross-modally fused to meet the requirements of diverse data forms, and the final multimodal advertising data set is obtained, and the output results suitable for multi-scenario applications are determined; based on the final multimodal advertising data set, a security mechanism module is embedded to detect potential privacy leakage risks. If reversible derivation features are found in the data, the reversible derivation features are desensitized a second time to obtain a final data set with higher security.
[0006] Preferably, the method of determining a preliminary synthetic data set that is suitable for the application scenario through preliminary feature mapping results includes performing feature decomposition on the text and image data respectively according to the requirements of multimodal data support through preliminary feature mapping results, obtaining corresponding multimodal feature vectors, and obtaining classified feature combinations.
[0007] Preferably, the method of determining a preliminary synthetic data set suitable for the application scenario through preliminary feature mapping results includes embedding a dynamic balance design module based on the classified feature combination, adjusting parameters for the conflict between data generation quality and privacy protection mechanism, and iteratively optimizing the model parameters to obtain an adjusted feature set if it is detected that the generated feature vector deviates from a preset statistical characteristic threshold.
[0008] Preferably, the method of determining a preliminary synthetic data set that is suitable for the application scenario through preliminary feature mapping results includes constructing a multimodal data generation submodule based on the adjusted feature set for complex analysis requirements, obtaining generated text and image data samples, and determining a preliminary synthetic data set that is suitable for the application scenario.
[0009] Preferably, the layered encryption method includes using homomorphic encryption for identification fields and using a differential privacy protection strategy for behavioral fields.
[0010] Preferably, the preset evaluation indicators include a distribution distance indicator, a feature diversity indicator, and an actual advertisement click-through rate simulation indicator.
[0011] Preferably, the cross-modal fusion processing includes aligning the generated text description with the image content using a shared semantic space, and enhancing cross-modal semantic consistency through an attention mechanism.
[0012] Preferably, if the detection result is lower than a preset threshold, the specific formula for performing secondary training on the generative adversarial network model is: in, represents the total loss function, Represents resistance to loss, represents the loss of statistical properties, represents the loss of diversity, represents the regularization term, λ1 represents the statistical feature loss weighting coefficient, λ2 represents the diversity loss weighting coefficient, and λ3 represents the regularization loss weighting coefficient.
[0013] Preferably, the specific calculation formula of the statistical characteristic loss weighting coefficient λ1 is:
[0014] Among them, λ1 represents the weighted coefficient of statistical feature loss, C represents the number of category coverage of synthetic data, and L represents the level of statistical distribution deviation.
[0015] Preferably, the specific calculation formula of the diversity loss weighted coefficient λ2 is:
[0016] Among them, λ2 represents the diversity loss weighting coefficient, D represents the feature distribution dispersion, and R represents the model output repetition rate; λ3 = 1-λ1-λ2; among them, λ3 represents the regularization loss weighting coefficient.
[0017] It can be seen from the above technical solution that the present invention has the following beneficial effects:
[0018] This privacy computing-driven cross-platform advertising joint modeling method first desensitizes the original advertising data, and then uses a generative adversarial network to extract features and learn data distribution patterns. The text and image data are then feature decomposed separately to obtain multimodal feature vectors. Parameters are adjusted through a dynamic balance design module to resolve the conflict between data generation quality and privacy protection. A multimodal data generation submodule is then constructed to generate a synthetic data set that meets the application scenario. Finally, cross-modal fusion processing is performed, and a security mechanism is embedded for secondary desensitization to obtain a more secure multimodal advertising data set. The present invention realizes privacy protection and multimodal synthesis of advertising data, and can be widely used in advertising analysis and multi-scenario applications that need to protect user privacy. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 Flow chart of the method of the present invention. DETAILED DESCRIPTION
[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0021] like Figure 1 As shown, the present invention provides a technical solution: a cross-platform advertising joint modeling method driven by privacy computing, the method comprising: performing preliminary desensitization processing on the original advertising data through a pre-established privacy protection mechanism, adopting a layered encryption method for the sensitive fields contained therein, obtaining a desensitized data set after the first stage of processing, and obtaining a data set that preliminarily shields sensitive information;
[0022] In this technical solution, raw advertising data typically originates from multiple platforms, encompassing both structured and unstructured fields. To ensure data security preprocessing prior to cross-platform joint modeling, a reusable privacy protection mechanism adaptable to diverse data sources is required. This mechanism consists of three components: a field sensitivity assessment system, encryption policy matching rules, and a key management and distribution module. The first step involves classifying field sensitivity levels. First, a field sensitivity grading criteria is established, with three levels defined based on the data's ability to identify users. Level 1 sensitive fields include data that directly identifies individuals, such as user IDs, national ID numbers, and mobile phone numbers. Level 2 sensitive fields include data that indirectly links to user identities, such as device IDs, IP addresses, location information, and operating versions. Level 3 sensitive fields include behavioral data that has no direct or indirect link to an individual but still requires protection, such as ad click history, ad display time, and ad category tags. Field classification is based on five factors: field name, field value examples, field value uniqueness analysis, correlation analysis with user IDs, and platform classification guidelines. Each factor is assigned a fixed weight, for a total of 100 points. Fields with a combined score greater than 80 are classified as Level 1 fields; those between 50 and 80 are classified as Level 2 fields; and those below 50 are classified as Level 3 fields. The field name recognition score is weighted at 30 points, based on the direct role of field names in identifying field sensitivity. In the sensitivity scoring system, field names are the first visible and most intuitive identifier. The semantics they convey often quickly reflect the data type and purpose of the field. Especially when the field is not masked or replaced, the name is highly recognizable. Therefore, field name recognition is designated as the primary factor in the scoring mechanism and given a maximum weight of 30 points, accounting for 30% of the total score. This weighting is based on two criteria: first, statistical analysis of the consistency between field names and field sensitivity levels in historical masking projects showed that field names correctly indicate sensitivity levels with an accuracy rate exceeding 85%; second, field names can be identified without parsing sample values, facilitating rapid screening during the initial processing phase. Therefore, field name recognition has a higher a priori value and processing priority than other scoring factors. Therefore, field name recognition has the highest weight among the five scoring factors and is set to 30 points. The field name recognition score is determined by analyzing whether the field name contains privacy-related keywords. A set of highly sensitive keywords, including "id", "user", "phone", "device", "ip", etc., is preset as a standard vocabulary for name matching. If the field name completely matches the terms in the vocabulary, for example, the field name is "user_id" or "device_id", it will be assigned a full score of 30 points; if the field name is a partial match, such as "phone_hash" or "ip_code", it will be assigned 20 points; if the field name does not contain any sensitive keywords, the score is 0 points.This method relies on standard field naming and has high recognition accuracy in well-structured data tables. The field value sample recognition score is assessed by performing semantic analysis and format recognition on the actual data content in the field. The first 1,000 non-null sample values in the field are randomly selected and tested for conformance to standard formats for common sensitive information, such as 11-digit mobile phone numbers, 18-digit ID numbers, 15-digit device numbers, or typical IP addresses. If more than 90% of the values meet a specific sensitive format, the field is assigned a score of 25; if the match rate is between 50% and 90%, the field is assigned a score of 15. If the sample data lacks structure, is natural language text, or is loosely formatted, it is considered to have no significant sensitivity and is given a score of 0. This method ensures that even if a field is poorly named, its sensitivity can still be identified based on its actual data characteristics. The field value uniqueness score reflects the strength of the field's recognition within the overall data. It calculates the percentage of unique values in the field across the entire dataset. For example, if a field has 95,000 unique values in 100,000 records, the unique value percentage is 95%. If the field's unique value ratio is greater than or equal to 95%, it is scored a full 20 points; if the unique value ratio is between 70% and 95%, it is scored 10 points; if the unique value ratio is less than 70%, the field's uniqueness is low and it is scored 0 points. Fields with high uniqueness generally have stronger user identification capabilities, so this scoring item is particularly critical for identifying private fields.
[0023] The correlation score between the field and the user identification field is based on the statistical association between the fields. A predefined user identification field (such as user_id) is used, and the frequent item set mining and association rule analysis algorithm is used to calculate the confidence index between the target field and the user identification field. If the confidence index exceeds 0.8, the target field is determined to be strongly associated with the user identification and the score is 15 points; if the confidence is between 0.5 and 0.8, the score is 10 points; if the confidence is lower than 0.5, the field correlation is not significant and the score is 0 points. This method can reveal hidden fields that are highly coupled to user information even if the field name and example are not obviously sensitive. The platform classification guide score is assigned in combination with the field classification document provided by the business party. The platform can classify its data fields into three levels in advance: high sensitivity, medium sensitivity, and non-sensitive, and upload them to through the classification rule configuration. After identifying the field name, consult the platform guidance document and directly assign points according to the guidance classification: if it is marked as a high-sensitivity field by the platform, it will get 10 points; if it is marked as a medium-sensitivity field, it will get 5 points; if it does not appear in the guidance or is marked as a non-sensitive field, it will get 0 points. This method improves the business adaptability of the scoring model and avoids missing sensitive data types customized by the enterprise. Encryption strategy and parameter setting basis, according to the field classification results, automatically match different encryption schemes: Level 1 field: highly sensitive, using integer homomorphic encryption, encryption key length: set to 1024 bits, generated by a secure entropy source, number of encryption rounds: set to 5 rounds, taking into account both encryption strength and computing performance, field expansion factor: the length of the encrypted ciphertext is approximately 8 times that of the plaintext, and a fixed field padding mechanism is used to maintain the data structure, operation compatibility: supports addition and multiplication operations in the encrypted state, which is used for subsequent distributed training tasks. Secondary fields: Medium sensitivity, using advanced symmetric encryption. Encryption algorithm: Standard symmetric encryption. Key length: Fixed to 256 bits. Initialization vector length: Set to 128 bits to prevent pattern recognition attacks. Block mode: Electronic codebook block encryption mode. Distribution cycle: Automatic key replacement every 24 hours. Third-level fields: Low sensitivity, using hash desensitization. Hash algorithm: Use the national standard hash algorithm, outputting a 256-bit fixed-length digest. Salting method: Introducing a 32-bit salt value, the salt value is generated by concatenating the "platform identification code" and the "current timestamp." Collision protection: Uniqueness is achieved by changing the salt value strategy to avoid hash collisions. Key generation and management process, key generation and distribution are completed by a dedicated module, the process is as follows: Entropy source reading: obtain a 128-bit random seed from the operating hardware random number source; Seed expansion: concatenate the random seed with the current time (accurate to milliseconds) and perform multiple rounds of expansion to generate a 1024-bit master key; Key numbering rule: The key is numbered in the format of "platform identification code + timestamp" to ensure global uniqueness; Distribution method: The key is pushed to each platform security module through an encrypted channel and is only accessible to encryption nodes, and no other access rights are allowed; Key expiration policy: The default validity period of the key is 24 hours, and it will be automatically cancelled and replaced after expiration.Data processing and output flow: Data processing follows these steps: read raw data, verify the number of fields and data format; call the field identification module and output the field classification results; match the encryption strategy field by field; pull the latest key from the key module; perform encryption operations: first-level fields undergo five rounds of encryption, replacing the original data with each round; second-level fields undergo group symmetric encryption, with each round being encrypted every 128 bits; third-level fields are salted and hashed, with the salt value generated independently for each record; complete data integrity check to confirm that no fields are missing or misplaced; generate the first-stage desensitized dataset, with the data structure unchanged and only the field values masked; and write the data to the cache for joint modeling. Computing resources and efficiency control: To ensure performance, each encryption node supports parallel processing, processing 5,000 records per second, with an average latency of less than 200 milliseconds per node. The data batch size is set to 100,000 records per batch, and the desensitization process completes in approximately 20 seconds. This step effectively shields sensitive fields from the original ad data by building a structured privacy protection mechanism and performing layered encryption. This protects user privacy while preserving the data's basic structure and statistical characteristics, allowing the data to be used for subsequent joint modeling without revealing user identities, thereby improving data availability, security, and compliance. Based on the desensitized dataset processed in the first phase, a generative adversarial network model is used to extract features from the data. Deep learning is performed on the data's statistical characteristics to obtain implicit distribution patterns and determine the resulting preliminary feature mapping results.
[0024] This step uses a generative adversarial network to perform deep modeling on the desensitized data set after the first stage of processing. Its core purpose is to extract the potential statistical structure information in the data and construct a unified high-dimensional feature map. The generative adversarial network used contains two sub-networks, namely the generator and the discriminator, which are continuously optimized through the adversarial training mechanism. The input of the generator is a random vector of length 128, which comes from the standard normal distribution sampling and represents the latent variable of the potential feature space. The generator has a total of 5 layers of fully connected networks, with the number of nodes being 256, 512, 1024, 2048 and the output dimension being consistent with the number of data fields. The specific output dimension is dynamically adjusted according to the number of fields in the desensitized data. Each layer uses linear transformation with the rectified linear unit activation function for nonlinear modeling, and finally outputs a floating-point vector to represent the generated sample features. The discriminator's inputs are real samples and pseudo samples output by the generator. Its structure is a symmetrical five-layer neural network with 2048, 1024, 512, 256, and 64 nodes, respectively. The output is a probability value between 0 and 1, used to determine whether the sample is real data. The discriminant result is output using a logistic regression activation function. The input vector length is set to 128, based on the finding from principal component analysis that over 95% of the data variance can be represented by the first 128 principal components. The optimization algorithm uses the adaptive moment estimation method, which automatically adjusts the learning rate, improving training convergence speed and accuracy. The learning rate is initially set to 0.0002 and is reduced to 90% of its original value every 200 training epochs to prevent overfitting in later training stages. The batch size is set to 64 samples to balance computational efficiency and training stability. Both the generator and the discriminator use cross entropy as the loss function. By optimizing the generator's deceptive ability and the discriminator's discriminative ability, the two networks are pushed to gradually approach the equilibrium state in the confrontation until the generator can output high-quality feature samples that are difficult to distinguish from real data.
[0025] The optimization algorithm uses the adaptive moment estimation method, which dynamically tracks and adaptively controls the gradient trends during training. During each training round, the gradient of each parameter to be updated is calculated, and the weighted moving averages of its first-order and second-order moments are recorded. The first-order moment represents the expected direction of the gradient, while the second-order moment reflects the magnitude of the gradient fluctuation. To eliminate bias in the initial estimation, these two values are corrected. Subsequently, these two metrics are used to assign a dynamically adjusted learning rate to each parameter, ensuring that parameters with large fluctuations receive smaller step sizes and stable parameters receive larger update steps, thereby improving the overall convergence speed and robustness of the model. This optimization method uses a first-order moment decay rate of 0.9 and a second-order moment decay rate of 0.999, standard values widely validated in generative adversarial network architectures. These values effectively adapt to different gradient scales and improve training efficiency. Both the generator and discriminator use cross-entropy as the loss function, implementing a discriminant mechanism based on binary probabilistic adversarial learning. During the discriminator training phase, real samples are labeled as label 1, and forged samples generated by the generator are labeled as label 0. The cross-entropy loss between the discriminator output probability and the actual label is calculated, and the discriminator parameters are updated through backpropagation to improve its ability to distinguish between real and forged samples. During the generator training phase, the discriminator parameters are frozen and not updated. Forged samples continue to be input into the discriminator, and the target label is set to 1, that is, it is expected that the discriminator will mistakenly believe them to be real samples. By minimizing the cross-entropy loss of forged samples that are judged as fake samples, the generator gradually learns the data distribution characteristics and thus generates more realistic samples. This alternating training mode of the generator and discriminator prompts the two networks to gradually reach a balance in the confrontation, making the feature results output by the final generator difficult for the discriminator to accurately distinguish, thereby achieving the goal of improving the authenticity of the samples and the quality of feature expression.
[0026] The entire training process consists of four steps: sample preparation, data generation, adversarial training, and model updating. Each training round begins by randomly selecting 64 samples from the desensitized dataset as real inputs. Then, 64 random vectors sampled from a standard normal distribution are fed into the generator to generate fake samples. A mixture of real and fake samples is fed into the discriminator, where the discriminator parameters are updated by calculating the accuracy and backpropagating the error. The discriminator is then frozen, and the generator is optimized backwards to increase the probability that its generated samples will be misclassified as real by the discriminator. This process continues for 2000 rounds, with the learning rate dynamically adjusted every 200 rounds. Intermediate models are saved to monitor the training status and select the best model as the final output. In each training round, the discriminator receives two inputs: 64 real samples sampled from the desensitized dataset and 64 fake samples generated by the generator. Real samples are assigned a label of 1, and fake samples are assigned a label of 0, forming a label vector of length 128. At the same time, the discriminator performs forward inference on these 128 input samples, outputting a probability value for each sample being "real," resulting in a predicted probability vector of length 128. Discrimination accuracy is calculated by comparing each predicted probability value with its corresponding label. A correct discrimination is considered achieved if the predicted probability for a real sample is greater than or equal to 0.5 and the label is 1, or if the predicted probability for a forged sample is less than 0.5 and the label is 0. The number of correctly classified samples is counted and divided by the total number of samples, 128, to obtain the discrimination accuracy for that round of training. For example, if 116 samples are correctly classified, the accuracy is 116 divided by 128, or approximately 90.6%. Error is calculated using the cross-entropy loss function. For each sample, a cross-entropy loss is calculated based on its label value and predicted probability. The loss for real samples is the negative logarithm of the predicted probability, while the loss for forged samples is the negative logarithm of (1 minus the predicted probability). The loss values for all samples are then averaged to obtain the overall loss for that round of training. Finally, a backpropagation operation is performed based on the loss value. The partial derivatives of the loss with respect to the parameters of each layer of the discriminator are calculated by the chain rule, and then the weights and bias terms are updated. The update method combines the previously set optimizer (such as the adaptive moment estimation method) to perform parameter iteration, adjust the direction and amplitude, and gradually enhance the discriminator's ability to distinguish between true and false samples. After training, the floating-point vector output by the generator is the preliminary feature mapping result, which can be used as the input for subsequent modeling tasks. This mapping retains the structure and correlation between the data, while eliminating sensitive information that may exist in the original field, thereby achieving the multiple goals of feature unification, privacy protection, and data availability. The final generator model is saved to the model service platform for subsequent cross-platform joint modeling calls, to achieve standardized representation and consistent processing of desensitized data, and to improve the data fusion capability and generalization performance of the entire modeling.
[0027] Based on the preliminary feature mapping results, and in response to the need for multimodal data support, the text and image data are separately decomposed to obtain corresponding multimodal feature vectors, resulting in a classified feature combination. After completing the preliminary feature mapping, to achieve unified modeling of the multimodal content information contained in the advertising data, this step utilizes a parallel feature decomposition mechanism to extract semantic and visual features from the text and image data, respectively. The feature vectors of the two modalities are then combined to generate a unified multimodal representation for subsequent classification and model input. The text feature decomposition implementation process and parameter settings are as follows: Text data types include natural language text such as user comments, ad headlines, keyword tags, and product descriptions. First, a tokenizer is used to segment the original sentences into word sequences at a granularity of 15 to 30 words per text entry. After token segmentation, a pretrained Chinese word embedding model is used to map each word to a vector of length 300, forming a two-dimensional tensor with a dimension equal to the number of words multiplied by 300. This vector matrix is then input into a bidirectional gated recurrent neural network for feature modeling. The neural network has two layers, each containing 256 hidden units. The bidirectional architecture models the text in both forward and reverse order at each layer, capturing both forward and backward context. After processing each piece of text, it outputs hidden states in both directions. These outputs are average-pooled, ultimately resulting in a semantic feature vector of length 256. This length was determined after balancing expressiveness with computational efficiency, ensuring that it covers key semantic features found in everyday advertising, such as sentiment, product intent, and user intent. The image feature decomposition process involves image data, including the main image of the ad, product images, brand logos, and user-uploaded display images. First, the image input is resized to a standard three-dimensional tensor of 224 pixels in height, 224 pixels in width, and 3 channels. After the image input, it enters the image convolution processing pipeline, employing a classic convolutional neural network architecture. This convolutional network consists of five convolutional layers, each with a kernel size of 3 pixels by 3 pixels, a stride of 1, and 1-pixel padding to maintain consistent input and output dimensions. Each convolutional layer is followed by a batch normalization layer and an activation function using a rectified linear unit. The number of channels is 64, 128, 256, 512, and 512, respectively. After the last convolutional layer, a global average pooling operation is performed to compress the three-dimensional tensor into a one-dimensional vector of length 512. This feature vector can accurately capture structured visual information such as the main contour, edge details, and color distribution in the image, and is suitable for describing product types, styles, brand differences, etc. The multimodal feature combination strategy concatenates the text feature vector (length 256) and the image feature vector (length 512) to obtain a preliminary multimodal feature combination vector of length 768. The concatenation order is text first and then image to ensure dimensional consistency and semantic controllability. The vector is then input into a fully connected network for unified mapping, and the fully connected layer outputs a vector result of length 128, which is the final unified multimodal feature vector.The fully connected layer is followed by a dropout layer with a dropout rate of 0.5 to enhance the model's robustness and generalization capabilities. After feature combination, the 768-dimensional vector is divided into three feature subsets based on modality source: the first 256 dimensions are text semantic features, the middle 512 dimensions are image visual features, and the last 128 dimensions are fused expression features. During the training phase, modality labels are added to each of the three subsets to supervise the learning process, and the minimum-maximum scaling method is used to compress all feature values to between 0 and 1 to improve the stability and comparability of feature expression. Classification labeling and output format: The fused feature vector is fed into a multi-task classification network as input, which outputs label values such as ad category, intent label, and target user group. The classification network contains two parallel output branches, corresponding to ad type and user preference label prediction, respectively. Each branch is a two-layer fully connected network with outputs set to 10 and 8 categories, respectively. The cross-entropy loss function is used for optimization. Ultimately, this step outputs a set of multimodal combined feature vectors with classification labels, which have uniform length and standardized processing and are suitable for use in subsequent modeling or model inference modules, ensuring that the model has good compatibility and discrimination capabilities for multimodal information.
[0028] Based on the classified feature combination, a dynamic balance design module is embedded to adjust parameters to address the conflict between data generation quality and privacy protection mechanism. If it is detected that the generated feature vector deviates from the preset statistical characteristic threshold, the model parameters are iteratively optimized to obtain the adjusted feature set.
[0029] The step introduces a dynamic balance design module based on the classified feature combination, aiming to coordinate the contradiction between data generation quality and privacy protection. Through statistical characteristic deviation analysis and privacy reconstruction risk assessment, feature deviation and privacy leakage signals are dynamically identified, and parameter adjustment strategies are triggered accordingly, finally realizing the adaptive optimization and update of the feature generation model, and outputting a high-quality feature set that meets the constraint conditions. The statistical characteristic deviation detection calculation process performs real-time statistical analysis on the feature set generated in each round, calculates its distribution characteristics in multiple dimensions, including mean, variance, skewness and kurtosis. The calculation of each index is based on all dimensions of the feature vector, and the process is as follows: mean calculation: calculate the average value of all dimensions in each feature vector, and then take the average of the result for the whole batch to get the batch mean; variance calculation: calculate the deviation square of each dimension from the mean, and take the average to get the variance, which reflects the dispersion degree of samples; skewness calculation: use the third central moment divided by the cube of standard deviation to measure the symmetry of distribution; kurtosis calculation: use the fourth central moment divided by the fourth power of standard deviation to judge the sharpness of the distribution. The above indexes are referred to the statistical baseline in the initial training stage as the reference standard. The following threshold judgment conditions are set: if the relative deviation of the current batch feature mean exceeds 5%, that is, the absolute difference between the mean of the generated feature and the reference value exceeds 5% of the reference value, it is considered as deviation; if the variance deviation exceeds 10%, that is, the variance value exceeds the range of plus or minus 10% of the reference value; if the skewness changes more than plus or minus 0.5, or the kurtosis changes more than plus or minus 0.5, it is also considered as abnormal distribution structure. Once any index triggers the deviation condition, the parameter adjustment module is started. Privacy protection evaluation mechanism, the privacy protection mechanism evaluates the reversibility of the generated features through the feature reconstruction attack model, that is, whether the original sensitive field can be restored based on the generated feature vector: a set of sensitive fields are preset as attack targets, such as user identification, device ID, etc.; use the auxiliary trained reconstruction neural network model to input the feature vector into the model, and the output is the predicted value of the sensitive field; statistical accuracy between predicted results and true fields, if the accuracy exceeds 20%, it means that the current feature has reversibility, and the privacy leakage risk is too high. This 20% threshold comes from the public data protection standard, combined with the "inference acceptable error range" in the actual modeling process to determine. The auxiliary reconstruction neural network model is an independent detection mechanism for privacy risk assessment, mainly used to judge whether the generated feature vector contains information that can be used to restore sensitive fields. The model uses a multi-layer feedforward neural network structure, the input is the feature vector output by the generator, usually with a dimension of 128 or an adjusted length, and the output is the predicted result of the target sensitive field, including user ID, device code or geographic location, etc.The network architecture consists of five layers, with the number of nodes in the input layer equal to the feature dimension. The first to third layers are fully connected hidden layers, with 256, 128, and 64 nodes, respectively. The activation function uses rectified linear units. The output layer is configured based on the prediction target type: for regression tasks, it outputs two continuous values using linear activation; for classification tasks, it outputs a classification probability distribution using normalized activation. The model uses the first-stage desensitized data to construct a training set, with the original sensitive fields as labels and the generated features as input. Training parameters include a learning rate of 0.001, a maximum number of training epochs of 500, and a batch size of 64. Adaptive moment estimation is used for optimization, with early stopping implemented to prevent overfitting. After training, the validation set is evaluated. Mean absolute error is used for regression output, and prediction accuracy is used for classification output. If the reconstruction accuracy of sensitive fields exceeds 20%, it is considered a significant privacy risk, triggering the dynamic balancing module to adjust the generator parameter structure. The model runs in an independent sandbox environment, decoupled from the main model, and does not participate in the formal modeling process. It is automatically activated and executed only after each round of generator training, and outputs a risk assessment report to guide model optimization decisions. Parameter adjustment and optimization strategy, if a statistical offset or privacy leakage warning is triggered, the parameter adjustment strategy is automatically matched according to the risk type: Generator structure adjustment, if the problem is caused by statistical offset, first adjust the generator output dimension, for example, when the original dimension is 128, increase it by 10% and set it to 140, and at the same time adjust the activation function parameters of the middle hidden layer, and reduce the slope of the corrected linear unit from 0.2 to 0.1 to enhance the expression of low-amplitude features; at the same time, adjust the activation function parameters of the middle hidden layer, and reduce the slope of the corrected linear unit from 0.2 to 0.1 to enhance the expression of low-amplitude features; training rhythm and learning rate adjustment Adjustment: If the issue is privacy leakage, the learning rate is reduced from the original setting of 0.0002 to 0.00005 to prevent rapid gradient updates and model overfitting. The training cycle ratio of the generator and discriminator is adjusted from the original 1:1 to 1.5:1, that is, the generator is trained once for every 1.5 rounds of the discriminator, to strengthen privacy suppression capabilities. Regularization and dropout are enhanced. The model regularization coefficient is adjusted from 0.01 to 0.1, introducing a stronger penalty term to control model complexity. A random dropout mechanism is also added to the feature fusion layer, with the dropout rate increased from 0.3 to 0.5. Half of the nodes are randomly dropped in each training round to reduce dependence on sensitive features. After implementing the above adjustments, the model is retrained, with a maximum number of training rounds set to 300. After every 50 rounds of training, statistical indicators and privacy risk double checks are re-performed, and indicator trends are recorded. If the results of three consecutive rounds of testing meet the following conditions: all statistical characteristic deviations are within the set thresholds and the reconstruction accuracy is less than 20%, the optimization is considered effective and the current model passes the verification.The model is solidified and output as an "optimized version", and a final feature set is generated for downstream calls. The final output feature set will be accompanied by a version number, adjustment history, and risk assessment report, and archived in the model library of the modeling platform for subsequent model training, evaluation, or deployment reuse.
[0030] Based on the adjusted feature set, a multimodal data generation submodule is constructed to meet complex analysis requirements. This module generates text and image data samples and determines a preliminary synthetic dataset suitable for the application scenario. Based on a dynamically privacy-optimized feature set, this solution designs and deploys a multimodal data generation submodule to meet the needs for text and image synthetic data in advertising modeling or behavioral analysis tasks. This module automatically generates structured vectors into natural language text and structured images, ultimately outputting a preliminary synthetic dataset that meets the requirements of the business scenario. The overall process consists of five stages: feature preparation, modality branch mapping, sample generation, consistency judgment, and dataset output. The input feature vector is prepared and processed. The input is a 128-dimensional feature vector set derived from the model output of the previous stage and has passed both statistical and privacy screening. First, the vector is normalized, adjusting the value range of each dimension to the range of 0 to 1 to eliminate scale differences between dimensions. The feature vector is then split into two and fed into the text generation branch and the image generation branch, respectively, forming the input tensors for the two modal pathways. The computational process and parameter settings for the text generation path are as follows: A sequential decoder based on gated recurrent units (GRUs) is used for text generation. The input is a 128-dimensional vector, which is projected into a 256-dimensional semantic representation through a linear mapping layer. This vector is then passed as the initial state to a two-layer GRU network. Specific parameters are as follows: each GRU layer contains 256 hidden units; the decoding length is capped at 40 steps, based on the average word count distribution of advertisement descriptions and product descriptions; the decoder word vectors are 300 long and initialized with pretrained word vectors to ensure semantic coherence; each decoding round selects the next word based on the maximum probability, using a 30,000-dimensional word list constructed using a pre-established word frequency dictionary; the decoding process stops when two consecutive stop tokens are generated or when the total length reaches 40. Training uses a cross-entropy loss function to optimize the difference between the decoded result and the reference text. Each round samples 64 feature vector combinations for 1000 rounds, with a learning rate of 0.0002. The computational process and parameter settings in the image generation path use a deconvolutional neural network to convert vectors into images. The 128-dimensional input vector is first mapped to an initial tensor of 64 by 64 by 8, which serves as the starting input for the generation network. This tensor is then upsampled layer by layer through a five-layer deconvolution module, ultimately outputting a 224 by 224 by 3 color image.The specific parameters are as follows: the transposed convolution kernel size is fixed at 4 per layer, with a stride of 2 and an edge padding of 1. The number of channels is set to 128, 64, 32, 16, and 3, respectively. The activation function is rectified linear units (RLUs), except for the output layer, which uses a normalization function to constrain pixel values between 0 and 1. Image quality is evaluated using a joint loss function consisting of mean squared error (MSE) and perceptual loss. The perceptual loss extracts intermediate image features from a pretrained classification network and compares them with a reference image. Each training round uses 32 samples, with a total of 1500 rounds. The optimizer uses the adaptive moment estimation method. In the image generation network, all intermediate layers use RLUs as the activation function. The specific calculation logic is: when the input value is greater than or equal to zero, the value is directly output; when the input value is less than zero, the output value is zero. The non-saturation property of this function helps accelerate training convergence and alleviate the vanishing gradient problem. This structure ensures the sparsity and stability of features across convolution channels. At the output layer, to generate content that conforms to image pixel standards, a normalization function is used to constrain the output values, forcing each pixel channel to be mapped between 0 and 1. Specifically, the output values are normalized to ensure a smooth distribution of pixel values within a certain range, maintaining consistent proportions across channels and avoiding overly bright or dark images. This mechanism ensures that the generated images can be directly used by image viewers or model loaders and are consistent with standard image input formats. Image quality is evaluated using a joint loss function consisting of two components: mean squared error (MSE) and perceptual loss. The MSE measures the pixel-level difference between the generated image and the reference image. For each pair of co-located pixels, the squared difference is calculated, and then the squared differences are averaged across all pixels in the entire image region to produce the final MSE value. A smaller MSE value indicates higher pixel-level similarity. The perceptual loss focuses on the similarity of the images at the semantic level. The calculation process is as follows: the generated image and the reference image are respectively input into a pre-trained image classification network (such as a classifier based on a convolutional neural network), and the feature map of a specified layer in the middle is extracted. Usually, the third or fourth layer in the middle is selected as the feature extraction location to obtain the structural and semantic information of the image. Then, an element-by-element difference operation is performed on the output results of the two images at this feature layer, and the mean of the squared differences is calculated as the perceptual loss. The final total loss value is the weighted sum of the mean squared error and the perceptual loss, where the mean squared error weight is set to 0.7 and the perceptual loss weight is set to 0.3 to ensure that the model takes into account both detail accuracy and overall structural quality when generating images.
[0031] After the text and image samples are generated, the semantic consistency judgment mechanism is executed, which uses the following method: the natural language understanding model is used to extract the key intent label for the text, and the image classification network is used to extract the theme label expressed by the image. If the semantic similarity score of the two labels is higher than the set threshold (set to 0.7), it is determined to be a consistent sample. At the same time, the scene label constraint mechanism is introduced, and the business label embedding (such as “e-commerce”, “game”, “education” and the like) is added in the feature input stage to guide the generated content style to match the specific application scene. The label control signal is merged into the generator input to change the decoding state initial vector direction, thereby affecting the type and structure of the generated content. The synthetic data set generation and structure organization: all samples marked as “semantic consistency” and “scene adaptation qualified” are classified as high-confidence multi-modal pairs and stored in the synthetic data pool. Organize according to the following structure: each sample contains a feature vector number, generated text content, generated image content, label classification result and generation timestamp, while recording the corresponding parameter configuration such as vocabulary version, model version and scene label. Perform quality re-inspection every 500 samples generated, and automatically exclude samples with low quality below the standard. The final output synthetic data set has high semantic consistency, structural integrity and application controllability, and is suitable for multi-modal modeling, content recommendation testing and simulation training and other downstream analysis scenarios.
[0032] According to the preliminary synthetic data set, in order to meet the demand of data practicality evaluation, the preset evaluation index is used to detect the reservation of the statistical characteristics of the generated data. If the detection result is lower than the preset threshold, the generative adversarial network model is trained again to obtain an optimized synthetic data set.
[0033] After the initial synthetic dataset construction is completed, to ensure the representativeness and practicality of the dataset in actual analysis and modeling tasks, a practicality evaluation module is embedded. The module quantifies the retention of the generated data in multiple statistical dimensions with reference to the original de-identified data, and judges whether the generation model needs to be optimized and adjusted accordingly. The statistical property retention evaluation index system includes three main indicators: distribution consistency indicator, category structure matching indicator, and covariance structure reproduction indicator. For each feature dimension, the mean and standard deviation of the original data and the generated data are calculated, and then the absolute difference is divided by the original data value to obtain the mean deviation rate and the standard deviation deviation rate. Assuming that this calculation is performed on all 128 feature dimensions, and then the average of all deviation rates is taken to form a unified distribution difference score. Set the distribution difference score less than fifty percent as the acceptable range. The category structure matching indicator, if the data contains a classification label, then the proportion of each category in the original data and the generated data is counted, the absolute value of the proportion difference is calculated, and the category deviation degree is formed. Set when the value is less than fifteen percent, it means good matching. The covariance structure reproduction indicator calculates the covariance matrix of the two sets of data, and uses the mean of the absolute difference of the main diagonal elements of each dimension as the covariance deviation indicator. The lower this indicator, the better the generated data reproduces the relationships between the features in the original data. Set less than twenty percent as the acceptable range. Each indicator is weighted and summarized according to the weight, with the default setting being fifty percent for distribution consistency, thirty percent for category matching, and twenty percent for covariance structure. The total score ranges from 0 to 1, with a higher score indicating that the generated data is closer to the original structure. Set the comprehensive score less than 0.8 to trigger the model optimization mechanism. The secondary training process of the generation model, if the detection result shows that the statistical property retention does not meet the preset requirements, will start the secondary training program of the generative adversarial network. The process includes the following steps: input construction and sample expansion, under the premise of keeping the original de-identified feature set unchanged, introduce the generated samples of the last round as auxiliary input to enhance the distribution fitting ability of the model. The input vector length is still 128 dimensions to ensure structural consistency. Loss function expansion, based on the original discriminator loss and generator loss, a new statistical deviation loss term is added, which calculates the sum of the mean, variance, and covariance deviations of the generated samples and the original samples. The weight of this loss is 0.3, and the weights of the discriminator and generator losses are 0.4 and 0.3 respectively, and the three are optimized together. Hyperparameter setting and training regulation, the initial learning rate is set to 0.00015, and it is reduced to 90% of the original value every 200 training rounds to avoid oscillation. The total training rounds are 2000, the batch sample size is 64, the optimization algorithm is adaptive moment estimation, and the training is stable. Dynamic adjustment mechanism, statistical property re-evaluation is performed every 100 training rounds, if the evaluation score is higher than 0.85 for three consecutive times, the training is terminated in advance and the model of that round is saved; otherwise, continue iteration until the upper limit of the number of rounds.Output data verification and archiving: Use the optimized model to regenerate synthetic data and perform the same indicator calculations as the preliminary evaluation. If the final score exceeds 0.85 and no single indicator is below the tolerance line (0.6, 0.7, and 0.8 respectively), the dataset is considered a "usable dataset" and written to a compliant data warehouse, and metadata such as the model version, optimization parameters, and evaluation results are recorded.
[0034] This technical approach ensures that synthetic data maintains high privacy compliance while significantly improving its statistical representativeness and analytical accuracy in real-world applications. The optimization process features automated triggering, evaluation-driven, and parameter-adaptive capabilities, meeting the dual requirements of data quality and structural alignment in complex business scenarios.
[0035] Using the optimized synthetic dataset, the generated text and image data are cross-modally fused to meet the diverse data formats required, obtaining a final multimodal advertising data set and determining output results suitable for multi-scenario applications. Based on the optimized synthetic dataset, to address the coexistence of text and images, the two main modal forms of advertising data, a cross-modal fusion processing mechanism is introduced to ensure the coordination and unity of the generated data in terms of semantic integrity and visual consistency, ultimately forming a multimodal advertising data set that can be used in multiple business scenarios. The entire processing process consists of four technical stages, covering feature encoding, modality alignment, fusion modeling, and result screening.
[0036] In the first stage, feature encoding, text data is first standardized through a natural language processing pipeline, including word segmentation, part-of-speech filtering, and stop word removal. It is then vectorized using a word embedding model. The word embedding dimension is set to 128, determined by evaluating semantic coverage and overfitting risk in three types of advertising material: product descriptions, campaign slogans, and user reviews. Image data is input into a convolutional neural network for feature extraction, using a structure containing five convolutional layers. The output feature dimension is set to 128, which maintains the semantic integrity of the image while controlling computational overhead. Image input preprocessing includes scaling to a fixed resolution of 128×128 and grayscale normalization. In the second stage, modality alignment, due to the differences in semantics and representational structure between text and images in the feature space, an alignment module is established to unify their encoding dimensions and distribution. The alignment process consists of two parts: first, the two 128-dimensional vectors are projected into a common embedding space through linear transformation. The projection matrix dimension is 128 times 128, and the parameters are obtained by maximizing the cosine similarity between synonymous sentences and corresponding images; second, batch normalization technology is used to adjust the consistency of mean and standard deviation between modalities. The normalization parameters are obtained by batch statistics on the entire dataset.
[0037] The third stage: Fusion modeling. The aligned text and image vectors are fed into the fusion network. The fusion strategy employs a gated attention mechanism to dynamically determine the contribution weight of each modality in the sample. These weights are estimated using a two-layer perceptron network. The first layer has a width of 64 and uses a rectified linear unit as the activation function. The second layer outputs two values, which are normalized to obtain the weights of the two modalities, controlling the ratio and weighting of the two vectors. The fusion result is a 128-dimensional multimodal vector, which serves as the basis for representing the ad sample. The fourth stage: Output screening. The fused vector is fed into a scenario adaptation discriminant network to identify specific scenarios for which the sample is suitable. This network uses a five-category classification architecture and outputs confidence scores for five application scenarios: e-commerce platforms, information recommendations, social media display, search advertising, and app recommendations. A confidence threshold of 0.6 is set; samples below this threshold are not output. Finally, the identified and classified multimodal samples are stored in a standard dataset with five fields: unique identifier, text content, image path, fusion vector, and scenario label. This processing chain ensures the extraction of effective semantics from multi-source heterogeneous modalities and the construction of a unified expression, providing content resources for advertising that meet business needs and have data consistency. The technical path is highly reproducible, and the parameter settings are supported by logic and data.
[0038] Based on the final multimodal advertising data set, a security mechanism module is embedded to detect potential privacy leakage risks. If reversible features are found in the data, the reversible features are desensitized twice to obtain a final data set with higher security.
[0039] After the multi-modal advertising data set is generated, the final stage of privacy protection process is performed, mainly including 5 steps: feature disintegration, multi-modal joint modeling, reversibility determination, secondary desensitization processing and result re-inspection. Each step adopts quantitative calculation standard and controllable parameter configuration to ensure that the technical scheme has engineering implementation ability. Feature disintegration, each record in the multi-modal advertising data is deconstructed. For the text part, the sentence is divided into word level units by using the word segmentation engine, the top 100 words with the highest frequency are counted and their context information in the corpus is retained to form a word frequency vector. The image part is extracted by the image processing module to form a multi-dimensional vector composed of image coordinate points, color histogram, boundary box area, etc. The vector length is 32. The behavior data includes timestamp, operation type and device identification, which are standardized to form a behavior feature vector with a uniform length of 16. The audio data is 1 frame with 256 sampling points, and the energy spectrum density in the main frequency range is extracted by performing time-frequency domain conversion on 20 consecutive frames to generate a frequency feature vector with a length of 24. Multi-modal joint modeling, after the uniformization of various features, the feature mapping is performed through a multi-modal fusion network. The network structure is a 3-layer perceptron, the input is a composite vector formed by splicing all features, and the total dimension is between 128 and 256. The number of neural units in each layer is 128, 64 and 32 respectively, the activation function is fixed as a rectified linear unit, the training round number is set to 300, the batch size is 128, the learning rate is 0.0002, and the optimizer uses an adaptive gradient descent algorithm. The output is a standard normalized feature vector, which is used for subsequent privacy sensitivity determination. Reversibility determination, the information entropy and mutual information of the above features are evaluated. For each feature dimension, the numerical range is divided into 10 equal width intervals, the frequency of each interval is counted and the information entropy value is calculated. If the information entropy is less than 0.85, the feature is a "high concentration" feature with reversible risk. At the same time, by constructing a joint frequency table, the joint distribution of the feature and the known sensitive label in history is counted, and the mutual information value is calculated according to the joint probability and marginal probability. If the mutual information value is greater than 0.75, the feature is further marked as a "high sensitivity" feature. The features meeting the above 2 conditions will enter the next step of processing. Secondary desensitization processing, different strategies are adopted for accurate desensitization according to the feature type. For high-risk keywords in text features, synonym library is used for replacement, and the replacement items are selected from 100000 word entries to ensure that the language fluency score is not less than 90%. The image pixels located in the center area of the boundary box in the image feature are processed by Gaussian filter with a blur radius of 10. The click sequence in the behavior feature inserts pseudo-behavior with a legal identity but not belonging to the original sequence at a proportion of 15%, and the interference points are randomly selected from the full behavior library to meet the context time sequence structure. If the frequency domain distribution overlap degree of the audio feature is more than 70%, the rhythm contour is retained but the high frequency and low frequency end information is removed, so that it can no longer be inversely deduced to the specific speaker identity by the speech model.
[0040] The results are rechecked and confirmed, and the processed features are re-evaluated for information entropy and mutual information to ensure that all information entropy values are not less than 0.85 and mutual information values are not higher than 0.75. After the evaluation passes, the data set is marked as a "high security data set" and allowed to output into the downstream training, transmission or display stage. Otherwise, return to step 3 for reprocessing until all indicators meet the standards. The basis for determining the lower limit of information entropy of 0.85: In the test, the distribution information entropy of each dimension is calculated for the features of multiple modalities such as text, image and behavior. Using information entropy interval analysis, it is found that when the information entropy is less than 0.85, the value distribution of this dimension generally presents a high concentration phenomenon, and the matching rate with the historical sensitive field is significantly increased, which inversely deduces the risk of identifying or interest category. Therefore, 0.85 is set as the safety lower limit of information entropy as the criterion for determining whether the concentration of the feature set is too high. The basis for determining the upper limit of mutual information of 0.75: In actual processing and analysis, it is found that if the mutual information exceeds 0.75, the feature dimension has strong coupling with the user's real identity or sensitive field, which can significantly improve the success rate of re-identification attack. In the simulation attack test, when the mutual information value is greater than 0.75, the attacker can restore the corresponding user's accuracy rate of more than 80% through known part of the feature. Therefore, the upper limit of mutual information is set to 0.75 as one of the conditions for identifying high-risk features. The basis for setting the behavior disturbance ratio to 15%, in practice, while keeping the recommendation accuracy not less than 90% of the original model, it is found that a disturbance ratio of 15% can maximize the reduction of the re-identification rate of behavior trajectory. Under this setting, the overlap between the behavior path of the desensitization data set and the original path is reduced to less than 60%, while the accuracy rate of the user portrait model is still maintained at more than 95% of the original model. Therefore, 15% is selected as the optimal disturbance ratio. The basis for setting the image blurring radius to 10, using the mainstream object detection model (such as the YOLO series) to test the recognition accuracy of images before and after desensitization, experiments were conducted under the conditions of blurring radius of 5, 10, 15 and 20 pixels, respectively. The results show that when the blurring radius is set to 10 pixels, the model recognition accuracy can be effectively suppressed to less than 20%, while the overall content of the image still has readability, and does not affect the aesthetic and layout requirements of the advertisement, therefore the blurring radius of the high-risk area of the image is determined to be 10 pixels. The lower limit of the language fluency score of the replaced text is set to 90%, in the language replacement experiment, the language model scores the original sentence and the replaced sentence respectively, and establishes a "language acceptability" scoring function. In the comparison sample, more than 90% of the scores can ensure that the replaced text will not be identified as harsh or illogical by the user. In the A / B test, when the language model score is less than 90% of the original sentence, the user complaint rate increases by more than 20% and the reading time significantly decreases. Therefore, 90% is set as the minimum score of the language fluency after replacement.
[0041] The layered encryption method includes homomorphic encryption for the identifying field and differential privacy protection strategy for the behavioral field.
[0042] In this implementation, the original advertisement data is first divided into two categories according to the privacy sensitivity level of the field: identifying field and behavioral field. The identifying field mainly includes user identity number, device unique identification code, IP address, MAC address and other sensitive information that can directly or indirectly identify the user's identity; the behavioral field includes user click behavior, dwell time, interest preference label, page browsing sequence and other information with behavioral pattern characteristics but not directly pointing to identity. For the identifying field, an additive homomorphic encryption mechanism is used for encryption processing. In implementation, an integer addition-based homomorphic encryption scheme such as Paillier encryption algorithm is selected, and the specific process is as follows: first, a pair of keys is generated, of which the public key is used for data encryption and the private key is kept in the controlled server and only used in the federal model aggregation or necessary verification stage. The original value of each identifying field is converted into ciphertext format after the encryption function, and the ciphertext can participate in addition operation and aggregation processing, ensuring that feature statistics can be completed without decryption in the modeling process. Taking user ID as an example, the original ID value is about 2048-bit binary string after encryption, and the ciphertext is stored in the model input pipeline and participates in the federal modeling process. For the behavioral field, a differential privacy protection strategy is used, and a Laplace mechanism is specifically used to add noise to numerical value features. The implementation steps are as follows: first, determine the value range of each behavioral field, for example, the maximum value of click count is 100 and the minimum value is 0, and the global sensitivity is 100. The preset privacy budget parameter ε is 0.5, and the standard deviation of the noise value calculated according to the Laplace distribution is the sensitivity divided by ε, i.e. 200. Before each data call, the corresponding Laplace random noise value is automatically generated for the behavioral field and added to the original behavioral data to form the desensitized behavioral features. For example, the page click count of user A is 20, and a random value with noise of ±15 is generated for it, and the final submission data value is 35 or 5. This noise mechanism is implemented at the data layer to ensure that a single query does not leak the user's real behavior characteristics. In addition, to balance privacy protection and modeling accuracy, a numerical boundary control strategy is set after the noise is added to clip the noise results that exceed the reasonable value range. Taking the above click count as an example, if the noise result exceeds 100 or is less than 0, the value is automatically limited within the range.
[0043] The preset evaluation indexes include a distribution distance index, a feature diversity index, and an actual advertisement click rate simulation index. The scheme comprehensively evaluates the performance of the generated data in statistical fidelity, sample structure complexity, and real business adaptability by constructing an evaluation index system in three dimensions to ensure that the generated de-identification data is both safe and practical. The distribution distance index: this index is used to quantify the difference between the synthetic data and the original data in the statistical distribution level. First, normalize the feature fields with comparable data in the two types of data sets, such as gender, age, region, interest category, etc. Then, divide each feature dimension into 10 equal intervals, and count the sample frequency of each interval to obtain the frequency vector of the original distribution and the synthetic distribution. Calculate the average absolute difference between the two distributions by summation, and the obtained value is the output result of the distribution distance index. If the value is less than 0.2, it is determined that the distribution consistency of the feature between the two data sets is good; if it exceeds 0.2, the data generation model needs to be retrained to improve the restoration effect. In addition, for numerical features such as click duration and dwell time, the cumulative probability difference method is used to calculate the distance between the distributions, and the result is normalized to 0 to 1. The acceptance threshold is set to 0.15. The feature diversity index: this index is used to evaluate the difference between the synthetic data samples to prevent the generated samples from being too similar and lacking modeling value. Randomly select 1000 synthetic samples, and calculate the Euclidean distance and direction angle between each pair of feature vectors. If the proportion of sample pairs with a Euclidean distance less than 0.3 exceeds 20%, or the proportion of sample pairs with a feature angle less than 15 degrees exceeds 25%, it is determined that the diversity of the synthetic data is insufficient. To further quantify the overall distribution of the data structure, the feature entropy calculation mechanism is introduced, and the frequency distribution entropy of each feature dimension is calculated. If the average feature entropy is higher than 3.5, it indicates that the data diversity is good; if it is lower than 3.0, the structure updating mechanism of the data generation module is triggered to improve the randomness and breadth of sample generation. The actual advertisement click rate simulation index: this index is used to detect the modeling adaptability of the generated data in the actual advertisement recommendation task. A set of advertisement click rate prediction model structures is constructed, and the original data set and the generated data set are used as training data, and a set of click behavior records from real business logs is used as a unified test set for verification. Calculate the click prediction accuracy of the two training models on the test set, and compare the results as the evaluation benchmark. If the accuracy of the generated data training model is not less than 90% of the accuracy of the original data training model, it is considered that the synthetic data has real business modeling capability; if it is lower than the threshold, it indicates that the modeling value of the synthetic data is insufficient, and the generation model needs to be reoptimized. At the same time, the fitting degree of the click prediction curve under each type of population label (such as gender, age group, interest label) is calculated, and the maximum deviation of the click rate curve between each sub-population should not exceed 10%. The part that exceeds the standard is recorded and returned to the synthesis strategy optimization module for difference compensation learning.
[0044] The cross-modal fusion processing includes aligning the generated text description and image content in a shared semantic space, and enhancing cross-modal semantic consistency through attention mechanism. After the generation of text and image data, the text and image are respectively encoded, and then a unified shared semantic space is introduced to realize alignment processing. The process includes the following steps: text feature encoding, first, each text description is encoded, and a pre-trained Chinese language model is used to encode the text content into a text feature vector with a length of 512. The model can process text input of up to 512 characters and output a semantic vector for each character. Finally, an average pooling operation is performed to obtain a unified representation vector for the entire text. Image feature encoding, the generated image data is input into an image feature extraction network for multi-level convolution operation. The network structure includes 5 convolution layers, and the number of convolution kernels is 32, 64, 128, 128 and 256 respectively, and the convolution kernel size is fixed at 3 by 3. After the maximum pooling layer, the key region features of the image are extracted, and finally compressed into an image feature vector with a length of 512 through a fully connected layer, which is consistent with the text feature dimension. Constructing a shared semantic space. The above text and image features are respectively sent into a dual mapping network, and the network structure is two layers of fully connected transformation modules, which project the two modal features into the same semantic space with a dimension of 256. The mapping process maintains structural symmetry, i.e. the text and image input paths share the same structure and parameter size. In this semantic space, the similarity score between the two modal features is calculated, and a set of alignment weights is obtained by using dot product matching method. If the score is higher than 0.8, it means that the alignment result has strong semantic consistency, and the result is retained; otherwise, re-matching or feedback adjustment will be performed.
[0045] Introducing attention mechanism to enhance semantic consistency, in order to further improve the accuracy of cross-modal information fusion, an attention mechanism module is introduced. The module includes three channels: self-attention channel, cross-attention channel and gate channel. The self-attention channel is used to enhance the internal semantic dependence of the text, especially to capture the subject-predicate structure, adjectival modification and other syntactic relationships in the description; the cross-attention channel is used to capture the significant alignment relationship between the key regions in the image and the key words in the text; the gate channel is used to adjust the attention weight according to the context information, to avoid semantic drift caused by forced alignment between modalities. The attention weight is calculated through a feedforward neural network, and the input is the embedding vector of the text and image in the shared semantic space, and the output is a weighted vector with a dimension of 256, representing the final cross-modal fusion representation.
[0046] The output fusion feature is used for downstream tasks. The fused feature will be sent to subsequent advertisement modeling modules or multi-task learning networks for tasks such as advertisement content recommendation, click rate prediction, and advertisement material review, ensuring the structural rationality, semantic consistency, and business adaptability of multi-modal content. By constructing a structure-symmetrical shared semantic space and an attention regulation mechanism, this method achieves deep coupling of different modal content, ensuring that the generated text description and image content maintain logical consistency and semantic accuracy in multi-scene applications, significantly improving the modeling quality and recommendation performance of multi-modal advertisement data. If the detection result is lower than the preset threshold, the generative adversarial network model is trained again, and the specific formula is: wherein, Ltotal represents the total loss function, Ladversarial represents the adversarial loss, Lstatistical represents the statistical property loss, Ldiversity represents the diversity loss, denotes a regularization term, λ1 denotes a statistical property loss weighting coefficient, λ2 denotes a diversity loss weighting coefficient, and λ3 denotes a regularization loss weighting coefficient. In the privacy computation driven cross-platform advertising joint modeling method proposed by the present application, when the synthetic data quality evaluation result is lower than the preset threshold, a secondary training mechanism based on loss function optimization is designed to improve the performance of the generative adversarial network model. The mechanism takes the weighted sum of multiple loss functions as the objective function, and the specific form is: the total loss function is equal to the adversarial loss plus the statistical property loss multiplied by the statistical property loss weighting coefficient, plus the diversity loss multiplied by the diversity loss weighting coefficient, plus the regularization term multiplied by the regularization loss weighting coefficient. The generative adversarial network model takes the total loss function as the optimization target during the training process, and updates the parameters of the generator and the discriminator through back propagation. The adversarial loss is responsible for controlling the discriminability between the generated samples and the real samples, the statistical property loss is used to measure the closeness of the generated data to the original data in terms of statistical indicators such as mean and variance, the diversity loss is used to limit the repetition rate between generated samples and maintain the diversity of data structure, and the regularization term is used to suppress the model complexity and prevent overfitting. Through this structured loss function system, the generation effect can be quickly iterated and optimized when the model evaluation is not up to standard, and the security and usability of the desensitized data can be improved. The total loss function represents the weighted total value of various losses in a forward propagation process, and is an important basis for calculating the gradient in the back propagation phase of the model. The goal of this function is to jointly minimize the difference between the generated samples and the real data, while balancing sample quality, structural diversity, and model stability. The adversarial loss is used to measure the difference between the generated samples and the real samples in the output of the discriminator. The smaller the value, the more realistic the generated samples. The cross-entropy loss of the discriminator in the generative adversarial network is used as the measurement standard. This loss directly participates in gradient calculation during generator training and is the main driving force for improving generation quality. The statistical property loss represents the deviation between the generated data and the original data in terms of overall statistical distribution. It mainly includes four indicators: mean, standard deviation, skewness, and kurtosis. In each training round, the average difference value of the above four statistical indicators between the original data and the generated data in each feature dimension is calculated, and this value is taken as the statistical property loss. This loss controls the global distribution consistency of the generated samples, avoiding problems such as distribution skew and sample polarization. The diversity loss is used to evaluate the difference between the generated samples, preventing the collapse of patterns. Its calculation method is: randomly select a number of sample pairs from the current batch of generated samples, calculate the average Euclidean distance between them in the feature space, and if the distance is too small, the diversity loss will increase. This indicator promotes the model to generate more extensive structural samples, improving sample coverage and modeling breadth. The regularization term is used to constrain the complexity of the generator parameters to prevent the model from overfitting. The L2 norm is used to sum all trainable parameters in the generator to obtain the regularization value.In training, the model weight is encouraged to keep a small value, which improves the generalization ability of the model and the stability after deployment.
[0047] The specific calculation formula of the statistical characteristic loss weighting coefficient λ1 is: Wherein, λ1 represents a statistical characteristic loss weighting coefficient, C represents a category coverage number of the synthetic data, and L represents a statistical distribution deviation level. The statistical characteristic loss weighting coefficient is equal to the category coverage number divided by the sum of the category coverage number and the statistical distribution deviation level. The formula can dynamically adjust the contribution value of the statistical characteristic loss in the total loss function according to the actual distribution of the current generated data, thereby realizing fine-grained model training control. The number of categories effectively covered in the current generated data, i.e., the category coverage number, is calculated after each round of training, and the deviation degree between the statistical distribution of each feature and the real data is evaluated to form the statistical distribution deviation level. Substitute the two parameters into the above calculation formula to obtain the statistical characteristic loss weighting coefficient in real time, which is used to guide the loss function optimization direction of the next round of training. The category coverage number represents the total number of effective categories contained in the current synthetic data. For example, if there are 10 dimensions of user interest categories, and only 8 of them are covered in the synthetic data, then the category coverage number is 8. The determination method of this value is as follows: after generating the data, the frequency of each key category is counted, and if the number of occurrences of a category in the synthetic data is not less than a certain threshold (such as 30 times), then the category is considered to be effectively covered. This setting can ensure that the statistical coverage is representative rather than incidental. The statistical distribution deviation level L is used to quantify the statistical consistency deviation between the generated data and the original real data in the key feature dimensions, and is a core evaluation index reflecting the overall fidelity of the generated data. First, select 5 types of key features including gender ratio, age distribution, user interest label distribution, page click behavior frequency and access time period distribution as the statistical distribution deviation analysis objects. For each feature dimension, the mean deviation, standard deviation and distribution coincidence difference between the generated data and the real data in that dimension are calculated, all the differences are normalized to 0 to 1, and finally the average of the three indexes represents the deviation degree of that dimension. Then, according to the influence of each feature on the modeling of the advertisement, the weights of gender and age distribution are set to 0.2, the weight of interest label distribution is set to 0.3, the weight of click frequency distribution is set to 0.2, and the weight of time period distribution is set to 0.1. A total deviation score S is calculated by weighted averaging. In order to use standardized values when calculating the statistical characteristic loss weighting coefficient, S is mapped to a discrete level L, specifically: when S is less than or equal to 0.1, L is 1; when S is greater than 0.1 and less than or equal to 0.2, L is 2; when S is greater than 0.2 and less than or equal to 0.35, L is 3; when S is greater than 0.35 and less than or equal to 0.5, L is 4; when S is greater than 0.5, L is 5. The level division standard is set based on a large number of cross-platform advertisement data measurement results, which can accurately reflect the fitting degree of the generated data to the statistical distribution of the real data, and ensure good interpretability and training guidance value when dynamically adjusting the statistical characteristic loss weighting coefficient in the loss function.
[0048] The specific calculation formula of the diversity loss weighting coefficient λ2 is: wherein, λ2 represents a diversity loss weighting coefficient, D represents a feature distribution dispersion degree, and R represents a model output repetition rate; λ3 = 1 - λ1 - λ2; wherein, λ3 represents a regularization loss weighting coefficient. The formula dynamically adjusts the proportion of diversity loss in the total loss function by considering the expansion range of the generated sample in the feature dimension and the deduplication performance of the generated model output. When the proportion of repeated samples in the model output is high, the denominator increases, and the value of λ2 relatively decreases, reducing the reward for diversity; when the dispersion degree of the generated sample distribution is large, it means that the generated content has strong distinguishability, the numerator increases, and λ2 increases accordingly, thereby encouraging the model to output more differentiated samples. On this basis, in order to ensure the normalization structure of the total loss function, the regularization loss weighting coefficient is automatically determined by "1 minus the statistical property loss weighting coefficient minus the diversity loss weighting coefficient", thereby realizing the balanced closed loop of the weighting structure of the overall loss function. D represents the dispersion degree of the synthetic sample in the feature space, which is used to measure whether the distribution range of the sample in different dimensions is wide enough. The following method is used to determine the value of D: first, select structural strong and classification strong feature dimensions, such as interest label encoding, behavior frequency, image content embedding vector, etc., and divide them into N intervals; then, the number of samples in each interval is counted, and the standard deviation and information entropy of each feature are calculated. The weighted average of the normalized standard deviation values and the information entropy values of multiple features is taken as the dispersion score D of this batch of data, and finally D is normalized to 0 to 1. If the value of D is close to 0, it means that the data distribution is highly concentrated; if the value of D is close to 1, it means that the data coverage is wide and the structure is rich. R represents the proportion of repeated samples output by the model in the current training round, reflecting the degree of degeneration of the diversity of the generated model. The calculation method of R value is as follows: the feature hash is performed on all sample vectors generated in the current batch, and the repeated statistics of the hash value is performed. The proportion of repeated samples is the number of repeated samples divided by the total number of samples, and the value range is 0 to 1. It is found that when R is higher than 0.4, the sample diversity decreases significantly, and the model training tends to converge and stagnate. Therefore, R is introduced as an adjusting factor to reduce the value of λ2 when the repetition rate is high, avoiding the solidification of the model. λ3 represents the proportion of the regularization term in the total loss function, and the purpose is to constrain the size of the generator parameters and suppress the risk of overfitting.
[0049] Although embodiments of the present application have been shown and described, it is to be understood that various modifications, substitutions, replacements and changes can be made to these embodiments without departing from the principles and spirit of the present application, and the scope of the present application is defined by the appended claims and their equivalents.
Claims
1. A cross-platform advertising joint modeling method driven by privacy computing, characterized by: The method comprises: Through the pre-established privacy protection mechanism, the original advertising data is initially desensitized. Layered encryption is used for the sensitive fields contained therein to obtain the desensitized data set after the first stage of processing, resulting in a data set that initially shields sensitive information. Based on the desensitized dataset processed in the first stage, a generative adversarial network model is used to extract features from the data. Deep learning is performed on the statistical characteristics of the data to obtain implicit distribution patterns and determine the generated preliminary feature mapping results. Based on the preliminary feature mapping results, a preliminary synthetic dataset that is suitable for the application scenario is determined. Based on the preliminary synthetic dataset and the requirements for data practicality evaluation, the statistical characteristics of the generated data are tested using preset evaluation indicators. If the test result is lower than the preset threshold, the generative adversarial network model is trained again to obtain an optimized synthetic dataset. Through the optimized synthetic dataset, the generated text and image data are cross-modally fused to meet the requirements of diverse data forms, obtaining the final multimodal advertising data set and determining the output results suitable for multi-scenario applications; Based on the final multimodal advertising data set, a security mechanism module is embedded to detect potential privacy leakage risks. If reversible features are found in the data, the reversible features are desensitized twice to obtain a final data set with higher security.
2. The privacy computing-driven cross-platform advertising joint modeling method according to claim 1, characterized in that: The preliminary synthetic data set that is determined to be suitable for the application scenario through the preliminary feature mapping results includes: Based on the preliminary feature mapping results, in order to meet the needs of multimodal data support, the text and image data are feature decomposed separately to obtain the corresponding multimodal feature vectors and the classified feature combination.
3. The privacy computing-driven cross-platform advertising joint modeling method according to claim 2, characterized in that: The preliminary synthetic data set that is determined to be suitable for the application scenario through the preliminary feature mapping results includes: Based on the classified feature combination, a dynamic balance design module is embedded to adjust parameters to address the conflict between data generation quality and privacy protection mechanism. If it is detected that the generated feature vector deviates from the preset statistical characteristic threshold, the model parameters are iteratively optimized to obtain the adjusted feature set.
4. The privacy computing-driven cross-platform advertising joint modeling method according to claim 3 is characterized by: The preliminary synthetic data set that is determined to be suitable for the application scenario through the preliminary feature mapping results includes: Through the adjusted feature set, a multimodal data generation sub-module is constructed to meet complex analysis requirements, obtain generated text and image data samples, and determine the preliminary synthetic data set that is suitable for the application scenario.
5. The privacy computing-driven cross-platform advertising joint modeling method according to claim 1, characterized in that: The layered encryption method includes using homomorphic encryption for identification fields and using a differential privacy protection strategy for behavioral fields.
6. The privacy computing-driven cross-platform advertising joint modeling method according to claim 1, characterized in that: The preset evaluation indicators include a distribution distance indicator, a feature diversity indicator, and an actual advertisement click-through rate simulation indicator.
7. The privacy computing-driven cross-platform advertising joint modeling method according to claim 1, characterized in that: The cross-modal fusion process includes aligning the generated text description with the image content using a shared semantic space and enhancing cross-modal semantic consistency through an attention mechanism.
8. The privacy computing-driven cross-platform advertising joint modeling method according to claim 1, characterized in that: If the detection result is lower than the preset threshold, the generative adversarial network model is trained twice. The specific formula is: in, represents the total loss function, Represents resistance to loss, represents the loss of statistical properties, represents the loss of diversity, represents the regularization term, λ1 represents the statistical feature loss weighting coefficient, λ2 represents the diversity loss weighting coefficient, and λ3 represents the regularization loss weighting coefficient.
9. The privacy computing-driven cross-platform advertising joint modeling method according to claim 8, characterized in that: The specific calculation formula of the statistical characteristic loss weighting coefficient λ1 is: Among them, λ1 represents the weighted coefficient of statistical feature loss, C represents the number of category coverage of synthetic data, and L represents the level of statistical distribution deviation.
10. The privacy computing-driven cross-platform advertising joint modeling method according to claim 8, characterized in that: The specific calculation formula of the diversity loss weighted coefficient λ2 is: Among them, λ2 represents the diversity loss weighting coefficient, D represents the feature distribution dispersion, and R represents the model output repetition rate; λ3=1-λ1-λ2; Among them, λ3 represents the regularization loss weight coefficient.
Citation Information
Cited By
Multi-modal training data desensitization and traceability management method
CN121278775A
User privacy regression modeling method fusing multi-platform data
CN121365381A