Cross-platform advertisement data fusion analysis method and system based on multi-source heterogeneity
By using a multi-source heterogeneous data fusion and analysis method, the problems of inconsistent data formats, difficulties in entity alignment, and inaccurate attribution analysis in cross-platform advertising have been solved. This method enables unified collection and analysis of cross-platform data, improving advertising efficiency and analysis accuracy.
Patent Information
- Application Number
- CN202511899870.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-16
- Publication Date
- 2026-03-20
AI Technical Summary
Existing technologies for cross-platform advertising suffer from issues such as inconsistent data formats, difficulties in entity alignment, asynchronous time sequences, and inaccurate attribution analysis, making it difficult for advertisers to conduct comprehensive analysis and evaluation from a holistic perspective.
Through the collaborative work of multi-source heterogeneous data acquisition modules, data standardization modules, entity alignment modules, time-series alignment modules, and cross-platform attribution analysis modules, unified acquisition, standardization, entity matching, and attribution analysis of cross-platform data are achieved, forming a closed-loop optimization system.
It enables unified collection and analysis of cross-platform data, eliminates differences in data format and time base, accurately identifies campaigns by the same advertiser, provides scientific budget allocation suggestions, and improves campaign efficiency and analysis accuracy.
Smart Images

Figure CN121707646A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data fusion and advertising technology, specifically to a method and system for cross-platform advertising data fusion and analysis based on multi-source heterogeneity. Background Technology
[0002] With the rapid development of short video platforms, advertising channels are becoming increasingly diversified. Advertisers typically need to run ads on multiple platforms simultaneously, such as ByteDance's Qianchuan, ByteDance's Xingtu, Local Promotion, DOU+, and ByteDance Ads. However, differences in data formats, inconsistent indicator systems, and data update delays among these platforms make it difficult for advertisers to comprehensively analyze and evaluate the effectiveness of their campaigns from a holistic perspective.
[0003] In the prior art, Chinese invention patent CN119863280A discloses a method and system for adjusting digital ad placement based on multi-platform data analysis. This method obtains platform ad placement information, categorizes ads according to product type, and acquires ad category information. Further, it obtains an ad performance index based on the weights of ad performance data, click-through rate, and conversion rate, as well as platform information. Then, it sorts ads of the same category according to their ad performance index based on the ad category information and adjusts the digital placement of ads based on this order information.
[0004] However, the aforementioned existing technologies have the following shortcomings: First, existing technologies mainly classify advertisements based on product type and target user age group, lacking in-depth analysis of the characteristics of the advertising materials themselves, and cannot accurately identify the same creative materials placed by the same advertiser on different platforms, resulting in inaccurate correlation analysis of cross-platform campaigns; Second, when calculating the advertising effectiveness index, existing technologies only consider limited dimensions such as click-through rate, conversion rate, and user overlap, failing to fully consider the user's conversion path across multiple platforms and the actual contribution of each platform to the final conversion, resulting in inaccurate attribution analysis; Third, existing technologies lack a mechanism to handle differences in data update delays across platforms, which may lead to deviations when comparing cross-platform data due to inconsistent time bases; Fourth, existing technologies are one-way data processing flows, lacking a closed-loop mechanism to feed the analysis results back to the data collection and processing stages, and cannot achieve dynamic optimization of data collection strategies. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a method and system for cross-platform advertising data fusion and analysis based on multi-source heterogeneity, aiming to solve technical problems such as inconsistent data formats, difficulties in entity alignment, asynchronous time sequences, and inaccurate attribution analysis in cross-platform advertising campaigns.
[0006] The technical solution of the present invention is as follows:
[0007] A cross-platform advertising data fusion and analysis method based on multi-source heterogeneous data includes: a multi-source heterogeneous data acquisition module that connects to the data services of ByteDance's Qianchuan, Xingtu, Local Push, DOU+, and ByteDance advertising platforms through standardized interfaces, collecting real-time advertising exposure, click, conversion, and consumption data from each platform to generate raw data streams; a data standardization module that receives the raw data streams and maps the differentiated field naming and indicator definitions of each platform to a preset standard data model to establish unified standardized indicator data across platforms; and an entity alignment module that receives the standardized indicator data, generates material fingerprint vectors based on the multimodal features of advertising materials, and combines them with the association features of the advertising account. The system performs cross-platform entity matching and outputs an entity association graph. The time-series alignment module receives standardized indicator data and aggregates it at a uniform time granularity based on the data update latency characteristics of each platform to generate a time-series alignment sequence. The cross-platform attribution analysis module receives the entity association graph and the time-series alignment sequence, tracks the conversion path of users after reaching multiple platforms, calculates the attribution weight vector of each platform for the final conversion based on the reach sequence, and outputs a cross-platform campaign overview analysis report. The feedback optimization module receives the attribution weight vector, generates optimization adjustment parameters, and feeds them back to the multi-source heterogeneous data acquisition module and the data standardization module to dynamically adjust the data acquisition frequency and standardization mapping rules.
[0008] The present invention also provides a cross-platform advertising data fusion and analysis system based on multi-source heterogeneity, including a multi-source heterogeneous data acquisition module, a data standardization module, an entity alignment module, a time-series alignment module, a cross-platform attribution analysis module, and a feedback optimization module. Each module works together to implement the various steps of the above method.
[0009] The beneficial effects of this invention are as follows: By standardizing the interface of the multi-source heterogeneous data acquisition module to connect with multiple advertising platforms, unified cross-platform data acquisition is achieved, solving the problem of scattered data sources; through the field mapping and caliber conversion of the data standardization module, a unified indicator system comparable across platforms is established, eliminating differences in data formats and calculation methods among platforms; through the material fingerprint comparison and account association analysis of the entity alignment module, the advertising activities of the same advertiser on different platforms are accurately identified, achieving precise association of cross-platform advertising activities; through the delay compensation and time granularity unification of the time sequence alignment module, the differences in data update delays between platforms are eliminated, ensuring a consistent time benchmark for cross-platform data comparison; through the conversion path tracking and attribution weight calculation of the cross-platform attribution analysis module, the actual contribution of each platform to the final conversion is accurately assessed, providing a scientific basis for budget allocation; and through the parameter feedback mechanism of the feedback optimization module, dynamic optimization of data acquisition and processing strategies is achieved, forming a complete closed-loop optimization system. Attached Figure Description
[0010] Figure 1 This is a flowchart of the cross-platform advertising data fusion and analysis method based on multi-source heterogeneity of the present invention.
[0011] Figure 2 This is an architecture diagram of the cross-platform advertising data fusion and analysis system based on multi-source heterogeneity of the present invention. Detailed Implementation
[0012] Please refer to the attached document. Figures 1-2 The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0013] Reference Figure 1 As shown, this invention provides a method for cross-platform advertising data fusion analysis based on multi-source heterogeneity. This method achieves fusion analysis of cross-platform advertising data through the collaborative work of a multi-source heterogeneous data acquisition module 1, a data standardization module 2, an entity alignment module 3, a time-series alignment module 4, a cross-platform attribution analysis module 5, and a feedback optimization module 6.
[0014] The multi-source heterogeneous data acquisition module 1 connects to the data services of ByteDance's Qianchuan, Xingtu, Local Push, DOU+, and ByteDance's advertising platform through standardized interfaces, and collects advertising exposure data, click data, conversion data, and consumption data from each platform in real time to generate raw data streams.
[0015] In one embodiment of the present invention, the multi-source heterogeneous data acquisition module 1 establishes an independent data acquisition channel for each advertising platform. For the ByteDance Qianchuan platform, the acquisition channel obtains platform API access permissions through the OAuth2.0 authorization mechanism and calls the advertising data interface provided by the platform to obtain performance data at three levels: advertising plan, advertising group, and advertising creative. The collected data fields include core indicators such as impressions, clicks, conversions, cost per thousand impressions, click-through rate, conversion rate, and conversion cost. The data acquisition frequency is dynamically adjusted according to the advertising status. For ads in the advertising process, the acquisition frequency is set to once every 15 minutes; for ads that have been paused or ended, the acquisition frequency is reduced to once every 6 hours.
[0016] For the ByteDance Star Chart platform, since it primarily serves influencer marketing scenarios, its data structure differs significantly from performance advertising platforms. The data collected through various channels includes multiple dimensions such as influencer information, task order information, content publishing information, and performance data. Performance data covers three categories: video views, interaction data, and conversion data. Video views include total views, number of completed views, and average playback duration; interaction data includes likes, comments, shares, and increase in followers; conversion data includes product clicks, product transaction amount, and number of completed orders.
[0017] For local promotion platforms, which target local service merchants, geographical location information plays a significant role in data collection. In addition to acquiring standard exposure, click, and conversion data, the data collection channels also need to collect in-store redemption data and geofence-triggered data. In-store redemption data includes the number of redeemed orders, redemption amount, and redemption rate; geofence-triggered data includes the number of users entering the geofence, dwell time within the geofence, and conversions within the geofence.
[0018] For the DOU+ platform, which is a content boosting tool, its data structure is relatively simple. The data collected mainly includes metrics such as boosting targets, spending amount, additional views, additional interaction, and increase in followers. Since DOU+'s campaign cycle is usually short, the data collection frequency is set to once every 30 minutes.
[0019] For the ByteDance advertising platform, which covers various advertising formats including feed ads, search ads, and splash screen ads, the data collection channel needs to call the corresponding data interface according to different ad formats. The data collection fields for feed ads are similar to those for ByteDance's Qianchuan platform; search ads additionally collect search term reports and quality score data; splash screen ads additionally collect exposure duration distribution and click heatmap data.
[0020] The multi-source heterogeneous data acquisition module 1 uses a message queue mechanism to buffer and distribute the acquired data. In this embodiment, Kafka is used as the message middleware, and an independent message topic is established for each delivery platform. The acquisition process of each platform encapsulates the acquired raw data into a standardized message format and sends it to the corresponding message topic. The data standardization module 2, as a consumer, subscribes to data from each message topic and performs subsequent processing. The message format includes five parts: platform identifier, data type, timestamp, raw data body, and acquisition metadata. The platform identifier is used to distinguish the data source; the data type indicates the business category to which the data belongs; the timestamp records the time information of data generation and acquisition; the raw data body is the raw JSON or XML format data returned by the platform; and the acquisition metadata records the running status and abnormal information of the acquisition process.
[0021] Data standardization module 2 receives the raw data stream and maps the different field names and indicator definitions of each platform to the preset standard data model to establish unified standardized indicator data across platforms.
[0022] In one embodiment of the present invention, the data standardization module 2 first establishes a platform field mapping table. This mapping table defines the correspondence between the original field names and standard field names of each platform. Taking the impression count metric as an example, the ByteDance Qianchuan platform uses the field name "show_cnt", the ByteDance Xingtu platform uses the field name "play_count", the local push platform uses the field name "impression", the DOU+ platform uses the field name "add_play", and the ByteDance Advertising platform uses the field name "stat_show". The standard data model uniformly maps the above fields to the standard field "standard_impression". Similarly, complete field mapping relationships are established for core metrics such as click count, conversion count, and cost.
[0023] Furthermore, Data Standardization Module 2 needs to address the differences in the calculation methods of metrics across different platforms. Taking conversion rate as an example, the conversion rate calculation formula for the ByteDance Qianchuan platform is the number of conversions divided by the number of clicks, while the conversion rate calculation formula for the ByteDance Xingtu platform is the transaction amount divided by the number of content views multiplied by 100. To eliminate these differences, Data Standardization Module 2 sets a conversion factor to normalize the original metric values. Specifically, for each metric requiring conversion, a source conversion expression and a target conversion expression are defined, and a conversion function is established. The conversion function extracts the basic metric values required for calculation from the original data based on the source conversion expression, and recalculates according to the target conversion expression to obtain the standardized metric value.
[0024] In one possible implementation, data standardization module 2 uses a unified advertising campaign identifier to associate and tag data from the same advertising campaign across different platforms. The advertising campaign identifier consists of three parts: advertiser ID, campaign creation time, and campaign sequence number, and is generated using the snowflake algorithm to create a globally unique identifier. When data from different platforms is detected to belong to the same advertising campaign, these data records are tagged with the same advertising campaign identifier, facilitating subsequent cross-platform correlation analysis.
[0025] Entity alignment module 3 receives standardized indicator data, generates material fingerprint vectors based on the multimodal features of advertising materials, performs cross-platform entity matching by combining the association features of the advertising account, and outputs an entity association graph.
[0026] In one embodiment of the present invention, the entity alignment module 3 first performs multimodal feature extraction on the advertising material. For video materials, keyframe image sequences are extracted, and a pre-trained visual encoder is used to extract visual feature descriptors for each keyframe. In this embodiment, the keyframe extraction uses a scene change detection algorithm. When the pixel difference between adjacent frames exceeds a preset threshold, it is determined to be a scene change point, and the frame at the scene change point is extracted as a keyframe. The visual encoder uses a ResNet-50 network structure and extracts the features of the penultimate layer as a 2048-dimensional visual feature descriptor. For the visual feature descriptors of multiple keyframes, a temporal average pooling method is used to merge them into a single visual feature descriptor.
[0027] For the text content in advertising creatives, including video titles, subtitles, and landing page text, a pre-trained language model is used to extract semantic feature descriptors. In this embodiment, the language model uses the Chinese version of BERT-Base, and the hidden layer output at the CLS position is taken as a 768-dimensional semantic feature descriptor after the text is input into the model. For multiple text segments, the semantic feature descriptors are extracted separately and then merged using a weighted average method, with the weights determined based on text length and positional importance.
[0028] For the audio content in the advertising materials, Mel-frequency cepstral coefficients are extracted as the basic representation of audio features. In this embodiment, 40-dimensional Mel-frequency cepstral coefficients are used, with one frame extracted every 25 milliseconds and a frame shift of 10 milliseconds. For the Mel-frequency cepstral coefficient sequence of the entire audio segment, the mean and standard deviation are calculated using statistical pooling, and then concatenated to form an 80-dimensional audio feature descriptor.
[0029] Entity alignment module 3 performs weighted fusion of visual feature descriptors, semantic feature descriptors, and audio feature descriptors to generate a material fingerprint vector. The calculation formula of the multimodal material fingerprint fusion algorithm proposed in this invention is as follows:
[0030] ,
[0031] in, The fingerprint vector of the merged material; This is a visual feature descriptor with a dimension of 2048; This is a semantic feature descriptor with a dimension of 768; This is an audio feature descriptor with a dimension of 80; , and These are projection functions for visual, semantic, and audio features, respectively, used to project features of different dimensions onto a unified 256-dimensional space. , and These are the modal weighting coefficients, satisfying the constraints. In this embodiment, based on the characteristics of short video ads, visual modal weights are used. Set to 0.5, semantic modality weight Set to 0.3, audio modal weights The value is set to 0.2. The projection function is implemented using a single-layer fully connected network, and the parameters are obtained through pre-training using a contrastive learning method.
[0032] After generating the material fingerprint vector, Entity Alignment Module 3 performs cross-platform entity matching by combining the associated features of the advertising accounts. The associated features of the advertising accounts include account name, contact information, enterprise qualification information, and historical advertising behavior characteristics. The account name is evaluated using a combination of edit distance and Jaccard similarity; the contact information is matched precisely; the enterprise qualification information is compared with business registration information; and the historical advertising behavior characteristics include advertising category preferences, advertising time period distribution, and budget range.
[0033] The cross-platform entity similarity calculation algorithm proposed in this invention comprehensively considers both material fingerprint similarity and account association similarity. Its calculation formula is as follows:
[0034] ,
[0035] in, For cross-platform entity similarity scores, the value range is: ; To measure the similarity of the fingerprints, cosine similarity is used to calculate the degree of similarity between two fingerprint vectors. The account association similarity is calculated by weighted similarity of each association feature of the account; This is a balancing coefficient used to adjust the weighting relationship between material similarity and account similarity. In this embodiment, The weight is set to 0.6, meaning that the fingerprint similarity of the material accounts for 60% of the weight, and the account association similarity accounts for 40% of the weight.
[0036] Material fingerprint similarity The calculation formula is:
[0037] ,
[0038] in, and These are the fingerprint vectors of the two ad creatives to be matched; This represents the vector dot product operation; This represents the L2 norm of a vector.
[0039] Account association similarity The calculation formula is:
[0040] ,
[0041] in, Account name similarity; Score based on contact information matching; Scoring based on enterprise qualifications; Similarity of delivery behavior characteristics; , , and The weight coefficients for each feature satisfy the constraints. In this embodiment, Set to 0.2, Set to 0.3, Set to 0.3, Set it to 0.2.
[0042] Entity alignment module 3 sets the matching threshold When cross-platform entity similarity score Greater than At that time, the corresponding entity pairs are marked as cross-platform campaigns of the same advertiser. In this embodiment, Set to 0.75. For all identified cross-platform entity pairs, construct an entity association graph. The entity association graph is stored in a graph structure, where nodes represent campaign activities on each platform, edges represent cross-platform relationships, and the weight of each edge is the entity similarity score.
[0043] The time-series alignment module 4 receives standardized indicator data and aggregates the standardized indicator data according to the delay characteristics of data updates on each platform, generating a time-series alignment sequence.
[0044] In one embodiment of the present invention, the timing alignment module 4 first detects the delay time of data updates on each platform. Delay detection is implemented using a probe mechanism; for each deployment platform, test requests are periodically sent and the timestamp differences of data updates are recorded. Specifically, at time... Initiate a data query request to obtain the timestamp of the latest data record returned by the platform. The estimated data update delay for this platform is... To obtain a stable delay estimate, a sliding window method is used to take a weighted average of the most recent N detection results, with the most recent detection results given a higher weight.
[0045] Based on actual measurement results, in this embodiment, the typical data update delays for each platform are as follows: Juxing Qianchuan platform: approximately 5 to 15 minutes; Juxing Xingtu platform: approximately 30 minutes to 2 hours; Local Push platform: approximately 10 to 30 minutes; DOU+ platform: approximately 5 to 20 minutes; Juxing Advertising platform: approximately 10 to 20 minutes. The delay differences between different platforms can reach tens of minutes or even hours. Without time-series alignment, directly comparing the real-time data of each platform will result in significant discrepancies.
[0046] The adaptive temporal alignment algorithm proposed in this invention corrects the timestamps of data from various platforms. The correction process includes two steps: delay compensation and time granularity unification. Delay compensation adjusts the timestamps of data from each platform forward by the corresponding delay time, aligning the data from different platforms in the time dimension. Time granularity unification aggregates the corrected data based on a preset time window.
[0047] The formula for calculating delay compensation is:
[0048] ,
[0049] in, For the platform The corrected timestamp; For the platform The original timestamp; For the platform The estimated data update delay.
[0050] The aggregation calculation formula with uniform time granularity is as follows:
[0051] ,
[0052] in, For the platform In the Aligned data for each time window; For the first The time range of each time window; For the platform At any moment The corrected data; This is the aggregation weighting function within the time window. In this embodiment, the time window size is set to 1 hour, and the aggregation weighting function uses uniform weighting, i.e. ,in For time window The number of data points within.
[0053] The time-series alignment module 4 also needs to handle data missing cases. When a platform has no data within a specific time window, interpolation methods are used to fill in the missing values. This embodiment uses a linear interpolation method to estimate the data value of the missing window based on the data values of adjacent time windows. For cases where multiple consecutive time windows are missing, a maximum interpolation span threshold is set; missing intervals exceeding the threshold are marked as invalid data and not included in subsequent analysis.
[0054] The cross-platform attribution analysis module 5 receives entity association maps and time-series alignment sequences, tracks the conversion path of users after reaching multiple platforms, calculates the attribution weight vector of each platform for the final conversion based on the reach sequence, and outputs a cross-platform campaign panoramic analysis report.
[0055] In one embodiment of the present invention, the cross-platform attribution analysis module 5 first obtains the set of advertising activities of the same advertiser on various platforms based on the entity association graph. For each connected component in the entity association graph, all advertising activity nodes contained therein are extracted to form a cross-platform advertising activity set. Each set represents a group of associated advertising activities of the same advertiser.
[0056] Furthermore, the cross-platform attribution analysis module 5 reconstructs the user's reach sequence within the campaign set based on user device identifiers and timestamp information. User device identifiers are matched using three identifiers: device ID, IP address, and user account. For each user, their advertising reach records across various platforms are collected, including impression events, click events, and conversion events. The reach records are sorted according to the timestamps of the events to form an ordered user reach sequence. Each element in the reach sequence contains four fields: platform identifier, event type, timestamp, and event attribute.
[0057] The cross-platform attribution analysis module 5 identifies the conversion node in the reach sequence where the final conversion occurs. A conversion node is defined as the reach record corresponding to the last conversion event in the reach sequence. For each conversion node, the reach sequence preceding it is traced back to extract all platform reach points involved in that conversion path. The time window for the reach path is set to 30 days, meaning only reach events within 30 days prior to the conversion are considered.
[0058] The multi-touchpoint attribution weight allocation algorithm proposed in this invention calculates the contribution of each platform to the final conversion based on a Markov chain model. The algorithm constructs a state transition matrix of the reach sequence and determines the attribution weight by calculating the removal effect value of each platform node.
[0059] The state transition matrix is constructed as follows: Each platform is defined as a state node in the state space, with two special states: the initial state and the transition completion state. The transition frequency between states is statistically analyzed from historical reach sequence data, and the state transition probability is calculated.
[0060] The formula for calculating the state transition probability is:
[0061] ,
[0062] in, From state Transition to state The probability of; From state Transition to state The number of observations; This is the state space, which includes all platform states and special states.
[0063] Based on the state transition matrix, the removal effect value of each platform node is calculated. The removal effect value is defined as the change in probability from the initial state to the transition completion state after removing a certain platform node.
[0064] The formula for calculating the removal effect size is:
[0065] ,
[0066] in, For the platform The removal effect value; Given the complete state transition matrix, the probability of reaching the completed state from the initial state. To remove the platform The probability of reaching the completed state from the initial state under the state transition matrix after the node.
[0067] The probability of reaching the completed transition state from the initial state is calculated using the matrix power series method. Let the state transition matrix be... The initial distribution vector of the initial state is Then after The state distribution vector after the step transition is:
[0068] ,
[0069] Conversion completion probability State distribution vector The components of the completed transformation state in The limit value as it approaches infinity.
[0070] The formula for calculating the attribution weight vector is:
[0071] ,
[0072] in, For the platform Attribution weights; For the platform The removal effect value; This refers to the set of all platforms participating in the attribution calculation. The attribution weights satisfy the normalization constraint. .
[0073] The platform contribution comprehensive evaluation algorithm proposed in this invention further calculates the platform contribution index, which comprehensively considers three dimensions: attribution weight, reach cost, and reach efficiency.
[0074] The formula for calculating the platform contribution index is:
[0075] ,
[0076] in, For the platform Contribution index; For the platform Attribution weights; For the platform Conversion rate; For the platform The cost per conversion; For the platform The reach efficiency coefficient is defined as the ratio of the number of effective reaches to the total number of reaches.
[0077] The cross-platform attribution analysis module 5 outputs a comprehensive cross-platform campaign analysis report based on the above calculation results. The report includes a summary of independent performance metrics for each platform, cross-platform user overlap analysis results, attribution contribution rankings for each platform, and budget allocation strategy recommendations generated based on attribution weight vectors. The budget allocation strategy recommendations calculate recommended budget allocation ratios based on the contribution index of each platform.
[0078] The formula for calculating the budget allocation ratio is:
[0079] ,
[0080] in, To be allocated to the platform Recommended budget; For the platform Contribution index; This is the total budget amount.
[0081] Feedback optimization module 6 receives the attribution weight vector, generates optimization adjustment parameters, and feeds them back to multi-source heterogeneous data acquisition module 1 and data standardization module 2 to dynamically adjust the data acquisition frequency and standardization mapping rules.
[0082] In one embodiment of the present invention, the feedback optimization module 6 evaluates the delivery efficiency of each platform based on the attribution weight vector. Delivery efficiency is defined as the attribution contribution value generated per unit cost. For platforms with delivery efficiency higher than a preset efficiency threshold, the system determines that the data of that platform has higher analytical value, and therefore increases the data collection frequency of the corresponding platform to obtain more granular data. For platforms with delivery efficiency lower than the preset efficiency threshold, the system adjusts the weight coefficients in the standardized mapping rules of the corresponding platform to reduce the influence weight of the platform's data in the fusion analysis.
[0083] The closed-loop feedback parameter optimization algorithm proposed in this invention uses gradient descent to dynamically optimize data acquisition and standardization parameters. The optimization objective function is defined as the accuracy index of attribution analysis, and the optimal acquisition frequency and standardization weights are determined by maximizing this index.
[0084] The objective function is defined as follows:
[0085] ,
[0086] in, To optimize the objective function value; For the first The actual attribution label for each conversion event; For the first Predicted attribution values for each conversion event; To assess the sample size; This is a regularization term for the data acquisition frequency, used to constrain the acquisition frequency to a reasonable range; This is a regularization term for the standardized weights, used to constrain the smoothness of the weight coefficients; and is the regularization coefficient.
[0087] The formula for updating the sampling frequency parameter is:
[0088] ,
[0089] in, For the platform In the The sampling frequency parameters at the next iteration; The learning rate is set to 0.01 in this embodiment; This is the gradient of the objective function with respect to the acquisition frequency parameter.
[0090] The formula for updating the standardized weight parameters is:
[0091] ,
[0092] in, For the platform In the The standardized weight parameters at the next iteration.
[0093] The feedback optimization module 6 sets the parameter update cycle, which in this embodiment is once every 24 hours. After each parameter update, the new acquisition frequency parameters are sent to the multi-source heterogeneous data acquisition module 1, and the new standardized weight parameters are sent to the data standardization module 2. Upon receiving the parameters, both modules immediately apply the new configuration, achieving dynamic adjustment of the data processing strategy.
[0094] Through the aforementioned closed-loop feedback mechanism, this invention achieves deep coupling and collaborative operation of six modules: data acquisition, data standardization, entity alignment, temporal alignment, attribution analysis, and feedback optimization. The output of the multi-source heterogeneous data acquisition module 1 serves as the input of the data standardization module 2. The output of the data standardization module 2 serves as the input of the entity alignment module 3 and the temporal alignment module 4. The outputs of the entity alignment module 3 and the temporal alignment module 4 together serve as the input of the cross-platform attribution analysis module 5. The output of the cross-platform attribution analysis module 5 serves as the input of the feedback optimization module 6. The output of the feedback optimization module 6 inversely influences the parameter configuration of the multi-source heterogeneous data acquisition module 1 and the data standardization module 2, forming a complete closed-loop optimization system.
[0095] Reference Figure 2 As shown, the present invention also provides a cross-platform advertising data fusion and analysis system based on multi-source heterogeneity. The system includes a multi-source heterogeneous data acquisition module 1, a data standardization module 2, an entity alignment module 3, a time-series alignment module 4, a cross-platform attribution analysis module 5, and a feedback optimization module 6.
[0096] The multi-source heterogeneous data acquisition module 1 is used to connect to the data services of ByteDance's Qianchuan, Xingtu, Local Push, DOU+, and ByteDance advertising platforms through standardized interfaces. It collects real-time ad exposure, click, conversion, and consumption data from each platform to generate raw data streams. In this embodiment, the multi-source heterogeneous data acquisition module 1 includes three sub-layers: an interface adaptation layer, a data acquisition layer, and a message distribution layer. The interface adaptation layer is responsible for connecting to the API interfaces of each advertising platform, handling authorization authentication and request encapsulation; the data acquisition layer is responsible for periodically triggering data acquisition tasks and parsing the response data returned by the platform; and the message distribution layer is responsible for encapsulating the collected raw data into a standard message format and sending it to a message queue.
[0097] Data standardization module 2 receives the raw data stream and maps the differentiated field naming and indicator definitions across platforms to a preset standard data model, establishing unified standardized indicator data across platforms. In this embodiment, data standardization module 2 includes three sub-modules: a field mapping engine, a definition conversion engine, and an identifier generation engine. The field mapping engine maintains a mapping table between platform fields and standard fields and performs field renaming operations; the definition conversion engine maintains the definition conversion rules for each indicator and performs numerical normalization calculations; the identifier generation engine generates unified advertising campaign identifiers for cross-platform associated data records.
[0098] Entity alignment module 3 receives standardized indicator data, generates material fingerprint vectors based on the multimodal features of advertising materials, performs cross-platform entity matching by combining the association features of the advertising account, and outputs an entity association graph. In this embodiment, entity alignment module 3 includes a feature extraction submodule, a fingerprint generation submodule, a similarity calculation submodule, and a graph construction submodule. The feature extraction submodule extracts features from visual, semantic, and audio content respectively; the fingerprint generation submodule fuses multimodal features to generate material fingerprint vectors; the similarity calculation submodule calculates the comprehensive similarity score between entity pairs; and the graph construction submodule constructs an entity association graph based on the similarity score.
[0099] The time-series alignment module 4 receives standardized indicator data and, based on the latency characteristics of data updates across different platforms, aggregates the standardized indicator data at a uniform time granularity to generate a time-series aligned sequence. In this embodiment, the time-series alignment module 4 includes a latency detection submodule, a time correction submodule, and an aggregation processing submodule. The latency detection submodule periodically detects the data update latency of each platform; the time correction submodule corrects the data timestamps based on latency estimates; and the aggregation processing submodule aggregates the corrected data according to a uniform time window.
[0100] The cross-platform attribution analysis module 5 receives entity association graphs and time-series aligned sequences, tracks user conversion paths after reaching multiple platforms, calculates attribution weight vectors for each platform's contribution to the final conversion based on the reach sequences, and outputs a cross-platform delivery panoramic analysis report. In this embodiment, the cross-platform attribution analysis module 5 includes a path tracing submodule, an attribution calculation submodule, a contribution evaluation submodule, and a report generation submodule. The path tracing submodule reconstructs the reach sequences based on user identifiers; the attribution calculation submodule calculates attribution weights based on a Markov chain model; the contribution evaluation submodule comprehensively evaluates the contribution index of each platform; and the report generation submodule summarizes the analysis results to generate a panoramic report.
[0101] The feedback optimization module 6 receives the attribution weight vector, generates optimization adjustment parameters, and feeds them back to the multi-source heterogeneous data acquisition module 1 and the data standardization module 2, dynamically adjusting the data acquisition frequency and standardization mapping rules. In this embodiment, the feedback optimization module 6 includes an efficiency evaluation submodule, a parameter optimization submodule, and a parameter distribution submodule. The efficiency evaluation submodule calculates the deployment efficiency index for each platform; the parameter optimization submodule uses the gradient descent algorithm to optimize the acquisition and standardization parameters; and the parameter distribution submodule sends the optimized parameters to the corresponding modules.
[0102] The data flow relationship between the above modules is as follows: the raw data stream of the multi-source heterogeneous data acquisition module 1 is output to the data standardization module 2; the standardized index data of the data standardization module 2 is output to the entity alignment module 3 and the time-series alignment module 4; the entity association graph of the entity alignment module 3 and the time-series alignment sequence of the time-series alignment module 4 are jointly output to the cross-platform attribution analysis module 5; the attribution weight vector of the cross-platform attribution analysis module 5 is output to the feedback optimization module 6; the optimization adjustment parameters of the feedback optimization module 6 are fed back to the multi-source heterogeneous data acquisition module 1 and the data standardization module 2, forming a closed-loop data flow.
[0103] The system of this invention can be deployed on a cloud server cluster, employing a microservice architecture to achieve independent deployment and elastic scaling of each module. In terms of hardware configuration, the server uses an Intel Xeon Gold 6248R processor, equipped with 256GB of memory and a 4TB solid-state drive. Regarding the software environment, the operating system is Ubuntu 22.04 LTS, the container orchestration is Kubernetes version 1.28, the message queue is Apache Kafka version 3.5, the graph database is Neo4j version 5.12, and the time-series database is InfluxDB version 2.7.
[0104] Through actual testing and verification, the system of this invention, in a scenario involving five delivery platforms and 10 million data records per day, exhibits a data acquisition latency of less than 5 seconds, a data standardization processing throughput of 100,000 records per second, an entity alignment accuracy of 92.6%, a temporal alignment error of less than 1 minute, an attribution analysis calculation time of less than 30 seconds, and an end-to-end processing latency of less than 2 minutes. Compared with existing technologies, this invention improves cross-platform entity recognition accuracy by approximately 15%, attribution analysis accuracy by approximately 12%, and budget allocation efficiency by approximately 18%.
[0105] The embodiments described above are merely illustrative of specific implementations of the present invention, and while the descriptions are detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention.
Claims
1. A method for cross-platform advertising data fusion and analysis based on multi-source heterogeneity, characterized in that, include: The multi-source heterogeneous data acquisition module connects to the data services of ByteDance's Qianchuan, Xingtu, Local Push, DOU+, and ByteDance's advertising platform through standardized interfaces, and collects advertising exposure data, click data, conversion data, and consumption data from each platform in real time to generate raw data streams; The data standardization module receives the raw data stream and maps the different field names and indicator definitions of each platform to a preset standard data model to establish unified standardized indicator data across platforms. The entity alignment module receives the standardized indicator data, generates a material fingerprint vector based on the multimodal features of the advertising material, performs cross-platform entity matching by combining the association features of the advertising account, and outputs an entity association graph. The time-series alignment module receives the standardized indicator data and, based on the delay characteristics of data updates on each platform, aggregates the standardized indicator data at a uniform time granularity to generate a time-series alignment sequence. The cross-platform attribution analysis module receives the entity association map and the time-series alignment sequence, tracks the conversion path of users after reaching multiple platforms, calculates the attribution weight vector of each platform for the final conversion based on the reach sequence, and outputs a cross-platform delivery panoramic analysis report. The feedback optimization module receives the attribution weight vector, generates optimization adjustment parameters, and feeds them back to the multi-source heterogeneous data acquisition module and the data standardization module to dynamically adjust the data acquisition frequency and standardization mapping rules.
2. The method according to claim 1, characterized in that, The data standardization module maps the differentiated field naming and indicator definitions of various platforms to a preset standard data model, specifically including: Establish a platform field mapping table to create a one-to-one correspondence between the original field names and the standard field names of each platform; Based on the differences in the calculation methods of indicators across different platforms, a conversion factor is set to normalize the original indicator values. For data from the same advertising campaign across different platforms, a unified advertising campaign identifier is used for association and tagging.
3. The method according to claim 1, characterized in that, The entity alignment module generates a material fingerprint vector based on the multimodal features of the advertising material, specifically including: Extract visual feature descriptors from the visual content of advertising materials; Extract semantic feature descriptors from the text content of advertising materials; Extract audio feature descriptors from the audio content of advertising materials; The visual feature descriptor, the semantic feature descriptor, and the audio feature descriptor are weighted and fused to generate the material fingerprint vector.
4. The method according to claim 1, characterized in that, The entity alignment module performs cross-platform entity matching by combining the association features of the advertising account, specifically including: Extract the basic attribute features of the advertising account, including account name, contact information, and enterprise qualification information; Calculate the overall similarity score between the pairs of entities to be matched; When the overall similarity score is greater than the preset matching threshold, the corresponding entity pair is marked as a cross-platform campaign by the same advertiser.
5. The method according to claim 1, characterized in that, The time-series alignment module aggregates the standardized indicator data according to a unified time granularity, specifically including: Detect the latency of data updates on each platform; The data from each platform is timestamped based on the aforementioned delay time. The corrected data is aggregated based on a preset time window to generate the time-aligned sequence with a uniform time granularity.
6. The method according to claim 1, characterized in that, The cross-platform attribution analysis module tracks the conversion path of users after being reached on multiple platforms, specifically including: Based on the entity association graph, obtain the set of advertising activities of the same advertiser on various platforms; Based on the user device identifier and timestamp information, reconstruct the user's reach sequence in the set of campaigns; Identify the conversion node in the reach sequence that ultimately undergoes conversion behavior.
7. The method according to claim 6, characterized in that, The cross-platform attribution analysis module calculates the attribution weight vector for each platform on the final conversion based on the reach sequence, specifically including: Construct the state transition matrix of the reach sequence; The removal effect value of each platform node is calculated based on the state transition matrix; The attribution weight vector for each platform is determined based on the removal effect value.
8. The method according to claim 1, characterized in that, The feedback optimization module generates optimization adjustment parameters, specifically including: The delivery efficiency of each platform is evaluated based on the attribution weight vector. For platforms whose delivery efficiency exceeds the preset efficiency threshold, increase the data collection frequency for the corresponding platforms; For platforms whose delivery efficiency is lower than the efficiency threshold, adjust the weight coefficients in the standardized mapping rules for the corresponding platforms.
9. The method according to claim 1, characterized in that, The cross-platform delivery overview analysis report includes: Summary of independent performance metrics for each platform; Cross-platform user overlap analysis results; Ranking of attribution contribution across platforms; Budget allocation strategy recommendations generated based on the attribution weight vector.
10. A cross-platform advertising data fusion and analysis system based on multi-source heterogeneity, used to implement the method described in any one of claims 1-9, characterized in that, include: The multi-source heterogeneous data acquisition module is used to connect to the data services of ByteDance's Qianchuan, Xingtu, Local Push, DOU+ and ByteDance's advertising platforms through standardized interfaces, and collect advertising exposure data, click data, conversion data and consumption data of each platform in real time to generate raw data streams; The data standardization module is used to receive the raw data stream, map the different field names and indicator definitions of each platform to a preset standard data model, and establish cross-platform unified standardized indicator data. The entity alignment module is used to receive the standardized indicator data, generate a material fingerprint vector based on the multimodal features of the advertising material, perform cross-platform entity matching in combination with the association features of the advertising account, and output an entity association graph. The time-series alignment module is used to receive the standardized indicator data, and aggregate the standardized indicator data according to the delay characteristics of data updates on each platform, thereby generating a time-series alignment sequence. The cross-platform attribution analysis module is used to receive the entity association map and the time-series alignment sequence, track the conversion path of users after reaching multiple platforms, calculate the attribution weight vector of each platform for the final conversion based on the reach sequence, and output a cross-platform delivery panoramic analysis report. The feedback optimization module is used to receive the attribution weight vector, generate optimization adjustment parameters, and feed them back to the multi-source heterogeneous data acquisition module and the data standardization module to dynamically adjust the data acquisition frequency and standardization mapping rules.
Citation Information
Patent Citations
Digital delivery adjustment method and system based on multi-platform data analysis
CN119863280A
Cited By
A cross-platform advertisement delivery anti-cheating and strategy optimization method
CN122155790A