Cross-platform crowd delivery data tracking and monitoring method and system based on information flow

By introducing data acquisition, integration, analysis and prediction modules into the cross-platform data analysis system, the problem of insufficient data integration and in-depth analysis capabilities in the existing technology is solved, and more accurate and efficient advertising delivery results are achieved.

CN120146929APending Publication Date: 2025-06-13GUANGZHOU BLOCK NETWORK TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510241748.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing cross-platform data analysis methods lack the ability to integrate and in-depth analysis of data from different sources, and it is difficult to deal with the problems of incompleteness, inconsistency and delay, which affects the accuracy of analysis and prediction.

Method used

Provide a cross-platform crowd delivery data tracking and monitoring method and system based on information flow, including data acquisition module, data integration module, data analysis and prediction module and data storage module. Through the collaborative work of these modules, cross-platform data collection, integration, cleaning, conversion and analysis can be achieved, and future delivery results can be predicted.

Benefits of technology

It improves the accuracy and efficiency of advertising delivery, provides advertisers with effective delivery optimization solutions and data support, and enhances the ability to integrate and in-depth analysis of data from different sources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146929A_ABST
    Figure CN120146929A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-platform crowd delivery data tracking and monitoring method and system based on information flow, relates to the technical field of big data, and is used for solving the problem of integration and deep analysis ability of data from different sources in an existing cross-platform data analysis method. Comprising a data acquisition module, a data integration module and a data analysis and prediction module, and the data acquisition module acquires advertisement putting related data from a plurality of platforms and covers user behavior data such as advertisement display, click and conversion. And the data integration module selects different data cleaning and conversion modes by analyzing the influence of different factors on data processing, uniformly integrates data from different platforms, eliminates redundancy and ensures data consistency. And the data analysis and prediction module carries out deep analysis on the integrated data and predicts a future delivery effect. Through cooperative work of the above modules, the accuracy and benefit of advertisement putting can be improved, and an effective putting optimization scheme and data support are provided for an advertiser.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data technology, and more specifically, to a cross-platform population placement data tracking and monitoring method and system based on information flow. Background Art

[0002] With the development of the Internet and the improvement of the degree of informatization, cross-platform population placement based on information flow has become an important data application scenario in various fields. Information flow placement data includes, but is not limited to, multi-dimensional data such as user clicks, exposures, conversions, access durations, interactions, etc. These data are generated through multiple channels and cover multiple ecosystems such as social media, search engines, e-commerce platforms, short video platforms, etc.

[0003] The prior art has the following deficiencies:

[0004] In existing cross-platform data analysis methods, most rely on data interfaces of a single platform and lack the ability to integrate and deeply analyze data from different sources. Especially when facing problems such as incomplete, inconsistent, and delayed data, traditional methods are difficult to effectively process the data, thus affecting the accuracy of subsequent analysis and prediction.

[0005] In view of the above problems, the present invention proposes a solution. Summary of the Invention

[0006] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a cross-platform population placement data tracking and monitoring method and system based on information flow to solve the problems raised in the above background art.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] In a preferred embodiment, it includes: a data acquisition module, a data integration module, a data analysis and prediction module, and a data storage module, and the modules are signal-connected to each other;

[0009] The data acquisition module extracts key behavior data related to placement from each platform and uploads it to the data storage module;

[0010] The data integration module determines the data feature coefficients of each platform by weighted calculation based on the key behavior data of different platforms, the repetition degree between the key behavior data, and the influence of different external conditions, performs different data cleaning and conversion methods on different key behavior data through the data feature coefficients, and summarizes the cross-platform key behavior data;

[0011] After the data analysis and prediction module obtains the key behavior data summarized by the data integration module from the data storage module, it analyzes the integrated key behavior data through differencing and the time series model ARIMA to predict the future delivery effect;

[0012] The data storage module is used to store all data during the platform processing.

[0013] In a preferred embodiment, the data acquisition module docks with the API interface provided by the platform, obtains the access token using the OAuth 2.0 protocol; according to the RESTful document provided by the platform, specifies the fields and the range of key behavior data to be collected, uses the paging mechanism to gradually pull a large amount of key behavior data; and sets the collection frequency;

[0014] And configures the callback address in the management interface of the platform, and listens to this callback address to receive the key behavior data pushed by the platform in real time.

[0015] In a preferred embodiment, by analyzing the key behavior data of different platforms, the duplication degree between the key behavior data, and the influence of different external conditions, the data characteristic coefficients of each platform are determined, and different data cleaning and conversion methods are performed on different key behavior data through the data characteristic coefficients;

[0016] According to the field completeness rate, error rate, and average time delay in the key behavior data of each platform, calculate the quality of the key behavior data of each platform. The specific implementation steps are as follows:

[0017] Step A1: The data integration module extracts the field data of each platform from the key behavior data of each platform obtained, including: the number of valid fields, the total number of fields, and calculates the field completeness rate of each platform. The specific formula is: field completeness rate = number of valid fields / total number of fields;

[0018] Step A2: The data integration module extracts the error data of each platform from the key behavior data of each platform obtained, including the number of errors, the total number of records, and calculates the error rate of each platform. The specific formula is: error rate = 1 - (number of errors / total number of records);

[0019] Step A3: Extract the time data of each platform from the key behavior data of each platform obtained, including: average delay time, maximum tolerable delay time, and calculate the average time delay from the generation to the completion of data collection of each platform. The specific formula is: average time delay = 1 - (average delay time / maximum tolerable delay time);

[0020] The data integration module calculates the quality of the key behavior data of each platform by weighted calculation based on the field completeness rate, error rate, and average time delay of the key behavior data of each platform;

[0021] The specific implementation steps for the duplicate rate of key behavior data on different platforms are as follows:

[0022] Step B1: The data integration module extracts the click behavior data of each platform from the key behavior data of each platform obtained;

[0023] Step B2: The data integration module calculates the cross-platform duplicate rate of the key behavior data of each platform through the duplicate record key behavior data set of each platform. Specifically, according to the formula: Duplicate rate = (Number of duplicate entries / Total number of entries) × 100%;

[0024] The data integration module determines the external condition coefficient of the key behavior data of each platform through weighted average calculation. The specific steps are as follows:

[0025] Step C1: The data integration module determines the dimensions affected by external conditions: geographical location, device type, and time period;

[0026] Step C2: The data integration module quantifies the external condition factors and determines the external condition values as follows: By analyzing the conversion rate of click behavior data in each region through historical key behavior data, the geographical location data is quantified; For the mobile terminal, since user operations are direct and the possibility of false clicks is low, a higher value is assigned; For the PC terminal, since the probability of script brushing is relatively high, a slightly lower value is assigned. Combining the time distribution law of clicks, the time period data is quantified;

[0027] According to the quantified data of each external condition, calculate the quantified values of different external conditions for each platform. Specifically, according to the formula: Quantity of each external condition data × Proportion of the corresponding external condition data volume;

[0028] Step C3: The data integration module calculates the overall external condition coefficient of each platform through the weight values of different external conditions;

[0029] The data integration module calculates the data feature coefficient of each platform through weighted calculation based on the quality of the key behavior data of each platform, the duplicate rate of the key behavior data, and the external condition coefficient;

[0030] The data integration module sets different feature coefficient thresholds according to the data feature coefficients of each platform, classifies the platform data into different categories, and applies different data cleaning and conversion methods to each category of data. Set the feature coefficient threshold Y. When the data feature coefficient of the platform is greater than or equal to Y, simple cleaning and conversion methods are used for such key behavior data. The specific steps are as follows:

[0031] Step S1: By checking the unique identifiers (such as ID, timestamp, etc.) in the key behavior data, delete the duplicate records;

[0032] Step S2: Perform unified format conversion on the dates, numbers, and strings in the key behavior data;

[0033] Step S3: Use mean filling to fill in the missing key behavior data;

[0034] When the data characteristic coefficient of the platform is less than or equal to Y, such key behavior data needs to be deeply cleaned and converted; the specific steps are as follows:

[0035] Step D1: The local outlier factor anomaly detection algorithm based on machine learning identifies abnormal key behavior data that does not conform to the normal pattern through multi-dimensional analysis, and marks it as the part that needs to be cleaned or eliminated;

[0036] Step D2: Use regression interpolation algorithm to infer missing values ​​based on the relationship between known key behavior data points and fill in the missing key behavior data.

[0037] In a preferred implementation, after the data integration module completes data cleaning and conversion, it is necessary to integrate and summarize the key behavior data from different platforms. The specific steps are as follows:

[0038] Step F1: The data integration module removes duplicate records in the key behavior data of different platforms through the identifiers contained in the key behavior data set of each platform;

[0039] Step F2: The data integration module maps fields of key behavior data that use different fields to represent the same meaning on different platforms;

[0040] Step F3: The data integration module aggregates the cross-platform user behavior data by user ID.

[0041] In a preferred embodiment, the data analysis and prediction module applies time series analysis to remove trend components of key behavior data through first-order differences;

[0042] The data analysis and prediction module uses the ARIMA model to model the key behavioral data after stabilization and calculate the model parameters.

[0043] In a preferred embodiment:

[0044] Step 1: Extract key behavioral data related to delivery from each platform;

[0045] Step 2: Determine the data characteristic coefficient of each platform through weighted calculation of the key behavior data of different platforms, the duplication between key behavior data, and the influence of different external conditions. Perform different data cleaning and conversion methods on different key behavior data through the data characteristic coefficient, and summarize the key behavior data across platforms;

[0046] Step 3: Analyze the integrated key behavior data through difference and time series to predict future delivery effects.

[0047] The present invention discloses a cross-platform population delivery data tracking and monitoring method and system based on information flow, which relates to the technical field of big data and is used to solve the problems of the integration and in-depth analysis capabilities of existing cross-platform data analysis methods for data from different sources; it includes a data acquisition module, a data integration module, and a data analysis and prediction module. The data acquisition module obtains advertisement delivery-related data from multiple platforms, covering user behavior data such as advertisement display, click, and conversion. The data integration module selects different data cleaning and conversion methods by analyzing the influence of different factors on data processing, and uniformly integrates the data from different platforms, eliminating redundancy and ensuring data consistency. The data analysis and prediction module deeply analyzes the integrated data to predict future delivery effects. Through the collaborative work of the above modules, the present invention can improve the accuracy and efficiency of advertisement delivery, and provide an effective delivery optimization plan and data support for advertisers. Description of the Drawings

[0048] Figure 1 It is a schematic structural diagram of the cross-platform population delivery data tracking and monitoring system based on information flow of the present invention.

[0049] Figure 2 It is an operation schematic diagram of the cross-platform population delivery data tracking and monitoring method based on information flow of the present invention. Detailed Embodiments

[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0051] Embodiment

[0052] The present invention discloses a cross-platform population delivery data tracking and monitoring system based on information flow, including: a data acquisition module, a data integration module, a data analysis and prediction module, and a data storage module, and the modules are connected by signals.

[0053] The data acquisition module extracts key behavior data related to delivery from each platform. In this embodiment, the collected key behavior data includes:

[0054] User ID (uniquely identifying the user's identity);

[0055] Advertisement ID (identifying specific advertisement content);

[0056] Click on the timestamp (records the time when the click occurred);

[0057] Platform source (indicates which platform the user's click comes from);

[0058] User device type (such as mobile or PC);

[0059] User's geographical location (such as country or city information);

[0060] Specifically, the data collection module docks with the API interface provided by the platform, uses the OAuth 2.0 protocol to obtain an access token to ensure that only authorized access to platform data is possible; according to the RESTful documentation provided by the platform, specifies the fields and the range of key behavior data to be collected, uses a paging mechanism to gradually pull a large amount of key behavior data to avoid performance issues or interface blockages caused by excessive single requests; and sets the collection frequency, for example: pull the latest click behavior data every 15 minutes to ensure the timeliness and integrity of the data.

[0061] At the same time, the API interface has deficiencies in terms of real-time performance. The data collection module configures a callback address (Webhook URL) in the platform's management interface and listens to this callback address to receive key behavior data pushed by the platform in real time. After the user clicks, the platform packages the relevant key behavior data into an event object and sends it to this address. After receiving the key behavior data, the data collection module extracts the key fields (user ID, click time) from the JSON format to parse the key behavior data, checks whether the fields are missing or the format is incorrect, and verifies the integrity and validity of the received key behavior data;

[0062] The collected data is as follows:

[0063]

[0064] After the data integration module obtains the key behavior data related to the delivery uploaded by the data collection module from the data storage module, it cleans, transforms, and summarizes the key behavior data from different platforms.

[0065] It should be noted that when the data collection module collects data from each platform, due to reasons such as data sources and the quality of key behavior data, the data is incomplete and inaccurate. In this embodiment, when the data integration module processes the key behavior data, considering the integrity and accuracy of the key behavior data, by analyzing the key behavior data of different platforms, the duplication degree between key behavior data, and the influence of different external conditions, the data characteristic coefficients of each platform are determined, and different data cleaning and transformation methods are used for different key behavior data through the data characteristic coefficients.

[0066] Specifically, the quality of key behavior data on different platforms (such as Weibo, Douyin, and Kuaishou) varies due to differences in collection mechanisms and audience characteristics. For example, Douyin mainly features short videos, with relatively high data interactivity, and its credibility is generally better than that of Weibo, which mainly focuses on pictures and texts. In this embodiment, the quality of key behavior data for each platform is calculated based on the field completeness rate, error rate, and average time delay in the key behavior data of each platform. The specific implementation steps are as follows:

[0067] Step A1: The data integration module extracts the field data of each platform from the key behavior data of each platform obtained, including the number of valid fields and the total number of fields, and calculates the field completeness rate of each platform. Specifically, according to the formula: field completeness rate = number of valid fields / total number of fields. For example, in the data of the Weibo platform, the total number of fields is 10, and the number of valid fields is 9, so the field completeness rate of the Weibo platform is 90%; in the data of the Douyin platform, the total number of fields is 10, and the number of valid fields is 9.5, so the field completeness rate of the Douyin platform is 95%; in the data of the Kuaishou platform, the total number of fields is 10, and the number of valid fields is 8.5, so the field completeness rate of the Kuaishou platform is 85%.

[0068] Step A2: The data integration module extracts the error data of each platform from the key behavior data of each platform obtained, including the number of errors and the total number of records, and calculates the error rate of each platform. Specifically, according to the formula: error rate = 1 - (number of errors / total number of records). For example, in the data of the Weibo platform, among 100 click records, 15 are false records, so the error rate of the Weibo platform is 85%; in the data of the Douyin platform, among 100 click records, 10 are false records, so the error rate of the Douyin platform is 90%; in the data of the Kuaishou platform, among 100 click records, 20 are false records, so the error rate of the Kuaishou platform is 80%.

[0069] Step A3: The data integration module extracts the time data of each platform from the data of each platform obtained, including the average delay time and the maximum tolerable delay time, and calculates the average time delay from the generation to the completion of data collection for each platform. Specifically, according to the formula: average time delay = 1 - (average delay time / maximum tolerable delay time). For example, in the data of the Weibo platform, the average delay time is 7 minutes, and the maximum tolerable delay time is 10 minutes, so the average time delay of the Weibo platform is 0.7; in the data of the Douyin platform, the average delay time is 5 minutes, and the maximum tolerable delay time is 10 minutes, so the average time delay of the Douyin platform is 0.5; in the data of the Kuaishou platform, the average delay time is 9 minutes, and the maximum tolerable delay time is 10 minutes, so the average time delay of the Kuaishou platform is 0.9.

[0070] The data integration module determines the quality of the key behavior data of each platform based on the field completeness rate, error rate, and average time delay of the key behavior data of each platform. Specifically, it is based on the formula: Q = w1×ZD + w2×WC - w3×PJT, where Q is the quality of the key behavior data of the platform, ZD is the field completeness rate of the key behavior data of the platform, WC is the error rate of the key behavior data of the platform, PJT is the average time delay of the key behavior data of the platform, and w1, w2, and w3 are weight distribution factors, which are adjusted according to the actual situation. For example: By calculation, w1 is 0.4, w2 is 0.4, and w3 is 0.2. Substituting the example data into the calculation, we get: The quality value of the key behavior data of the Weibo platform is 0.56; The quality value of the key behavior data of the Douyin platform is 0.64; The quality value of the key behavior data of the Kuaishou platform is 0.48.

[0071] The specific implementation steps for the duplicate rate of the key behavior data of different platforms are as follows:

[0072] Step B1: The data integration module extracts the click behavior data of each platform from the data of each platform obtained. The example is as follows:

[0073]

[0074]

[0075] Step B2: The data integration module regards each record as a unique identifier composed of the user ID, advertisement ID, and click timestamp, identifies duplicates from the cross-platform data, and obtains the duplicate record dataset after deduplicating the example data: {(1001, 001, 08:30:00), (1002, 002, 09:00:00), (1003, 003, 10:00:00)), (1004, 004, 11:00:00)}. In the example data obtained, (1001, 001, 08:30:00) of Weibo is repeated with Douyin; (1001, 001, 08:30:00) of Weibo itself is repeated; (1004, 004, 11:00:00) of Douyin is repeated with Kuaishou; (1002, 002, 09:00:00) of Weibo is repeated with Kuaishou.

[0076] Step B3: The data integration module calculates the cross-platform duplication rate of each platform through the duplicate record datasets of each platform. Specifically, according to the formula: Duplication rate = (Number of duplicate entries / Total number of entries) × 100%. For example, in this embodiment, the click behavior dataset of the Weibo platform is: {(1001,001,08:30:00), (1002,002,09:00:00), (1001,001,08:30:00)}, the total number of entries: 3, the duplicate item: (1001,001,08:30:00) duplicates itself once. After deduplication, the entries are: {(1001,001,08:30:00), (1002,002,09:00:00)}, the number of entries after deduplication: 2, the number of duplicate entries: 1. Substituting the data into the formula, the click behavior data duplication rate of the Weibo platform is calculated to be 33.33%; the click behavior dataset of the Douyin platform is, the total number of entries: 3, the duplicate items: (1001,001,08:30:00) duplicates with Weibo once, (1004,004,11:00:00) duplicates with Kuaishou once. After deduplication, the entries are: {(1001,001,08:30:00), (1003,003,10:00:00), (1004,004,11:00:00)}, the number of entries after deduplication: 3, the number of duplicate entries: 2. Substituting the data into the formula, the click behavior data duplication rate of Douyin is calculated to be 66.67%; the click behavior dataset of the Kuaishou platform is: {(1002,002,09:00:00), (1004,004,11:00:00)}, the total number of entries: 3, the duplicate items: (1002,002,09:00:00) duplicates with Weibo once, (1004,004,11:00:00) duplicates with Douyin once. After deduplication, the entries are: {(1002,002,09:00:00), (1004,004,11:00:00)}, the number of entries after deduplication: 2, the number of duplicate entries: 2. Substituting the data into the formula, the click behavior data duplication rate of the Kuaishou platform is calculated to be 100%.

[0077] Different external conditions will directly affect the credibility and representativeness of the key behavior data. In this embodiment, the data integration module determines the external condition coefficients of each platform through the analysis of three factors: geographical location, device type, and time period. The specific steps are as follows:

[0078] Step C1: Determine the dimensions affected by external conditions. The example data is as follows:

[0079]

[0080] Step C2: Quantify external condition factors and determine external condition values. Specifically, analyze the conversion rate of click behavior data in each region through historical key behavior data to quantify geographical location data. For example: City A (high consumption capacity): high true click-through rate, assigned a value of 1.0; City B (medium consumption capacity): medium true click-through rate, assigned a value of 0.8; City C (low consumption capacity): low true click-through rate, assigned a value of 0.6. For the mobile terminal, since user operations are direct and the probability of false clicks is low, a higher value is assigned; for the PC terminal, since the probability of script-based click farming is relatively high, a slightly lower value is assigned. For example: Mobile terminal: 0.9; PC terminal: 0.8. Combine the time distribution pattern of clicks to quantify time period data. For example: Morning (high-efficiency working hours, high authenticity of clicks): 1.0; Afternoon (medium authenticity of clicks): 0.7; Evening (high probability of false clicks): 0.5. Specific to the following example data:

[0081] Douyin platform:

[0082]

[0083]

[0084] Weibo platform:

[0085]

[0086]

[0087] Kuaishou platform:

[0088]

[0089]

[0090] Calculate the quantization values of different external conditions for each platform according to the specific data of each external condition, specifically based on the formula: the amount of data of each external condition × the proportion of the amount of data of the corresponding external condition. For example: for Douyin, Eld = (0.4×1.0)+(0.3×0.8)+(0.3×0.6) = 0.82, Edd = 0.8×0.9+0.2×0.7 = 0.86, Etd = 0.3×1.0+0.5×0.8+0.2×0.6 = 0.82, where Eld is the quantization value of the geographical location of Douyin, Edd is the quantization value of the device type of Douyin, and Etd is the quantization value of the time period of Douyin; for Weibo, Elw = 0.5×1.0+0.333×0.7+0.167×0.5 = 0.833, Edw = 0.75×0.8+0.25×0.7 = 0.775, Etw = 0.25×0.9+0.5×0.8+0.25×0.6 = 0.775, where Eld is the quantization value of the geographical location of Douyin, Edd is the quantization value of the device type of Douyin, and Etd is the quantization value of the time period of Douyin; for Kuaishou, Elk = 0.6×0.9+0.3×0.7+0.1×0.5 = 0.8, Edk = 0.7×0.85+0.3×0.75 = 0.82, Etk = 0.25×1.0+0.6×0.8+0.15×0.5 = 0.8, where Elk is the quantization value of the geographical location of Kuaishou, Edk is the quantization value of the device type of Kuaishou, and Etk is the quantization value of the time period of Kuaishou.

[0091] It should be noted that the quantization of external conditions is specifically set according to actual needs.

[0092] Step C3: Different external conditions have different degrees of influence on platform data. In this embodiment, calculate the overall external condition coefficient of each platform through the weight values of different external conditions, specifically based on the formula: external condition coefficient = Qd×(Eld / Elw / Elk)+Qs×(Edd / Edw / Edk)+Qt×(Etd / Etw / Etk), where Qd is the geographical location weight, Qs is the device type weight, and Qt is the time period weight. For example: through calculation, the weights of each external condition are obtained as follows: geographical location Qd = 0.5, device type Qs = 0.3, time period Qt = 0.2. Substituting the example data into the calculation, the external condition coefficient of the Douyin platform is 0.836, the external coefficient of the Weibo platform is 0.806, and the external condition coefficient of the Kuaishou platform is 0.816.

[0093] The data integration module determines the data characteristic coefficient of each platform through the key behavior data quality, key behavior data repetition rate and external condition coefficient of each platform. It should be noted that the degree of influence of each factor on the data characteristic coefficient is different. It is necessary to perform weighted average calculation according to the weight of each factor to obtain a more accurate and scientific data characteristic coefficient. The specific formula is: Data characteristic coefficient = Qzl×zl+Qcf×cf+Qwb×wb, where zl is the key behavior data quality value, cf is the key behavior data repetition value, wb is the external condition quantification value, Qzl is the key behavior data quality weight value, Qcf is the key behavior data repetition rate weight value, and Qwb is the external condition weight value. For example, by calculation, Qzl is 0.45, Qcf is 0.35, and Qwb is 0.2. The data characteristic coefficient of the Douyin platform is 0.688 when the sample data is brought into the calculation; the data characteristic coefficient of the Weibo platform is 0.530; and the data characteristic coefficient of the Kuaishou platform is 0.729.

[0094] The data integration module sets different characteristic coefficient thresholds according to the data characteristic coefficients of each platform, divides the platform data into different categories, and applies different data cleaning and conversion methods to each category of data. Specifically, the characteristic coefficient threshold Y is set. When the data characteristic coefficient of the platform is greater than or equal to Y, it means that the key behavior data quality is high, the repetition rate is low, and the external environment is relatively stable. This type of data uses a simple cleaning and conversion method; the specific steps are as follows:

[0095] Step S1: Delete duplicate records by checking unique identifiers (such as ID, timestamp, etc.) in key behavior data;

[0096] Step S2: convert the dates, numbers, and character strings in the key behavior data into a unified format;

[0097] Step S3: Use mean filling to fill in the missing key behavior data;

[0098] When the data characteristic coefficient of the platform is less than or equal to Y, it means that the quality of key behavior data is poor, the repetition rate is high, and the external environment fluctuates greatly. This type of data needs to be deeply cleaned and converted; the specific steps are as follows:

[0099] Step D1: The local outlier factor anomaly detection algorithm based on machine learning identifies abnormal data that does not conform to the normal pattern through multi-dimensional data analysis, and marks it as the part that needs to be cleaned or eliminated;

[0100] Step D2: Using regression interpolation algorithm to infer missing values ​​based on the relationship between known key behavior data points, and fill in the missing key behavior data;

[0101] In this embodiment, the threshold value Y of the feature coefficient is set to 0.7. When the data feature coefficient of the platform is less than or equal to 0.8, the deep cleaning and conversion method is adopted for such data. For example, the data feature coefficient of the Kuaishou platform is 0.729, which is greater than the feature coefficient threshold of 0.7. Therefore, the simple cleaning and conversion method needs to be adopted for the data of the Kuaishou platform; the data feature coefficient of the Douyin platform is 0.688, which is less than the feature coefficient threshold of 0.7. Therefore, the deep cleaning and conversion method needs to be adopted for the data of the Douyin platform; the data feature coefficient of the Weibo platform is 0.530, which is less than the feature coefficient threshold of 0.7. Therefore, the deep cleaning and conversion method needs to be adopted for the data of the Weibo platform.

[0102] It should be noted that both simple data cleaning and conversion and deep data cleaning and conversion are conventional technical means and will not be elaborated here.

[0103] After the data integration module completes data cleaning and conversion, it is necessary to integrate and summarize the key behavior data from different platforms to ensure that all key behavior data can be analyzed and processed in a unified format. The specific steps are as follows:

[0104] Step F1: The data integration module removes duplicate records in the key behavior data of different platforms through the identifiers (such as user ID, advertisement ID, click time) contained in the key behavior data sets of each platform, ensuring that multiple click data of the same user or advertisement are only counted once. For example, if there are the same user ID and advertisement ID in the data records of the Douyin and Kuaishou platforms, the latest record can be judged as valid data according to the timestamp, and the rest of the duplicate records need to be removed.

[0105] Step F2: The data integration module maps the fields of the key behavior data that represent the same meaning using different fields on different platforms. For example, the "device type" field in the Douyin platform is "device_type", while the "device type" field in the Kuaishou platform is "platform_device". The data integration module unifies these two fields into a standard field "device_type" for subsequent analysis and statistics. At the same time, for the same type of data (such as "click count", "conversion count") from different platforms, the data integration module combines them into the same column and ensures that the units are consistent to avoid affecting the analysis results due to different units.

[0106] Step F3: The data integration module summarizes the cross-platform user behavior data by user ID. For example, the number of clicks, purchase behaviors, etc. of the same user on different platforms are integrated into a complete user behavior record. Finally, the integrated data needs to be output in a unified format, usually in tabular form, with each row representing a record for subsequent analysis, report generation, or storage. The output format should be flexibly designed according to requirements.

[0107] It should be noted that data integration and summarization are conventional technical means and will not be elaborated here.

[0108] After the data analysis and prediction module obtains the key behavior data summarized by the data integration module from the data storage module, it analyzes the integrated key behavior data to predict the future delivery effect.

[0109] The data analysis and prediction module applies time series analysis and processes the integrated key behavior data through techniques such as differencing to make the integrated key behavior data stationary.

[0110] Specifically, after being summarized by the data integration module, a historical key behavior dataset summarized by week is generated, forming a data table with the following example:

[0111]

[0112]

[0113] The data analysis and prediction module removes the trend component of the key behavior data through first-order differencing to ensure that the data becomes stationary. For example, the change in the click-through rate (CTR) on the Douyin platform is as follows:

[0114] Week 1 to Week 2: 0.06 - 0.05 = 0.01

[0115] Week 2 to Week 3: 0.07 - 0.06 = 0.01

[0116] This indicates that the change in CTR is stationary and there is no significant trend. The same applies to the key behavior data of other platforms.

[0117] The data analysis and prediction module uses the ARIMA model to model the stationary key behavior data and calculates the model parameters.

[0118] Specifically, the ARIMA model consists of three main parts: autoregressive (AR), differencing (I), and moving average (MA). To build an ARIMA model, it is necessary to determine p (the order of autoregression), d (the order of differencing), and q (the order of moving average) to ensure that the model can correctly fit the key behavior data. Based on the differenced key behavior data, the differencing order d = 1 of the ARIMA model is determined. Further, the data analysis and prediction module selects the ARIMA(1,1,1) model by analyzing the ACF and PACF graphs, that is, p = 1 (the order of autoregression), d = 1 (the order of differencing), and q = 1 (the order of moving average). The data analysis and prediction module uses historical data (differenced click behavior data) to estimate the coefficients of the AR (autoregressive) and MA (moving average) parts through the ARIMA(1,1,1) model. Fit the model. For example: The model fitting results show that the AR coefficient is 0.8 and the MA coefficient is 0.3. Based on the already fitted ARIMA(1,1,1) model, the click-through rate (CTR) for the next few weeks can be predicted. Specifically, the autoregressive (AR) part reflects the relationship between the current value and the previous value. In the ARIMA model, the coefficient of the AR part (0.8) multiplies the differenced value of the previous week (0.01) by the coefficient to obtain the prediction of the differenced value for the next week. For example: To predict the click-through rate (CTR) of the fourth week through the ARIMA(1,1,1) model, the coefficient of the AR part (0.8) multiplies the differenced value of the previous week (0.01) by the coefficient to obtain the prediction of the differenced value for the next week. The difference in the previous week (the 3rd week) is 0.01 (from 0.06 to 0.07), so the AR part predicts an increment of 0.008 in the differenced value for the next week; the MA part reflects the impact of errors. If we make a prediction for the click-through rate of the 3rd week and calculate the error (assumed to be 0.002), then the MA part will multiply this error by the MA coefficient (0.3). For example: The prediction error for the 3rd week is 0.002, so the MA part predicts a differenced increment of 0.0006. Add the predicted values of the AR part and the MA part to obtain the total differenced prediction value for the 4th week as 0.0086. It should be noted that due to the first-order differencing, the differenced prediction value needs to be restored to the original click-through rate value. Specifically, according to the formula: The predicted CTR value for next week = this week's CTR + total differenced prediction. For example: The predicted CTR value for the 4th week = the CTR of the 3rd week + total differenced prediction, and the predicted click-through rate of the Douyin platform for the 4th week is 0.0786.

[0119] The present invention also relates to a cross-platform population placement data tracking and monitoring method based on information flow;

[0120] The specific process is as follows:

[0121] Step 1: Extract key behavior data related to placement from each platform.

[0122] Step 2: Clean and transform the key behavior data from different platforms, and summarize the cross-platform key behavior data.

[0123] Step 3: Analyze the integrated key behavior data to predict future placement effects.

[0124] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0125] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product.

[0126] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and invention constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0127] In addition, in each embodiment of this application, the functional modules can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.

[0128] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claimed rights.

[0129] Finally: The above is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A cross-platform crowd delivery data tracking and monitoring system based on information flow, characterized by: include: Data acquisition module, data integration module, data analysis and prediction module and data storage module, and signal connection between each module; The data collection module extracts key behavior data related to delivery from each platform and uploads it to the data storage module; The data integration module determines the data characteristic coefficient of each platform by weighted calculation of the key behavior data of different platforms, the duplication between key behavior data, and the influence of different external conditions. It performs different data cleaning and conversion methods on different key behavior data through the data characteristic coefficient, and summarizes the key behavior data across platforms. After the data analysis and prediction module obtains the key behavior data summarized by the data integration module from the data storage module, it analyzes the integrated key behavior data through the difference and time series model ARIMA to predict the future delivery effect; The data storage module is used to store all data during the platform processing process.

2. The cross-platform crowd delivery data tracking and monitoring system based on information flow according to claim 1 is characterized by: The data collection module connects to the API interface provided by the platform and uses the OAuth2.0 protocol to obtain an access token. According to the RESTful document provided by the platform, the data collection module specifies the fields and key behavior data ranges to be collected, and uses the paging mechanism to gradually pull large quantities of key behavior data. The collection frequency is also set. And configure the callback address in the platform's management interface, and receive key behavior data pushed by the platform in real time by monitoring the callback address.

3. The cross-platform crowd delivery data tracking and monitoring system based on information flow according to claim 2 is characterized by: By analyzing the key behavior data of different platforms, the duplication between key behavior data and the influence of different external conditions, the data characteristic coefficients of each platform are determined, and different key behavior data are cleaned and converted in different ways according to the data characteristic coefficients; The quality of key behavior data of each platform is calculated based on the field completeness rate, error rate and average time delay in the key behavior data of each platform. The specific implementation steps are as follows: Step A1: The data integration module extracts the field data of each platform from the key behavior data of each platform, including: the number of valid fields, the total number of fields, and calculates the field integrity rate of each platform, specifically according to the formula: field integrity rate = number of valid fields / total number of fields; Step A2: The data integration module extracts the error data of each platform from the key behavior data of each platform, including the number of errors and the total number of records, and calculates the error rate of each platform according to the formula: error rate = 1-(number of errors / total number of records); Step A3: extract the time data of each platform from the key behavior data of each platform, including: average delay time, maximum tolerable delay time, and calculate the average time delay from the generation to the completion of the collection of each platform data, specifically according to the formula: average time delay = 1-(average delay time / maximum tolerable delay time); The data integration module calculates the quality of key behavior data on each platform based on the field completeness rate, error rate, and average time delay of key behavior data on each platform; The specific steps to implement the repetition rate of key behavior data on different platforms are as follows: Step B1: The data integration module extracts the click behavior data of each platform from the key behavior data of each platform obtained; Step B2: The data integration module calculates the cross-platform repetition rate of the key behavior data of each platform through the repeated record key behavior data set of each platform, specifically according to the formula: repetition rate = (number of repeated entries / total number of entries) × 100%; The data integration module determines the external condition coefficient of key behavior data of each platform through weighted average calculation. The specific steps are as follows: Step C1: The data integration module determines the dimensions of external conditions: geographic location, device type, and time period; Step C2: The data integration module quantifies external condition factors and determines the external condition values, as follows: Analyze the conversion rate of click behavior data in each region through historical key behavior data and quantify the geographic location data; the mobile terminal has a higher value because the user operation is direct and the possibility of false clicks is low; the PC terminal has a higher probability of script brushing and the value is slightly lower. Combined with the time distribution law of clicks, the time period data is quantified; The quantitative values ​​of different external conditions of each platform are calculated according to the quantitative data of each external condition, specifically according to the formula: the amount of data of each external condition × the proportion of the corresponding external condition data; Step C3: The data integration module calculates the overall external condition coefficient of each platform through the weight values ​​of different external conditions; The data integration module calculates the data characteristic coefficient of each platform by weighting the key behavior data quality, key behavior data repetition rate and external condition coefficient of each platform; The data integration module sets different characteristic coefficient thresholds according to the data characteristic coefficients of each platform, divides the platform data into different categories, and applies different data cleaning and conversion methods to each category of data; sets the characteristic coefficient threshold Y, when the data characteristic coefficient of the platform is greater than or equal to Y, this type of key behavior data adopts a simple cleaning conversion method; the specific steps are as follows: Step S1: Delete duplicate records by checking unique identifiers (such as ID, timestamp, etc.) in key behavior data; Step S2: convert the dates, numbers, and character strings in the key behavior data into a unified format; Step S3: Use mean filling to fill in the missing key behavior data; When the data characteristic coefficient of the platform is less than or equal to Y, such key behavior data needs to be deeply cleaned and converted; the specific steps are as follows: Step D1: The local outlier factor anomaly detection algorithm based on machine learning identifies abnormal key behavior data that does not conform to the normal pattern through multi-dimensional analysis, and marks it as the part that needs to be cleaned or eliminated; Step D2: Use regression interpolation algorithm to infer missing values ​​based on the relationship between known key behavior data points and fill in the missing key behavior data.

4. The cross-platform crowd delivery data tracking and monitoring system based on information flow according to claim 3 is characterized in that; After the data integration module completes data cleaning and conversion, it is necessary to integrate and summarize the key behavior data from different platforms. The specific steps are as follows: Step F1: The data integration module removes duplicate records in the key behavior data of different platforms through the identifiers contained in the key behavior data set of each platform; Step F2: The data integration module maps fields of key behavior data that use different fields to represent the same meaning on different platforms; Step F3: The data integration module aggregates the cross-platform user behavior data by user ID.

5. The cross-platform crowd delivery data tracking and monitoring system based on information flow according to claim 4 is characterized by: The data analysis and prediction module applies time series analysis to remove the trend component of key behavioral data through first-order differences; The data analysis and prediction module uses the ARIMA model to model the key behavioral data after stabilization and calculate the model parameters.

6. A cross-platform crowd delivery data tracking and monitoring method based on information flow, characterized by: Step 1: Extract key behavioral data related to delivery from each platform; Step 2: Determine the data characteristic coefficient of each platform through weighted calculation of the key behavior data of different platforms, the duplication between key behavior data, and the influence of different external conditions. Perform different data cleaning and conversion methods on different key behavior data through the data characteristic coefficient, and summarize the key behavior data across platforms; Step 3: Analyze the integrated key behavior data through differential analysis and time series to predict future delivery effects.

Citation Information

Cited By

  • Third-party advertisement plan management method and system for multiple delivery channels

    CN121280092A