E-commerce live broadcast data drainage management optimization method and system based on big data
By collecting user behavior data in e-commerce live broadcasts, identifying abnormal traffic using DTW and Louvain algorithms, and optimizing advertising delivery with Bayesian formulas, the problem of false traffic interference is solved, and the accuracy and efficiency of advertising delivery is improved.
Patent Information
- Application Number
- CN202510491166.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-04-18
AI Technical Summary
There are fake traffic such as robot brushing volume and IP address fake in e-commerce live broadcasts, resulting in reduced advertising delivery efficiency and inability to accurately distinguish between real users and brushing volume users.
By obtaining the user click behavior sequence, using the DTW algorithm to match the trajectory, building the user feature matrix and clustering, using the Louvain algorithm to analyze IP region jumps, combining Bayesian formulas to calculate the path probability value, and optimizing advertising delivery traffic.
Accurately identify and eliminate abnormal traffic, distinguish real users from flash users, improve the accuracy and ROI of advertising delivery, and optimize the advertising delivery effect.
Smart Images

Figure CN120434458A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of live broadcast data management optimization, and more specifically, to an e-commerce live broadcast data diversion management optimization method and system based on big data. Background Art
[0002] In recent years, with the popularization of mobile Internet, the rise of social e-commerce and the digital transformation of consumer shopping behavior, e-commerce live streaming has become a core channel for interaction between brands and consumers. E-commerce live streaming data diversion refers to the process of guiding potential consumers to the live broadcast room or product page by analyzing user behavior data generated during the live broadcast, such as viewing time, interaction frequency and click-through conversion, and combining it with algorithmic marketing strategies, thereby increasing the traffic scale and conversion efficiency of the live broadcast room and achieving sales growth.
[0003] For example, the invention patent announcement with announcement number: CN117714722A discloses a data analysis method and system for live-streaming shopping on an e-commerce platform. The system receives verification requests and initial bonuses from target anchors sent by willing sellers, compares them with a pre-stored blacklist, and conducts a potential risk assessment on the target anchor based on live-streaming record data. A relationship map is established for target anchors with potential risk tags, and whether there are abnormalities is determined based on the relationship map. A model for monitoring the authenticity of live-streaming traffic is established, and first and second judgment criteria are established for anchors with potential risk tags and dangerous tags, respectively, for live-streaming monitoring. If false traffic is detected, the target anchor is added to the blacklist, and a warning report is sent to the willing seller. This achieves efficient and accurate identification and response to false traffic, thereby improving the credibility and user satisfaction of the live-streaming shopping platform.
[0004] For example, the invention patent announcement with announcement number: CN112019871B discloses an intelligent management platform for live e-commerce content based on big data, including a live content segmentation and classification module, a video frame position matching module, a playback rule database, a playback rule selection module, a video screening and playback module, and a playback rule intelligent recommendation module. The present invention segments the complete live playback video of the product according to the content and performs keyword annotation to form a product feature keyword video list, and obtains the video frame position matching the keyword from the complete live playback video of the product according to the manually input product keyword through the video frame position matching module, and manually selects a certain playback rule for playback, thereby realizing intelligent management of the live video content of the product, having the characteristics of strong operability, making up for the poor operability, low efficiency and low matching degree problems caused by manually adjusting the video progress bar, improving the adjustment efficiency, and enhancing the consumer's viewing experience of watching live broadcasts.
[0005] The above disclosed technical solutions have at least the following technical problems:
[0006] E-commerce live streaming has become a core means of product promotion, but the live streaming room is exposed to fake traffic such as robots and forged IP addresses, making it impossible to accurately distinguish between real users and fake users, resulting in reduced advertising efficiency. Summary of the Invention
[0007] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides an e-commerce live broadcast data traffic management optimization method and system based on big data, which solves the problem of false traffic interference in the existing live broadcast traffic management by integrating user behavior feature extraction, abnormal traffic detection, community brushing analysis and path probability correction technology.
[0008] To achieve the above object, the present invention provides the following technical solutions:
[0009] The big data-based e-commerce live streaming data diversion management optimization method includes the following steps: obtaining user click behavior sequences and using the DTW algorithm to perform trajectory matching on the sequences to obtain a first abnormal traffic set; constructing a user feature matrix based on the first abnormal traffic set, and screening it according to a clustering algorithm to obtain a second abnormal traffic set; constructing a graph model based on the second abnormal traffic set, and using the Louvain algorithm to analyze user IP regional jumps to obtain a list of brushing community markers; constructing a user conversion path set, and calculating the first probability value of each path based on the brushing community marker list; and correcting the advertising delivery traffic according to the first path probability to obtain optimized diversion decision instructions.
[0010] In a preferred embodiment, the method of obtaining a user click behavior sequence and using a DTW algorithm to perform trajectory matching on the sequence to obtain a first abnormal traffic set is specifically as follows: obtaining a user click behavior sequence and constructing a trajectory time series graph; performing feature extraction on the trajectory time series graph based on a preset neural network model to obtain a user behavior density feature sequence; using a DTW algorithm to match the user behavior density feature sequence with a preset baseline trajectory; performing data analysis on the matching results, identifying abnormal sequences, and constructing a first abnormal traffic set.
[0011] In a preferred embodiment, the method of constructing a user feature matrix based on the first abnormal traffic set and screening it according to a clustering algorithm to obtain a second abnormal traffic set is specifically as follows: constructing a user feature matrix based on the first abnormal traffic set, and using a clustering algorithm to density cluster the user feature matrix based on a preset density threshold; merging clusters whose spatial distance is less than a preset first threshold, and marking clusters containing a number of users greater than a preset second threshold as second abnormal traffic to obtain a second abnormal traffic set.
[0012] In a preferred embodiment, a graph model is constructed based on the second abnormal traffic set, and the user IP regional jump is analyzed using the Louvain algorithm to obtain a list of markers for brushing communities. Specifically, each user in the second abnormal traffic set is used as a node to construct a graph model; the Louvain algorithm is used to divide the graph model into communities, and the communities are identified as abnormal to obtain a list of markers for brushing communities.
[0013] In a preferred embodiment, the Louvain algorithm is used to divide the graph model into communities, and the communities are identified as abnormal to obtain a list of community markers for inflated traffic. Specifically, the Louvain algorithm is used to perform initial community division on the graph model to obtain several communities; the geographical distances between all IPs in the community are obtained, and the maximum geographical distance and the average geographical distance within the community are calculated, and the ratio of the maximum geographical distance to the average geographical distance is used as the regional jump index; the modularity of each community is calculated, and the modularity of each community is corrected according to the regional jump index to obtain the corrected modularity; according to the corrected modularity, the community merging strategy is adjusted to obtain several adjusted communities; the adjusted communities are identified as abnormal to obtain a list of community markers for inflated traffic.
[0014] In a preferred embodiment, the method of constructing a user conversion path set and calculating the first probability value of each path based on the brushing community tag list is as follows: obtaining the timestamp of the user's behavior in the live broadcast room, and constructing a behavior chain sequence according to a preset time sequence; extracting the key nodes in the behavior chain sequence to obtain a user conversion path set; dividing the user conversion path set based on the brushing community tag list to obtain a real path subset and a brushing path subset; using a preset kernel density calculation formula to perform feature calculation on the real path subset and the brushing path subset to obtain a real path feature value and a brushing path feature value; calculating the ratio of the real path feature value to the brushing path feature value, and calculating the first probability value of each path according to the Bayesian formula.
[0015] In a preferred embodiment, the advertising delivery traffic is screened according to the first probability value to obtain an optimized diversion decision instruction, specifically: obtaining advertising delivery channel data, and calculating the first traffic quality evaluation value of all paths under each advertising delivery channel according to the first path probability; screening the delivery traffic according to the first traffic quality evaluation value to obtain a screened advertising traffic pool; optimizing the diversion decision according to the screened advertising traffic pool to obtain an optimized diversion decision instruction.
[0016] Also provided is an e-commerce live broadcast data diversion management optimization system based on big data, including a data acquisition module, an anomaly screening module, a community identification module, a probability calculation module and a traffic screening module: the data acquisition module is used to obtain the user click behavior sequence and use the DTW algorithm to match the sequence trajectory to obtain a first abnormal traffic set; the anomaly screening module is used to build a user feature matrix based on the first abnormal traffic set, and screen it according to the clustering algorithm to obtain a second abnormal traffic set; the community identification module is used to build a graph model based on the second abnormal traffic set, and use the Louvain algorithm to analyze the user IP regional jump to obtain a list of brushing community tags; the probability calculation module is used to build a user conversion path set, and calculate the first probability value of each path based on the brushing community tag list; the traffic screening module is used to screen the advertising delivery traffic according to the first probability value to obtain the optimized diversion decision instruction. The technical effects and advantages of the e-commerce live broadcast data diversion management optimization method and system based on big data of the present invention:
[0017] 1. The present invention collects user click behavior sequences and uses the dynamic time warping (DTW) algorithm for trajectory matching to accurately identify and eliminate abnormal traffic; secondly, by constructing a user feature matrix and using a clustering algorithm for screening, the quality of traffic is further optimized. By density clustering abnormal traffic, it can effectively distinguish between real users and inflated users, helping advertisers better understand the characteristics of user groups, thereby improving the accuracy of target audiences. In addition, the Louvain algorithm is used for community identification, and the user IP regional jump is analyzed to further accurately mark the inflated communities, effectively avoiding the impact of inflated behavior on the effectiveness of advertising.
[0018] 2. This method constructs a set of user conversion paths, calculates the first probability of each path based on a list of social media markers for inflated traffic, and accurately predicts the path using the Bayesian formula. This in-depth analysis not only helps identify high-quality conversion paths but also effectively evaluates the traffic quality of each path, optimizing the effectiveness of advertising. Furthermore, by filtering advertising traffic based on the first path probability, optimized traffic flow decision instructions can accurately match high-efficiency traffic, thereby improving the ROI of advertising. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is a flow chart of the e-commerce live streaming data traffic management optimization method based on big data of the present invention.
[0020] Figure 2 This is a structural diagram of the e-commerce live broadcast data traffic management optimization system based on big data of the present invention. DETAILED DESCRIPTION
[0021] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0022] It should be noted that the acquisition and processing of all information or data in the present invention are carried out in compliance with the relevant national data protection laws and policies and based on the authorization given by the authorized administrator.
[0023] Example 1, Figure 1 The present invention provides an e-commerce live broadcast data diversion management optimization method based on big data, which includes the following steps:
[0024] S1, obtain the user click behavior sequence and use the DTW algorithm to match the sequence trajectory to obtain the first abnormal traffic set;
[0025] In this example, we obtain the user click behavior sequence and use the DTW algorithm to match the sequence trajectory to obtain the first abnormal traffic set, which is specifically:
[0026] Get the user click behavior sequence and construct a trajectory time series graph:
[0027] Based on the preset neural network model, the trajectory time series graph is extracted to obtain the user behavior density feature sequence;
[0028] The DTW algorithm is used to match the user behavior density feature sequence with the preset benchmark trajectory;
[0029] The matching results are analyzed to identify abnormal sequences and construct a first abnormal traffic set.
[0030] It should be noted that the user's click, stay, and jump behavior sequences in the live broadcast room are collected, and a trajectory time series diagram is generated according to the timestamp. For example, a user's behavior sequence in the live broadcast room is [enter the live broadcast room (10:00) → click on the product link (10:02) → collect the product (10:05) → leave the live broadcast room (10:08)]; the input layer of the preset neural network model is the trajectory time series diagram, and the hidden layer includes a time series convolution layer, an attention mechanism layer, and a fully connected layer. The time series convolution layer extracts click frequency features; the attention mechanism layer is used to extract high-frequency operation nodes, such as flash sale buttons, repeated clicks, etc.; the fully connected layer outputs a user behavior density feature sequence.
[0031] The specific calculation formula for matching the user behavior density feature sequence with the preset benchmark trajectory using the DTW algorithm is as follows:
[0032]
[0033] Among them, D is the matching distance, π is the regular path, S i is the user behavior density feature sequence, R j is the preset reference trajectory.
[0034] Furthermore, by constructing user click behavior sequences and generating trajectory time series graphs, we provide detailed and organized user behavior data for subsequent analysis. This step ensures comprehensive capture of user behavior and lays the foundation for in-depth data analysis.
[0035] Next, feature extraction is performed on the trajectory time series graph based on a pre-set neural network model, further enhancing the ability to identify user behavior patterns. By extracting feature sequences of user behavior intensity, the intensity and behavioral patterns of user interactions during live broadcasts can be accurately captured, thereby better reflecting their true interests and activity levels. Compared to traditional single-behavior analysis methods, this feature extraction method can effectively improve the ability to identify complex user behaviors and reduce misjudgments.
[0036] One of the core advantages of this method is its use of the DTW algorithm to match user behavior intensity feature sequences with pre-set baseline trajectories. The DTW algorithm can handle nonlinear alignment issues in time series, enabling precise matching and comparison of user behaviors and identifying anomalous sequences that significantly deviate from standard behavior patterns. This approach effectively distinguishes between real users and fraudulent users, eliminates anomalous traffic that doesn't conform to expected behavior patterns, and ensures the authenticity and validity of traffic data.
[0037] Ultimately, through in-depth data analysis of the matching results, we can identify abnormal behavior sequences and construct the first abnormal traffic set, thereby providing reliable data support for subsequent traffic screening and conversion path optimization, accurately identifying abnormal traffic, minimizing the interference of invalid traffic, and providing a solid data basis for advertising decisions, ensuring the effectiveness of advertising and the improvement of ROI.
[0038] S2, constructing a user feature matrix based on the first abnormal traffic set, and screening it according to the clustering algorithm to obtain a second abnormal traffic set;
[0039] In this example, a user feature matrix is constructed based on the first abnormal traffic set, and is screened using a clustering algorithm to obtain a second abnormal traffic set, specifically:
[0040] Constructing a user feature matrix based on the first abnormal traffic set, and using a clustering algorithm to perform density clustering on the user feature matrix based on a preset density threshold;
[0041] The clusters whose spatial distance is less than a preset first threshold are merged, and the clusters whose number of users is greater than a preset second threshold are marked as second abnormal traffic, to obtain a second abnormal traffic set.
[0042] It should be noted that the user feature matrix includes behavior density, operation type distribution and device fingerprint, among which behavior density is the number of user operations per unit time; operation type distribution is the ratio of clicks / collections / purchases; and device fingerprint is the mark of multiple account logins on the same device.
[0043] When screening according to the clustering algorithm, the parameters are set as follows: the neighborhood radius is 0.5, the minimum number of samples is 5, if the cluster center distance is less than 0.3, they are merged into the same cluster; when the number of users in the cluster is >100, it is determined to be a large-scale brushing cluster, and the second abnormal traffic set is generated.
[0044] It's important to note that by constructing a user feature matrix based on the first abnormal traffic set, the process integrates the behavioral characteristics of different users, providing a comprehensive view of user characteristics for subsequent cluster analysis. A clustering algorithm is then used to perform density clustering on the user feature matrix, grouping users based on a preset density threshold. This density clustering method effectively identifies groups of users with similar behavioral patterns and distinguishes these groups from other groups. Compared to traditional classification methods, density clustering better handles irregular data distributions, avoiding misclassification issues caused by improper threshold settings or uneven data distribution. Furthermore, by merging clusters with a spatial distance below a preset first threshold, the user group segmentation can be further refined, ensuring that users with similar behavioral patterns and frequent interactions are correctly classified. This strategy helps avoid confusing legitimate users with users engaged in fraudulent activity, ensuring accurate traffic analysis. Furthermore, clusters containing users greater than a preset second threshold are labeled as second abnormal traffic, effectively identifying fraudulent activity. These clusters usually have obvious characteristic differences, such as user concentration and abnormal behavior. Through this screening method, invalid traffic can be further eliminated and more accurate traffic data can be provided for advertising delivery.
[0045] S3: Build a graph model based on the second abnormal traffic set and use the Louvain algorithm to analyze the user IP regional jumps to obtain a list of markers for inflated traffic communities.
[0046] In this example, a graph model is constructed based on the second abnormal traffic set, and the Louvain algorithm is used to analyze the user IP regional jumps to obtain a list of markers for brushing communities, specifically:
[0047] Build a graph model by taking each user in the second abnormal traffic set as a node;
[0048] The Louvain algorithm is used to divide the graph model into communities, identify abnormalities in the communities, and obtain a list of marked communities with fake traffic.
[0049] Among them, each user in the second abnormal traffic set is used as a node to build a graph model, and the edges of the graph model are constructed based on the association relationships such as shared IP, equipment, payment account, etc. between users.
[0050] In this example, the Louvain algorithm is used to divide the graph model into communities and identify anomalies in the communities to obtain a list of community markers for fake traffic. Specifically:
[0051] Use the Louvain algorithm to perform initial community division on the graph model and obtain several communities;
[0052] Obtain the geographical distance between all IPs in the community, calculate the maximum geographical distance and average geographical distance within the community, and use the ratio of the maximum geographical distance to the average geographical distance as the regional jump index;
[0053] Calculate the modularity of each community and correct it according to the regional jump index to obtain the corrected modularity;
[0054] According to the revised modularity, the community merging strategy is adjusted to obtain several adjusted communities;
[0055] Identify abnormalities in the adjusted communities and obtain a list of labeled communities with fake traffic.
[0056] The modularity of each community is corrected according to the regional jump index to obtain the corrected modularity. The specific calculation formula is as follows:
[0057] Q new =Q-λ·T
[0058] Among them, Q new is the modified modularity, Q is the modularity of each community, λ is the preset weight coefficient, and T is the regional jump index.
[0059] It should be noted that, first, each user in the second abnormal traffic set is used as a node in the graph model, and the connections between users can be constructed based on their behavior patterns or other relevant features. In this way, the relationships and interactions between users can be systematically captured, thereby providing strong data support for community division and anomaly identification; in addition, the Louvain algorithm is an efficient community detection algorithm that can automatically divide communities based on the connection strength of nodes in the graph, ensuring a detailed division of user groups. Through this community division, potential brushing communities can be identified, that is, those user groups with abnormal behavior patterns and frequent regional hopping. Compared with traditional rule screening methods, this method has stronger adaptability and flexibility and can cope with complex and changing user behaviors.
[0060] Furthermore, by calculating the geographic distance between IP addresses within a community, we obtain the maximum and average geographic distances, and use the ratio of these distances as the regional jump index, effectively identifying unnatural patterns in user behavior. The regional jump index accurately reflects the IP hopping phenomenon in brush traffic, helping to identify malicious users who attempt to evade monitoring by changing their geographic location. By correcting the modularity and adjusting the community merging strategy based on the regional jump index, we can optimize the community segmentation results and further improve the accuracy of identifying brush traffic communities.
[0061] S4, constructing a user conversion path set and calculating the first probability value of each path based on the list of brushing community tags;
[0062] In this example, a user conversion path set is constructed, and the first probability value of each path is calculated based on the list of brushing community tags, specifically:
[0063] Obtain the timestamp of the user's behavior in the live broadcast room and construct a behavior chain sequence according to the preset time sequence;
[0064] Extract key nodes in the behavior chain sequence to obtain a set of user conversion paths;
[0065] Divide the user conversion path set based on the list of fake traffic community tags to obtain the real path subset and the fake traffic path subset;
[0066] The preset kernel density calculation formula is used to calculate the characteristics of the real path subset and the brush volume path subset to obtain the real path characteristic value and the brush volume path characteristic value;
[0067] Calculate the ratio of the true path characteristic value to the brush path characteristic value, and calculate the first probability value of each path according to the Bayesian formula.
[0068] Among them, the preset kernel density calculation formula is:
[0069]
[0070] Among them, f s (x) is the true path feature value, n is the number of true paths in the true path subset, h is the preset bandwidth parameter, K is the kernel function, x i is the eigenvalue of the i-th true path in the true path subset, and x is the feature dimension.
[0071] It's important to note that constructing a set of user conversion paths and calculating the first probability value for each path based on a list of fake traffic community markers can achieve precise traffic screening and optimized advertising delivery in e-commerce live streaming. This method, through detailed behavioral chain sequence analysis, combined with a list of fake traffic community markers and a Bayesian algorithm, can effectively identify the conversion paths of real users and fake traffic users, thereby improving the authenticity, effectiveness, and precision of traffic delivery.
[0072] First, by capturing the timestamps of each user's actions within the livestream and constructing a behavior chain sequence according to a preset chronological order, this step accurately records all of each user's interactive behaviors during the livestream. Building a behavior chain by following the timestamp sequence comprehensively reflects the user's behavioral trajectory, capturing every important milestone in the livestream, such as clicks, viewing time, and interactions, providing rich data support for building conversion paths.
[0073] Next, by extracting key nodes from the behavioral chain sequence, we can effectively transform user behaviors into a set of conversion paths. Extracting key nodes helps accurately identify the behaviors that are crucial to a user's conversion process, thereby constructing representative conversion paths. These conversion paths not only help identify the entire user journey from viewing to purchasing, but also distinguish between different user groups. This is particularly true in e-commerce live streaming, where accurate identification of conversion paths is crucial for advertising.
[0074] On this basis, combined with the list of markers for fake traffic communities, we can segment user conversion paths into a subset of real paths and a subset of fake traffic paths. By segmenting conversion paths, we can effectively filter out fake traffic.
[0075] Then, a pre-set kernel density calculation formula is used to calculate the characteristics of the subset of real and fake paths. This method can further improve the accuracy of traffic analysis. The kernel density estimation method can estimate the probability density based on user behavior characteristics and calculate the characteristic value of each path. The characteristic values of real and fake paths can reflect the differences between different paths, thus better identifying the boundary between normal and fake paths.
[0076] Ultimately, by calculating the ratio of the true path characteristic value to the inflated path characteristic value and then calculating the first probability value for each path using the Bayesian formula, we can assign a precise conversion probability value to each user path based on traffic characteristics. This step uses the Bayesian algorithm to combine prior knowledge and path characteristic values to determine the authenticity probability of each path in the current data environment. This method effectively distinguishes truly effective user paths from inflated paths, thereby optimizing advertising placement decisions.
[0077] S5: Modify the advertising delivery traffic according to the first probability value to obtain an optimized traffic diversion decision instruction.
[0078] In this example, the advertising traffic is filtered according to the first path probability to obtain the optimized traffic diversion decision instructions, specifically:
[0079] Acquire advertisement delivery channel data, and calculate the first traffic quality evaluation value of all paths under each advertisement delivery channel based on the first path probability;
[0080] Filtering the delivered traffic according to the first traffic quality evaluation value to obtain a filtered advertising traffic pool;
[0081] The traffic diversion decision is optimized based on the screened advertising traffic pool to obtain the optimized traffic diversion decision instruction.
[0082] It should be noted that the advertising channel data is pulled in real time from the API of advertising platforms (such as ByteDance and Tencent Advertising), including channel names (such as Douyin Live, Kuaishou Store, and Taobao Direct); the collection of user conversion paths under each channel (such as "homepage → product details page → payment completed"); and indicators such as exposure, clicks, conversion rate, and average stay time corresponding to the path.
[0083] The first flow quality evaluation value formula is as follows:
[0084]
[0085] Among them, r c is the first flow quality evaluation value of channel c, t is the number of paths in channel c, p i is the first probability value of the i-th path, l i is the preset historical conversion volume of the ith path under channel c, l c is the total conversion volume of all paths under channel c.
[0086] It's important to note that by obtaining ad delivery channel data and calculating a traffic quality assessment based on the first-path probability for each path, the effectiveness of each ad path can be quantified. This step helps advertisers identify traffic paths with higher conversion potential from massive amounts of traffic. Because each path's probability reflects its authenticity and conversion performance, advertisers can accurately determine which paths will deliver the best return on investment based on the probability. Furthermore, by calculating and analyzing the first-path quality assessment, advertisers can filter the delivered traffic and effectively eliminate low-quality traffic. This screening process ensures that every path in the ad traffic pool has a high conversion potential, effectively avoiding unnecessary advertising costs. This precise screening not only increases advertising targeting precision but also improves ad exposure, allowing advertising budgets to be more efficiently invested in potential, high-quality user groups. Optimizing traffic diversion decisions based on this filtered ad traffic pool ensures more effective advertising decisions.
[0087] Example 2, Figure 2 The present invention provides an e-commerce live streaming data diversion management and optimization system based on big data, including a data acquisition module, an anomaly screening module, a community identification module, a probability calculation module, and a traffic screening module:
[0088] A data acquisition module is used to obtain the user click behavior sequence and use the DTW algorithm to perform trajectory matching on the sequence to obtain the first abnormal traffic set;
[0089] An abnormality screening module, used to construct a user feature matrix based on the first abnormal traffic set, and screen it according to a clustering algorithm to obtain a second abnormal traffic set;
[0090] The community identification module is used to build a graph model based on the second abnormal traffic set and use the Louvain algorithm to analyze the user IP regional jump to obtain a list of community markers for inflating traffic;
[0091] A probability calculation module is used to construct a set of user conversion paths and calculate the first probability value of each path based on the list of community tags for inflated traffic;
[0092] The traffic screening module is used to screen the advertising delivery traffic according to the first probability value to obtain an optimized traffic diversion decision instruction.
[0093] The above embodiments may be implemented in whole or in part through software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments may be implemented in whole or in part in the form of a computer program product.
[0094] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0095] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0096] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0097] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. An e-commerce live streaming data diversion management optimization method based on big data, characterized by: The following steps are involved: Obtain the user click behavior sequence and use the DTW algorithm to match the sequence trajectory to obtain the first abnormal traffic set; Constructing a user feature matrix based on the first abnormal traffic set, and screening it according to a clustering algorithm to obtain a second abnormal traffic set; A graph model is constructed based on the second abnormal traffic set, and the Louvain algorithm is used to analyze the user IP regional jumps to obtain a list of markers for inflated traffic communities. Construct a set of user conversion paths and calculate the first probability value of each path based on the list of fake community tags; The advertising delivery traffic is screened according to the first probability value to obtain an optimized traffic diversion decision instruction.
2. The e-commerce live streaming data diversion management optimization method based on big data according to claim 1 is characterized in that: The user click behavior sequence is obtained and the DTW algorithm is used to perform trajectory matching on the sequence to obtain the first abnormal traffic set, which is specifically: Get the user click behavior sequence and construct a trajectory time series graph: Based on the preset neural network model, the trajectory time series graph is extracted to obtain the user behavior density feature sequence; The DTW algorithm is used to match the user behavior density feature sequence with the preset benchmark trajectory; The matching results are analyzed to identify abnormal sequences and construct a first abnormal traffic set.
3. The e-commerce live streaming data diversion management optimization method based on big data according to claim 2 is characterized in that: The user feature matrix is constructed based on the first abnormal traffic set, and is screened according to the clustering algorithm to obtain the second abnormal traffic set, which is specifically: Constructing a user feature matrix based on the first abnormal traffic set, and using a clustering algorithm to perform density clustering on the user feature matrix based on a preset density threshold; The clusters whose spatial distance is less than a preset first threshold are merged, and the clusters whose number of users is greater than a preset second threshold are marked as second abnormal traffic, to obtain a second abnormal traffic set.
4. The e-commerce live streaming data diversion management optimization method based on big data according to claim 3 is characterized in that: The graph model is constructed based on the second abnormal traffic set, and the Louvain algorithm is used to analyze the user IP regional jump to obtain a list of markers for the brushing community, specifically: Build a graph model by taking each user in the second abnormal traffic set as a node; The Louvain algorithm is used to divide the graph model into communities, identify abnormalities in the communities, and obtain a list of marked communities with fake traffic.
5. The e-commerce live streaming data diversion management optimization method based on big data according to claim 4 is characterized in that: The Louvain algorithm is used to divide the graph model into communities and identify abnormalities in the communities to obtain a list of community markers for inflated traffic, specifically: Use the Louvain algorithm to perform initial community division on the graph model and obtain several communities; Obtain the geographical distance between all IPs in the community, calculate the maximum geographical distance and average geographical distance within the community, and use the ratio of the maximum geographical distance to the average geographical distance as the regional jump index; Calculate the modularity of each community and correct it according to the regional jump index to obtain the corrected modularity; According to the revised modularity, the community merging strategy is adjusted to obtain several adjusted communities; Identify abnormalities in the adjusted communities and obtain a list of labeled communities with fake traffic.
6. The e-commerce live streaming data diversion management optimization method based on big data according to claim 5 is characterized in that: The user conversion path set is constructed, and the first probability value of each path is calculated based on the list of brushing community tags, specifically: Obtain the timestamp of the user's behavior in the live broadcast room and construct a behavior chain sequence according to the preset time sequence; Extract key nodes in the behavior chain sequence to obtain a set of user conversion paths; Divide the user conversion path set based on the list of fake community tags to obtain the real path subset and the fake path subset; The preset kernel density calculation formula is used to calculate the characteristics of the real path subset and the brush volume path subset to obtain the real path characteristic value and the brush volume path characteristic value; Calculate the ratio of the true path characteristic value to the brush path characteristic value, and calculate the first probability value of each path according to the Bayesian formula.
7. The e-commerce live streaming data diversion management optimization method based on big data according to claim 6 is characterized in that: The advertisement delivery traffic is screened according to the first probability value to obtain an optimized traffic diversion decision instruction, specifically: Acquire advertisement delivery channel data, and calculate the first traffic quality evaluation value of all paths under each advertisement delivery channel based on the first path probability; Filtering the delivered traffic according to the first traffic quality evaluation value to obtain a filtered advertising traffic pool; The traffic diversion decision is optimized based on the screened advertising traffic pool to obtain the optimized traffic diversion decision instruction.
8. An e-commerce live streaming data diversion management optimization system based on big data, applied to an e-commerce live streaming data diversion management optimization method based on big data as described in any one of claims 1-7, characterized in that: It includes data acquisition module, anomaly screening module, community identification module, probability calculation module and traffic screening module: A data acquisition module is used to obtain the user click behavior sequence and use the DTW algorithm to perform trajectory matching on the sequence to obtain the first abnormal traffic set; An abnormality screening module, used to construct a user feature matrix based on the first abnormal traffic set, and screen it according to a clustering algorithm to obtain a second abnormal traffic set; The community identification module is used to build a graph model based on the second abnormal traffic set and use the Louvain algorithm to analyze the user IP regional jump to obtain a list of community markers for inflating traffic; A probability calculation module is used to construct a set of user conversion paths and calculate the first probability value of each path based on the list of community tags for inflated traffic; The traffic screening module is used to screen the advertising delivery traffic according to the first probability value to obtain an optimized traffic diversion decision instruction.
Citation Information
Patent Citations
User group identification method and device, electronic equipment and storage medium
CN112839027A
Intelligent channel operation data analysis method, system and device and storage medium
CN115408586A
Sales planning method and system based on computer assistance
CN117557299A
Digital media advertisement effect evaluation system
CN117829914A
System for providing advertising campaign recommendation service based on unsupervised learning
KR102662258B1
Cited By
Data processing method and system based on business management platform
CN120952880A