A method and system for storing big data
By analyzing ad views and user behavior, and dynamically adjusting storage priorities, the problem of resource waste caused by malicious access in the ad delivery system was solved, achieving efficient utilization and rapid response of storage resources.
Patent Information
- Application Number
- CN202510399109.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-04-01
AI Technical Summary
Existing technologies cannot accurately identify malicious access in advertising delivery systems, leading to unreasonable allocation of storage resources and affecting system stability and efficiency.
By analyzing the time-series growth trend of ad views, page jump sequences, and user behavior, we can calculate potential anomaly indicators, distinguish between normal access and malicious behavior, dynamically adjust storage priorities, and optimize storage resource allocation.
Effectively identify and exclude malicious access, improve storage resource utilization, reduce resource waste, and enhance advertising data response speed and system stability.
Smart Images

Figure CN120335720B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of big data storage technology, specifically to a big data storage service method and system. Background Technology
[0002] Advertising placement requires real-time recording and analysis of user advertising interaction behaviors, such as ad access behavior and conversion behavior, and real-time adjustment of the ad content delivered to users based on user access behavior, thereby optimizing the accuracy of ad placement.
[0003] To efficiently process and maintain the stability of advertising delivery systems, existing methods are typically based on distributed storage principles. Storage resources are allocated according to user access habits, such as recent ad interaction frequency and dwell time. For example, relevant ad data with high recent user interaction frequency is stored on edge storage nodes, while other data with lower interaction frequency is stored in the cloud, thereby reducing access latency and improving the effective utilization of storage resources.
[0004] In the actual operation of big data-based storage service systems, there are sudden surges in access volume caused by sudden traffic and malicious attacks. Sudden traffic may be caused by sudden events, and compared with malicious access, the access behavior has storage value. For example, before and after shopping festivals, the number of ad views may increase. However, existing methods cannot accurately identify malicious access during surges in browsing, which affects the rational allocation of storage resources and brings the potential risk of abnormal storage service response. Summary of the Invention
[0005] To address the technical problem in existing technologies that fail to accurately identify malicious access during surges in browsing, thereby affecting the rational allocation of storage resources, the present invention aims to provide a big data storage service method and system, the specific technical solution of which is as follows:
[0006] This invention provides a method for storing big data, the method comprising:
[0007] Get the page jump sequence, page browsing time and access depth, and payment mark for each ad in each period of the historical time period when each user visits it once;
[0008] Analyze the growth trend of the number of visits to each ad in the current period to determine the growth period; based on the stability of the increase in the number of visits during the growth period, and combined with the deviation from the number of visits during the same period in historical periods, obtain the anomaly probability index; and determine the suspected anomaly period for each ad based on the anomaly probability index.
[0009] During the suspected abnormal period of each advertisement, the jump anomaly index for a single visit is obtained by analyzing the deviation of the page jump sequence length between a single visit and other visits in non-suspected abnormal periods, as well as the correlation of each jump. Based on the possible abnormal indicators of each visit of the user corresponding to a single visit in the current period, and the similarity between the page jump sequence and non-suspected abnormal periods, combined with the visit frequency, the user anomaly index for a single visit is obtained. Based on the user anomaly index and the jump anomaly index, the normal visits of each advertisement in the current period are determined.
[0010] In each storage node, the storage priority of each advertisement is obtained by combining the frequency of each advertisement being accessed by each user and the number of payment tokens, as well as the page browsing time and access depth when the advertisement is normally accessed; the storage content is then reallocated according to the storage priority of the advertisement.
[0011] Furthermore, the method for obtaining the growth period includes:
[0012] For any given ad, the total number of times the ad is accessed at each time step in the time sequence is taken as the access count at each time step;
[0013] In the current period, calculate the difference in visits to the ad at each time point compared to the previous time point, and use this difference as the growth index of the ad at each time point; filter out the time points where the growth index is less than 0, and record the remaining time points as the target time points;
[0014] The period consisting of consecutive target moments in the current cycle is designated as the growth period for this advertisement.
[0015] Furthermore, the method for obtaining the possible anomaly indicators includes:
[0016] For any given advertisement, each growth period is used as the analysis period in turn;
[0017] For any period outside the current period in the historical time period, by analyzing the distribution of the time period in the current period, the same distributed time period in that period is obtained; the difference between the average number of visits to the advertisement in the analysis period and the average number of visits to the advertisement in the same distributed time period is taken as the visit deviation between the analysis period and that period.
[0018] The average of the visit deviations between the analysis period and all other periods in the historical data is used as the historical visit deviation of the ad in the analysis period.
[0019] After calculating the difference in the growth index of the advertisement between every two adjacent moments during the analysis period, the sum of all the differences in the growth index is calculated and a negative correlation mapping is performed to obtain the stability of the advertisement's fluctuations during the analysis period.
[0020] By combining the average growth index, historical visit deviation, and volatility stability of the advertisement at all times during the analysis period, possible indicators of anomalies for the advertisement during the analysis period can be obtained.
[0021] Furthermore, the method for obtaining the jump anomaly indicator includes:
[0022] In a single visit, each pair of adjacent pages in the page jump sequence is taken as a jump group, and the page with the larger sequence number in the jump group is taken as the jump page; based on the distribution order of the jump groups in the page jump sequence, the jump pages of the jump groups are arranged to obtain the jump behavior sequence of a single visit.
[0023] For any visit to each ad during a suspected abnormal period, all visits to the ad corresponding to that visit during all non-suspected abnormal periods are taken as reference visits.
[0024] Each number in the redirection sequence of the access is taken as the target number; the total number of accesses with the target number in the reference access redirection sequence is taken as the number of times the target number is redirected; the total number of accesses with the target number in the reference access redirection sequence and the target number corresponds to the same redirect page is taken as the number of times the target number is redirected; the ratio of the number of times the target number is redirected to the number of times the target number is redirected is taken as the redirection correlation index of the target number.
[0025] By performing negative correlation mapping on the possible abnormal indicators of the suspected abnormal time period of the visit, we can obtain the possible normal analysis indicators of the visit.
[0026] The product of the jump association index of the target sequence number and the normal analysis possible index of the visit is used as the association feature index of the target sequence number in the visit; the mean of the association feature index of the sequence number in the jump behavior sequence of the visit is calculated and negative correlation mapping is performed to obtain the jump association abnormal index of the visit.
[0027] Calculate the difference between the length of the page jump sequence during this visit and the average length of all page jump sequences in the reference visits to obtain the abnormal jump count index for this visit;
[0028] By combining the redirection-related anomaly indicators and the redirection count anomaly indicators for this access, the redirection anomaly indicators for this access can be obtained.
[0029] Furthermore, the method for obtaining the user anomaly indicators includes:
[0030] The first two pages in the page jump sequence during a single visit are treated as one analysis group, and the subsequent pages in the page jump sequence are successively merged into the analysis group. Each time they are merged, a new analysis group is obtained.
[0031] For any ad during a suspected abnormal period, calculate the frequency of each analysis group of that ad in all non-suspected abnormal periods to obtain the fixed probability index of each analysis group; use the maximum fixed probability index of the ad as the similarity index of that ad.
[0032] The user corresponding to this access is designated as the analysis user; in the current period, the ratio of the total number of times the analysis user accesses this advertisement to the total number of times all users access this advertisement is used as the access frequency index of the analysis user; the mean of the possible abnormal indicators in all accesses of the analysis user is calculated to obtain the access abnormality index of the analysis user; the mean of the similar jump indicators in all accesses of the analysis user is calculated and negative correlation mapping is performed to obtain the jump order abnormality index of the analysis user.
[0033] By combining the analysis of user access frequency indicators, access anomaly indicators, and redirection order anomaly indicators, the user anomaly indicators for this access are obtained.
[0034] Furthermore, the method for determining normal access includes:
[0035] The product of the redirection anomaly indicator and the user anomaly indicator for a single visit is normalized and used as the malicious behavior indicator for a single visit; visits with malicious behavior indicators greater than the preset malicious threshold are considered malicious visits for each advertisement.
[0036] In the current period of each advertisement, all visits except malicious visits are treated as normal visits for each advertisement.
[0037] Furthermore, the method for obtaining the storage priority includes:
[0038] For any normal visit to any advertisement, the ratio of the number of times the user corresponding to the normal visit visits the advertisement in the current period to the total number of times the advertisement is displayed to the user is used as the visit interest index for the normal visit.
[0039] Multiply the average page browsing time during a normal visit by the average page visit depth to obtain the behavioral interest index of the normal visit; perform a negative correlation mapping on the malicious behavior index of the normal visit to obtain the behavioral reliability of the normal visit.
[0040] The effective index of the normal access is obtained by multiplying the reliability of the normal access behavior, the access interest index, and the behavior interest index.
[0041] In each storage node, the total number of times each advertisement is accessed by each user is taken as the node access score; the ratio of the total number of times the advertisement is accessed normally to the node access score is taken as the node access frequency of the advertisement.
[0042] Calculate the average of all valid metrics for normal access to the advertisement, the product of the number of paid tokens and the node access frequency, and obtain the storage priority of the advertisement in the storage node.
[0043] Furthermore, the reallocation of storage content based on the storage priority of the advertisement includes:
[0044] In each storage node, the storage priority of the advertisement is normalized to obtain the storage retention metric;
[0045] When the storage retention index of an ad exceeds the preset storage threshold, the relevant data of the corresponding ad will be retained in the storage node; otherwise, the relevant data of the ad will be transferred to cloud storage.
[0046] Furthermore, the method for determining the suspected abnormal time period includes:
[0047] Within all growth periods of each advertisement, the growth periods where the abnormality probability index exceeds the preset abnormality threshold are designated as suspected abnormal periods.
[0048] The present invention also provides a big data storage service system, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the big data storage service method described above.
[0049] The present invention has the following beneficial effects:
[0050] This invention analyzes the temporal growth trend of ad views within the current period, combined with the deviation of views and the stability of the increase in views during the same period in historical periods, to calculate possible anomaly indicators and initially screen for suspected abnormal periods with significant traffic surges, providing a time-based basis for subsequent differentiation of malicious traffic behavior. Within suspected abnormal periods, further correlation analysis of page jump sequences and length differences with non-abnormal period sequences are used to quantify the jump anomaly indicators of a single visit. Simultaneously, combined with abnormal user access behavior within the current period, a comprehensive judgment of normal access behavior is made, eliminating interference from malicious behavior on data storage resource allocation. Based on normal access behavior data, the invention integrates ad access status by each user, the number of payment tokens, page browsing duration, and access depth to dynamically calculate the storage priority of each ad, achieving real-time optimized allocation of edge storage node resources. This invention identifies abnormal access based on temporal behavior analysis and jump logic differences, and dynamically adjusts ad storage strategies based on the depth of user interaction during normal access, reducing resource waste caused by malicious behavior, improving the rapid response of effective ad data, and enhancing the effectiveness of storage resource utilization. Attached Figure Description
[0051] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0052] Figure 1 This is a flowchart of a big data storage service method provided in one embodiment of the present invention;
[0053] Figure 2 This is a schematic diagram of a distributed storage structure provided in one embodiment of the present invention. Detailed Implementation
[0054] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a big data storage service method and system proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0056] The following description, in conjunction with the accompanying drawings, details a specific solution for a big data storage service method and system provided by the present invention.
[0057] A distributed big data storage service based on advertising data, utilizing a distributed architecture that coordinates edge storage nodes with the cloud, please refer to [link / reference needed]. Figure 2 This illustration shows a schematic diagram of a distributed storage structure provided by an embodiment of the present invention. Edge storage nodes, such as CDN nodes and local servers, are deployed close to user terminals, forming the first layer of storage in the distributed network. When a user initiates an ad access request, the edge node can directly access high-priority data cached locally, reducing the response time to milliseconds and avoiding network latency when accessing the cloud across regions. Meanwhile, cloud storage carries massive amounts of low-frequency data, supports long-term analysis and global strategy generation, and reduces bandwidth costs by decreasing the frequency of cloud data transmission.
[0058] Dynamically prioritizing storage resources can improve resource utilization. Please refer to [link / reference]. Figure 1 The diagram illustrates a flowchart of a big data storage service method according to an embodiment of the present invention, which includes the following steps:
[0059] S1: Obtain the page jump sequence, page browsing time and access depth, and payment mark for each advertisement during each period in the historical time period when each user visits it once.
[0060] User data is transmitted to the nearest edge storage device for storage. Edge storage nodes store user access logs and advertising materials. In this embodiment, a 24-hour cycle is used, and user advertising interaction behavior logs are recorded at each edge storage point at each sampling time. This mainly includes the page jump sequence when each ad is accessed by the user, the browsing time on each page, the corresponding access depth, and the payment page where the ad appears is marked based on tags. Historical data is analyzed over a one-week period to facilitate subsequent screening of malicious traffic.
[0061] S2: Analyze the growth trend of the number of times each ad is accessed in the current period to determine the growth period; based on the stability of the increase in the number of times accessed during the growth period, and combined with the deviation from the number of times accessed during the same period in historical periods, obtain the abnormality probability index; based on the abnormality probability index, determine the suspected abnormal period for each ad.
[0062] During the storage service operation, there are malicious attacks that forge massive requests with abnormal traffic, and situations where users' normal access behavior causes a sudden surge in the number of visits to advertising pages due to factors such as promotional activities. Malicious attacks usually lead to the occupation and waste of system storage resources.
[0063] There are some differences between the two types of sudden traffic surges. When a surge in visits to a particular advertisement is caused by normal user behavior, the behavior still aligns with the user's normal browsing habits. For example, normal visits tend to occur during users' leisure time, such as lunch breaks, after get off work, or weekends. The growth rate is relatively slow and the time distribution is relatively even. In contrast, malicious attacks have irregular access times, and the growth rate is faster and the growth period is more concentrated.
[0064] First, by analyzing the growth trend of website traffic, preliminary screening is performed to identify abnormal periods of traffic surges. Preferably, in this embodiment of the invention, the method for obtaining growth periods includes:
[0065] For any given advertisement, the total number of times the advertisement is accessed at each time point in the time series is taken as the access volume at that time point. Within the current period, the difference between the access volume of the advertisement at each time point and the previous time point is calculated as the growth index of the advertisement at that time point. A growth index less than zero indicates a decreasing trend in access volume, and is not a period of malicious traffic. Therefore, times with a growth index less than 0 are filtered out, and the remaining times are recorded as the target times.
[0066] The target times are merged. If there are adjacent times that are both target times, they are merged to reflect the continuous possible access anomalies. The time period consisting of consecutive target times in the current period is used as the growth period of the advertisement. There may be malicious attack phases in the growth period.
[0067] Anomaly attacks differ from normal user behavior. If the number of visits during a single growth period differs significantly from the number of visits during the same period in historical cycles, and the increase in visits is rapid with low volatility, then the likelihood of an anomaly attack occurring during that period is higher. Preferably, in this embodiment of the invention, the method for obtaining anomaly probability indicators includes:
[0068] For any given advertisement, each growth period is sequentially used as an analysis period, and all growth periods are analyzed. For any period outside the current period in the historical timeframe, the distribution of the analysis period within the current period is used to obtain the co-distributed periods within that period. This means that the time series of both the historical and current periods are analyzed end-to-end, and the portions of the same time segment within the analysis period are selected as the co-distributed periods. For example, if the analysis period is from 0:00 to 1:00 in the current period, then the co-distributed segments are from 0:00 to 1:00 in each historical period.
[0069] Furthermore, the difference between the average number of visits to the ad during the analysis period and the average number of visits to the ad during the same period is taken as the visit deviation between the analysis period and that cycle. The difference in visit volume between weeks reflects the difference in visit patterns. Combining all historical cycles, the average of the visit deviations between the analysis period and all other cycles in the historical period is taken as the historical visit deviation degree of the ad during the analysis period. The higher the historical visit deviation degree, the more serious the deviation between the visit situation during the analysis period and the historical situation, and the higher the possibility of anomalies.
[0070] Furthermore, by analyzing the growth rate, the difference in the growth index of the advertisement between each two adjacent moments during the analysis period is calculated. The sum of all growth index differences is then applied using a negative correlation mapping to obtain the stability of the advertisement's fluctuations during the analysis period. This stability characterizes the degree of growth rate stability at different times; the smaller the overall difference, the more stable the growth rate. It should be noted that the negative correlation mapping is a technique well-known to those skilled in the art; however, the use of an inverse proportional form is not restricted here.
[0071] Finally, by combining the average growth index, historical visit deviation, and volatility stability of the advertisement across all moments in the analysis period, the possible anomaly indicators for the advertisement during the analysis period are obtained. In this embodiment of the invention, the average growth index, historical visit deviation, and volatility stability across all moments in the analysis period are multiplied together and normalized to obtain the possible anomaly indicators for the advertisement during the analysis period. It should be noted that normalization is a technique well-known to those skilled in the art, and the choice of normalization can be linear normalization or standard normalization, etc., and the specific normalization method is not limited here.
[0072] The higher the average growth index, the greater the increase in visits during the analysis period. The greater the historical visit deviation, the higher the degree of abnormality of the analysis period compared to historical cycles. The greater the stability fluctuation, the more stable the increase at different times, thus the higher the possibility of anomaly attacks.
[0073] Ultimately, malicious attack periods can be identified through threshold judgment. Suspected abnormal periods are then filtered based on anomaly probability indicators. In this embodiment, among all growth periods for each advertisement, growth periods with anomaly probability indicators greater than a preset anomaly threshold are considered suspected abnormal periods, while non-suspected abnormal periods are considered normal periods. The preset anomaly threshold can be set to 0.5, and the specific value can be adjusted by the implementer according to the specific implementation scenario; no restrictions are imposed here.
[0074] S3: During the suspected abnormal period of each advertisement, obtain the jump abnormality index for a single visit by analyzing the deviation of the page jump sequence length between a single visit and other visits in non-suspected abnormal periods, as well as the correlation of each jump; obtain the user abnormality index for a single visit by analyzing the possible abnormality index of each visit of the user corresponding to the single visit in the current period, the similarity between the page jump sequence and non-suspected abnormal periods, and the visit frequency; determine the normal visit of each advertisement in the current period based on the user abnormality index and the jump abnormality index.
[0075] During suspected abnormal periods, malicious behavior is analyzed and filtered out based on access patterns. Malicious access behavior is usually executed automatically by pre-programmed procedures, exhibiting a high degree of regularity and significantly differing from the access behavior of normal users. When normal users access advertisements, their page navigation process typically follows a certain flow or access habits; for example, a user's access habits might be: from the advertisement campaign page to the product details page, then to the order page, and finally to the payment page, etc. Malicious access usually targets a specific page frequently, and even if page navigation occurs, the navigation behavior differs significantly from the access habits of normal users. Therefore, by combining the differences in navigation behavior, normal user access behavior and malicious traffic can be distinguished.
[0076] First, anomaly analysis is performed on the page navigation behavior. Although user page navigation behavior has randomness, the order of page navigation in each visit follows a certain pattern, such as from the product details page to the add-to-cart page. Therefore, the higher the correlation between each navigation behavior in the visit to this advertisement, the more consistent the navigation behavior is with the user's browsing habits. Furthermore, the similarity in the number of navigations to normal time periods is considered to obtain anomaly indicators for the navigation analysis.
[0077] Preferably, in this embodiment of the invention, the method for obtaining the jump anomaly indicator includes:
[0078] First, the page jump sequence is broken down to facilitate subsequent analysis of each jump. Each pair of adjacent pages in the page jump sequence during a single visit is considered a jump group, and each jump group represents one jump process. The page with the higher sequence number in the jump group is designated as the jump page, which is also the destination page for each jump. Based on the distribution order of the jump groups in the page jump sequence, the jump pages of the jump groups are arranged to obtain the jump behavior sequence for a single visit.
[0079] For example, an advertisement might have 5 pages from its homepage to payment completion. A user might visit all pages sequentially from beginning to end, with a page jump sequence of [1.2.3.4.5], jumping 4 times in total, with jump groups of 1-2, 2-3, 3-4, 4-5, and each jump landing on pages 2, 3, 4, and 5 respectively. Alternatively, a user might repeatedly jump between the product details page and the review page, with a page jump sequence of [1.2.3.2.3.4.5], jumping 6 times, and each jump landing on pages 2, 3, 2, 3, 4, and 5 respectively. In this case, the two users have the same purpose in the first two jumps, indicating a behavioral correlation. Notably, there are also cases where a user exits immediately after clicking on the advertisement, in which case the page jump sequence is [1], with 0 jumps.
[0080] Then, for any single visit to each ad during a suspected abnormal period, we first analyze all visits to the ad corresponding to that visit during all non-suspected abnormal periods, using the ad's visits during normal periods as a reference. Then, we analyze each sequence number in the redirection behavior sequence of that visit as a target number, and sequentially analyze the destination page for each redirect.
[0081] The total number of accesses with the target sequence number in the reference access jump behavior sequence is taken as the same jump count for the target sequence number. In other words, it is the number of jump situations with the same target sequence number in all reference access jump sequences. For example, when the target sequence number is 2, the number of accesses with sequence number 2 in the reference access jump behavior sequence is taken as the same jump count.
[0082] The total number of visits with the same target number and corresponding redirect page within the jump behavior sequence of the statistical reference access is taken as the same behavior count for the target number, which is the number of visits with the same target number that also redirect to the same page within the jump behavior sequence. Finally, the ratio of the same behavior count for the target number to the same redirect count is taken as the jump correlation index for the target number, reflecting the degree of correlation between each jump situation when compared with normal access.
[0083] Furthermore, a negative correlation mapping was performed on the possible abnormal indicators of the suspected abnormal time period of the visit to obtain the possible normal analysis indicators of the visit. Combined with the time period anomaly analysis, if the possibility of anomaly in the time period corresponding to a single visit is smaller, that is, the normal analysis possible indicator is larger, it indicates that the contribution of the visit behavior to the analysis of normal visit behavior is higher.
[0084] Therefore, the product of the jump association index of the target sequence number and the normal analysis possibility index of the access is used as the association feature index of the target sequence number in the access, reflecting the association characteristics of each jump. Combining the association of all jump situations, the mean of the association feature index of the sequence number in the jump behavior sequence of the access is calculated, and a negative correlation mapping is performed to obtain the jump association anomaly index of the access. The smaller the association feature of multiple jump situations in the access, the lower the degree of association with normal jumps, and the more likely it is malicious behavior. Therefore, the larger the jump association anomaly index is.
[0085] Furthermore, the difference between the length of the page jump sequence during this visit and the average length of all page jump sequences in the reference visit is calculated to obtain the jump number anomaly index for this visit. The higher the overall difference in jump length, the more significant the degree of jump anomaly.
[0086] Finally, by combining the redirect association anomaly index and the redirection count anomaly index of the access, the redirection anomaly index of the access is obtained. In this embodiment of the invention, the product of the redirect association anomaly index and the redirection count anomaly index is used as the redirection anomaly index of the access. The larger the redirection anomaly index, the higher the likelihood that the access is malicious based on the redirection behavior analysis.
[0087] Secondly, anomalies in user access are analyzed. If a user with a single access visits the advertisement more frequently in the current period, and the degree of anomaly is higher during each access period, then the user is more likely to have abnormal access behavior in the current period. Furthermore, the greater the difference between the redirected page when the user accesses the advertisement in the current period and the access behavior when accessing the advertisement in normal periods, the greater the likelihood that the user considers this access behavior to be malicious.
[0088] Preferably, in this embodiment of the invention, the method for obtaining user anomaly indicators includes:
[0089] Considering that the ad interface settings are relatively fixed, the order in which users jump to certain pages is highly consistent. For example, there is only a single jump between the ad introduction page and the ad details page. In this case, the similarity of the user's page sequence should be high. Therefore, we consider that the order of page jumps during user visits is similar to that during normal periods, and analyze the situation where a single visit conforms to the user's access habits.
[0090] First, the first two pages in the page jump sequence during a single visit are grouped together as an analysis group. Subsequent pages in the page jump sequence are then merged into the analysis group in turn. Each merge results in a new analysis group. For example, if the page jump sequence is "abcd", then the analysis groups are "ab", "abc", and "abcd".
[0091] For any ad visit during a suspected anomaly period, the higher the frequency of the analysis group appearing during normal periods, the more normal the user's visit behavior is. Calculate the frequency of each analysis group appearing during all non-suspected anomaly periods for that visit, obtaining a fixed probability index for each analysis group's redirection. The highest fixed probability index for redirection in that visit is used as the similarity index for that visit. For multiple analysis groups, the highest frequency is taken as the probability of a fixed order of redirection. The higher the similarity index, the higher the similarity of user redirection habits.
[0092] The user corresponding to this access is designated as the analysis user. In the current period, the ratio of the total number of times the analysis user accesses this advertisement to the total number of times all users access this advertisement is used as the access frequency index of the analysis user. The higher the frequency of users accessing this advertisement, the higher the possibility of abnormal access by that user.
[0093] Furthermore, within the current period, the average of the possible abnormal indicators across all user accesses is calculated and analyzed to obtain the user's access anomaly indicators. Combining these access anomaly indicators, the more periods within the current period where a user's access is likely to be subject to abnormal attacks, the higher the likelihood of the user being abnormal.
[0094] Furthermore, within the current period, the mean of similar redirection metrics for all user visits is calculated and negatively correlated to obtain anomaly indicators for the redirection order of users. Combined with the analysis of similar redirection metrics for all user visits, the smaller the overall similar redirection metrics, the less similar the user's access behavior is to normal times, and the higher the possibility of anomaly.
[0095] Finally, by combining the analysis of the user's access frequency index, access anomaly index, and redirection order anomaly index, the user anomaly index for that access is obtained. In this embodiment of the invention, the product of the analyzed user's access frequency index, access anomaly index, and redirection order anomaly index is used as the user anomaly index for that access. The larger the user anomaly index, the higher the likelihood that the access is malicious based on user behavior analysis.
[0096] By combining malicious behavior analysis from both aspects, and based on user anomaly indicators and redirection anomaly indicators, the normal access to each advertisement in the current period is determined. In this embodiment of the invention, the product of the redirection anomaly indicator and the user anomaly indicator for a single access is normalized and used as the malicious behavior indicator for that single access. The higher the indicator, the more abnormal the access behavior. Accesses with malicious behavior indicators greater than a preset malicious threshold are considered malicious accesses to each advertisement. The preset malicious threshold is set to 0.6, and the specific value can be adjusted by the implementer.
[0097] In the current period of each advertisement, all visits except malicious visits are treated as normal visits for each advertisement, thus distinguishing between normal access behavior and malicious traffic attacks.
[0098] S4: In each storage node, the storage priority of each advertisement is obtained by combining the frequency of each advertisement being accessed by each user and the number of payment tokens, as well as the page browsing time and access depth when the advertisement is normally accessed; storage content is reallocated according to the storage priority of the advertisement.
[0099] The data cached on each edge storage node meets the requirement for rapid response, thus enabling priority analysis of all ads based on the efficiency generated by normal ad access. Higher-priority ads are cached on edge storage nodes closest to traffic police, allowing for rapid response to their ad delivery needs. Lower-priority ad-related cached data is moved from edge storage nodes to cloud storage, thereby fully utilizing storage resources.
[0100] The more frequently an ad is viewed by users and the higher its effectiveness, the greater the benefit of the ad. Inefficient ad views are typically short-lived, with users usually exiting immediately after viewing and exhibiting minimal redirection. Conversely, if an ad view involves multiple page navigations, and the user scrolls deeply and stays on the page for a long time, it indicates a higher level of interest in the ad's content and a greater likelihood of a purchase, meaning the ad view is more effective.
[0101] Preferably, in this embodiment of the invention, the method for obtaining storage priority includes:
[0102] For any normal visit to any advertisement, the ratio of the number of times the user accesses the advertisement in the current period to the total number of times the advertisement is displayed to the user is used as the access interest index for that normal visit. The more consistent the number of times the user clicks on the displayed advertisement, the higher the user's access interest in that advertisement.
[0103] Then, the average page browsing time during a normal visit is multiplied by the average page exploration depth to obtain the behavioral interest index of the normal visit. The longer the page is browsed and the greater the page exploration depth, the higher the level of interest in the advertisement.
[0104] By negatively mapping the malicious behavior indicators of the normal access, the reliability of the normal access behavior is obtained. The degree of malicious behavior reflects the reliability of each access in effective analysis. The smaller the malicious behavior indicator, the higher the reliability of the normal behavior.
[0105] Furthermore, the reliability of the normal visit, the visit interest index, and the behavior interest index are multiplied together to obtain the effective index of the normal visit. The higher the reliability of the visit, the visit interest index, and the behavior interest index, the higher the interest value generated by the user in this visit to the advertisement, and the greater the effectiveness of the advertisement.
[0106] In each storage node, the total number of times each advertisement is accessed by each user is used as the node access score. The ratio of the total number of normal visits to the node access score is used as the node access frequency of the advertisement. The proportion of the number of visits to all advertisements reflects the high frequency of visits to each user. The higher the node access frequency, the higher the possibility of being accessed efficiently again in the future.
[0107] Finally, the storage priority of the advertisement in the storage node is obtained by calculating the average of the effective metrics of all normal visits to the advertisement, the product of the number of paid tokens and the node access frequency. The higher the interest level that the corresponding advertisement in the node can generate, the higher the frequency of access and payment, and the higher the frequency among all advertisement visits, the higher the storage priority of the advertisement.
[0108] Based on the storage priority of each advertisement, the cached content of each storage node is adjusted. In this embodiment of the invention, the storage priority of the advertisement is normalized in each storage node to obtain a storage retention index. When the storage retention index of the advertisement is greater than the preset storage threshold, the relevant data of the corresponding advertisement is retained in the storage node. Otherwise, the relevant data of the advertisement is transmitted to the cloud storage. The storage content of the edge storage point can be dynamically allocated periodically, such as updating once per minute.
[0109] In summary, this invention analyzes the temporal growth trend of ad views within the current period, combines the deviation of views in the same period of historical periods and the stability of the increase in views, calculates possible anomaly indicators, and initially screens suspected abnormal periods with relatively abnormal traffic surges, providing a time-based basis for subsequent differentiation of malicious traffic behavior. Within suspected abnormal periods, further correlation analysis of page jump sequences and length differences with non-abnormal period sequences are used to quantify the jump anomaly indicators of a single visit. Simultaneously, combined with abnormal user access behavior within the current period, normal access behavior is comprehensively determined, eliminating interference from malicious behavior on data storage resource allocation. Based on normal access behavior data, the access status of ads by each user, the number of payment tokens, page browsing duration, and access depth are integrated to dynamically calculate the storage priority of each ad, achieving real-time optimized allocation of edge storage node resources. This invention identifies abnormal access based on temporal behavior analysis and jump logic differences, and dynamically adjusts the storage strategy for ads based on the user interaction depth of normal access, reducing resource waste caused by malicious behavior, improving the rapid response of effective ad data, and enhancing the effectiveness of storage resource utilization.
[0110] The present invention also provides a big data storage service system, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the big data storage service method described above.
[0111] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0112] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
Claims
1. A method for storing big data, characterized in that, The method includes: Get the page jump sequence, page browsing time and access depth, and payment mark for each ad in each period of the historical time period when each user visits it once; Analyze the growth trend of the number of visits to each ad in the current period to determine the growth period; based on the stability of the increase in the number of visits during the growth period, and combined with the deviation from the number of visits during the same period in historical periods, obtain the anomaly probability index; and determine the suspected anomaly period for each ad based on the anomaly probability index. During the suspected abnormal period of each advertisement, the jump anomaly index for a single visit is obtained by analyzing the deviation of the page jump sequence length between a single visit and other visits in non-suspected abnormal periods, as well as the correlation of each jump. Based on the possible abnormal indicators of each visit of the user corresponding to a single visit in the current period, and the similarity between the page jump sequence and non-suspected abnormal periods, combined with the visit frequency, the user anomaly index for a single visit is obtained. Based on the user anomaly index and the jump anomaly index, the normal visits of each advertisement in the current period are determined. In each storage node, the storage priority of each advertisement is obtained by combining the frequency of each advertisement being accessed by each user and the number of payment tokens, as well as the page browsing time and access depth when the advertisement is normally accessed; and the storage content is reallocated according to the storage priority of the advertisement. The method for obtaining the growth period includes: For any given ad, the total number of times the ad is accessed at each time step in the time sequence is taken as the access count at each time step; In the current period, calculate the difference in visits to the ad at each time point compared to the previous time point, and use this difference as the growth index of the ad at each time point; filter out the time points where the growth index is less than 0, and record the remaining time points as the target time points; The period consisting of consecutive target moments in the current cycle is designated as the growth period for this advertisement. The methods for obtaining the possible indicators of anomalies include: For any given advertisement, each growth period is used as the analysis period in turn; For any period outside the current period in the historical time period, by analyzing the distribution of the time period in the current period, the same distributed time period in that period is obtained; the difference between the average number of visits to the advertisement in the analysis period and the average number of visits to the advertisement in the same distributed time period is taken as the visit deviation between the analysis period and that period. The average of the visit deviations between the analysis period and all other periods in the historical data is used as the historical visit deviation of the ad in the analysis period. After calculating the difference in the growth index of the advertisement between every two adjacent moments during the analysis period, the sum of all the differences in the growth index is calculated and a negative correlation mapping is performed to obtain the stability of the advertisement's fluctuations during the analysis period. By combining the average growth index, historical visit deviation, and volatility stability of the advertisement at all times during the analysis period, possible indicators of anomalies for the advertisement during the analysis period can be obtained.
2. The big data storage service method according to claim 1, characterized in that, The method for obtaining the jump anomaly indicator includes: In a single visit, each pair of adjacent pages in the page jump sequence is taken as a jump group, and the page with the larger sequence number in the jump group is taken as the jump page; based on the distribution order of the jump groups in the page jump sequence, the jump pages of the jump groups are arranged to obtain the jump behavior sequence of a single visit. For any visit to each ad during a suspected abnormal period, all visits to the ad corresponding to that visit during all non-suspected abnormal periods are taken as reference visits. Each number in the redirection sequence of the access is taken as the target number; the total number of accesses with the target number in the reference access redirection sequence is taken as the number of times the target number is redirected; the total number of accesses with the target number in the reference access redirection sequence and the target number corresponds to the same redirect page is taken as the number of times the target number is redirected; the ratio of the number of times the target number is redirected to the number of times the target number is redirected is taken as the redirection correlation index of the target number. By performing negative correlation mapping on the possible abnormal indicators of the suspected abnormal time period of the visit, we can obtain the possible normal analysis indicators of the visit. The product of the jump association index of the target sequence number and the normal analysis possible index of the visit is used as the association feature index of the target sequence number in the visit; the mean of the association feature index of the sequence number in the jump behavior sequence of the visit is calculated and negative correlation mapping is performed to obtain the jump association abnormal index of the visit. Calculate the difference between the length of the page jump sequence during this visit and the average length of all page jump sequences in the reference visits to obtain the abnormal jump count index for this visit; By combining the redirection-related anomaly indicators and the redirection count anomaly indicators for this access, the redirection anomaly indicators for this access can be obtained.
3. The big data storage service method according to claim 1, characterized in that, The methods for obtaining the user anomaly indicators include: The first two pages in the page jump sequence during a single visit are treated as one analysis group, and the subsequent pages in the page jump sequence are successively merged into the analysis group. Each time they are merged, a new analysis group is obtained. For any ad during a suspected abnormal period, calculate the frequency of each analysis group of that ad in all non-suspected abnormal periods to obtain the fixed probability index of each analysis group; use the maximum fixed probability index of the ad as the similarity index of that ad. The user corresponding to this access is designated as the analysis user; in the current period, the ratio of the total number of times the analysis user accesses this advertisement to the total number of times all users access this advertisement is used as the access frequency index of the analysis user; the mean of the possible abnormal indicators in all accesses of the analysis user is calculated to obtain the access abnormality index of the analysis user; the mean of the similar jump indicators in all accesses of the analysis user is calculated and negative correlation mapping is performed to obtain the jump order abnormality index of the analysis user. By combining the analysis of user access frequency indicators, access anomaly indicators, and redirection order anomaly indicators, the user anomaly indicators for this access are obtained.
4. The big data storage service method according to claim 1, characterized in that, The method for determining normal access includes: The product of the redirection anomaly indicator and the user anomaly indicator for a single visit is normalized and used as the malicious behavior indicator for a single visit; visits with malicious behavior indicators greater than the preset malicious threshold are considered malicious visits for each advertisement. In the current period of each advertisement, all visits except malicious visits are treated as normal visits for each advertisement.
5. The big data storage service method according to claim 4, characterized in that, The method for obtaining the storage priority includes: For any normal visit to any advertisement, the ratio of the number of times the user corresponding to the normal visit visits the advertisement in the current period to the total number of times the advertisement is displayed to the user is used as the visit interest index for the normal visit. Multiply the average page browsing time during a normal visit by the average page visit depth to obtain the behavioral interest index of the normal visit; perform a negative correlation mapping on the malicious behavior index of the normal visit to obtain the behavioral reliability of the normal visit. The effective index of the normal access is obtained by multiplying the reliability of the normal access behavior, the access interest index, and the behavior interest index. In each storage node, the total number of times each advertisement is accessed by each user is taken as the node access score; the ratio of the total number of times the advertisement is accessed normally to the node access score is taken as the node access frequency of the advertisement. Calculate the average of all valid metrics for normal access to the advertisement, the product of the number of paid tokens and the node access frequency, and obtain the storage priority of the advertisement in the storage node.
6. The big data storage service method according to claim 1, characterized in that, The reallocation of storage content based on the storage priority of advertisements includes: In each storage node, the storage priority of the advertisement is normalized to obtain the storage retention metric; When the storage retention index of an ad exceeds the preset storage threshold, the relevant data of the corresponding ad will be retained in the storage node; otherwise, the relevant data of the ad will be transferred to cloud storage.
7. The big data storage service method according to claim 1, characterized in that, The method for determining the suspected abnormal time period includes: Within all growth periods of each advertisement, the growth periods where the abnormality probability index exceeds the preset abnormality threshold are designated as suspected abnormal periods.
8. A big data storage service system, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the big data storage service method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Abnormal traffic detection method and device, electronic equipment and storage medium
CN113271322A
Authentication method of data service interface
CN119299227A