Serverless traffic protection method based on feature extraction
By adopting a traffic protection method based on feature extraction in the Serverless architecture, the problem of single protection capabilities and difficulty in adapting to business scenarios in the existing technology is solved, and lightweight, customizable and efficient traffic protection is achieved to meet the security needs of different business scenarios.
Patent Information
- Application Number
- CN202510120643.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing Serverless architecture traffic protection solution has problems such as single protection capabilities, difficulty in adapting to business scenarios, and high deployment costs. Especially when facing complex attack scenarios and customized crawler attacks, the protection effect is not ideal.
A Serverless traffic protection method based on feature extraction is adopted, including a feature extraction engine, feature storage module, decision programming interface and runtime support module. Through systematic feature extraction and flexible programming interfaces, a lightweight and customizable traffic protection solution is realized.
It realizes lightweight feature extraction and storage, reducing system resource consumption; supports multi-dimensional feature analysis, improving protection accuracy; it realizes flexible customization of protection strategies through programmable interfaces to meet the needs of different business scenarios; ensures the security and stability of the system; feature storage adopts efficient compression algorithms to optimize storage efficiency.
Smart Images

Figure BDA0005258848000000051 
Figure BDA0005258848000000052 
Figure BDA0005258848000000061
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer network security technology, and in particular to a method and device for protecting serverless architecture traffic based on feature extraction. The method is applied to a serverless function computing platform in a cloud computing environment to defend against network crawlers and abnormal traffic attacks. Background Art
[0002] With the development of cloud computing technology, Serverless architecture has been widely used in various application scenarios due to its on-demand use and maintenance-free management. In the Serverless architecture, function computing, as a core service, carries a large number of business processing tasks. However, this architecture also brings new security challenges, especially in network traffic protection. Due to the characteristics of Serverless functions being started on demand and billed on a per-use basis, malicious traffic and web crawlers may cause a large number of function instances to be triggered, resulting in resource waste and cost increases. Therefore, effective traffic protection is of great significance in the Serverless architecture.
[0003] At present, the Serverless platform mainly adopts two traffic protection solutions: one is based on the traditional Web Application Firewall (WAF), which identifies and filters malicious traffic by deploying WAF; the other is based on the API gateway, which implements request flow control and simple attack identification at the API gateway level. These solutions usually identify abnormal traffic through preset rules and feature libraries and provide basic protection capabilities.
[0004] However, these traditional protection solutions have obvious shortcomings under the Serverless architecture. First, traditional WAFs are relatively heavyweight, with high deployment and maintenance costs, and are difficult to adapt to the elastic characteristics of the Serverless architecture; second, the protection capabilities of the API gateway are relatively simple and difficult to cope with complex attack scenarios; most importantly, these solutions lack an understanding of specific business scenarios and cannot provide accurate protection strategies based on the characteristics of different businesses, resulting in unsatisfactory protection results.
[0005] To solve the above problems, the industry has proposed some lightweight traffic detection solutions, such as protection components based on function runtime and simplified WAF. Although these solutions have reduced deployment costs to a certain extent, they still have problems such as insufficient protection capabilities and difficulty in adapting to business changes. Especially when facing customized crawler attacks, due to the lack of in-depth understanding of business logic, the protection effect of these solutions is often not ideal.
[0006] Therefore, there is an urgent need for a traffic protection solution that can maintain lightweight features while supporting customized business protection logic. This solution should be able to effectively extract traffic features and allow users to flexibly customize protection strategies according to business needs, thereby achieving precise protection of the Serverless architecture. This can not only improve the protection effect, but also reduce the misjudgment rate, and better meet the security needs of different business scenarios. Summary of the invention
[0007] The purpose of the present invention is to solve the problems existing in the existing Serverless architecture traffic protection solution, such as single protection capability, difficulty in adapting to business scenarios, and high deployment cost. In addition, the present invention also aims to provide a programmable protection framework, which enables users to customize protection strategies according to specific business needs and achieve precise protection.
[0008] To achieve the above objectives, the present invention provides a serverless traffic protection method based on feature extraction, which is characterized by comprising a feature extraction engine, a feature storage module, a decision programming interface, and a runtime support module. The method implements a lightweight and customizable traffic protection solution through systematic feature extraction and flexible programming interfaces.
[0009] Specifically, the feature extraction engine is responsible for extracting and calculating traffic features from multiple dimensions, including time dimension features, space dimension features, behavior dimension features, and business dimension features. Among them, time dimension features include request frequency and time distribution features, space dimension features include IP distribution and network features, behavior dimension features include request mode and parameter features, and business dimension features are customized according to specific application scenarios.
[0010] Furthermore, the feature storage module adopts a sliding bitmap compression algorithm for feature storage, represents the time window features through a fixed-length bitmap, and uses bit operations to update the feature status, thereby realizing an efficient feature data access mechanism.
[0011] Preferably, the decision programming interface provides feature access methods and rule programming capabilities, supporting users to customize protection decision logic based on extracted features. Through this interface, users can flexibly combine various features according to business needs to implement precise protection strategies.
[0012] In some embodiments, the runtime support module ensures the safe operation of user-defined code and prevents excessive consumption of resources through a code isolation mechanism and a resource limitation mechanism.
[0013] In addition, the present invention can also integrate an incremental calculation mechanism in the feature extraction engine, reduce calculation overhead through feature data caching, and improve system performance.
[0014] In a preferred implementation, the system first deploys a feature extraction engine to continuously collect and calculate multi-dimensional traffic features; then, these features are efficiently stored through a sliding bitmap compression algorithm; finally, users can write custom protection logic through a decision programming interface, and the system can securely execute these logics in the runtime environment to achieve precise protection.
[0015] By adopting the above scheme, the present invention has the following beneficial effects: 1. It realizes lightweight feature extraction and storage, reducing system resource consumption; 2. Support multi-dimensional feature analysis to improve the accuracy of protection; 3. Flexible customization of protection strategies is achieved through programmable interfaces to meet the needs of different business scenarios; 4. The code isolation and resource limitation mechanism are adopted to ensure the security and stability of the system; 5. Feature storage uses an efficient compression algorithm to optimize storage efficiency.
[0016] In summary, the present invention provides a Serverless traffic protection solution that maintains lightweight characteristics and supports flexible customization, effectively solves the problems in the prior art, and has important practical value. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention, and for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative work.
[0018] Figure 1 It is the overall architecture diagram of the Serverless traffic protection system, showing the four core modules of the system (feature extraction engine, feature storage module, decision programming interface, and runtime support module) and the hierarchical relationship of their internal components, and clearly describing the overall structure of the system and the interaction between modules.
[0019] Figure 2 It is a feature extraction pipeline diagram, which describes the complete processing flow from network request input to feature extraction, including preprocessing of the input layer, analysis of three dimensions (time dimension, space dimension, behavior dimension) of the feature extraction layer, and aggregation and compression process of the feature fusion layer.
[0020] Figure 3It is a class diagram of the feature storage structure, showing the core classes of the feature storage module and their relationships, including the specific methods and interdependencies of the bitmap manager (BitMapManager), storage engine (StorageEngine), index manager (IndexManager) and compression manager (CompressionManager).
[0021] Figure 4 It is a state flow diagram of the decision programming interface, which shows the complete process from rule parsing to execution, including key stages such as rule parsing, syntax analysis, AST generation, code optimization, bytecode generation and execution engine, as well as the processing steps within each stage.
[0022] Figure 5 It is an organizational chart of the runtime protection mechanism, which describes the four main components in the runtime environment (container isolation, resource monitoring, fault handling, and security policy) and their internal implementation details, and shows how the system implements a secure and reliable code execution environment. DETAILED DESCRIPTION
[0023] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0024] In the description of the present invention, it should be noted that the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside" and the like indicate directions or positional relationships based on the directions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore cannot be understood as a limitation on the present invention.
[0025] Example 1: Infrastructure Example
[0026] Reference Figure 1 As shown in the figure, the serverless traffic protection system adopts a modular design, including a feature extraction engine, a feature storage module, a decision programming interface, and a runtime support module. The system is deployed at the access layer of the serverless platform to analyze and protect all incoming requests.
[0027] As the core component of the system, the feature extraction engine adopts a streaming processing architecture and processes request data in real time in an event-driven manner. The engine implements a lightweight feature calculation framework and connects multiple feature extractors in series through a data pipeline. Each feature extractor is responsible for feature calculation of a specific dimension and transfers data through a shared memory queue, ensuring the efficiency of the feature extraction process.
[0028] The feature storage module uses an innovative sliding bitmap compression algorithm for feature storage. This algorithm uses a fixed-size bitmap to represent the feature state within a time window. Each bitmap size is 2 n bits, where n is determined according to the actual monitoring time window size. The bitmap implements time window sliding through circular shift operations and uses bit operations to update feature states, significantly reducing storage overhead. Feature data is organized in the format of {timestamp, feature_type, feature_value}, where timestamp uses Unix timestamps, accurate to milliseconds.
[0029] The decision programming interface provides a lightweight domain-specific language (DSL) that supports users to write protection rules. The interface defines the feature access syntax: feature.get(dimension,name,window), where dimension specifies the feature dimension, name specifies the feature name, and window specifies the time window size. The rule execution engine uses the interpreter mode to convert the rules written by users into optimized bytecode to ensure execution efficiency.
[0030] The runtime support module implements a container-based code isolation mechanism. Each user-defined protection rule is executed in an independent container environment. The container limits CPU and memory usage through cgroup. The preset CPU usage limit is 10% per core and the memory limit is 64MB. The module uses a token bucket algorithm for resource control. The token generation rate is dynamically adjusted according to the system load to ensure system stability.
[0031] The system's workflow is as follows: 1. After a request enters the system, it is first processed by the feature extraction engine. The engine starts multiple feature extractors at the same time according to the preset feature extraction configuration, and calculates the features of each dimension in parallel. 2. The extracted features are compressed and stored through the feature storage module. The storage module maintains a feature index table and supports feature queries with O(1) time complexity. When the amount of feature data exceeds the preset threshold, the data compression mechanism is triggered. 3. The decision programming interface loads user-defined protection rules, and the rule compiler converts them into optimized execution plans. The execution plan contains feature access paths and judgment logic, and supports complex condition combinations. 4. The runtime support module creates an isolated execution environment for each rule, and ensures that rule execution does not affect the overall system performance through resource restrictions. The module implements a fault isolation mechanism, and the exception of a single rule will not affect the execution of other rules.
[0032] The innovation of this architecture design is to realize a lightweight but fully functional feature extraction and protection framework. Through the optimized feature storage algorithm and rule execution mechanism, the additional overhead is controlled at a low level while ensuring functional flexibility. The system supports horizontal expansion and can dynamically adjust the number of instances according to the load situation.
[0033] The specific parameters and implementation methods in this embodiment are only for illustration, and those skilled in the art may make appropriate adjustments according to actual needs. For example, the configuration of the feature extractor may be adjusted according to business characteristics, or the resource restriction parameters of the container may be adjusted according to system resource conditions.
[0034] Example 2: Feature extraction example
[0035] Based on Example 1, this example describes in detail the specific implementation method of the feature extraction engine. Figure 2 As shown in the figure, the feature extraction engine adopts a multi-dimensional feature extraction architecture and realizes efficient traffic analysis through an original feature calculation method.
[0036] The time dimension feature extraction uses the sliding time window technology. The window size is configurable and the default is 60 seconds. The request frequency feature is obtained by calculating the number of requests per unit time. The calculation formula is:
[0037] where F req Indicates the request frequency, N req Indicates the number of requests in the window, T window Indicates the time window size. Time distribution features are analyzed by Fourier transform to extract periodic features from the request time series:
[0038] Where S(ω) represents the power spectral density and x(t) represents the request time series.
[0039] The spatial dimension feature extraction includes two parts: IP distribution features and network features. The IP distribution features use the cardinality estimation algorithm and the HyperLogLog data structure to record the IP address distribution. The error range is set to Where m is the number of registers. In practice, m = 2048 can meet the accuracy requirement. The network characteristics are obtained by analyzing the characteristics of TCP connections, including connection establishment time, data transmission rate and other parameters.
[0040] The behavior dimension feature extraction focuses on request patterns and parameter features. The request pattern features are obtained by analyzing the Markov characteristics of the request sequence and constructing the state transition matrix: P ij =P(X t+1 =j|X t =i)(3)
[0041] Where P ij Indicates the probability of transitioning from state i to state j. Parameter characteristics analyze the distribution characteristics of request parameters through information entropy calculation method:
[0042] Where H represents the parameter entropy value, p i represents the probability of occurrence of parameter value i.
[0043] The feature extraction process uses a pipeline processing mechanism, and triggers feature calculation in an event-driven manner. The feature extractor of each dimension implements incremental computing capabilities. When a new request arrives, only the relevant counters and status values need to be updated, without recalculating all features. The feature calculation results are temporarily stored in a ring buffer with a buffer size of 4096 records. A double buffer mechanism is used to avoid read-write conflicts.
[0044] To improve the accuracy of feature extraction, the system implements an adaptive sampling mechanism. When the request traffic is large, the computational load is reduced through a stratified sampling method:
[0045] where p sample represents the sampling probability, C base Indicates the benchmark processing capability, F req Indicates the current request frequency.
[0046] The feature extraction results are output in the form of vectors with a dimension of 32, including the feature values of each dimension. The feature vector is stored in a compressed format, and the sparse features are stored in the CSR (Compressed Sparse Row) format, which can effectively reduce storage overhead.
[0047] The innovation of this embodiment is to design a complete set of multi-dimensional feature extraction methods, which accurately characterizes traffic characteristics through mathematical models. This method has the characteristics of high computational efficiency and low storage overhead, and is suitable for real-time traffic analysis in a Serverless environment.
[0048] The feature extraction method can be expanded according to actual needs, such as adding new feature dimensions or adjusting feature calculation parameters. In specific implementation, parameters such as sampling rate and buffer size can be adjusted according to system resource conditions and performance requirements.
[0049] Embodiment 3: Feature storage embodiment
[0050] Based on the feature extraction results of Example 1 and Example 2, this embodiment describes in detail the specific implementation method of the feature storage module. Figure 3 As shown, the feature storage module adopts an innovative sliding bitmap compression algorithm to achieve an efficient feature data storage and retrieval mechanism.
[0051] The core of the sliding bitmap compression algorithm is the fixed-length bitmap representation. Each feature uses a bitmap array to represent its state within the time window, and the bitmap length L is determined by the following formula:
[0052] Where T window Indicates the monitoring window size (seconds), F sample Indicates the sampling frequency (times / second). The bitmap adopts a ring structure and implements time window sliding through displacement operation. The bitmap update operation is defined as: B new =(B old <<s)|mask (7)
[0053] Among them B new and B old They represent the bitmaps before and after the update, s represents the displacement, and mask represents the bit pattern of the newly added features.
[0054] To improve storage efficiency, feature data uses a hierarchical storage structure. Hot data is stored in a bitmap array in memory, and cold data is stored in persistent storage after compression. The compression uses the run-length encoding (RLE) method to encode consecutive identical bits in the bitmap: RLE(B)={(v i ,l i )|i=1,2,...,n} (8)
[0055] where v i Indicates the bit value (0 or 1), l i Indicates the run-on length.
[0056] The feature index adopts a two-level structure. The first level is the feature type index, which is implemented using a balanced tree (B+ tree); the second level is the time index, which is implemented using a skip list (Skip List). The index item structure is:
[0057] The bitmap update operation uses atomic operations to ensure concurrency safety. When you need to update the bitmap, first acquire the write lock: lock=atomic_compare_exchange(bitmap_lock,0,1)(9)
[0058] After the update is completed, the lock is released through atomic operations. To avoid occupying the lock for a long time, the update operation is broken down into multiple small batches.
[0059] To handle the overflow of bitmap capacity, an adaptive bitmap expansion mechanism is implemented. When the bitmap usage exceeds the threshold (default 75
[0060] Where L new and L old Respectively represent the length of the bitmap before and after expansion, U current Indicates the current usage rate, U threshold Indicates the expansion threshold.
[0061] The cleaning of feature data adopts a lazy deletion strategy. When the time window slides, the expired data is not deleted immediately, but marked as deleted. The cleaning operation is performed asynchronously when the system load is low:
[0062] Where P clean represents the probability of triggering cleaning, L current Indicates the current load, L max Indicates the maximum load.
[0063] The innovation of this embodiment is to design a bitmap compression algorithm suitable for traffic feature storage, which realizes efficient feature storage and retrieval through bit operations and compression coding. This method has the characteristics of low storage overhead and fast access speed, and is particularly suitable for real-time feature analysis in a Serverless environment.
[0064] The specific implementation of the feature storage module can be adjusted according to actual needs, such as modifying the compression algorithm, adjusting the index structure, or changing the cleanup strategy. In actual deployment, parameters such as bitmap size and compression ratio can be adjusted according to system resource conditions and performance requirements.
[0065] Embodiment 4: Decision-making programming embodiment
[0066] Based on the feature extraction and storage capabilities in the previous embodiments, this embodiment details the design and implementation of the decision programming interface. Figure 4As shown, the decision programming interface is designed using a domain specific language (DSL) and implements a complete feature access and rule programming framework.
[0067] The core of the decision programming interface is the feature access syntax, which is defined as follows: feature.get(dimension,name,window)->FeatureValue feature.combine(features,operator)->FeatureValue feature.transform(value,method)->FeatureValue
[0068] Among them, dimension specifies the feature dimension, name specifies the feature name, and window specifies the time window size. The calculation of feature values supports multiple aggregation operations, and the expression is as follows: V feature =Agg(F(d,n,w))(12)
[0069] Where V feature represents the eigenvalue, F(d,n,w) represents the feature extraction function, and Agg represents the aggregation operation (such as sum, avg, max, etc.).
[0070] Rule programming uses declarative syntax, and the rule definition format is:
[0071] The rule compiler converts the rules into an abstract syntax tree (AST) with the following optimized structure: AST={Node(type,value,children)}(13)
[0072] The rule execution engine adopts an event-driven model. The execution process includes the following steps: 1. Parse rule definition and build AST 2. Optimize AST and remove redundant nodes 3. Generate bytecode instruction sequence 4. Execute instructions in the virtual machine
[0073] The bytecode instruction set is designed as follows: LOAD_FEATURE; Load feature values COMPARE; conditional comparison JUMP_IF; Conditional jump EXECUTE_ACTION; execute action
[0074] In order to improve the efficiency of rule execution, a rule optimizer is implemented. The optimization strategies include:
[0075] Cost rule represents the rule execution cost, W i represents the operation weight, C i Represents the cost of the operation. Reorder the rules based on the cost estimation result.
[0076] The feature access interface implements a cache mechanism, and the cache structure is defined as:
[0077] The cache elimination strategy uses the LRU algorithm, and the elimination probability is calculated as follows:
[0078] Where T current Indicates the current time, T last Indicates the last access time, and TTL indicates the survival time.
[0079] The innovation of this embodiment is to design a feature-driven rule programming language, which realizes efficient protection strategy customization through declarative syntax and optimized execution engine. This method has the characteristics of strong expression ability and high execution efficiency, allowing users to flexibly define protection rules according to business needs.
[0080] The decision programming interface can be expanded according to actual needs, such as adding new feature operators, supporting more complex rule logic, or optimizing execution performance. In specific implementation, parameters such as cache size and optimization strategy can be adjusted according to system resource conditions and performance requirements.
[0081] Example 5: Runtime protection example
[0082] Following the decision-making programming function in the previous embodiment, this embodiment describes in detail the implementation method of the runtime support module. Figure 5 As shown in the figure, the runtime support module ensures the safe execution of user-defined code through code isolation and resource limitation mechanisms.
[0083] The code isolation mechanism is implemented using lightweight container technology. The container configuration is defined as follows:
[0084] The container runtime uses cgroup to implement resource isolation. The key resource limit calculation formula is: R limit =min(R base α,R max)(16)
[0085] Where R limit Indicates the resource limit value, R base represents the reference value, α represents the adjustment factor, and R max Indicates the maximum limit.
[0086] Resource monitoring adopts a hierarchical monitoring architecture and defines a set of monitoring indicators: M={CPU usage ,MEM usage ,IO rate ,NET rate}(17)
[0087] Each indicator has a corresponding alarm threshold: T alert =T base +kσ(18)
[0088] Where T base represents the benchmark threshold, k represents the alarm level, and σ represents the standard deviation.
[0089] Code execution security is ensured by the sandbox mechanism, which implements system call filtering:
[0090] System call risk level assessment formula:
[0091] Where P i represents the call probability, S i Indicates the security threat level.
[0092] Resource limitation uses a token bucket algorithm, and the token generation rate is dynamically adjusted:
[0093] Rate token represents the token generation rate, L current Indicates the current load, L max Indicates the maximum load.
[0094] The exception handling mechanism implements multi-level fault isolation:
[0095] The fault recovery strategy is determined based on the fault level:
[0096] Where T recovery Indicates the recovery time, Levelfault Indicates the fault level.
[0097] The innovation of this embodiment is to design a lightweight but safe and reliable runtime protection mechanism, which ensures the safe execution of user-defined code through multi-level isolation and restriction means. This method effectively controls security risks while ensuring flexibility.
[0098] The runtime protection module can be adjusted according to actual needs, such as adding new resource restriction dimensions, optimizing monitoring strategies, or improving fault handling mechanisms. In specific implementations, resource restriction parameters, alarm thresholds, and other configurations can be adjusted according to the deployment environment and security requirements. In this way, the flexibility of the protection strategy is guaranteed while ensuring the security and reliability of the system.
Claims
1. A traffic protection method for a Serverless architecture, characterized in that: include: (a) a feature extraction engine, used to extract and calculate traffic features, wherein the traffic features include time dimension features, space dimension features, behavior dimension features, and business dimension features; (b) Feature storage module, which uses a sliding bitmap compression algorithm to store features and achieve efficient access to feature data; (c) Decision programming interface, which provides feature access methods and rule programming capabilities, and supports users to customize protection decision logic based on extracted features; (d) Runtime support module, responsible for the isolated operation and resource management of custom code.
2. The flow protection method according to claim 1, characterized in that: The time dimension features include: (a) Request frequency feature, which is used to represent the number of requests per unit time; (b) Time distribution features, used to characterize the timing pattern of requests.
3. The flow protection method according to claim 1, characterized in that: The spatial dimension features include: (a) IP distribution features, used to characterize the geographical distribution of request sources; (b) Network characteristics, used to characterize the network transmission characteristics of the request.
4. The flow protection method according to claim 1, characterized in that: The behavioral dimension features include: (a) Request pattern features, used to characterize the behavior pattern of requests; (b) Parameter features, used to characterize the distribution patterns of request parameters.
5. The flow protection method according to claim 1, characterized in that: The implementation of the feature storage module includes: (a) Fixed-length bitmap represents time window features; (b) Bitwise operations are performed to update feature states.
6. The flow protection method according to claim 1, characterized in that: The decision programming interface includes: (a) Feature access interface, which provides methods for reading extracted features; (b) Rule programming interface, which supports writing feature-based judgment logic.
7. The flow protection method according to claim 1, characterized in that: The runtime support module includes: (a) Code isolation mechanism to ensure the safe operation of custom code; (b) Resource limitation mechanism to prevent excessive consumption of resources.
8. A serverless traffic protection device for implementing the method according to any one of claims 1 to 7, characterized in that: include: (a) a feature extraction unit, used to perform feature extraction and calculation; (b) a feature storage unit, used to store feature data; (c) Decision execution unit, used to run user-defined protection logic; (d) Runtime management unit, used to ensure the safe and stable operation of the system.
Citation Information
Cited By
Digitalized marketing anti-disturbing method and system
CN121684996A