Method and apparatus for optimizing representations of traffic flow patterns
Patent Information
- Application Number
- CN202610447615.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-02-17
- Filing Date
- 2026-04-07
- Publication Date
- 2026-08-18
AI Technical Summary
例如,为了阻止或转发某种类型的业务流,可能很难在大量签名或规则中确定能够管理该业务流的签名或规则
[0009]Examples of this technology offer numerous advantages, including providing methods, non-transitory computer-readable media, apparatus, and systems for optimizing and deploying optimized representations on network devices to detect traffic flow patterns in a network. Therefore, such optimized representations can improve the performance of one or more networks and achieve a better user experience. In some examples, by iteratively performing the operations described in this disclosure to combine or discard original or inefficient representations, the number of representations can potentially be reduced to a manageable value. Furthermore, by allowing users to specify preferred metric values, flexibility is introduced into this improved solution depending on which network device is used to detect patterns. The above and other aspects and advantages, and embodiments thereof, are described in more detail in the following figures, specification, and claims.
Smart Images

Figure CN122601501A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to optimized representations, and more particularly to optimized representations and their deployment on network devices for detecting traffic flow patterns in a network. Background Technology
[0002] Detecting patterns in network traffic can be used for various purposes, such as traffic management, data analytics, security, and load balancing. Typically, when processing incoming traffic, this pattern detection is performed using various network devices based on a set of signatures or rules. While this can be efficient, the sheer number of signatures or rules can make the process very challenging and inefficient. For example, to block or forward a certain type of traffic flow, it can be difficult to identify the signature or rule that manages that flow from a large pool of signatures or rules.
[0003] Another issue is that network devices used for pattern detection may have varying constraints in terms of hardware or software resources (such as memory, network capabilities, etc.), limiting their capabilities and effectiveness. For example, in scenarios involving a large number of traffic flows (e.g., requests to application servers in the network), this number may be so large that the network device cannot apply some pattern matchers to each request. This situation is not uncommon in today's typical network environments, and therefore there may not be enough opportunity to spend sufficient time effectively performing pattern detection on each traffic flow without causing response timeouts or significant delays. Summary of the Invention
[0004] This disclosure relates to methods and apparatus for optimizing representations used to detect traffic flow patterns. Related non-transitory computer-readable media and network traffic management systems are also disclosed.
[0005] According to one aspect of this disclosure, a method can be implemented by a network traffic management system, wherein the network traffic management system may include one or more network traffic management devices, client devices, or server devices. The method may include retrieving a representation and a dataset associated with the representation from a storage device, and prompting a natural language processing model to convert the retrieved representation into a first candidate representation based on the dataset, wherein the first candidate representation differs from the retrieved representation. Next, the method inputs the first candidate representation into a simulator to generate one or more first candidate metrics corresponding to one or more preferred metrics. The method further evaluates whether to replace the retrieved representation with the first candidate representation based on the generated one or more first candidate metrics and one or more preferred metrics. In response to the evaluation providing an indication to replace the retrieved representation with the first candidate representation, the method deploys the first candidate representation on a network device of a network managed by the network traffic management system, wherein the network device is configured to use the deployed first candidate representation to detect network traffic flow patterns.
[0006] According to another aspect of this disclosure, an apparatus may include a memory and one or more processors. The memory includes programming instructions stored in the memory. The one or more processors are configured to execute the programming instructions stored in the memory to: retrieve a representation and a dataset associated with the representation from the memory, and prompt a natural language processing model to convert the retrieved representation into a first candidate representation based on the dataset, the first candidate representation being different from the retrieved representation. The one or more processors may also input the first candidate representation into a simulator to generate one or more first candidate metrics corresponding to one or more preferred metrics, and evaluate whether to replace the retrieved representation with the first candidate representation based on the generated one or more first candidate metrics and one or more preferred metrics. In response to the evaluation providing an indication to replace the retrieved representation with the first candidate representation, the one or more processors may deploy the first candidate representation on network devices of a network managed by a network traffic management system, wherein the network devices are configured to use the deployed first candidate representation to detect network traffic flow patterns.
[0007] According to another aspect of this disclosure, a non-transitory computer-readable medium may have instructions stored thereon for protecting network service devices, including executable code that, when executed by one or more processors, causes one or more processors to retrieve a representation and a dataset associated with the representation from storage, and prompts a natural language processing model to convert the retrieved representation into a first candidate representation based on the dataset, the first candidate representation being different from the retrieved representation. The executable code may also cause one or more processors to input the first candidate representation into a simulator to generate one or more first candidate metrics corresponding to one or more preferred metrics, and to evaluate whether to replace the retrieved representation with the first candidate representation based on the generated one or more first candidate metrics and one or more preferred metrics. In response to the evaluation providing an instruction to replace the retrieved representation with the first candidate representation, the executable code may also cause one or more processors to deploy the first candidate representation on network devices of a network managed by a network traffic management system, wherein the network devices are configured to use the deployed first candidate representation to detect network traffic flow patterns.
[0008] According to another aspect of this disclosure, a network traffic management system includes one or more traffic management devices, server devices, or client devices. The network traffic management system may include a memory and one or more processors. The memory includes programming instructions stored thereon. The processors are configured to execute the stored programming instructions to retrieve a representation and a dataset associated with the representation from the storage device, and to prompt a natural language processing model to convert the retrieved representation into a first candidate representation based on the dataset. The first candidate representation differs from the retrieved representation. The one or more processors may also input the first candidate representation into a simulator to generate one or more first candidate metrics corresponding to one or more preferred metrics, and evaluate whether to replace the retrieved representation with the first candidate representation based on the generated one or more first candidate metrics and one or more preferred metrics. In response to the evaluation providing an indication to replace the retrieved representation with the first candidate representation, the one or more processors may deploy the first candidate representation on network devices of a network managed by the network traffic management system, wherein the network devices are configured to use the deployed first candidate representation to detect network traffic flow patterns.
[0009] Examples of this technology offer numerous advantages, including providing methods, non-transitory computer-readable media, apparatus, and systems for optimizing and deploying optimized representations on network devices to detect traffic flow patterns in a network. Therefore, such optimized representations can improve the performance of one or more networks and achieve a better user experience. In some examples, by iteratively performing the operations described in this disclosure to combine or discard original or inefficient representations, the number of representations can potentially be reduced to a manageable value. Furthermore, by allowing users to specify preferred metric values, flexibility is introduced into this improved solution depending on which network device is used to detect patterns. The above and other aspects and advantages, and embodiments thereof, are described in more detail in the following figures, specification, and claims. Attached Figure Description
[0010] The above and other aspects of this disclosure are best understood from the following detailed description when read in conjunction with the accompanying drawings. Specific examples are shown in the drawings to illustrate the technology; however, it should be understood that the examples of the technology are not limited to the specific means disclosed. The drawings include the following illustrations:
[0011] Figure 1 An exemplary network traffic management system is shown;
[0012] Figure 2 An exemplary execution environment for a network traffic management device is shown;
[0013] Figure 3 An exemplary block diagram of a network traffic management device is shown;
[0014] Figure 4 A flowchart is shown of an exemplary method for optimizing representation executed at a network traffic management device;
[0015] Figure 5 It shows the method for execution Figure 4 An exemplary data source for the exemplary method shown;
[0016] Figure 6 It shows the use of Figure 5 The data source shown is used as input data for execution. Figure 4 An exemplary flowchart of the method shown;
[0017] Figure 7 An exemplary flowchart for prompting a natural language processing model to generate candidate representations is shown;
[0018] Figure 8 An example preference selection is shown, which allows the user to specify a preferred metric.
[0019] Figure 9 An exemplary simulation and evaluation process is shown;
[0020] Figure 10 Another exemplary simulation and evaluation process is shown;
[0021] Figure 11 This demonstrates yet another exemplary simulation and evaluation process; and
[0022] Figure 12 An exemplary application scenario for performing the operations described in this disclosure is shown. Detailed Implementation
[0023] This disclosure can be more readily understood by referring to the detailed description of the following exemplary examples. Before disclosing and describing exemplary implementations and examples of methods, apparatus, and systems according to this disclosure, it should be understood that the implementations are not limited to those described in this disclosure. Many modifications and variations therein will be apparent to those skilled in the art and remain within the scope of this disclosure. It should also be understood that the terminology used herein is for describing particular implementations only and is not intended to be limiting. Some implementations of the disclosed technology will be described more fully below with reference to the accompanying drawings. However, the disclosed technology may be embodied in many different forms and should not be construed as limited to the implementations set forth herein.
[0024] Many specific details are set forth in the following description. However, it should be understood that examples of the disclosed techniques can be practiced without these specific details. In other instances, well-known components, structures, and techniques have not been shown in detail so as not to obscure the understanding of this description. References to “one implementation,” “one example,” “some examples,” etc., indicate that implementations of the disclosed techniques described may include specific features, structures, or characteristics, but not every implementation must include that specific feature, structure, or characteristic. Furthermore, the repeated use of the phrase “in some examples” does not necessarily refer to the same implementation, although it may. Moreover, it should be understood that specific features, structures, or characteristics described in different examples, implementations, etc., can be further combined in various ways and implemented in one or more implementations.
[0025] A network traffic management system may involve a set of tools, processes, devices, and related technologies for controlling and optimizing data flow within a computer network. Such a system monitors, analyzes, controls, and balances network traffic to maintain the performance and reliability of the computer network. Network traffic management systems can be implemented in a variety of network topologies. The devices used and the topology designed for the network environment may depend on specific requirements and the size of the network. Factors may include, for example, network size, geographical distribution, the type of applications and services provided, and the organization's traffic management requirements. For instance, a network traffic management system can be implemented in a variety of network topologies, including centralized, distributed, or cloud-based topologies. Network traffic management systems can be implemented in a variety of networks, including but not limited to Local Area Networks (LANs), Wide Area Networks (WANs), Metropolitan Area Networks (MANs), data center networks, cloud networks, hybrid networks, or any suitable existing or future network. Depending on the specific network and topology used, a variety of devices may be involved in a network traffic management system. For example, edge routers or switches, firewalls, proxies, load balancers, content delivery network (CDN) servers, application servers, etc., may be included in a network traffic management system.
[0026] A network traffic management device can refer to an apparatus that performs one or more operations as described below in various examples optimized according to this disclosure. The network traffic management device may reside on any network device (e.g., a router, switch, smart NIC, or a combination of these functions such as a BIG-IP device) or component having the ability to intercept, analyze, and process traffic flows transmitted between client devices and network service devices, or reside on any network device or component communicatively connected thereto, to implement the operations of this disclosure.
[0027] A network service device can be any network device that provides services to client devices. Network service devices can be implemented in various ways, such as hardware, software, firmware, or any combination thereof. For example, a network service device can be a server for a network traffic management system (e.g., a network application server, such as...). Figure 1 The servers shown are one of 30(1)-30(n), which will be described below, or virtual machines, virtual servers, containers, engines, instances, etc. residing on servers or other network elements.
[0028] A client device can refer to any end-user device that can send or initiate requests to a network service device to establish or continue a communication connection with the network service device. Similar to a network service device, a client device can be implemented in various ways, including but not limited to hardware, software, firmware, or any combination thereof.
[0029] A representation may include one or more detection instructions for detecting patterns in a traffic flow. In this example, a pattern can refer to any characteristic of a traffic flow, including but not limited to attacks or specific types of attacks, characteristic use of a client device for a specific network function (e.g., a specific network service or network application or APP), utilization of a specific type of network resource by a specific network function, or user-defined patterns (e.g., detecting the number and / or frequency of times a specific API is used in a specific manner). One or more detection instructions included in a representation may individually or collectively describe one or more characters in a traffic flow whose presence indicates a match between the traffic flow and a pattern that the representation should embody and detect. The difference between representations lies in the different detection instructions that constitute them. Therefore, executing detection instructions within a representation enables network devices to perform pattern detection. Depending on the specific network environment and the tools used, detection instructions can be in any suitable form, as long as they can be used by network devices to detect various patterns of interest in a traffic flow. For example, detection instructions can be, but are not limited to, regular expressions, custom rules (irule), Python, or DEX programs. Regular expressions can be used in a variety of programming languages and tools to specify patterns in various tasks (e.g., search tasks). Python functions or code can execute on devices with more resources, but at the cost of latency. Irules can execute on routers to route traffic, during which a series of different pattern-recognition statements can be executed against the incoming traffic flow. Therefore, irules can be used for pattern detection, where the destination to be routed can indicate a match for the pattern being detected (e.g., if the destination is blocked, it might indicate an attack). Therefore, in the following description, in some cases, the detection instructions might be referred to as regular expressions, or, if coded and represented in Python, as Python functions.
[0030] A traffic flow can refer to one or more packets (e.g., data packets, control packets) transmitted in a network. A traffic flow can be a single packet or data stream that matches one or more patterns that a user might be interested in. In this example, "user" refers to an individual or entity that values pattern detection for various purposes as described above (e.g., an operator of an entity). It should be understood that such a user can be a user of a client device, a network service provider, a server administrator, a firewall, an entity providing security, management, or analysis services, or an entity playing any such role in network traffic flows, or any combination thereof.
[0031] The network device used in deployment optimization can refer to any physical or virtual network device or apparatus located between client devices and network service devices and processing data packets or traffic flows. For example, a network traffic management device or the device or apparatus on which a network traffic management device resides can be such a network device. As another example, any device or apparatus directly or indirectly connected to a network traffic management device can also be such a network device (e.g., an Internet of Things (IoT) device).
[0032] Figure 1 An exemplary simplified network traffic management system 100 according to examples of this disclosure is shown. Figure 1 As shown, the network traffic management system 100 may include multiple client devices 10(1)-10(n), a communication network 40, and multiple servers 30(1)-30(n) that provide services to these client devices 10(1)-10(n). The client devices 10(1)-10(n) and the servers 30(1)-30(n) can communicate with each other via the communication network 40.
[0033] refer to Figure 1 As an exemplary implementation of the aforementioned client devices, one of the client devices 10(1)-10(n) may (e.g., via a web browser installed on one of the client devices 10(1)-10(n)) send a service request to one of the servers 30(1)-30(n). The client devices 10(1)-10(n) may also be referred to as “clients,” “user devices,” or “user device devices,” and may include, but are not limited to, mobile phones, smartphones, tablets, laptops, smart electronic devices, wearable devices, video surveillance equipment, industrial wireless or wired sensors, or appliances including air conditioners, televisions, refrigerators, ovens, IoT devices, etc., or other devices capable of wireless communication via a network. Furthermore, one or more of the client devices 10(1)-10(n) may also be proxies or servers or any network element or device that can forward the aforementioned request, thereby initiating a service flow to one of the servers 30(1)-30(n) on behalf of its internal user devices. For example, one or more of client devices 10(1)-10(n) can be proxies (e.g., forward proxies) of a private network that forward request messages received from client devices isolated within the private network. In this way, the proxy sends request messages on behalf of the isolated devices and allows them to be served by one of servers 30(1)-30(n). In this case, the proxy, as... Figure 1 The network traffic management system 100 shown plays the role of one of the client devices 10(1)-10(n).
[0034] Continue to refer to Figure 1As an exemplary implementation of the aforementioned network service device, one of the servers 30(1)-30(n) may, in response to one of the client devices 10(1)-10(n) and in response to receiving a request from one of the client devices 10(1)-10(n) via the communication network 40, interact with one of the client devices 10(1)-10(n) once or more to provide the requested service or data. The servers 30(1)-(n) may be any type of server serving the client devices. For example, the servers 30(1)-(n) may be application servers that run applications and manage and perform various tasks related to the processing of requests from client devices in the network environment. The servers 30(1)-(n) may provide a variety of services.
[0035] like Figure 1 As shown, the communication network 40 may include multiple network elements 42(1)-42(n) that provide connectivity and data processing and transmission. Depending on the topology and characteristics of the communication network 40, various types of network elements 42(1)-42(n) (e.g., routers, proxies, load balancers, firewalls, etc.) may be used to perform specified functions. Figure 1 As shown, one of the client devices 10(1)-10(n) can communicatively connect to the communication network 40. When one of the client devices 10(1)-10(n) sends a message to request a service from one of the servers 30(1)-30(n), the message may pass through some of the network elements 42(1)-42(n) before reaching its destination. Therefore, such network elements 42(1)-42(n) can be the network device types described above for detecting service flow patterns, acting as intermediate devices between the client devices 10(1)-10(n) and the servers 30(1)-30(n). Therefore, such network elements 42(1)-42(n) or any suitable devices connected to network elements 42(1)-42(n) can be deployed using the optimized representation discussed in this disclosure. It should be understood that different network technologies can be applied to the communication network 40. For example, communication network 40 can be one or more wired or wireless public or private networks based on any industry standard protocol, such as Ethernet, Wi-Fi, satellite networks, 4G / LTE (Long Term Evolution), 5G, and various Internet protocols such as TCP / IP. Communication network 40 can also be formed by connecting an appropriate number of network connections as needed.
[0036] exist Figure 1In the network environment shown, to protect servers 30(1)-30(n) from attacks, or for anti-fraud (e.g., anti-bot) purposes, or for management purposes (e.g., load balancing, analysis such as statistics on the usage of certain network services or applications, resource usage statistics of certain network services and applications, etc.), detection of a pattern or patterns of traffic flows can be performed on appropriate devices. For such detection, sophisticated machine learning models can provide accurate identification, thereby accurately determining whether a given traffic flow matches a pattern of interest (e.g., an attack). However, due to the limited hardware and / or software resources of a given network element, resulting in limited processing capacity of that given network element, the traffic flow entering one of network elements 42(1)-(n) may be too large to handle, thus making the use of machine learning models on such network devices infeasible. Even if hardware upgrades or the deployment of additional equipment are tolerable from a cost perspective, this is still not a practical solution due to the long or sometimes huge latency introduced by performing pattern detection.
[0037] As an alternative solution, a representation can be generated, consisting of a series of regular expressions or signatures that can identify or characterize a given pattern (e.g., a specific type of attack). This representation can then be applied on network devices to detect and filter traffic matching the pattern. Implementing this representation is relatively more cost-effective (e.g., consuming fewer resources and introducing less latency) compared to machine learning model solutions. One problem with this solution is that a large number of regular expressions or signatures are available, and at least a certain number of these expressions and signatures are independent of each other. As a result, redundancy is common in the representation. For example, including... The representation can describe attacks in web request logs. However, another, smaller representation with fewer regular expressions, such as a single regular expression, is also possible. Similarly, the same attack can be described, and therefore web requests containing such attacks can be detected with similar or comparable accuracy. In other words, different regular expressions or different combinations of regular expressions can detect and identify the same important data that reflects business flow patterns. This means that different representations can be used to perform pattern detection with similar results but different processing performance (e.g., resource consumption, throughput, and latency). In this example, the number of regular expressions included in the representation may affect the processing performance of detecting business flow patterns using that representation. This is because even if executing a single regular expression is inexpensive, executing a representation consisting of a bunch or a large number of regular expressions for a large number of business flows is no longer cheap. Therefore, when performing pattern detection on business flows, an optimized representation can potentially improve processing performance. Various examples and operations for optimizing representations are described below. References Figure 1These operations can be performed on network traffic management device 20, which is deployed on any suitable device or component located between the client device and the network service device along the network communication connection established between the client device and the network service device (e.g., residing on an intermediate device, such as a firewall device, router or load balancer located between one of the client devices 10(1)-10(n) and one of the servers 30(1)-30(n)).
[0038] It should be understood that Figure 1 An exemplary simplified network traffic management system 100 is shown, which can be varied in many ways. For example, other types and numbers of systems, devices, components, and elements from other topologies can be used to add to or replace any part of the illustrated system. Furthermore, one or more components described in the network traffic management system 100 (e.g., network traffic management device 20) can be configured to run as virtual instances on the same or different physical machines. In some scenarios, the network traffic management device 20 can operate as multiple separate devices on different physical devices and communicate with each other as needed via communication network 40 or other related networks, rather than as... Figure 1 As shown, it runs on the same physical device.
[0039] Figure 2 An exemplary execution environment 200 for a network traffic management device 20 is illustrated. In the execution environment 200, the network traffic management device 20 may include a processor 22, a memory 24, a communication interface 26, and / or other circuitry coupled together via a bus 202 or other communication link. It should be understood that the network traffic management device 20 may include other types and / or numbers of other configured elements. The processor 22 of the network traffic management device 20 can execute programming instructions stored in the memory 24 of the network traffic management device 20 for any number of operations or tasks identified in this disclosure. For example, the processor 22 of the network traffic management device 20 may include one or more central processing units (CPUs) or a general-purpose processor with one or more processing cores, but other types of processors may also be used. The communication interface 26 may support wireless (e.g., Bluetooth, Wi-Fi, WLAN, cellular (4G, LTE / A, 5G)) and / or wired, Ethernet, Gigabit Ethernet, and optical network protocols. Communication interface 26 may also include a serial interface, such as Universal Serial Bus (USB), Serial ATA, IEEE 1394, Lightning port, I2C, SlimBus, or other serial interfaces. In some examples, execution environment 200 may also include power functions and various input interfaces. Figure 2(Not shown in the image). In some examples, the execution environment 200 may also include a user interface, which may include a human-machine interface device and / or a graphical user interface (GUI).
[0040] The memory 24 of the network traffic management device 20 may store non-transitory computer-readable instructions for programming one or more aspects of the techniques described and illustrated herein, although some or all of the programming instructions may be stored elsewhere. Various types of memory storage devices, such as random access memory (RAM), read-only memory (ROM), hard disk drive (HDD), solid-state drive, flash memory, erasable programmable read-only memory (EPROM), or other computer-readable media (e.g., disks or optical discs, e.g., CD-ROMs) that are read and written by magnetic, optical, or other machine-readable media coupled to the processor 22, may be used as memory 24. Therefore, the memory 24 of the network traffic management device 20 may store applications that may include computer-executable instructions that, when executed by the network traffic management device 20, cause the network traffic management device 20 to perform actions or operations, such as sending, receiving, or otherwise processing messages, and performing other actions or operations described and illustrated below with reference to the accompanying drawings. Applications may be implemented as units, modules, components, instances, or engines of other applications and / or operating system extensions, plug-ins, etc. Applications can run within virtual machines or virtual servers managed in a cloud-based computing environment, or as such virtual machines or virtual servers, without being tied to one or more specific physical network devices.
[0041] The methods, apparatus, processes, circuits, and logic described below can be implemented in many different ways and in many different combinations of hardware, software, firmware, or combinations thereof. For example, all or part of the implementation may be a circuit including an instruction processor, such as a central processing unit (CPU), microcontroller, or microprocessor; or as an application-specific integrated circuit (ASIC), programmable logic device (PLD), or field-programmable gate array (FPGA); or as a circuit including discrete logic or other circuit components, including analog circuit components, digital circuit components, or both; or any combination thereof. For example, the circuit may include discrete interconnected hardware components, or may be combined on a single integrated circuit die, distributed among multiple integrated circuit dies, or implemented in a multi-chip module (MCM) of multiple integrated circuit dies in a common package.
[0042] Therefore, the circuit can store or access instructions for execution, or its function can be implemented solely in hardware. Instructions can be stored in a tangible storage medium (e.g., memory 24) other than transient signals. Products such as computer program products may include storage media and instructions stored in or on the media, and when the instructions are executed by circuitry in the device, the device can perform any of the processes shown above or in the accompanying drawings.
[0043] The implementations discussed herein can be distributed. For example, the circuit may include multiple different system components, such as multiple processors and memories, and may span multiple distributed processing systems. Parameters, databases, and other data structures may be stored and managed separately, may be merged into a single memory or database, may be logically and physically organized in many different ways, and may be implemented in many different ways. Example implementations include linked lists, program variables, hash tables, arrays, records (e.g., database records), objects, and implicit storage mechanisms. Instructions may form parts of a single program (e.g., subroutines or other code segments), may form multiple separate programs, may be distributed across multiple memories and processors, and may be implemented in many different ways. Example implementations include standalone programs, as well as shared libraries as part of a library, such as dynamic link libraries (DLLs). For example, the library may contain shared data and one or more shared programs that include instructions that, when executed by the circuit, perform any of the processes shown above or in the accompanying drawings.
[0044] Reference Figure 3 An exemplary block diagram of a network traffic management device 20 for optimizing representation is shown. Figure 3 In this context, the network traffic management device 20 may include a transceiver unit 240, a candidate generation unit 242, a simulator 244, and an evaluator 246. This will be combined with... Figure 4 The flowcharts shown describe the operations performed by these units. These units described herein can be implemented using various available or appropriate programming APIs, such as JavaScript, Python, etc.
[0045] The term "unit" (and other similar terms, such as module, submodule, etc.) can refer to computing software, firmware, hardware, and / or various combinations thereof. However, at a minimum, a unit should not be construed as software not implemented on hardware, firmware, or recorded on a non-transitory, processor-readable, and recordable storage medium. In fact, a "unit" should be construed as including at least some physical, non-transitory hardware, such as a processor, circuitry, or part of a computer. Two different units may share the same physical hardware (e.g., two different units may use the same processor and network interface). Units described herein can be combined, integrated, separated, and / or replicated to support a variety of applications. Furthermore, functions described in this example as performing in a particular unit may be performed in one or more other units and / or by one or more other devices, in place of or complementing the functions performed in that particular unit. Furthermore, these units may be implemented across multiple devices and / or other components, locally or remotely to each other. Furthermore, these units may be removed from one device and added to another, and / or contained in two devices. These units may be implemented in software stored in memory or a non-transitory, computer-readable medium. Software stored in memory or media can run on a processor or circuit capable of executing computer instructions or computer code (e.g., ASIC, PLA, DSP, FPGA, or any other integrated circuit). These units can also be implemented in hardware using processors or circuits on the same or different integrated circuits.
[0046] Figure 4 A flowchart of an exemplary process 400 for optimizing representation is shown, which can be implemented or executed by a network traffic management device 20. As mentioned above, the network traffic management device 20 can reside on and be implemented on any suitable device. Furthermore, the network traffic management device 20 can be distributed across different devices in the network. The following will combine... Figure 3 The logic of the network traffic management device 20 shown is described. Figure 4 The steps are shown.
[0047] In step 401, the transceiver unit 240 of the network traffic management device 20 can access the data from the storage device 302 (e.g., Figure 5 The data source 502 retrieves a representation and the dataset associated with that representation, although representations can be stored and retrieved from other locations. It should be understood that "one (a)" representation does not limit the number of representations to be retrieved. That is, in step 401, one, two, or any appropriate number of representations can be retrieved.
[0048] The dataset associated with the retrieved representation may include various records or events related to the representation. For example, the dataset may be records or events captured from a business flow that matches a detected pattern. As a non-limiting example for illustration only, a record may be a line from one or more log files, the name and path of a program executing on some host, the CPU and memory usage of these programs, files accessed by the program, etc. The dataset stored in storage device 302 may be preprocessed manually, semi-manually, or automatically to facilitate subsequent operations. Figure 5 An exemplary data source 502 is shown, from which representations and datasets can be retrieved. User 504 can preprocess the data stored or maintained in data source 502. Here, user 504 can be a user of the network traffic management device 20 that generates the optimized representation, or a user of the network device 304 that deploys the generated optimized representation, or both. For example, user 504 can label or tag a set of dominant representations, dominant or popular detection instructions, or any combination thereof (e.g., a set of regular expressions or rules), and organize them into one or more clusters. In this example, a cluster may be associated with one or more similar patterns.
[0049] User 504 can also label or annotate datasets associated with the representations and / or detection instructions associated with these labels. In this regard, User 504 can label the data manually or automatically, for example, using a classification model. Alternatively, User 504 can choose to generate related data (e.g., characters similar to those in the labeled datasets or related detection instructions or representations), which can be, for example... Figure 5 The labeled dataset shown may be all or part of it. This can be implemented by a generator that utilizes machine learning techniques, such as large language models (LLMs) or hidden Markov models, to generate data of a specific type and / or with specific characters, which can be specified by the user. Furthermore, Figure 5 User preferences that can be optionally stored in data source 502 are also shown, which will be described in more detail below.
[0050] It should be understood that, such as Figure 5The data maintained in data source 502 shown may be historical data and may be updated from time to time. Therefore, over time, the number of clustering and label detection instructions may increase significantly, as may the number of labeled datasets. As mentioned above, the number of detection instructions may be very large, resulting in a large number of related datasets, possibly several times the number of detection instructions. The same applies to the label detection instructions or representations and the associated labeled datasets. Therefore, it may be impossible to provide all labeled datasets associated with a given retrieval representation. Therefore, in this case, sampling techniques can be used when retrieving related datasets. Sampling techniques may include sampling rules, such as random sampling. Sampling rules may also involve other aspects, such as relevance, where labeled datasets with high deterministic relevance (i.e., highly relevant to the retrieval representation) may be sampled first or earlier than other labeled datasets with lower deterministic relevance. Similarly, sampling operations can be applied when transceiver unit 240 retrieves representations. For example, representations that have been determined (e.g., labeled by user 504) to consume fewer resources from network device 304 (e.g., below a corresponding predetermined threshold), introduce less latency (e.g., below a corresponding predetermined threshold), or have high accuracy (e.g., above a corresponding predetermined threshold) may be sampled first. Here, sampling rules can be determined or selected based on user 504's preferences (e.g., user 504's dominant or interested detection instructions or representations).
[0051] In step 402, the candidate generation unit 242 of the network traffic management device 20 can prompt the natural language processing model to convert the retrieved representation into a candidate representation. In this example, "one (a)" candidate representation does not limit the number of generated candidate representations to one. Instead, the candidate generation unit 242 can generate one or more candidate representations. For example, when the transceiver unit 240 retrieves multiple representations from the storage device, the candidate generation unit 242 can perform a prompting operation on each retrieved representation sequentially or in parallel. Alternatively, the candidate generation unit 242 can run the retrieved multiple representations together in a single prompting operation and generate one or more candidate representations for these retrieved multiple representations. In the case of retrieving multiple representations, as described above, these representations can be retrieved through a random sampling operation. Alternatively, for example, if a sampling rule for sampling representations with high relevance is used, all or part of the retrieved multiple representations can be a set of highly relevant or similar representations.
[0052] The transformation in step 402 can be based on a dataset associated with the retrieval representation, detection instructions, or both. The generated candidate representation differs from the representation input to the candidate generation unit in the detection instructions included in the candidate representation. For example, the generated candidate representation could be a more compact representation with fewer detection instructions, or a completely different representation with no general detection instructions, or with some general detection instructions but also one or more new detection instructions, etc.
[0053] like Figure 3 and Figure 6 As shown, the candidate generation unit 242 itself may include one or more natural language processing models (e.g., large language models), or be communicatively connected to these models. Here, multiple natural language processing models can be utilized. For example, in the process of generating candidates from... Figure 5 In the case where data source 502 retrieves multiple representations and performs the prompting operation of step 402 on each of these representations, there may be multiple LLMs, each processing a portion of the retrieved representation and generating a corresponding candidate representation. In some other examples, when a representation includes a series of detection instructions, multiple LLMs can be used, each processing one or more detection instructions for that representation and generating a portion of the candidate representation (e.g., one or more candidate detection instructions). Next, candidate generation unit 242 can combine these candidate portions into a single candidate representation. Alternatively, each candidate portion can be processed individually and... Figure 6 The evaluators 246 at the mid-to-end are combined together. In some examples, depending on the complexity of the model, the computing power of the device, or other factors, a single LLM can be used to perform step 402.
[0054] Figure 6 An exemplary flowchart is shown, in which, from Figure 5 The output data 506 of the data source 502 shown is used as input data 602 for optimized representation. Figure 6 In the candidate generation unit 242, multiple LLMs are deployed to generate candidates based on the input data 602.
[0055] Figure 7 A prompt was shown Figure 6 The two LLMs in the example generate candidate exemplary flowcharts as Python functions. Figure 7 As shown, it can be automatically input (e.g., pre-entered by user 504 and stored in network traffic management device 20, or automatically generated by network traffic management device 20) or by user (e.g., Figure 5The user's manual input (in the implementation of the prompting operation, 504) can come from multiple regular expressions of one or more representations to be included in the prompt. Furthermore, relevant datasets are also included in the prompt. Here, relevant datasets include not only positive datasets but also negative datasets. Here, a positive dataset refers to a dataset that matches the retrieval representation or any detection instructions contained within the retrieval representation, while a negative dataset refers to a dataset that does not match. For example, if the representation is used to detect attacks, positive records include data that has been identified as related to a real attack, and negative records include data that has been identified as not being attacked. It should be understood that providing both positive and negative datasets (e.g., including positive and negative records and events) may be beneficial for LLM candidate generation, but it is not necessary. In other examples, providing only positive datasets to the LLM is also an option for candidate generation unit 242 to generate candidates. Similarly, in other examples, only negative datasets may be provided to the LLM so that candidate generation unit 242 can generate candidates. By performing the prompting operation in step 402, candidate representations different from the input representation are generated, which can capture and combine all the insights from the detection instructions of the input LLM and the relevant datasets. In this example, the generated candidates are not limited to producing the same detection instructions included in the retrieved representation or stored in data source 502. Instead, in other examples, candidate generation unit 242 may include one or more new detection instructions in the candidate representation. Here, as an example, candidate generation unit 242 can generate a candidate representation by including a different set of detection instructions in the candidate representation (e.g., combining regular expressions into a Python program and then translating it into irule). In this way, the generated candidate representation can represent the semantics of the retrieved representation in an improved manner. For example, the generated candidate representation can represent a set of regular expressions as a single regular expression, or with a more efficient syntax. In some other examples, the generated candidate representation can convert the semantics of a retrieved representation in one language into a representation with similar semantics in a different language (e.g., converting a regular expression into a Python function), which can also be called a heterogeneous representation.
[0056] exist Figure 7In this example, user 504 can specify the form of the generated candidate representation (e.g., Python function, regular expression, irule, etc.). As another example, user 504 can specify the form as a Python program. With this input from user 504, candidate generation unit 242 can generate candidates in the form of a Python program using some high-level C code or library. It should be understood that this is an additional option available to user 504. However, the input from user 504 specifying the form is not required to perform the operations discussed herein. Instead, the default form of the generated candidate representation can be predetermined or set in network traffic management device 20. In other examples, candidate generation unit 242 can be configured to automatically generate candidate representations of appropriate forms without requiring input from user 504.
[0057] In step 403, the candidate representation generated by candidate generation unit 242 can be input into simulator 244 to generate one or more candidate metrics corresponding to one or more preferred metrics specified by the user. In this example, simulator 244 can provide one or more metrics and generate candidate metric values based on each metric, i.e., metric by metric. Each metric can measure the processing performance of the representation used to detect business flow patterns from a different perspective.
[0058] As an example, Figure 8 An exemplary preference selection interface 800 for specifying preferred metrics is shown for user 504. Figure 8 The system provides user 504 with six metrics, including false positive rate 802-1, true positive rate 802-2, false negative rate 802-3, latency 802-4, CPU utilization 802-5, and memory utilization 802-6. For each of these metrics, user 504 can select a metric value that the user prefers for an optimized representation or a batch of optimized representations (e.g., representations of the same or similar types or categories) to have (e.g., selectable preferred metric values 804-1, 804-2, 804-3, 804-4, 804-5, and 804-6). It should be understood that user 504 is not required to select a metric value for every metric provided. Instead, user 504 can enter only the preferred metric values for one or more metrics that are important to the user, leaving the rest blank (e.g., user 504 selecting "N / A" in the optional preferred metric value ranges 804-5 and 804-6). As a non-limiting example, a metric for user preference might be twice as strong as a metric for user dislike (e.g., metric values of 10 and 5 respectively), and a metric that the user doesn't care about at all could be set to zero or N / A. Figure 8The interface presents users with descriptive metrics such as "high" and "medium" for selection; understandably, other descriptive metrics are also appropriate (e.g., "low"). Alternatively, users could be allowed to manually set metrics (e.g., coefficients), or provided with an additional box to add extra metrics not offered in the preference selection interface 800. For example, users could enter a normalized metric (e.g., normalizing CPU utilization to between 0 and 1) or the current number of CPU cores (e.g., from 0 to an integer). Therefore, Figure 8 The exemplary preference selection interface 800 provides user 504 with a wide range of options. This allows user 504 to create a range of different preferences (e.g., minimizing latency at a lower precision cost, reducing false positive rate and latency) by combining these metrics 802-1 to 802-6 with different values 804-1 to 804-6. In this regard, user 504 can specify or input their preferences, taking into account which network device is used to perform pattern detection, or the actual network environment (e.g., characteristics and requirements in a real environment). There may be a balance between different metrics so that user 504 can decide which metric is more important and which is less important in a particular scenario. Furthermore, the preferred metric values 804-1 to 804-6 specified by user 504 can be based on representation or detection patterns. For example, user 504 can specify the same preferred metric value applicable to all optimized representations used to detect the same or a set of similar patterns. User 504 can pre-enter such preference information, which may be stored in the network traffic management device or entered in real time when performing the operations described herein.
[0059] In this example, the false positive rate 802-1 can refer to a misclassification rate or inaccuracy rate: when performing pattern detection representation, the detection result indicates that a flow matches a pattern, but a false positive occurs because the flow does not actually match the pattern (e.g., indicating an attack in the flow during attack detection, but it turns out not to be an attack). Similarly, the false negative rate 802-3 can refer to a misclassification rate or inaccuracy rate: the detection result indicates a mismatch between the flow and the pattern, but a match actually exists (e.g., indicating no attack in the flow, but it turns out to be an example of an attack). The false positive rate 802-1 can be calculated as FP / (FP+TN), that is, the ratio of FP to (FP+TN). Here, PF is the number of negative events that are misclassified as positive (false positive), TN is the number of true negative events, and (FP+TN) is the total number of actual negative events. It should be understood that, in contrast to the false positive rate 802-1, the true positive rate 802-2 can refer to an accuracy rate: when executing a pattern detection representation, the detection result indicates that the traffic flow matches the pattern and that a match actually exists between them (i.e., a true positive event correctly classified as positive). The metrics for latency 802-4, CPU utilization 802-5, and memory utilization 802-6 refer to the time required to execute the pattern detection representation, the latency introduced into the traffic flow transmission, and the amount of CPU and memory used during execution.
[0060] It should be understood that Figure 8 The metrics 802-1 to 802-6 shown are for illustrative purposes only, and various other metrics (e.g., classification accuracy) may be provided. These metrics are not limited to processing performance, as long as they provide performance indicators that user 504 may be interested in. In some examples, multiple preference selection interfaces may be provided to user 504, each corresponding to a specific one of multiple simulators. This means that in some examples, multiple simulators may be provided to user 504.
[0061] As an example, in Figure 6 Simulator 244 illustrates multiple simulators performing simulations, corresponding to the number of LLMs included in candidate generation unit 244. These simulators may be identical to each other or may differ based on different preferences specified by user 504. It is also understood that user-specified metrics (e.g., Figure 8The 804-1 to 804-6 in the candidate representation can be any suitable form, such as numbers (e.g., the number of detection instructions that can be included in the candidate representation, such as one, two, or “equal to or less than” an integer), natural language descriptions (e.g., “high,” “medium,” “low,” “better performance” of the candidate representation compared to the retrieval representation or detection instructions), “fewer” detection instructions included in the candidate representation (i.e., more compact detection instructions compared to the retrieval representation), such as trade-offs that reduce latency at the cost of a higher false positive or false negative rate), thresholds, etc. Figure 8 The objective function 806 is also shown, which will be described below in conjunction with step 404.
[0062] Return to reference Figure 3 and Figure 6 The simulator 244 can be implemented in various ways. For example, the simulator 244 can be a virtual machine that measures the performance of the candidate representations generated by the candidate generation unit 242 and generates corresponding candidate metrics. In this example, the measurement is performed by measuring the metrics when candidate representations are executed on simulated data (i.e., data used for simulation purposes, to which candidate representations are performed to generate metrics). For example, the simulator 244 can measure how much latency 802-4 exists when candidate representations are executed on simulated data and generate candidate metrics for the latency metric. As another example, the simulator 244 can monitor whether the candidate representation can detect the pattern and its accuracy (e.g., false positive rate 802-1, false negative rate 802-3, true positive rate 802-2, or...). Figure 8 (True negative rate not shown). In this example, performance and user 504 specify one or more corresponding preferred metrics (e.g., Figure 8 This relates to one or more metrics (804-1 to 804-4) in the dataset. This means that the simulator 244 does not necessarily measure all metrics provided to the user (e.g., ...). Figure 8 The performance of 802-1 to 802-6 in the standard is measured, while only a subset of the metrics that the user is interested in (e.g., in the standard) is measured. Figure 8 (The middle part is 804-1 to 804-4).
[0063] In some examples, simulator 244 can use natural language processing models (such as LLM) for simulation. In this regard, simulator 244 can simply use storage devices (e.g., Figure 3 Storage device 302 or Figure 5The LLM samples a certain amount of real data from data source 502 as simulation data. In some other examples, the LLM can generate synthetic data from real or test environments to simulate complex scenarios or behaviors that occur in historical business flows, as supplementary simulation data. For example, the prompt could be "This is a candidate representation, and this is the original representation. Please generate synthetic data that helps me distinguish between the two representations."
[0064] Next, the generated synthetic data can facilitate simulation, thereby facilitating subsequent evaluation of whether the generated candidate representations are optimized from certain perspectives. As another example, if the pattern to be detected is attack-related, the LLM can simulate complex benign and malicious behaviors based on the generated synthetic data (e.g., business flows initiated by attackers and non-attackers traversing a website, respectively). In some examples, during synthesis, the simulator 244 can also utilize the candidate representations to generate relevant synthetic data as supplementary simulation data. Alternatively, in some examples, the simulator 244 can retrieve data from the real environment that is associated with the candidate representations (e.g., data from...). Figure 5 Data from data sources (502 or other sources).
[0065] Next, simulator 244 can randomly combine all these related data to perform the simulation. Figure 6 In the example shown, the method is implemented iteratively, as will be described below, and simulator 244 can also utilize any candidate representations generated in previous iterations. Taking the network traffic management device 20 residing on a virtual BIG-IP device as an exemplary application scenario, simulator 244 can be implemented as an irule simulator. This irule simulator can generate synthetic data with similar characteristics to production data. Furthermore, the irule simulator can also use LLM to generate additional data that includes additional behaviors that enhance the simulated data.
[0066] In some examples, simulator 244 can generate candidate metrics based on the representation. This means that if a given candidate representation includes multiple detection instructions, the candidate metric for a particular metric indicates the overall performance of the representation, regardless of the individual contributions of each detection instruction. Next, the entire representation is evaluated in step 405, which will be described in detail below. In some other examples, simulator 244 can instead generate candidate metric values for each detection instruction included in the representation. In other words, simulator 244 generates candidate metrics based on detection instructions or levels. Then, when proceeding to step 405, the detection instructions can be evaluated individually.
[0067] In some examples, the original representation retrieved in step 401 is also fed into simulator 244 to generate metrics that display the performance difference between the original representation and the generated candidate representation. This provides a relatively straightforward comparison to show whether a candidate representation is an optimized representation compared to the original representation, and by how much. However, this is not necessary for performing the operations discussed in this example. For example, only the generated candidate representations might be fed into simulator 244 to obtain the performance indicated by the generated candidate metrics. These candidate metrics can then be compared to predetermined criteria or rules (e.g., acceptable value ranges or thresholds) to check whether the candidate is optimized.
[0068] In step 404, the evaluator 246 of the network traffic management device 20 can evaluate whether to replace the retrieval representation with the generated candidate representation. For example, this can be based on the candidate metric generated by the simulator 244 in step 403 and the preferred metric of the user 504 (e.g., by inputting the preferred metric). Figure 8 The evaluation is performed using the preference selection interface 800. Specifically, if the generated candidate metric satisfies the specified preferred metric (e.g., based on the user's preferred metric, the generated candidate metric indicates lower latency, higher true positive rate, and lower false positive rate), it means that the candidate representation is optimized from at least some perspective. In this case, the evaluator 246 can indicate a replacement. As another example, if the user 504 specifies a preferred range, the generated candidate can replace the original representation retrieved in step 401 as long as the generated candidate metric falls within that range. If the original representation is also input into the simulator in step 403, as described above, the evaluator 246 can also consider whether the candidate is superior to the original representation from any perspective when making a decision (e.g., whether the candidate is optimized in at least one perspective / relative to one metric, or whether the candidate is better based on user preferences). Based on the preferences specified by the user 504, for example, if the candidate representation is optimized in one or a number of perspectives (which may be reflected by one or more metrics), the evaluator 246 can decide whether to perform a replacement if the candidate representation can replace the original representation. If user 504 specifies that their preference is for candidates to be superior to the original representation, evaluator 246 can make a decision in a similar manner. In other examples, evaluator 246 can perform evaluations in various different ways, such as by comparing each generated candidate metric with a corresponding predetermined value. The predetermined value can be a default value pre-set at network traffic management device 20 or pre-entered by user 504.
[0069] In the example described above where simulator 244 generates candidate metrics based on detection instructions, evaluator 246 can also perform evaluation based on detection instructions, i.e., evaluating each detection instruction individually. In this case, different detection instructions can utilize separate evaluation rules or criteria. These separate evaluation rules can be the same as or different from each other.
[0070] In some examples, optionally, one or more preferred metrics specified by the user and the generated candidate metrics can be quantized using an objective function. Figure 8 An exemplary objective function 806 is illustrated. Objective function 806 encodes user preferences for various metrics. By inputting generated candidate metric values into objective function 806, more than one candidate metric value can be combined into a single result (e.g., an overall numerical score indicating the degree or range of optimization). In some examples that also simulate the original representation, the corresponding metric values of the original representation can also be input into objective function 806 to obtain numerical scores. Therefore, a comparison of two numerical scores can give an indication of the degree of optimization of the candidate representation. In some other examples, one or more preferred metric values 804-1 to 804-6 specified by user 504 can be input into objective function 806 to obtain preferred numerical scores quantifying one or more preferred metric values 804-1 to 804-6. For example, objective function 806 can be used to calculate a final score by considering user-inputted or selected metrics, where metrics whose values are set to zero are considered irrelevant to the final calculation result. Figure 8 As shown, the user selects a relative value of "high" (804-1) for metric 802-1, "medium" (804-4) for metric 802-4, and N / A (804-5 and 804-6) for metrics 802-5 and 802-6. Objective function 806 converts "high" (804-1) to a coefficient of "10" representing higher priority in the calculation, while converting "medium" (804-4) to a smaller coefficient of "5" representing lower priority, and assigns "0" to N / A (804-5 and 804-6). It should be understood that... Figure 8The allocation coefficients in objective function 806 are for illustrative purposes only. In practical applications of the examples described in this disclosure, other appropriate values can be assigned to different metrics, thereby assigning them weights. The benefit of objective function 806 to the user is that it directly presents a total score after calculating the correlation of various variables represented by the metrics, which may be more objective and can be implemented automatically. In this case, evaluator 246 can also assess the degree of optimization by comparing the numerical scores of candidate representations with the preferred numerical score. In the example of calculating the preferred numerical score, evaluator 246 can also set this score as the target value for user 504, reaching which indicates that maximum optimization has been achieved. In some other examples, the results obtained by inputting metric values into the objective function may not be a single result. Instead, the results can be metric-based, with each metric having a numerical score. Such results can indicate any improvement or optimization in a more direct way.
[0071] It should be understood that the evaluator 246 is deployed to evaluate whether and to what extent the candidate representations generated by the candidate generation unit 244 meet the user's goals. In other words, it is used to evaluate whether there is optimization in the candidate representation or detection instruction, the degree of optimization, and whether it meets the user's goals or preferences. Therefore, in some examples where the candidate generation unit 242 generates multiple candidate representations, the evaluator 246 can determine which candidate to replace the original representation with the best one by evaluating the individual candidate metrics generated by the simulator 244. Figure 6 As shown, the best candidate can be stored in storage device 604 and labeled as a new cluster. If the generated candidate representation is not optimized in any way, or is not optimized based on the specified preferred metric as expected by user 504, evaluator 246 can decide to store the original representation retrieved in step 401 in storage device 604 (e.g., in a separate cluster that can be used for feedback or simulation). Alternatively, as... Figure 5 As shown, evaluator 246 can include the raw representation in the feedback and send the feedback to data source 502. This raw representation can be stored in data source 502 and retrieved again when performing step 401.
[0072] Figures 9-11 Exemplary simulation and evaluation processes are shown respectively. Figure 9 In this context, LLM 902-1 and 902-2 are used to generate synthetic data as simulation data, which is stored in database 904. For example... Figure 9 As shown, the evaluation function 908, as a non-limiting example of the objective function described above, is used for the evaluation performed by the evaluator 246. Figure 9In this process, candidate representations are generated as Python functions. The candidate representations are simulated, during which candidate metrics are generated, and then these candidate metrics are fed into the result evaluator 906 to perform the aforementioned evaluation.
[0073] exist Figure 10 In this process, the data used to perform the simulation is captured from the real environment; that is, requests to web server 1002 are intercepted and then manually or automatically tagged. The tagged data is then stored as simulation data in database 1004. Next, with... Figure 9 Similarly, the generated candidate representations are input into simulator 244. Next, the generated candidate metrics are input into result evaluator 1006 for evaluation using evaluation function 1008. For example... Figure 10 As shown, candidate representations are generated into one or more regular expressions.
[0074] exist Figure 11 In the simulation data, with Figure 10 The generation process is similar. The difference is that the candidate representations are generated using an nginx configuration, as simulator 244 and evaluator 246 run on an nginx server. (Details omitted here.) Figure 11 Zhongyu Figures 9-10 Similar operations.
[0075] In some examples, if it is determined in step 404 that the generated candidate is optimized compared to the representation retrieved in step 401, steps 402-404 can be performed on the generated candidate. The new candidate representation generated in step 402 can be referred to as the second candidate representation, while the original candidate on which the prompting operation was performed can be referred to as the first candidate representation. Specifically, a dataset associated with the first candidate representation can be retrieved from a storage device (e.g., storage device 302 or data source 502), and the prompting operation of step 402 can be performed on the first candidate representation based on the retrieved dataset. As mentioned above, a large amount of relevant datasets may be stored in the storage device. Therefore, when performing the prompting operation to generate the first candidate representation, there may be a limitation on the amount of data that can be input into the natural language processing model, which may be a common situation in real-world environments. Therefore, by performing step 402 on the first candidate representation based on the relevant dataset retrieved from the storage device to additionally generate the second candidate representation, the first candidate representation can be refined or optimized.
[0076] In some other examples, alternatively, additional representations may be retrieved from a storage device, for example, through random retrieval such as random sampling, or through retrieval based on relevance rules, thereby retrieving representations related to the first candidate representation. Those additionally retrieved representations may include the basis for the cueing operations of the first candidate representation (e.g., including in...). Figure 7(As indicated in the prompt). Thus, the number of representations stored in the storage device (e.g., storage device 302, data source 502) that have not yet been processed by the operations described in this disclosure may decrease over time. On the other hand, for example, by combining multiple representations into an optimized representation, discarding redundant representations, etc., the number of optimized representations (e.g., stored in storage device 604) may be less than the number of representations initially stored in the storage device (e.g., storage device 302, data source 502). Similarly, it should be understood that the number of clusters and the size of clusters can also be reduced in the optimized representations compared to the representations initially stored in the storage device (e.g., storage device 302, data source 502).
[0077] Next, the simulation operation of step 403 can be performed on the newly generated second candidate representation to generate one or more second candidate metrics. Subsequently, the newly generated second candidate representation is evaluated in step 404. In this case, the evaluation is performed based on the generated first and second candidate representation metrics. In this regard, the two sets of metrics can be compared, and if the evaluator 246 determines that the second candidate representation is optimized, then the second candidate representation can replace the first candidate representation. This optimized second candidate representation can then be considered the best candidate and stored. Figure 6 The storage device 604 is used in the evaluation. It should be understood that, optionally, preferred metrics of user 504 may also be considered during the evaluation.
[0078] In some examples, steps 402-404 can be performed iteratively (e.g., repeated for the second candidate representation described above). Therefore, the first and second candidate representations as used herein do not necessarily refer to candidate representations generated exactly in the first and second iterations. Rather, they can refer to two candidates generated consecutively at any stage of the iteration, where the candidate representation generated first is called the first candidate representation, and the candidate representation generated in the next iteration based on the first candidate representation can be called the second candidate representation. This iterative strategy may be suitable in practical application environments for various reasons. For example, as mentioned above, due to the limited size of the dataset, the generated candidate representations can be further refined, and the representation can be used during the prompting operation in step 402. As another example, even if a candidate representation satisfies or conforms to preferred metrics, it can be determined during evaluation that the candidate representation can be further optimized. This determination can be based on predetermined upper and lower limits (e.g., user 504 input or default settings in network traffic management device 20) or on a maximized / minimized value calculated using an objective function (e.g., objective function 806). As a non-limiting example, if the objective is to maximize a metric (e.g., true positive rate) or the overall function, a candidate representation is considered optimized or better than the input retrieval representation if the calculated value after simulation is larger. Similarly, if the objective is to minimize a metric (e.g., latency) or the overall function, then the lower value wins. By iteratively performing these steps, the calculated values of the generated candidates can be continuously increased or decreased until no better representation is available. Therefore, it should be understood that at each iteration, if the simulation and evaluation described above are performed on a representation-based basis, and an overall result is generated during the evaluation, the candidate representation can be optimized overall (e.g., by removing or changing certain detection instructions). In some other examples, if the simulation and evaluation described above are performed on a metric-based basis, the candidate representation can be optimized within a specific metric. For example, optimization could include a smaller representation size, lower latency, a lower false positive / false negative rate, a higher true positive / true negative rate, lower CPU utilization, or any combination thereof.
[0079] The iteration can be repeated until, for example, no progress is made in any aspect of user preference (e.g., a certain level, threshold, or criterion has been reached), maximization has been achieved (e.g., calculated using the objective function), and all labeled datasets or representations in the storage device have been exhausted. For example, the preferences specified by user 504 may include, but are not limited to, a specified number of iterations, a specific performance threshold, a minimum or maximum number of detection instructions included in the candidate representations, and a time period for performing the iterations (e.g., overnight, week, month, etc.).
[0080] In some examples, the form or format of the generated candidate representation may change with each iteration. For example, the candidate representation may first be generated as a Python function, and then in subsequent iterations as an irule or a regular expression represented as an irule. This change may be caused by different preferences specified by user 504 (e.g., entered at the time of operation or entered in advance), such as user 504 specifying different forms at different stages. Alternatively, it may be changed automatically by network traffic management device 20, such as evaluator 246 or objective function determining the change.
[0081] It should be understood that in some examples, Figure 4 The entire process illustrated can be performed sequentially or in parallel iteratively on the representations stored in the storage device (e.g., storage device 302). For example, when the retrieved representation can no longer be optimized or refined, a new round can be performed on the remaining representations stored in the storage device. Figure 4 The operation shown involves retrieving the additional representation from the storage device by repeating step 401. In some examples, such as... Figure 5 As shown, the generated candidate representations that have been determined to be optimized can be included in the feedback of evaluator 246. The feedback is sent to data storage device 502 and can be stored in the relevant clusters. Subsequently, the candidate representations can be combined with other labeled representations in the clusters, and if retrieved by transceiver unit 240, can later proceed to iteration. In this respect, even if it differs from the refinement of a specific candidate representation as described above, the overall refinement of the representations stored in data source 502 can be implemented through iterative execution process 400.
[0082] In step 405, in response to the evaluator 246 deciding to replace the retrieved representation with a candidate representation, the transceiver unit 240 may deploy the candidate representation on the network device 304. Next, the network device 304 may perform pattern detection for incoming traffic flows by deploying the candidate representation generated by the candidate generation unit 242 as a candidate.
[0083] Back Figure 6 As described above, after evaluation by evaluator 246, the best candidate representation is stored in a new cluster in storage device 604. This may occur if more than one candidate representation is generated in step 402, or if steps 402-404 are performed iteratively on the generated candidates. In some examples, all generated candidate representations that satisfy the preferred but not optimal metric specified by user 504 may also be maintained (e.g., stored separately and used to generate synthetic data for simulation, or to generate feedback to simulator 244, candidate generation unit 242, or to obtain data from...). Figure 5(Data source retrieval information). In some other examples, generated candidate representations that are determined not to be the optimal representation are also stored in the storage device (e.g., stored separately from the best candidate representation in storage device 604). In this respect, these generated candidate representations can be used as counterexamples in the training or simulation data of the LLM.
[0084] like Figure 6 As shown, feedback from evaluator 246 can be sent to simulator 244 to improve the simulation, to candidate generation unit 242 to improve the quality of its generated candidates (e.g., refine the process of generating candidates with lower latency), and to data source 502 to improve data retrieval (e.g., retrieve more dominant representations and / or detection instructions, highly relevant datasets). Similarly, even Figure 6 As not shown in the diagram, simulator 244 can also provide candidate generation unit 242 and Figure 5 Feedback is provided by data source 502. Similarly, feedback from candidate generation unit 242 can also be sent to data source 502.
[0085] Figure 12 Exemplary application scenarios for performing the operations described in this disclosure are illustrated. For example... Figure 12 As shown, the optimized representation can be deployed on network device 304, i.e., local computer 1202, which communicatively connects to and executes LLM 1204 located in the cloud. Figure 12 In this context, a low false positive rate is preferred, thus reducing the amount of data requiring further analysis. A lower false positive rate, in turn, also reduces the latency introduced by further examining the data used to generate false positive detections. Therefore, the optimized representation output as result 1210 is expected to have the lowest false positive rate, which can be achieved through iterative execution based on the dataset stored in storage device 1208. Figure 12 Implemented using procedure 1206 within the dashed box. Procedure 1206 is... Figure 4 A specific example of process 400. It should be understood that... Figure 12 The environment shown is a simplified one for illustrative purposes; the actual environment may be more complex.
[0086] Based on the above description of various operations and examples, representations can be optimized. In some examples, an optimized representation may be a simplified and compact representation with a smaller number of detection instructions, which may consume fewer resources and introduce lower latency. This can be particularly advantageous when the total volume of traffic flows targeted by executing one or more representations is large. The operations in this disclosure also maintain flexibility for the user by allowing the user to input their own preferences. In this way, the user can specify one or more angles of optimization of the representation, depending on which metric(s) are more important or have higher priority to the user. The operations in this disclosure can be adapted to various deployment environments, where optimization can be tuned to different directions or angles (e.g., latency, false positive / false negative rate, true positive / true negative rate, etc.) by inputting different preferences. Different preferences can be reflected in preferred metric values input into the system performing the operations described herein. User preferences can be determined based on the different network devices on which the optimized representation will be deployed. In some other examples, an optimized representation may have improved processing performance at the expense of performance degradation at other angles.
[0087] Throughout the specification and claims, terms may have subtle meanings suggested or implied in the context that go beyond their expressly stated meanings. Further, this means that: unless expressly stated otherwise, the word “or” can be inclusive or exclusive; the term “set” can include zero, one, two, or more elements; the terms “some,” “another,” and “specific” are used as naming conventions to distinguish elements from each other and, unless otherwise stated, do not imply the order, time, or any feature of the referenced items; the terms “for example,” “e.g.,” etc., describe one or more examples, but are not limited to the described examples; the terms “comprising” and / or “such as” specify the presence of the stated feature but do not exclude the presence or addition of one or more other features.
[0088] References to features, advantages, or similar language in this specification do not imply that all features and advantages can be implemented using this solution and should or must be included in any implementation. Rather, references to features and advantages are understood to mean that a particular feature, advantage, or characteristic described in conjunction with the example is included in at least one example of this solution. Therefore, throughout this specification, discussions of features and advantages, as well as similar language, may but not necessarily refer to the same examples.
[0089] Furthermore, the features, advantages, and characteristics described herein can be combined in any suitable manner in one or more embodiments or examples. Based on the description herein, those skilled in the art will recognize that this solution can be practiced without one or more specific features or advantages of a particular embodiment or example. In other cases, additional features and advantages may be recognized in certain embodiments or examples that may not exist in all embodiments of this disclosure.
Claims
1. A method implemented by a network traffic management system, the network traffic management system comprising one or more network traffic management devices, client devices, or server devices, the method comprising: Retrieve representations and datasets associated with said representations from the storage device; The natural language processing model converts the retrieved representation into a first candidate representation based on the dataset, and the first candidate representation is different from the retrieved representation; The first candidate representation is input into the simulator to generate one or more first candidate metrics corresponding to one or more preferred metrics; Based on one or more first candidate metrics and one or more preferred metrics, evaluate whether to replace the retrieved representation with the first candidate representation; as well as In response to an evaluation providing an indication to replace the retrieved representation with the first candidate representation, the first candidate representation is deployed on network devices of the network managed by the network traffic management system, the network devices being configured to use the deployed first candidate representation to detect traffic flow patterns of the network.
2. The method according to claim 1, wherein, The assessment also includes: The one or more preferred metric values are input into the objective function to obtain preferred values for quantifying the one or more preferred metric values; The generated one or more first candidate metric values are input into the objective function to obtain the first candidate values for quantifying the generated one or more first candidate metric values; The evaluation of whether to replace the retrieved representation with the first candidate representation is based on a comparison between the preferred value and the first candidate value.
3. The method according to claim 1, wherein, The storage device stores multiple representations, and the retrieval further includes: Randomly sample the plurality of representations to obtain sampled representations as the retrieved representations; and Retrieve the dataset associated with the retrieved representation from the storage device.
4. The method according to claim 1, wherein, In response to the evaluation providing an indication to replace the retrieved representation with the first candidate representation, the method further includes: Retrieve additional datasets associated with the first candidate representation from the storage device; The natural language processing model is described in the prompt that it converts the first candidate representation into a second candidate representation based on the additional dataset, and the second candidate representation is different from the first candidate representation. The second candidate representation is input into the simulator to generate one or more second candidate metrics corresponding to one or more preferred metrics; Based on one or more first candidate metrics and one or more second candidate metrics, evaluate whether to replace the first candidate representation with the second candidate representation; and In response to an evaluation providing an indication to replace the first candidate representation with the second candidate representation, the second candidate representation is deployed on the network device.
5. The method according to claim 1, wherein, In response to the evaluation providing an indication not to replace the retrieved representation with the first candidate representation, the method further includes: The first candidate representation is stored in the storage device.
6. An apparatus comprising a memory and one or more processors, the memory including programming instructions stored in the memory, the one or more processors being configured to execute the programming instructions stored in the memory to: Retrieve representations and datasets associated with said representations from the storage device; The natural language processing model converts the retrieved representation into a first candidate representation based on the dataset, and the first candidate representation is different from the retrieved representation; The first candidate representation is input into the simulator to generate one or more first candidate metrics corresponding to one or more preferred metrics; Based on one or more first candidate metrics and one or more preferred metrics, evaluate whether to replace the retrieved representation with the first candidate representation; as well as In response to an evaluation providing an indication to replace the retrieved representation with the first candidate representation, the first candidate representation is deployed on network devices in a network managed by a network traffic management system, the network devices being configured to use the deployed first candidate representation to detect traffic flow patterns in the network.
7. The apparatus according to claim 6, wherein, The assessment also includes: The one or more preferred metric values are input into the objective function to obtain preferred values for quantifying the one or more preferred metric values; The generated one or more first candidate metric values are input into the objective function to obtain the first candidate values for quantifying the generated one or more first candidate metric values; The evaluation of whether to replace the retrieved representation with the first candidate representation is based on a comparison between the preferred value and the first candidate value.
8. The apparatus according to claim 6, wherein, The storage device stores multiple representations, and the retrieval further includes: Randomly sample the plurality of representations to obtain sampled representations as the retrieved representations; and Retrieve the dataset associated with the retrieved representation from the storage device.
9. The apparatus according to claim 6, wherein, In response to the evaluation providing an indication to replace the retrieved representation with the first candidate representation, the one or more processors are configured to execute programming instructions stored in the memory to: Retrieve additional datasets associated with the first candidate representation from the storage device; The natural language processing model is described in the prompt that it converts the first candidate representation into a second candidate representation based on the additional dataset, and the second candidate representation is different from the first candidate representation. The second candidate representation is input into the simulator to generate one or more second candidate metrics corresponding to one or more preferred metrics; Based on one or more first candidate metrics and one or more second candidate metrics, evaluate whether to replace the first candidate representation with the second candidate representation; as well as In response to an evaluation providing an indication to replace the first candidate representation with the second candidate representation, the second candidate representation is deployed on the network device.
10. The apparatus according to claim 6, wherein, In response to the evaluation providing an indication not to replace the retrieved representation with the first candidate representation, the one or more processors are configured to execute programming instructions stored in the memory to: The first candidate representation is stored in the storage device.
11. A non-transitory computer-readable medium having instructions stored thereon, the instructions including executable code, which, when executed by one or more processors, causes the one or more processors to: Retrieve representations and datasets associated with said representations from the storage device; The natural language processing model converts the retrieved representation into a first candidate representation based on the dataset, and the first candidate representation is different from the retrieved representation; The first candidate representation is input into the simulator to generate one or more first candidate metrics corresponding to one or more preferred metrics; Based on one or more first candidate metrics and one or more preferred metrics, evaluate whether to replace the retrieved representation with the first candidate representation; as well as In response to an evaluation providing an indication to replace the retrieved representation with the first candidate representation, the first candidate representation is deployed on network devices in a network managed by a network traffic management system, the network devices being configured to use the deployed candidate representation to detect traffic flow patterns in the network.
12. The non-transitory computer-readable medium according to claim 11, wherein, The assessment also includes: The one or more preferred metric values are input into the objective function to obtain preferred values for quantifying the one or more preferred metric values; The generated one or more first candidate metric values are input into the objective function to obtain the first candidate values for quantifying the generated one or more first candidate metric values; The evaluation of whether to replace the retrieved representation with the first candidate representation is based on a comparison between the preferred value and the first candidate value.
13. The non-transitory computer-readable medium according to claim 11, wherein, The storage device stores multiple representations, and the retrieval further includes: Randomly sample the plurality of representations to obtain sampled representations as the retrieved representations; and Retrieve the dataset associated with the retrieved representation from the storage device.
14. The non-transitory computer-readable medium according to claim 11, wherein, In response to the evaluation providing an indication to replace the retrieved representation with the first candidate representation, the one or more processors are also made to: Retrieve additional datasets associated with the first candidate representation from the storage device; The natural language processing model is described in the prompt that it converts the first candidate representation into a second candidate representation based on the additional dataset, and the second candidate representation is different from the first candidate representation. The second candidate representation is input into the simulator to generate one or more second candidate metrics corresponding to one or more preferred metrics; Based on one or more first candidate metrics and one or more second candidate metrics, evaluate whether to replace the first candidate representation with the second candidate representation; as well as In response to an evaluation providing an indication to replace the first candidate representation with the second candidate representation, the second candidate representation is deployed on the network device.
15. The non-transitory computer-readable medium according to claim 11, wherein, Also, the one or more processors: The first candidate representation is stored in the storage device.
16. A network traffic management system, the network traffic management system comprising one or more traffic management devices, server devices, or client devices, the network traffic management system comprising a memory and one or more processors, the memory comprising programming instructions stored thereon, the one or more processors being configured to execute the stored programming instructions to: Retrieve representations and datasets associated with said representations from the storage device; The natural language processing model converts the retrieved representation into a first candidate representation based on the dataset, and the first candidate representation is different from the retrieved representation; The first candidate representation is input into the simulator to generate one or more first candidate metrics corresponding to one or more preferred metrics; Based on one or more first candidate metrics and one or more preferred metrics, evaluate whether to replace the retrieved representation with the first candidate representation; as well as In response to an evaluation providing an indication to replace the retrieved representation with the first candidate representation, the first candidate representation is deployed on network devices of the network managed by the network traffic management system, the network devices being configured to use the deployed first candidate representation to detect traffic flow patterns of the network.
17. The network traffic management system according to claim 16, wherein, The assessment also includes: One or more preferred metric values are input into the objective function to obtain preferred values for quantifying the one or more preferred metric values; The generated one or more first candidate metric values are input into the objective function to obtain the first candidate values for quantifying the generated one or more first candidate metric values; The evaluation of whether to replace the retrieved representation with the first candidate representation is based on a comparison between the preferred value and the first candidate value.
18. The network traffic management system according to claim 16, wherein, The storage device stores multiple representations, and the retrieval further includes: Randomly sample the plurality of representations to obtain sampled representations as the retrieved representations; and Retrieve the dataset associated with the retrieved representation from the storage device.
19. The network traffic management system according to claim 16, wherein, In response to the evaluation providing an indication to replace the retrieved representation with the first candidate representation, the one or more processors are further configured to execute stored programming instructions to: Retrieve additional datasets associated with the first candidate representation from the storage device; The natural language processing model is described in the prompt that it converts the first candidate representation into a second candidate representation based on the additional dataset, and the second candidate representation is different from the first candidate representation. The second candidate representation is input into the simulator to generate one or more second candidate metrics corresponding to one or more preferred metrics; Based on one or more first candidate metrics and one or more second candidate metrics, evaluate whether to replace the first candidate representation with the second candidate representation; as well as In response to an evaluation providing an indication to replace the first candidate representation with the second candidate representation, the second candidate representation is deployed on the network device.
20. The network traffic management system according to claim 16, wherein, The one or more processors are further configured to execute stored programming instructions to: The first candidate representation is stored in the storage device.