A method of identifying interface assets and a data security monitoring system
By splitting URL parameters and introducing interface components, the interface asset identification process is optimized, solving the problems of invalid and duplicate identification in existing technologies, and achieving more efficient interface asset identification and recording.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU DPTECH TECH
- Filing Date
- 2023-04-06
- Publication Date
- 2026-07-21
AI Technical Summary
Existing interface asset identification methods result in the identification of a large number of invalid and duplicate interface assets, affecting identification quality and efficiency.
By parsing the URL from the interface messages of mirrored traffic, splitting it into a second URL and URL parameters, identifying interface assets based on the second URL, and introducing the use of interface components and system memory during the identification process, the identification process of interface assets is optimized.
It effectively reduced the generation of invalid and duplicate interface assets, improved asset identification quality and processing efficiency, shortened query time, reduced the number of interactions with the database, and improved overall operational efficiency.
Smart Images

Figure CN116471075B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data security, and in particular to a method for identifying interface assets and a data security monitoring system. Background Technology
[0002] Ensuring data security on terminal devices requires clearly defining the data to be protected and identifying and recording it. In related technologies, data security on terminal devices is achieved through a data security monitoring system. Packets carrying source data from the terminal device are forwarded to the data security monitoring system through designated ports on a switch; this process is called mirror traffic forwarding. After receiving the interface packets from the mirror traffic forwarding, the data security monitoring system analyzes these packets to identify interface assets corresponding to a specific terminal device in the current environment; this process is called interface asset identification. By recording these interface assets, the data security monitoring system identifies and records the monitored data on the terminal device, enabling data security industry applications such as sensitive behavior detection and access trajectory detection.
[0003] The current method of identifying interface assets results in the identification of a large number of invalid and duplicate interface assets. Summary of the Invention
[0004] In view of this, this application provides a method for identifying interface assets and a data security monitoring system to address the aforementioned deficiencies in related technologies.
[0005] The first aspect of this application provides a method for identifying interface assets, wherein the interface assets are identified by a data security monitoring system based on interface packets forwarded by mirrored traffic; mirrored traffic is the traffic of interface packets of the monitored device forwarded by a switch to the data security monitoring system; the method includes:
[0006] Parse the first URL from the interface message of the mirrored traffic, split the first URL into a second URL and URL parameters, and the second URL points to the monitored data of the monitoring device;
[0007] The interface asset is determined based on the second URL.
[0008] A second aspect of this application provides a data security monitoring system for identifying interface assets, wherein the interface assets are identified based on interface packets forwarded by mirrored traffic; mirrored traffic is the traffic of interface packets forwarded by a switch to the monitored device of the data security monitoring system; the system includes:
[0009] Traffic probes are used to monitor mirrored traffic. They parse the first URL from the interface packets of the mirrored traffic, split the first URL into a second URL and URL parameters, and the second URL points to the monitored data of the monitoring device.
[0010] The data processing engine receives the second URL and determines the interface assets based on the second URL.
[0011] This application parses the Uniform Resource Locator (URL) of the interface from the interface packets of mirrored traffic, and then splits the interface URL of the mirrored traffic into the URL pointing to the monitoring data and the URL parameters, thus obtaining the interface URL of the monitored device after removing the interference of URL parameters. In related technologies, the approach is to receive and parse all the interface packets forwarded by the mirrored traffic, and then directly identify the interface assets based on the parsed interface URLs. Due to the presence of URL parameters, interface URLs of the same monitored device but with different URL parameters will ultimately be identified as different interface assets. However, splitting the URL allows for effective filtering of URL parameters, ensuring that interface URLs of the same monitored device are identified as the same interface URL. This significantly reduces the generation of invalid and duplicate interface assets, thereby effectively improving the quality of asset identification and increasing the aggregation of asset identification results. Because the amount of identification data is large, the effective reduction of invalid and duplicate interface assets also greatly reduces the number of subsequent asset identification and recording operations, effectively improving the efficiency of asset identification processing.
[0012] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0013] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0014] Figure 1 This is a flowchart illustrating a method for identifying interface assets according to an exemplary embodiment;
[0015] Figure 2 This is a flowchart illustrating a method for determining interface assets based on a parsed URL, according to an exemplary embodiment.
[0016] Figure 3 This is a flowchart illustrating a method for determining interface assets based on a parsed URL, according to an exemplary embodiment.
[0017] Figure 4 This is a flowchart illustrating a method for determining interface assets based on a parsed URL, according to an exemplary embodiment.
[0018] Figure 5 This is a flowchart illustrating a method for identifying interface assets according to an exemplary embodiment;
[0019] Figure 6This is a block diagram illustrating a data monitoring system according to an exemplary embodiment;
[0020] Figure 7 This is a block diagram illustrating a data monitoring system according to an exemplary embodiment. Detailed Implementation
[0021] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0022] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0023] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0024] To facilitate understanding, some of the concepts involved in this application are explained:
[0025] A data security monitoring system refers to products and / or services centered on data security, covering data protection needs in various scenarios, such as resident information data, examinee information data, and enterprise asset data. It integrates multiple technologies to achieve data security protection, such as data access control, data anonymization, and data encryption. The prerequisite for data security protection is data discovery; for example, in the protection of enterprise asset data, discovering the enterprise's asset data is the foundation of data protection.
[0026] Mirroring traffic refers to replicating real online traffic, such as data packets and response parameters, to a mirror service through specific configurations. By analyzing the forwarded data traffic in the mirror service, data security can be effectively monitored. Furthermore, mirroring traffic forwarding allows for detailed analysis of traffic and / or request content without impacting online services. Mirroring traffic services can be implemented through switch ports. By forwarding data traffic from the source ports of one or more connected terminal devices on a switch to a designated port, analysis and other operations can be performed on the forwarded data traffic on that designated port, thereby achieving data security monitoring of the source ports of the aforementioned one or more connected terminal devices.
[0027] Interface messages serve as the medium for exchanging and transmitting data between online systems or within a network. These messages contain complete information about the exchanged data and must adhere to a predefined format. For example, in the operation of an enterprise system, multiple systems communicate and / or exchange and transmit data by forwarding interface messages to perform corresponding operations. The format of interface messages is not unique and is configured by the developers; it may include elements such as URLs, request headers / body, response headers / body, cookies, and destination IP addresses.
[0028] A URL (Uniform Resource Locator) is a unique identifier for a resource, much like a national identity card number, which is the address of a standard resource on the internet. A resource can be understood as any file accessible via the internet, such as web pages, images, audio, video, scripts, etc. URL resources within an enterprise serve as its information assets and play a crucial role in its development. A URL typically consists of the following parts: protocol, host (domain name), port, path, and query parameters. Generally, if a resource is accessible via the internet, it will have a corresponding URL. One URL corresponds to one asset, but the same asset may correspond to multiple URLs. This situation presents a challenge to the discovery of data resources in the aforementioned data security platform.
[0029] Interface asset identification involves the data security platform receiving interface packets from mirrored traffic. By analyzing these packets, the platform identifies interface assets corresponding to specific terminal devices within the current environment. This process is called interface asset identification. Identified interface assets can be applied to data security applications such as sensitive behavior detection and access tracking. They can also be used in other industries that require interface assets.
[0030] Current interface asset identification methods identify a large number of invalid and duplicate interface assets. Therefore, an efficient and reliable interface asset identification method is needed to ensure the quality of identification results, reduce the duplication rate of identification results, and increase the readability of identification results while processing large-scale mirror traffic data packets.
[0031] The inventors discovered through research that addressing the issue of the same resource potentially corresponding to multiple URLs is a key focus for improving current API asset identification methods. In current API asset identification methods, the parsed URLs may contain string parameters and / or path parameters and / or other forms of parameters, leading to different URLs representing the same asset due to variations in the parameters passed.
[0032] This issue arises from how we access resources: for example, if we want to find a specific book in an online library, we can search by title, filter by category, or search by keyword within the category. All three methods ultimately yield the book, but the URL generated by the first method includes a query string parameter; the second method, filtering by category, generates a URL with a path parameter; and the third method generates a URL with both query string and path parameters. As this example shows, all three methods ultimately lead to the same book, but generate three different URLs.
[0033] Therefore, the inventors conceived of splitting the parameters of the directly generated URL, separating the parameters passed in string form and / or path form and / or other forms, to obtain the interface URL of the monitored device. This would greatly reduce the generation of invalid and duplicate URLs, thereby reducing the identification of invalid and duplicate interface assets and effectively improving the identification quality of interface assets.
[0034] The technical solutions described in this specification will now be introduced in conjunction with specific embodiments.
[0035] Figure 1 This application is a flowchart illustrating an interface asset identification method according to an exemplary embodiment. The interface assets are identified by a data security platform based on interface packets forwarded by mirrored traffic; mirrored traffic is the traffic of interface packets of the monitored device forwarded by the switch to the data security monitoring system; the method includes:
[0036] Step 102: Parse the first URL from the interface message of the mirrored traffic, and split the first URL into the second URL and URL parameters;
[0037] Step 104: Determine the interface assets based on the second URL.
[0038] The second URL points to the monitored data of the monitoring device.
[0039] Because of the existence of URL parameters, a single interface asset can generate different URLs, each carrying different URL parameters, resulting in differences in form and thus making them "different." However, these URLs essentially all point to the same interface asset. This embodiment can solve the above problem; one of the key improvements in interface asset identification is the optimization of the interface asset itself. By splitting the different URLs carrying URL parameters, an interface URL for the monitored device is generated. The interface URL for the monitored device, after removing the parameters, no longer has the specific limitations of the parameters, allowing different URLs that essentially point to the same interface asset to be grouped into a single URL, i.e., the second URL in the method. However, "second" here does not imply any restriction on this URL; it is merely for descriptive convenience and can also be referred to as the third URL, fourth URL, etc.
[0040] URL parameter passing can take many forms, such as including only string parameters, only path parameters, or both. Of course, it may also include other parameters. Since the technical solutions in this specification may have multiple embodiments, the URL parameters may differ in different embodiments, meaning the URL parameter type is not necessarily limited to one. Those skilled in the art can make actual additions, subtractions, or selections as needed, and this specification does not impose any limitations on this.
[0041] This embodiment uses string-based URL parameter passing as an example. When there is only one string parameter, the URL will include a parenthesis (?key=value), where value is the specific parameter. When there are more than one string parameter, different parameters are connected by the '&' symbol, as shown in the parenthesis (?key=value&key=value&key=value). Without splitting the parameter, the parenthesis-like portion will be mixed into the URL. For example, when there is only one parameter, the URL with mixed string parameters might be: / url?key=1 or / url?key=1, etc.; when there are more than one parameter, the URL with mixed string parameters might be: / url?key1=1&key2=2 or / url?key3=3&key4=4, etc. This embodiment does not impose a specific limit on the number of parameters. All the URLs mentioned above add string parameters to the original interface URL of the monitored device. Therefore, although each URL is essentially the same, they differ in form. Related technologies will identify these interface asset URLs as different URLs, and thus as different interface assets. This embodiment can split and remove the string parameters that cause this problem, retaining the original interface URL of the monitored device. Removing unnecessary restrictions on the string parameters before identification will reduce the identification of a large number of duplicate interface assets, thus effectively improving the quality of interface asset identification.
[0042] This embodiment uses URL parameter passing in path form as an example. Path-based parameter passing refers to URL paths containing unnecessary port numbers, interface numbers, or other forms of path IDs. The presence of path IDs does not affect the current interface asset pointer in the URL. The specific interface asset pointer is configured by the user. For example, if the user needing to identify interface assets is in the zoo industry, the user can configure all animals in the current zoo as interface assets, such as / zoos / 1 / animals, / zoos / 2 / animals, etc. Then, when recording interface assets, all animals in different zoos within this industry are recorded as interface assets. Different animals within the same zoo belong to that zoo, i.e., the specific asset data under that interface asset, such as / zoos / 1 / Animals / elephants, / zoos / 1 / animals / tigers, etc., no longer need to be recorded separately. Alternatively, a specific animal category within a particular zoo can be recorded as an interface asset, such as / zoos / 1 / animals / elephants, / zoos / 1 / animals / tigers, etc. In this case, different individuals within the same category no longer need to be recorded, such as / zoos / 1 / animals / elephants / 1, / zoos / 1 / animals / elephants / 2, etc. Furthermore, only the animal category can be recorded as an interface asset, without restrictions on the specific zoo, such as / zoos / {id} / animals / elephants. Here, {id} represents the collection of all zoos, and the specific zoos are no longer distinguished; all zoos are treated as a single entity. For example, recording elephants no longer requires recording in the form of "Elephant 1 in Zoo 1, Elephant 2 in Zoo 1, Elephant 1 in Zoo 2," etc.
[0043] The specific splitting method and degree of the interface asset can be configured according to user needs, and this specification does not impose any restrictions on this. Therefore, path-based parameters can also be our splitting objects. Regardless of the splitting method and degree, splitting this parameter can achieve the same technical effect as the aforementioned example of splitting string parameters.
[0044] This embodiment uses URL parameters passed in both string and path formats as examples. This embodiment is a combination of the two embodiments mentioned above. Using the examples from the above embodiments, the scenario in this embodiment is as follows: the specific interface asset that needs to be identified is a juvenile elephant 11 in Zoo 1. Here, "juvenile" is the keyword for the query, in the form of / zoos / 1 / animals / elephants / 11?key=young. However, the interface asset that the user needs to record is all elephants in all zoos, in the form of / zoos / {id} / animals / elephants. Therefore, we need to separate the path parameter representing 1 (the specific zoo) and the path parameter representing 11 (the specific individual elephant) from the string parameter ?key=young.
[0045] The above embodiments of splitting URL parameters are intended to enable those skilled in the art to better understand the technical solution of this application. This technical solution does not impose specific restrictions on the type, number, or representation of URL parameters, and the above embodiments are not intended to limit this application.
[0046] Another key improvement in interface asset identification lies in the identification process itself, specifically step 104 mentioned above. After acquiring the interface asset, we need to compare it with the recorded interface assets. If the current interface asset is not found among the recorded assets, it should be recorded accordingly. Related technologies typically start the comparison directly from the first interface asset each time, and each query retrieves the complete URL. Furthermore, these technologies do not perform parameter splitting on the URL, making the URL itself lengthy. Given the already large amount of recorded interface asset data, the comparison process becomes extremely time-consuming.
[0047] Therefore, to address the issue of excessively long asset query times, the inventors added an interface component level query before querying interface assets. The addition of interface components creates a hierarchical query process, eliminating the need to start with the first specific interface asset and then query for corresponding interface assets under that component. This significantly reduces the time required to identify a large number of interface assets and optimizes the identification process.
[0048] For ease of understanding, the interface components described above are explained below. The interface components are generated based on the domain name of the parsed interface asset URL. The specific form of the interface component is not unique; for example, an algorithm can be used to generate key-value pairs based on the parsed domain name to represent the interface component. Alternatively, the parsed domain name can be directly used as the interface component of the current interface asset URL. This application does not restrict the specific form of the interface component.
[0049] The following example illustrates the role of the interface component in the interface asset identification process. For instance, a school has one thousand students in a certain grade. Each student has a unique number, but this number is not a direct number starting from 1; instead, it's a string of characters, equivalent to the interface asset URL mentioned above. These one thousand students are divided into 20 classes, each with a corresponding class number. Similarly, this class number is not a direct number starting from 1; instead, it's a string of characters, which is the interface component mentioned above. We currently know the specific student information of a student, for example, the student with the twentieth student ID in the tenth class. We need to query whether this student is already a student of this school. The approach of related technologies is that we only know the student ID, and to specifically query this student, we have to start searching from the first student ID in the first class and continue until we find the student ID we need. However, the solution in this application is that we first know the class ID of the student we need to query, then query the class ID, and then specifically query the student ID within that class. In this example, querying in this way eliminates the need to query and compare student information other than that in the tenth class, thus reducing a large number of unnecessary queries, greatly saving time, and improving query efficiency. The solution adopted in this embodiment has the same effect as the example: first, we obtain the interface component of the URL of the interface asset that needs to be identified; then, we query whether this interface component is recorded, and then continue to query whether there is an interface asset that needs to be identified under this interface component. Interface asset information under interface components other than this interface component no longer needs to be queried and compared, greatly reducing query and identification time and improving the efficiency of interface asset identification.
[0050] Adding an interface component-level query facilitates the identification of interface components and assets. A corresponding primary key can be generated using an algorithm, and the identification process can be achieved by querying the corresponding primary key. Algorithm-generated primary keys ensure the uniqueness of the corresponding interface component and asset, and these primary keys are often much shorter than direct interface asset representations. While the generation format and length of primary keys are generally fixed based on different generation rules, interface assets vary in length and format, making primary key queries relatively easier. There are various methods for generating primary keys; algorithms that can satisfy a low probability of collisions can generally generate primary keys. Examples include MD-series message digest algorithms and SHA-series secure hash algorithms. This embodiment does not limit the specific primary key generation method.
[0051] This embodiment uses the MD series message digest algorithm to generate primary keys as an example. The MD (Message Digest) algorithm, with the MD5 message digest algorithm being the most widely used, is the fifth version of the MD series, an improvement upon MD2 and MD4. The core of all MD series digest algorithms is to generate a fixed-length message digest, such as a 128-bit binary digest, to ensure the integrity and consistency of the data. Each data message generates a unique message digest, thus guaranteeing the uniqueness of the corresponding data. The MD2 algorithm is slower but more accurate; the MD4 algorithm has significantly improved speed but decreased accuracy; MD5 offers a significant improvement in both accuracy and speed. This application can use the MD5 digest algorithm to generate primary keys for interface components and interface assets. Because the MD5 values generated by MD5, i.e., the aforementioned primary keys, are unique, during the interface asset identification process, the primary key information can be queried to determine whether the current interface component or asset has been recorded. The uniqueness of the generated primary key effectively improves the identification accuracy of interface assets. Based on this, instead of directly querying specific interface assets, querying the primary key corresponding to the interface assets can also ensure the security of the recorded interface assets to a certain extent.
[0052] This embodiment uses the SHA series of secure hash algorithms, also known as secure hash algorithms, to generate primary keys as an example. SHA (Secure Hash Algorithm) evolved from the MD4 algorithm in the previous embodiment. Currently, there are SHA-1 and SHA-2 series, with the SHA-2 series specifically including SHA-256, SHA-384, and SHA512. The commonly used SHA algorithm is SHA-256. Taking SHA-256 as an example, its effect is similar to the MD algorithm; both generate unique values corresponding to data information through algorithms. However, the specific algorithm calculation method is different from the MD algorithm, resulting in different unique values. The SHA algorithm can also achieve the function of the MD algorithm, ensuring the integrity and consistency of data information and guaranteeing the uniqueness of corresponding data, improving the accuracy of asset identification, and to a certain extent ensuring the security of recorded interface assets.
[0053] Figure 2 This application illustrates a flowchart of a method for determining interface assets based on a parsed URL according to an exemplary embodiment. The method includes:
[0054] Step 202: Generate the interface component and the component primary key bound to the interface component based on the domain name of the second URL;
[0055] Step 204: Query the interface components that are bound to the component's primary key from the interface components already recorded in memory;
[0056] Step 206: Generate the interface asset primary key based on the second URL and domain name;
[0057] Step 208: Based on the component primary key, retrieve the interface assets that match the interface component from the interface assets already recorded in memory;
[0058] Step 210: Determine the interface assets that match the primary key of the interface assets from the interface assets that match the interface components.
[0059] Among them, the domain names of interface assets corresponding to the same interface component are the same;
[0060] This embodiment is merely an exemplary embodiment, and the specific method steps are not strictly limited to the order of steps in this embodiment, as long as the generation of the corresponding primary key precedes the query and identification of the corresponding interface component or the corresponding interface asset, and the identification of the interface component precedes the identification of the interface asset. For example, the method steps can also be in the following order: swapping steps 204 and 206, or swapping steps 206 and 208, etc. Among the two conditions that need to be met above, generating the corresponding primary key before querying and identifying the corresponding interface component or the corresponding interface asset is to enable the primary key to play its role in interface asset identification, as described in the above embodiments, and will not be described again in this embodiment; identifying the interface component before identifying the interface asset is to enable the interface component to play its role in interface asset identification, as described in the above embodiments, and will not be described again in this embodiment. This embodiment mainly describes the situation where the interface asset to be identified has already been recorded, and will eventually be found in the memory record. At this time, the next interface asset identification operation can be performed directly, or the recorded interface asset information and / or interface component information and / or corresponding primary key information can be returned first, and then the next interface asset identification operation can be performed. This embodiment does not limit this.
[0061] Figure 3 This application illustrates a flowchart of a method for determining interface assets based on a parsed URL according to an exemplary embodiment. The method includes:
[0062] Step 302: Record the interface components that are not yet recorded in memory and are bound to the component's primary key into memory;
[0063] Step 304: Record the interface assets that are not yet recorded in memory and match the interface components into memory.
[0064] This embodiment mainly describes the situation where the currently identified interface components and / or interface assets are not recorded. In the technical solution of this application, an interface component identification process is added to the interface asset identification process. In practical applications, the interface component identification process and the interface asset identification process can be carried out independently, but the interface component identification should be prioritized, also to give full play to the role of interface components in interface asset identification. The two processes can also be integrated together. When querying and identifying independently, it is first determined whether the current interface component or interface asset has been recorded. If it has been recorded, the next step is performed; if it has not been recorded, the current interface component or interface asset is recorded, and then the next step is performed. Before proceeding to the next step, the information of the currently queried or recorded interface components or interface assets can be returned, or the next step can be performed directly. This embodiment does not limit this. When integrating them for query and identification, first determine whether the current interface component has been recorded. If it has been recorded, continue to identify the interface asset under the current interface component. If it has not been recorded, it proves that there is no corresponding interface asset recorded under this interface component. Then directly record the interface component and interface asset. Similarly, before proceeding to the next step, the information of the currently queried or recorded interface component or interface asset can be returned. This embodiment does not impose any restrictions on this.
[0065] In the previous embodiment, if the current interface component or interface asset is not found, it needs to be recorded. In related technologies, both querying and recording involve direct interaction with a database. However, this database is a public resource on an external server, and many devices need to interact with it. The database type could be Redis or HBase, etc., and this embodiment does not limit this. Furthermore, because interface asset queries require a large amount of data interaction with the database and involve numerous interactions, even if a single interface asset identification process only requires a small number of database queries, it will still result in frequent database record queries for identification and comparison, consuming a large amount of database server resources and affecting overall operating efficiency. The interface asset information in the database ultimately needs to be stored in the corresponding storage device, such as an external hard disk drive or solid-state drive, which needs to be able to permanently record the interface asset information for relevant industry applications. However, the reading and recording speed of storage devices is limited, so the interface asset identification process, which directly interacts with the database, is slow, affecting overall operating efficiency. The inventors have significantly improved this problem by using internal system memory. In one example, the "internal" refers to the system memory within the aforementioned data security monitoring system. First, the interface asset information already recorded in the database is loaded into the system memory. The query and identification of interface assets is then transferred to this internal system memory. If no interface asset is found that needs to be recorded, it is recorded in memory, and then the database is interacted with for recording. Alternatively, after the current round of interface asset identification is completed, the updated interface asset information in the system memory can be updated in the database; this embodiment does not impose this limitation. In another example, the aforementioned data security monitoring system can use external memory modules. Any system that can transfer the interface asset query and identification process to these memory modules is acceptable; this embodiment does not impose this limitation either. Since query operations no longer interact with the database, only recording operations do. This solution significantly reduces the number of interactions with the database, and the interaction speed with the internal system memory is much faster than with an external database, greatly improving asset identification efficiency and reducing database resource consumption.
[0066] In the above embodiment, step 304 can be specifically divided into the following steps, see below. Figure 4 , Figure 4 This application illustrates a flowchart of a method for determining interface assets based on a URL according to an exemplary embodiment. The method includes:
[0067] Step 402: Query the memory for interface assets that match the interface component bound to the primary key of the component. If no match is found, record the interface assets that match the interface component bound to the primary key in memory and the database.
[0068] Step 404: Query the interface asset bound to the primary key of the interface asset in memory. If no interface asset is found, record the interface asset bound to the primary key in memory and database.
[0069] This embodiment primarily describes how to record interface assets when the interface component is found but not the interface asset, and describes the query method. Utilizing system memory can significantly improve the efficiency of interface asset identification and reduce the consumption of database server resources. The query and recording process for interface components and interface assets can be performed independently or integrated. This embodiment is a specific elaboration based on the conditions of the aforementioned embodiments, and does not limit these specific conditions.
[0070] Taking the process of querying and recording interface components and interface assets independently as an example, if an interface component has been found, the system first checks in memory whether there is a matching interface asset record under the current interface component bound to the component's primary key. If not, the current interface asset is recorded in memory at the corresponding interface component location, and simultaneously interacts with the database to record it in the corresponding location. If it is already recorded, proceed to the next step. Then, the system checks in memory whether the interface asset bound to the current interface asset's primary key has been recorded. If not, the current interface asset is recorded in memory at the corresponding interface component location, and simultaneously interacts with the database to record it in the corresponding location. If it is already recorded, the identification of the current interface asset is complete. Before each of the above operations, the information of the currently found or recorded interface components or interface assets can be returned, or the next operation can be performed directly; this embodiment does not impose any restrictions on this. Separating the query and recording process of interface components from that of interface assets allows the queries of interface components and interface assets to proceed independently without interference. Furthermore, this separation can be implemented on two different devices, which can reduce the computational burden on each device during the interface asset identification process to some extent.
[0071] Taking the integration of the query and recording process for interface components with that for interface assets as an example, the process begins by checking in system memory whether the interface component of the current interface asset has been recorded. If not, the current interface component is recorded in system memory, and interaction with the database is performed to record it as well. After finding a recorded component or recording an unrecorded component, the system continues to check in memory whether the interface asset matching the current interface component bound to the component's primary key has been recorded. If not, the current interface asset is recorded in memory at the corresponding interface component location, and interaction with the database is performed to record it in the corresponding location in the database. If recorded, the system continues to check in system memory whether the interface asset bound to the current interface asset's primary key has been recorded. If not, the current interface asset is recorded in memory at the corresponding interface component location, and interaction with the database is performed to record it in the corresponding location in the database. If recorded, the identification of the current interface asset is complete. After finding a record or recording an unrecorded component, the system can either return the currently found or recorded interface component or interface asset information, or proceed directly to the next step. This embodiment does not impose any restrictions on this. Integrating the two processes together can enhance the cohesion of the interface asset identification process.
[0072] Figure 5 This is a flowchart illustrating an interface asset identification method according to an exemplary embodiment of this application, the method comprising:
[0073] Step 502: Parse the IP address of the monitored data from the interface message;
[0074] Step 504: Discard the second URL whose IP address is not within the monitoring range;
[0075] The monitoring range is set through a data security monitoring system.
[0076] During the identification of interface assets, the scope of identification varies depending on the needs of different users. For example, using the zoo example in one of the embodiments above, a user may only want to record zoos within a certain area as interface assets, while zoos outside this area are considered invalid records for that user, i.e., invalid interface assets. In other words, during the interface asset identification process, the interface asset URLs that users need to record must also be within a specified range. If the recorded interface assets include those outside this range, users must manually determine whether the recorded interface assets fall within their required range when using them. This significantly reduces the quality of interface asset identification and also affects the user experience to some extent. The inventors noticed this problem and solved the problem of recording invalid interface assets by adopting the solution in this embodiment. Each interface asset URL has its corresponding IP address for monitored data. By configuring the required range of these IP addresses, unwanted invalid interface assets can be filtered out. These IP addresses are obtained when parsing the interface packets. Parsing the IP addresses of the monitored data and filtering out unnecessary IP ranges can be done separately or combined. The IP addresses of the monitored data can be obtained by parsing interface packets during traffic probe analysis, or they can be parsed later during IP range filtering. Filtering out unnecessary IP ranges can be performed directly by the data processing engine, or a filtering module can be added. This filtering module can also parse the interface packets to obtain the IP addresses of the monitored data; this embodiment does not impose any limitations on this. Filtering out unnecessary IP ranges avoids recording invalid interface assets, thereby improving the quality of interface asset identification. Because the generation of invalid interface assets is reduced, the number of interface asset identification attempts is also reduced to some extent, decreasing the number of interactions with memory and the database, thus improving the efficiency of interface asset identification.
[0077] Corresponding to the embodiments of the aforementioned interface asset identification method, this specification also provides embodiments of a data security monitoring system.
[0078] refer to Figure 6 , Figure 6 This application illustrates a block diagram of a data security monitoring system according to an exemplary embodiment. The system is used to identify interface assets; interface assets are identified based on interface packets forwarded by mirrored traffic; mirrored traffic is the traffic of interface packets forwarded by a switch to the monitored device of the data security monitoring system; the system includes:
[0079] Traffic probe 601 is used to monitor mirrored traffic, parse the first URL from the interface packets of mirrored traffic, and split the first URL into the second URL and URL parameters;
[0080] Data processing engine 602 is used to receive the second URL and determine the interface assets based on the second URL.
[0081] Among them, interface assets are based on the identification of interface packets forwarded by mirrored traffic; mirrored traffic is the traffic of interface packets of the monitored device forwarded by the switch; and the second URL points to the monitored data of the monitoring device.
[0082] In one embodiment, the data processing engine 602 determines the interface asset based on the second URL, including:
[0083] The interface component and the component primary key bound to the interface component are generated based on the domain name of the second URL; wherein, the domain names of the interface assets corresponding to the same interface component are the same;
[0084] Retrieve the interface components that are bound to the component's primary key from the interface components already recorded in memory;
[0085] Generate the primary key of the interface asset based on the second URL and domain name;
[0086] Based on the component primary key, retrieve the interface assets that match the interface component from the interface assets already recorded in memory;
[0087] Identify the interface assets that match the primary key of the interface assets from the interface assets that match the interface components.
[0088] In one embodiment, the data processing engine 602 determines the interface asset based on the second URL, and further includes:
[0089] Record the interface components that are not currently in memory and are bound to the component's primary key into memory; and / or
[0090] Record interface assets that match the interface component but are not currently recorded in memory into memory.
[0091] refer to Figure 7 , Figure 7 This is a block diagram illustrating a data security monitoring system according to an exemplary embodiment of this application. In one embodiment, the data security monitoring system further includes:
[0092] Filtering module 603 is used to parse the IP of the monitored data from the interface message and discard the second URL whose IP of the monitored data is not within the monitoring range;
[0093] The monitoring range is set by the data security monitoring system.
[0094] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for identifying interface assets, characterized in that, The interface assets are identified by the data security monitoring system based on the interface packets forwarded by the mirrored traffic; The mirrored traffic is the traffic of interface packets of the monitored device forwarded by the switch to the data security monitoring system; the method includes: The first URL is parsed from the interface message of the mirrored traffic, and the first URL is split into a second URL and URL parameters. The second URL points to the monitored data of the monitoring device. Determining the interface asset based on the second URL includes: generating an interface component and a component primary key bound to the interface component based on the domain name of the second URL; wherein the domain names of interface assets corresponding to the same interface component are the same; querying the interface components bound to the component primary key from the interface components already recorded in memory; generating an interface asset primary key based on the second URL and the domain name; querying the interface assets matching the interface component from the interface assets already recorded in memory based on the component primary key; and determining the interface asset matching the interface asset primary key from the interface assets matching the interface component.
2. The method according to claim 1, characterized in that, The method further includes: Record the interface components that are not currently recorded in memory and are bound to the primary key of the component into the memory; and / or Record the interface assets that do not match the interface component but are not recorded in the memory into the memory.
3. The method according to claim 1, characterized in that, The URL parameters include at least one of string-based URL parameters and path-based URL parameters.
4. The method according to claim 3, characterized in that, The step of splitting the first URL into a second URL and URL parameters includes: The first URL is split into a third URL and the string-format URL parameters, and the third URL is encapsulated based on the second URL and the path-format URL parameters.
5. The method according to claim 4, characterized in that, The step of splitting the first URL into a second URL and URL parameters also includes: The third URL is split into the second URL and the path-form URL parameters.
6. The method according to claim 1, characterized in that, The method further includes: The IP address of the monitored data is parsed from the interface message; The second URL whose IP address is not within the monitoring range is discarded; wherein, the monitoring range is set by the data security monitoring system.
7. A data security monitoring system for identifying interface assets, characterized in that, The interface assets are identified based on the interface packets forwarded by mirrored traffic; The mirrored traffic is the traffic of interface packets of the monitored device forwarded by the switch to the data security monitoring system; The data security monitoring system includes: A traffic probe is used to monitor mirrored traffic. It parses a first URL from the interface packets of the mirrored traffic, splits the first URL into a second URL and URL parameters, and the second URL points to the monitored data of the monitored device. A data processing engine is configured to receive the second URL and determine the interface asset based on the second URL, including: generating an interface component and a component primary key bound to the interface component based on the domain name of the second URL; wherein the domain names of interface assets corresponding to the same interface component are the same; querying the interface components bound to the component primary key from the interface components already recorded in memory; generating an interface asset primary key based on the second URL and the domain name; querying the interface assets matching the interface component from the interface assets already recorded in memory based on the component primary key; and determining the interface asset matching the interface asset primary key from the interface assets matching the interface component.
8. The system according to claim 7, characterized in that, The data processing engine is also used to perform the method of any one of claims 2-5.
9. The system according to claim 7, characterized in that, The system also includes a filtering module, used to parse the IP of the monitored data from the interface message and discard the second URL whose IP is not within the monitoring range; wherein, the monitoring range is set by the data security monitoring system.